Mohsen Seyedkazemi Ardebili

dblp:282/6179 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-1166-6559ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum
abstract
This paper presents the DECICE project (Device Edge Cloud Intelligent Collaboration framEwork), a Horizon Europe Research and Innovation Action (Grant No. 101092582, December 2022 to November 2025) that developed an open-source framework for intelligent workload scheduling across the cloud-HPC-edge compute continuum. A consortium of 12 partners across 6 European countries organized the work into six work packages covering AI-driven scheduling, digital twin infrastructure, system architecture and integration, monitoring, use case validation, and dissemination. The two core technical contributions are an Integrated AI Scheduler (IAIS) employing RNN-based prediction and formal workflow modeling for constraint-aware workload mapping, and a Digital Twin aggregating real-time metrics with carbon intensity and anomaly prediction for energy-aware scheduling. The framework operates within Kubernetes environments, supports unified workflow ingestion from multiple formats, and bridges cloud-native and HPC orchestration through a Slurm integration layer. We present the project vision, the overall architecture, contributions from each work package, quantitative evaluation results, and the open-source release.
Aasish Kumar Sharma, Felix Stein, Mirac Aydin, Michael Bidollahkhani, Sachin P. Nanavati, Mohsen Seyedkazemi Ardebili, Giorgi Mamulashvili, Mojtaba Akbari, Jonathan Decker, Zoya Masih, Julian M. Kunkel
COMPSAC6
2026 Elevating Datacenter Resilience with ThermADNet: A Thermal Anomaly Detection System
abstract
In the era of digital transformation, datacenters and High Performance Computing (HPC) Systems have emerged as the backbone of global technology infrastructure, powering essential services across various industries, including finance and healthcare. Therefore, ensuring the uninterrupted service of these datacenters has become a critical challenge. Thermal anomalies pose a significant risk to datacenter operation, potentially leading to hardware deterioration, system downtime, and catastrophic failures. This threat is exacerbated by the growing number of datacenters, increased power density, and heat waves fostered by global warming. Detecting thermal anomalies in datacenters involves several challenges. Large-scale data collection is difficult, requiring diverse monitoring signals from thousands of nodes over long periods. The absence of labeled data complicates the identification of normal and abnormal states. Establishing accurate classification thresholds to minimize false positives and negatives is another significant hurdle. Traditional statistical methods often fail to capture temporal dependencies and complex correlations in monitoring signals. Additionally, finding anomalies at both the system and subsystem levels adds to the complexity. Deploying machine learning models in production environments presents technical and operational challenges, making real-time anomaly detection a demanding task. This paper introduces ThermADNet, a Thermal Anomaly Detection framework that combines statistical rules-based methods with Deep Neural Network (DNN) techniques for thermal anomaly detection in datacenters. ThermADNet utilizes a semi-supervised learning approach by training on a ”semi-normal” dataset, addressing the challenges of large-scale data collection, semi-normal dataset identification, and classification threshold establishment. This framework’s efficacy is validated by its success in identifying real physical thermal failure events within a Tier-0 datacenter, pinpointing anomalies at both the system and subsystem levels, including compute nodes and datacenter infrastructure. In the critical evaluation window covering the July 28 failure, ThermADNet achieves precision and recall up to 0.97, with F1-scores as high as 0.97. By providing detailed information about anomalies, the framework clarifies the characteristics and reasoning behind the DNN outputs, thereby building trust in the AI model and ensuring that users can understand and rely on the system’s decisions. By offering a sophisticated method for thermal anomaly detection, ThermADNet significantly contributes to enhancing datacenter reliability and efficiency. This advancement supports the uninterrupted operation of critical HPC systems, averting considerable economic and societal losses.
Mohsen Seyedkazemi Ardebili, Andrea Acquaviva, Luca Benini, Andrea Bartolini
Future Gener. Comput. Syst.1
2026 KubeIntellect: A Modular LLM-Orchestrated Agent Framework for End-to-End Kubernetes Management
abstract
Abstract Kubernetes has become the foundation of modern cloud-native infrastructure, yet its operational complexity remains a persistent barrier. Administrators must navigate a vast API surface, manage heterogeneous workloads, and coordinate tasks across disconnected tools—often requiring precise commands, declarative configuration files, and deep domain expertise. This paper presents KubeIntellect , a Large Language Model (LLM)-powered system for end-to-end Kubernetes management through natural language. KubeIntellect spans all major categories of Kubernetes operations—read, write, delete, exec, access control, and lifecycle management—through a supervisor-coordinated set of domain-specialized agents, with human-in-the-loop (HITL) confirmation on all mutating operations. Operations outside the static tool set are handled by the Code Generator Agent, which synthesizes, validates, and registers new Kubernetes tools at runtime. The Code Generator Agent executes synthesized tools in an in-process Python REPL rather than a separate process or container; process-level isolation between synthesized code and the host runtime is therefore not enforced, and a defense-in-depth model comprising static analysis, API-call validation, and mandatory human-in-the-loop review constitutes the primary mitigation. Migration to pod-level isolation is a planned hardening step. Evaluation on a live four-node Kubernetes cluster (170 pods across 18 namespaces) shows: a 75% pass rate (12/16; 95% CI: 51%–91%; mean rubric score 31.2/40) on a 16-scenario controlled fault-injection corpus scored on an 8-dimension LLM-judge rubric; a +25 percentage-point improvement over a tool-less GPT-4o baseline on the same scenarios (75% vs. 50%); a 93% query resolution rate (186/200) with an 81.8% synthesis success rate (63/77 novel tool requests) on a 200-query operational corpus; and end-to-end latency in the 7–10 s range at a mean API cost of $0.036/query for read-only workloads and $0.039/query overall. A reproducible demo environment is available on a public managed-Kubernetes service, with a local single-node option for readers without cloud access. These results demonstrate that domain-specific multi-agent orchestration, structured HITL confirmation, and runtime tool synthesis together yield substantially higher task completion than general-purpose LLM reasoning on Kubernetes operations.
Mohsen Seyedkazemi Ardebili, Andrea Bartolini
J. Grid Comput.1
2024 HazardNet: A thermal hazard prediction framework for datacenters
Mohsen Seyedkazemi Ardebili, Andrea Acquaviva, Luca Benini, Andrea Bartolini
Future Gener. Comput. Syst.1
2024 GRAAFE: GRaph Anomaly Anticipation Framework for Exascale HPC systems
Martin Molan, Mohsen Seyedkazemi Ardebili, Junaid Ahmed Khan, Francesco Beneventi, Daniele Cesarini, Andrea Borghesi, Andrea Bartolini
Future Gener. Comput. Syst.2
2022 Multi-level anomaly prediction in Tier-0 datacenter: a deep learning approach
abstract
Modern scientific discoveries are driven by an unsatisfiable demand for computational resources. To solve large problems in science, engineering, and business, data centers provide High-Performance Computing (HPC) systems with aggregation of the computing capacity of thousand of computing nodes. Anomaly prediction is critical in order to preserve the continuity of the service of HPC systems and prevent hardware deterioration. In the datacenter, a thermal anomaly occurs when the balance of cooling capacity and computational demand is disturbed. Moreover, this is identifiable from a suspicious/abnormal pattern in the monitoring signals.
Mohsen Seyedkazemi Ardebili, Andrea Bartolini, Luca Benini
CF1
2021 Prediction of Thermal Hazards in a Real Datacenter Room Using Temporal Convolutional Networks
abstract
Datacenters play a vital role in today's society. At large, a datacenter room is a complex controlled environment composed of thousands of computing nodes, which consume kW of power. To dissipate the power, forced air/liquid flow is employed, with a cost of millions of euros per year. Reducing this cost involves using free-cooling and average case design, which can create a cooling shortage and thermal hazards. When a thermal hazard happens, the system administrators and the facility manager must stop the production to avoid IT equipment damage and wear-out. In this paper, we study the thermal hazards signatures on a Tier-0 datacenter room's monitored data during a full year of production. We define a set of rules for detecting the thermal hazards based on the inlet and outlet temperature of all nodes of a room. We then propose a custom Temporal Convolutional Network (TCN) to predict the hazards in advance. The results show that our TCN can predict the thermal hazards with an Fl-score of 0.98 for a randomly sampled test set. When causality is enforced between the training and validation set the F1-score drops to 0.74, demanding for an in-place online re-training of the network, which motivates further research in this context.
Mohsen Seyedkazemi Ardebili, Marcello Zanghieri, Alessio Burrello, Francesco Beneventi, Andrea Acquaviva, Luca Benini, Andrea Bartolini
DATE1