EDBT 2026 Demo / reviewers in the wild / expert
Jackson Clark
dblp:399/7175
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 59% Cloud and datacenter computing · 41% | |
| Artificial intelligence
1 paper |
Reinforcement learning · 77% Trustworthy machine learning · 23% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
agent evaluation |
0.9 | 1 | 2025 | ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks · ICML 2025 |
Distributed systems › fault tolerance
fault detection and diagnosis |
0.9 | 1 | 2025 | STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025 |
Distributed systems
fault tolerance |
0.9 | 1 | 2025 | STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
AI safety |
0.3 | 1 | 2025 | ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks · ICML 2025 |
Cloud and datacenter computing › datacenter operations
AIOps |
0.3 | 1 | 2025 | STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025 |
Cloud and datacenter computing
cloud service management |
0.3 | 1 | 2025 | STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 1.7state machine · 0.9multi-agent system · 0.9large language model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksabstractRealizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench. Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri |
ICML | 6 |
| 2025 | STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsabstractIn cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, as existing human-in-the-loop practices can hardly keep up with the scale of modern clouds. This paper presents STRATUS, an LLM-based multi-agent system for realizing autonomous Site Reliability Engineering (SRE) of cloud services. STRATUS consists of multiple specialized agents (e.g., for failure detection, diagnosis, mitigation), organized in a state machine to assist system-level safety reasoning and enforcement. We formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iteration. We show that TNR can effectively improve autonomous failure mitigation. STRATUS significantly outperforms state-of-the-art SRE agents in terms of success rate of failure mitigation problems in AIOpsLab and ITBench (two SRE benchmark suites), by at least 1.5 times across various models. STRATUS shows a promising path toward practical deployment of agentic systems for cloud reliability. Yinfang Chen, Jackson Clark, Yiming Su, Noah Zheutlin, Bhavya, Rohan R. Arora, Yu Deng 0004, Saurabh Jha, Tianyin Xu |
NeurIPS | 3 |