EDBT 2026 Demo / reviewers in the wild / expert
Suryanarayana Reddy Yarrabothula
dblp:408/3910
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0004-6401-9082ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Multi-agent systems · 100% |
Topics — the 1 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems › multi-agent systems engineering
multi-agent evaluation |
1.0 | 1 | 2026 | AssetOpsBench-Live: Privacy-Aware Online Evaluation of Multi-Agent Performance in Industrial Operations · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
containerized execution · 1.0clustering · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AssetOpsBench-Live: Privacy-Aware Online Evaluation of Multi-Agent Performance in Industrial OperationsabstractIndustrial automation increasingly relies on multi-agent AI, yet evaluation remains difficult due to task complexity and data confidentiality. We present AssetOpsBench-Live, a demo of a competition-ready platform for real-time, privacy-preserving evaluation of multi-agent AI in industrial contexts. The platform integrates AssetOpsBench, which measures six dimensions of multi-agent performance and performs automated failure-mode discovery, with Codabench, which supports reproducible, code-oriented competitions. End users first validate agents locally, then submit containerized code for execution on hidden industrial scenarios. Instead of raw trajectories, the system provides quantitative scores and clustered failure modes (e.g., reasoning--action mismatch, step repetition), enabling participants to identify failures, apply targeted improvements, and iteratively resubmit. By combining competition-based engagement with actionable diagnostics, AssetOpsBench-Live delivers reproducible, real-time insights reflecting real-world industrial constraints. Dhaval Patel 0002, Nianjun Zhou, Shuxin Lin, James T. Rayfield, Chathurangi Shyalika, Suryanarayana Reddy Yarrabothula |
AAAI | 6 |