EDBT 2026 Demo / reviewers in the wild / expert
Mudit Verma
dblp:192/7474
· DBLP profile ↗
15ranked-venue papers
1as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksabstractRealizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench. Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri |
ICML | 8 |
| 2024 | Intent-Driven Multi-Engine Observability Dataflows for Heterogeneous Geo-Distributed CloudsabstractWith the growth of multi-cloud computing across a heterogeneous substrate of public cloud, edge, and on-premise sites, observability has been gaining importance in compre-hending the state of availability and performance of large-scale geo-distributed systems. Collecting, processing and analyzing observability data from multiple geo-distributed clouds can be naturally modelled as dataflows comprising chained functions. These observability flows pose a unique set of challenges in-cluding ($a$) keeping cost budgets, resource overheads, network bandwidth consumed and latency low, (b) scaling to a large number of clusters, (c) adapting the volume of observability data to satisfy resource constraints and service level objectives (d) supporting diverse engines per dataflow depending on each processing function of the flow and (e) automating and optimizing placement of observability processing functions including closed-loop orchestration. Towards this end, we propose Octopus, a multi-cloud multi-engine observability processing framework. In Octopus, declarative observability dataflows (DODs) serve as an intent-driven abstraction for site reliability engineers (SREs) to specify self-driven observability dataflows. A dataflow engine in Octopus, then orchestrates these DODs to automatically deploy and self-manage observability dataflows over large fabrics spanning multiple clouds and clusters. Octopus supports a mix of streaming and batch functions and supports pluggable run-time engines, thereby enabling flexible composition of multi-engine observability flows. Our early deployment experience with Octopus is promising. We have successfully deployed production-grade metrics analysis and log processing data flows in an objective-optimized fashion across 1 cloud and 10 edge clusters spanning continents. Our results indicate data volume and WAN bandwidth savings of$2.3\mathrm{x}$and 56 %, respectively, the ability to support auto-scaling of DODs as input load varies, and the ability to flexibly relocate functions across clusters without hurting latency targets. Aishwariya Chakraborty, Anand Eswaran, Pankaj Thorat, Mudit Verma, Pranjal Gupta, Praveen Jayachandran |
CLOUD | 4 |
| 2024 | Self Adjusting Log Observability for Cloud Native ApplicationsabstractWith the increasing complexity of modern applications, particularly those relying on microservices architectures, the volume of observability data, encompassing logs, metrics, traces, etc., has surged significantly. This is further exacerbated by extensive cloud deployments, where observability is crucial for comprehending the health and performance of these systems, leading operations teams to collect as much data as possible for the “fear of missing out”. However, the collection, storage, and analysis of observability data entail significant costs, both in terms of resources and finances. Specifically, logs comprise the most substantial portion of observability data volume, thus exerting the greatest impact on observability cost. Moreover, logs also exhibit unstructured and noisy characteristics, where the efficacy of downstream AI for IT operations (AIOps tasks or Day-2 operations), such as fault classification, fault diagnosis, log anomaly detection etc., can be negatively impacted by log data volume. Hence, striking a balance between the verbosity of log observability and its impact on day-2 operations and debuggability is essential. In this paper, we introduce an autonomous system named SALO, which stands for Self-Adjusting Log Observability. SALO selectively collects logs based on real-time necessity, location, and granularity, as opposed to the conventional practice of collecting indiscriminately from all the components continuously. Our experiments show that SALO drastically decreases the log volume, by as much as 95 %, while still maintaining data quality for downstream AIOps usage, especially for post-hoc diagnosis tasks. Operating on a reduced volume of log data not only decreases storage, transfer, and retention costs but also streamlines observability pipelines, making them leaner, more efficient, and less resource-hungry. Divya Pathak, Mudit Verma, Aishwariya Chakraborty |
CLOUD | 2 |
| 2024 | Hindsight PRIORs for Reward Learning from Human PreferencesabstractPreference based Reinforcement Learning (PbRL) removes the need to hand specify a reward function by learning one from preference feedback over policy behaviors. Current approaches to PbRL do not address the credit assignment problem inherent in determining which parts of a behavior most contributed to a preference resulting in data intensive approaches and subpar reward models. We address such limitations by introducing a credit assignment strategy (PRIOR) that uses a forward dynamics world model to approximate state importance within a trajectory and then guides rewards to be proportional to state importance through an auxiliary predicted return redistribution objective. Incorporating state importance into reward learning improves the speed of policy learning, overall policy performance, and reward recovery on both locomotion and manipulation tasks. For example, PRIOR achieves 80% success rate with half the amount of data compared to baselines. The performance gains and our ablations demonstrate the benefits even a simple credit assignment strategy can have on reward learning and that state importance in forward dynamics prediction is a strong proxy for a state's contribution to a preference decision. Mudit Verma, Katherine Metcalf |
ICLR | 1 |
| 2024 | Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo FrameworksabstractWe argue that auto-regressive LLMs cannot, by themselves, do planning or self-verification (which is after all a form of reasoning), and shed some light on the reasons for misunderstandings in the literature. We will also argue that LLMs should be viewed as universal approximate knowledge sources that have much more meaningful roles to play in planning/reasoning tasks beyond simple front-end/back-end format translators. We present a vision of LLM-Modulo Frameworks that combine the strengths of LLMs with external model-based verifiers in a tighter bi-directional interaction regime. We will show how the models driving the external verifiers themselves can be acquired with the help of LLMs. We will also argue that rather than simply pipelining LLMs and symbolic components, this LLM-Modulo Framework provides a better neuro-symbolic approach that offers tighter integration between LLMs and symbolic components, and allows extending the scope of model-based planning/reasoning regimes towards more flexible knowledge, problem and preference specifications. Subbarao Kambhampati, Karthik Valmeekam, Lin Guan 0003, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Saldyt, Anil Murthy |
ICML | 4 |
| 2023 | Trust-Aware Planning: Modeling Trust Evolution in Iterated Human-Robot InteractionabstractTrust between team members is an essential requirement for any successful cooperation. Thus, engendering and maintaining the fellow team members' trust becomes a central responsibility for any member trying to not only successfully participate in the task but to ensure the team achieves its goals. The problem of trust management is particularly challenging in mixed human-robot teams where the human and the robot may have different models about the task at hand and thus may have different expectations regarding the current course of action, thereby forcing the robot to focus on the costly explicable behavior. We propose a computational model for capturing and modulating trust in such iterated human-robot interaction settings, where the human adopts a supervisory role. In our model, the robot integrates human's trust and their expectations about the robot into its planning process to build and maintain trust over the interaction horizon. By establishing the required level of trust, the robot can focus on maximizing the team goal by eschewing explicit explanatory or explicable behavior without worrying about the human supervisor monitoring and intervening to stop behaviors they may not necessarily understand. We model this reasoning about trust levels as a meta reasoning process over individual planning tasks. We additionally validate our model through a human subject experiment. Zahra Zahedi, Mudit Verma, Sarath Sreedharan, Subbarao Kambhampati |
HRI | 2 |
| 2022 | Network Aware Container Orchestration for Telco WorkloadsabstractIn recent years, with the maturation of container orchestration platforms like Kubernetes, containers are now becoming the default way to deploy cloud-native applications, designed as microservices, on public and private clouds. These trends have also spread to the field of Telecommunications, boosted by the onset of 5G. Network functions processing millions of packets per second, earlier run as proprietary physical boxes, are now being realized as disaggregated container based microservices (CNFs) running on commodity clusters managed by orchestrators, like Kubernetes, on Telco clouds. While container orchestrators have evolved to meet the needs of enterprise applications, Telco workloads still remain a second class citizen, as the orchestrator is presently unaware of the networking needs of CNFs and cannot guarantee QoS of network intensive functions. In this work, we examine orchestration of network sensitive functions and identify the key networking requirements of containerized Telco workloads from the orchestration platform. We design and propose NACO - Network Aware Container Orchestration, a minimal, cloud-native and scalable extension to the Kubernetes platform to address these requirements and provide first class lifecycle management of CNFs used in Telco workloads. We implement a prototype of the system and demonstrate that we can achieve network aware container orchestration with minimal operation times. Kavya Govindarajan, Chander Govindarajan, Mudit Verma |
CLOUD | 3 |
| 2022 | Towards More Effective and Explainable Fault Management Using Cross-Layer Service TopologyabstractAs microservice architecture becomes prominent, existing fault management techniques to deal with service disruption become limiting mainly due to the amount of data needed to be analyzed. This paper emphasizes the need to consider the cross-layer topology of the cloud service to intelligently identify and correlate the observability data and assist in implementing efficient and more accurate fault management techniques that can provide better explainability. Towards this goal, the paper presents a tool that discovers the cross-layer topology for a cloud microservice application and discusses the benefits of using cross-layer service topology to implement effective fault management. Dhanya R. Mathews, Mudit Verma, J. Lakshmi, Pooja Aggarwal |
CLOUD | 2 |
| 2022 | A Guided Approach Towards Complex Chaos Selection, Prioritisation and InjectionabstractThough Chaos Engineering is a popular method to test reliability and performance assurance, available tools can only inject random or manually curated faults into a target system. Given the vast array of faults that can be injected, it is crucial to a.) intelligently pick the faults that can have tangible effects, b.) increase the test coverage, and c.) reduce the overall time needed to assess the reliability of a system under adverse conditions. To the effect, we are proposing to learn from past major outages and use genetic algorithm-based meta-heuristics to evolve complex fault injections. Ojaswa Sharma, Mudit Verma, Saumya Bhadauria, Praveen Jayachandran |
CLOUD | 2 |
| 2022 | Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI SystemsabstractDespite the surprising power of many modern AI systems that often learn their own representations, there is significant discontent about their inscrutability and the attendant problems in their ability to interact with humans. While alternatives such as neuro-symbolic approaches have been proposed, there is a lack of consensus on what they are about. There are often two independent motivations (i) symbols as a lingua franca for human-AI interaction and (ii) symbols as (system-produced) abstractions use in its internal reasoning. The jury is still out on whether AI systems will need to use symbols in their internal reasoning to achieve general intelligence capabilities. Whatever the answer there is, the need for (human-understandable) symbols in human-AI interaction seems quite compelling. Symbols, like emotions, may well not be sine qua non for intelligence per se, but they will be crucial for AI systems to interact with us humans--as we can neither turn off our emotions not get by without our symbols. In particular, in many human-designed domains, humans would be interested in providing explicit (symbolic) knowledge and advice--and expect machine explanations in kind. This alone requires AI systems to at least do their I/O in symbolic terms. In this blue sky paper, we argue this point of view, and discuss research directions that need to be pursued to allow for this type of human-AI interaction. Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, Lin Guan 0003 |
AAAI | 3 |
| 2022 | Modeling the Interplay between Human Trust and MonitoringabstractIn this work, we investigate and model how human trust affects monitoring. We present a web-based human subject study in which the robot is a worker and the human plays the role of a supervisor. First, we evaluate the correlation between the human trust and monitoring by using statistical tests, and then we learn probabilistic models of the behavioral data collected through our user studies. These models can provide us with the likelihood of a human user monitoring a system given their level of trust. Such models can be leveraged in many systems including the ones designed to be resilient to automation bias and complacency. Zahra Zahedi, Sarath Sreedharan, Mudit Verma, Subbarao Kambhampati |
HRI | 3 |
| 2022 | Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations
Sarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava 0001, Subbarao Kambhampati |
ICLR | 3 |
| 2021 | Insights into Multi-Layered Fault Propagation and Analysis in a Cloud StackabstractEmerging application modernisation efforts are pushing new application services to be built and existing monoliths to be refactored as loosely coupled distributed components (e.g, microservices) for independent scaling and management in cloud. With dynamic operating conditions, component failures, complex component interconnections across the cloud stack, etc., it becomes a challenge to develop effective fault management techniques at the granularity of a multi-layered cloud application service. This paper emphasises on considering faults, errors and failure across the components in different layers of a cloud stack for effective fault management. Dhanya R. Mathews, Mudit Verma, Pooja Aggarwal, J. Lakshmi |
CLOUD | 2 |
| 2021 | Konveyor Move2Kube: Automated Replatforming of Applications to KubernetesabstractWe present Move2Kube, a replatforming framework that automates the transformation of the deployment specification and development pipeline of an application from a non-Kubernetes platform to a Kubernetes-based one, minimizing changes to the application's functional implementation and architecture. Our contributions include: (1) a standardized intermediate representation to which diverse application deployment artifacts could be translated, (2) an extension framework for adding support for new source platforms, and target artifacts while allowing customization as per organizational standards. We provide initial evidence of its effectiveness in terms of effort reduction, and highlight the current research challenges and future lines of work. Move2Kube is being developed as an open source community project and it is available at https://move2kube.konveyor.io/ Padmanabha Venkatagiri Seshadri, Harikrishnan Balagopal, Pablo Loyola, Akash Nayak, Chander Govindarajan, Mudit Verma, Ashok Pon Kumar, Amith Singhee |
CLOUD | 6 |
| 2021 | Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data AugmentationabstractHuman explanation (e.g., in terms of feature importance) has been recently used to extend the communication channel between human and agent in interactive machine learning. Under this setting, human trainers provide not only the ground truth but also some form of explanation. However, this kind of human guidance was only investigated in supervised learning tasks, and it remains unclear how to best incorporate this type of human knowledge into deep reinforcement learning. In this paper, we present the first study of using human visual explanations in human-in-the-loop reinforcement learning (HIRL). We focus on the task of learning from feedback, in which the human trainer not only gives binary evaluative "good" or "bad" feedback for queried state-action pairs, but also provides a visual explanation by annotating relevant features in images. We propose EXPAND (EXPlanation AugmeNted feeDback) to encourage the model to encode task-relevant features through a context-aware data augmentation that only perturbs irrelevant features in human salient information. We choose five tasks, namely Pixel-Taxi and four Atari games, to evaluate the performance and sample efficiency of this approach. We show that our method significantly outperforms methods leveraging human explanation that are adapted from supervised learning, and Human-in-the-loop RL baselines that only utilize evaluative feedback. Lin Guan 0003, Mudit Verma, Sihang Guo, Subbarao Kambhampati |
NeurIPS | 2 |