Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Bhavya

dblp:245/8864 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 59% Cloud and datacenter computing · 41%
Artificial intelligence
2 papers
Reinforcement learning · 36% Information extraction and text analysis · 27% Language models and text generation · 27%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
agent evaluation
0.912025
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks · ICML 2025
Distributed systems › fault tolerance
fault detection and diagnosis
0.912025
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025
Distributed systems
fault tolerance
0.912025
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › lexical semantics
analogy detection
0.712023
CAM: A Large Language Model-based Creative Analogy Mining Framework · WWW 2023
Natural language and speech › Language models and text generation › text generation › large language model generation
prompt-based generation
0.712023
CAM: A Large Language Model-based Creative Analogy Mining Framework · WWW 2023
Machine learning › Trustworthy machine learning
AI safety
0.312025
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks · ICML 2025
Cloud and datacenter computing › datacenter operations
AIOps
0.312025
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025
Cloud and datacenter computing
cloud service management
0.312025
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

benchmarking · 1.7state machine · 0.9multi-agent system · 0.9large language model · 0.9scoring function · 0.7prompt engineering · 0.7large pre-trained language model · 0.7
YearPublicationVenuePosition
2025 ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
abstract
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench.
Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri
ICML7
2025 STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
abstract
In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, as existing human-in-the-loop practices can hardly keep up with the scale of modern clouds. This paper presents STRATUS, an LLM-based multi-agent system for realizing autonomous Site Reliability Engineering (SRE) of cloud services. STRATUS consists of multiple specialized agents (e.g., for failure detection, diagnosis, mitigation), organized in a state machine to assist system-level safety reasoning and enforcement. We formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iteration. We show that TNR can effectively improve autonomous failure mitigation. STRATUS significantly outperforms state-of-the-art SRE agents in terms of success rate of failure mitigation problems in AIOpsLab and ITBench (two SRE benchmark suites), by at least 1.5 times across various models. STRATUS shows a promising path toward practical deployment of agentic systems for cloud reliability.
Yinfang Chen, Jackson Clark, Yiming Su, Noah Zheutlin, Bhavya, Rohan R. Arora, Yu Deng 0004, Saurabh Jha, Tianyin Xu
NeurIPS6
2024 AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction
abstract
Textual analogies that make comparisons between two concepts are often used for explaining complex ideas, creative writing, and scientific discovery.In this paper, we propose and study a new task, called Analogy Detection and Extraction (AnaDE), which includes three synergistic sub-tasks: 1) detecting documents containing analogies, 2) extracting text segments that make up the analogy, and 3) identifying the source and target concepts being compared.To facilitate the study of this new task, we create a benchmark dataset by scraping Metamia.com and investigate the performances of state-of-the-art models on all sub-tasks to establish the first-generation benchmark results for this new task.We find that the Longformer model achieves the best performance on all three sub-tasks demonstrating its effectiveness for handling long texts.Moreover, smaller models fine-tuned on our dataset perform better than non-fine-tuned ChatGPT, suggesting high task difficulty.Overall, the models achieve a high performance on document detection suggesting that it could be used to develop applications like analogy search engines.Further, there is a large room for improvement on the segment and concept extraction tasks 1 .
Bhavya, Shradha Sehgal, Jinjun Xiong, ChengXiang Zhai
EACL (1)1
2023 CAM: A Large Language Model-based Creative Analogy Mining Framework
abstract
Analogies inspire creative solutions to problems, and facilitate the creative expression of ideas and the explanation of complex concepts. They have widespread applications in scientific innovation, creative writing, and education. The ability to discover creative analogies that are not explicitly mentioned but can be inferred from the web is highly desirable to power all such applications dynamically and augment human creativity. Recently, Large Pre-trained Language Models (PLMs), trained on massive Web data, have shown great promise in generating mostly known analogies that are explicitly mentioned on the Web. However, it is unclear how they could be leveraged for mining creative analogies not explicitly mentioned on the Web. We address this challenge and propose Creative Analogy Mining (CAM), a novel framework for mining creative analogies, which consists of the following three main steps: 1) Generate analogies using PLMs with effectively designed prompts, 2) Evaluate their quality using scoring functions, and 3) Refine the low-quality analogies by another round of prompt-based generation. We propose both unsupervised and supervised instantiations of the framework so that it can be used even without any annotated data. Based on human evaluation using Amazon Mechanical Turk, we find that our unsupervised framework can mine 13.7% highly-creative and 56.37% somewhat-creative analogies. Moreover, our supervised scores are generally better than the unsupervised ones and correlate moderately with human evaluators, indicating that they would be even more effective at mining creative analogies. These findings also shed light on the creativity of PLMs 1.
Bhavya, Jinjun Xiong, ChengXiang Zhai
WWW1
2021 Scaling Up Data Science Course Projects: A Case Study
abstract
Large-scale, online Data Science (DS) courses and degree programs are becoming increasingly common due to the global rise in popularity and demand for data scientists. Although project-based learning is integral to gaining hands-on experience in DS education, providing fair, timely, and high-quality feedback on varied projects for a large number of diverse students is challenging. To address those challenges in scaling up the assessment of DS group projects, we integrated multiple techniques, such as rapid feedback, peer grading, graders as meta-reviewers, etc. We present a case study of deploying those strategies for group projects in a large online DS course titled Text Information Systems offered in Fall, 2020. We synthesize our findings from analyzing student and grader survey responses, and share useful lessons and future work.
Bhavya, Jinfeng Xiao, ChengXiang Zhai
L@S1
2020 Explanation Mining
abstract
Explanations are used to provide an understanding of a concept, procedure, or reasoning to others. Although explanations are present online ubiquitously within textbooks, discussion forums, and many more, there is no way to mine them automatically to assist learners in seeking an explanation. To address this problem, we propose the task of Explanation Mining. To mine explanations of educational concepts, we propose a baseline approach based on the Language Modeling approach of information retrieval. Preliminary results suggest that incorporating knowledge from a model trained on the ELI5 (Explain Like I'm Five) dataset in the form of a document prior helps increase the performance of a standard retrieval model. This is encouraging because our method requires minimal in-domain supervision, as a result, it can be deployed for multiple online courses. We also suggest some interesting future work in the computational analysis of explanations.
Bhavya, ChengXiang Zhai
L@S1
2020 Collective Development of Large Scale Data Science Products via Modularized Assignments: An Experience Report
abstract
Many universities are offering data science (DS) courses to fulfill the growing demands for skilled DS practitioners. Assignments and projects are essential parts of the DS curriculum as they enable students to gain hands-on experience in real-world DS tasks. However, most current assignments and projects are lacking in at least one of two ways: 1) they do not comprehensively teach all the steps involved in the complete workflow of DS projects; 2) students work on separate problems individually or in small teams, limiting the scale and impact of their solutions. To overcome these limitations, we envision novel synergistic modular assignments where a large number of students work collectively on all the tasks required to develop a large-scale DS product. The resulting product can be continuously improved with students' contributions every semester. We report our experience with developing and deploying such an assignment in an Information Retrieval course. Through the assignment, students collectively developed a search engine for finding expert faculty specializing in a given field. This shows the utility of such assignments both for teaching useful DS skills and driving innovation and research. We share useful lessons for other instructors to adopt similar assignments for their DS courses.
Bhavya, Assma Boughoula, Aaron Green 0002, ChengXiang Zhai
SIGCSE1
2019 Web of Slides: Automatic Linking of Lecture Slides to Facilitate Navigation
abstract
Lecture slides covering many topics are becoming increasingly available online, but they are scattered, making it a challenge for anyone to instantly access all slides relevant to a learning context. To address this challenge, we propose to create links between those scattered slides to form a Web of Slides (WOS). Using the sequential nature of slides, we present preliminary results of studying how to automatically create a basic link based on similarity of slides as an initial step toward the vision of WOS. We also explore interesting future research directions using different link types and the unique features of slides.
Sahiti Labhishetty, Bhavya, Kevin Pei, Assma Boughoula, ChengXiang Zhai
L@S2
2019 WOSView Demo: A Tool to Explore the Web of Slides
abstract
We will demonstrate a prototype system WOSView built based on the vision of the Web of Slides(WOS), which aims to link all the lectures slides so as to facilitate navigation over all the slides. The links can be created at the slide level or at the level of phrases inside a slide, and many types of links can be created. The prototype system we built implements the most basic type of links, which link slides that have similar content and integrates lectures from four different MOOCs. WOSView also supports keyword search, which generates virtual links dynamically. We will demonstrate how the graphical interface of the WOSView enables students to flexibly navigate into slides from different courses and explore related slides using both static and dynamic links and solicit feedback from the community about the vision of WOS.
Sahiti Labhishetty, Bhavya, Kevin Pei, Assma Boughoula, ChengXiang Zhai
L@S2