Oishik Chatterjee

dblp:191/1189 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0007-9543-1514ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ScriptSmith: A Unified LLM Framework for Enhancing IT Operations via Automated Bash Script Generation, Assessment, and Refinement
abstract
In the rapidly evolving landscape of site reliability engineering (SRE), the demand for efficient and effective solutions to manage and resolve issues in site and cloud applications is paramount. This paper presents an innovative approach to action automation using large language models (LLMs) for script generation, assessment, and refinement. By leveraging the capabilities of LLMs, we aim to significantly reduce the human effort involved in writing and debugging scripts, thereby enhancing the productivity of SRE teams. Our experiments focus on Bash scripts, a commonly used tool in SRE, and involve the CodeSift dataset of 100 tasks and the InterCode dataset of 153 tasks. The results show that LLMs can automatically assess and refine scripts efficiently, reducing the need for script validation in an execution environment. Results demonstrate that the framework shows an overall improvement of 7-10% in script generation.
Pooja Aggarwal, Oishik Chatterjee, Suranjana Samanta, Prateeti Mohapatra, Debanjana Kar, Ruchi Mahindru, Steve Barbieri, Eugen Postea, Brad Blancett, Arthur De Magalhaes
AAAI2
2025 ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
abstract
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench.
Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri
ICML20
2024 CodeSift: An LLM-Based Reference-Less Framework for Automatic Code Validation
abstract
The advent of large language models (LLMs) has greatly facilitated code generation, but ensuring the functional correctness of generated code remains a challenge. Traditional validation methods are often time-consuming, error-prone, and impractical for large volumes of code. We introduce CodeSift, a novel framework that leverages LLMs as the first-line filter of code validation without the need for execution, reference code, or human feedback, thereby reducing the validation effort. We assess the effectiveness of our method across three diverse datasets encompassing two programming languages. Our results indicate that CodeSift outperforms state-of-the-art code evaluation methods. Internal testing conducted with subject matter experts reveals that the output generated by CodeSift is in line with human preference, reinforcing its effectiveness as a dependable automated code validation tool.
Pooja Aggarwal, Oishik Chatterjee, Prateeti Mohapatra, Brent Paulovicks, Brad Blancett, Arthur De Magalhaes
CLOUD2
2024 CLA-RA: Collaborative Active Learning Amidst Relabeling Ambiguity
abstract
Obtaining diverse and high-quality labeled data for training efficient classifiers remains a practical challenge. Crowdsourcing, which involves employing multiple weak labelers, is a popular approach to address this issue. However, crowd labelers often introduce noise, inaccuracies, and possess limited domain knowledge. In this paper, we propose a novel framework CLA-RA to optimize the labeling process by determining what to label next and assigning tasks to the most suitable annotators. Our technique aims to optimize classifier efficiency by utilizing the collective wisdom of various annotators while limiting the influence of error-prone annotations. Key contributions include an instance selection mechanism based on annotator disagreement and an instance-dependent annotator confidence model. Experimental results over 9 datasets demonstrate significant improvements over state-of-the-art multi-annotator active learning methods, highlighting the effectiveness of our approach in obtaining high-quality labeled data for training classifiers with minimal labeling costs and errors.
Oishik Chatterjee, Kaizer Rahaman, Pooja Aggarwal
SSE1
2024 Efficient Incident Summarization in ITOps: Leveraging Entity-Based Grouping
abstract
An incident which is created due to a fault in an Application Monitoring System, gather large amount of diverse information, which helps in effective remediation of the fault. A Site Reliability Engineer (SRE) should resolve the outage quickly, for which all the fault related information should be presented to her in a crisp and summarized form. In this paper, we address this problem by summarizing an incident and presenting important details to the SRE. We group the list of related events, which is a part of the incident payload, and use Large Language Models (LLMs) to summarize each of these groups separately. The grouping is driven by the entities and symptoms occurring due to the fault, and used to design efficient prompts for LLM. Our approach addresses the known issue of LLM hallucination, and remove any false symptoms or facts from the generated summary. Our proposed method creates a resource-entity driven summarization, giving a bird's eye view of the entire outage in a cost and time efficient way, thus aiding an SRE to understand and resolve the incident at a faster pace.
Suranjana Samanta, Oishik Chatterjee, Hiten Gupta, Prateeti Mohapatra, Arthur De Magalhaes, Ameet Rahane, Marc Palaci-Olgun, Ragu Kattinakere
SSE2
2023 InsightsSumm - Summarization of ITOps Incidents Through In-Context Prompt Engineering
abstract
AI has been extensively used to help Site Reliability Engineers (SREs) to resolve faults in cloud services and applications. It helps to accelerate resolution time by navigating through the vast amount of heterogeneous data (logs, metrics, alerts, etc) related to a fault. A good ITOps system should help SREs by giving precise and meaningful insights for a quick understanding of the data at hand. In this paper, we design a framework to summarize the context or insight present in the heterogeneous data related to a fault. The proposed framework constructs queries/prompts, specific to the ITOps domain, which helps us to generate more insightful abstractive summaries using state-of-the-art text generator models. Initial study on simulated faults shows promising results, which can be expanded to accommodate other datatype, providing summaries for real-world cases.
Suranjana Samanta, Oishik Chatterjee, Neil Boyette, Guangya Liu, Prateeti Mohapatra
CLOUD2
2022 WARM: A Weakly (+Semi) Supervised Math Word Problem Solver
abstract
Solving math word problems (MWPs) is an important and challenging problem in natural language processing. Existing approaches to solving MWPs require full supervision in the form of intermediate equations. However, labeling every MWP with its corresponding equations is a time-consuming and expensive task. In order to address this challenge of equation annotation, we propose a weakly supervised model for solving MWPs by requiring only the final answer as supervision. We approach this problem by first learning to generate the equation using the problem description and the final answer, which we subsequently use to train a supervised MWP solver. We propose and compare various weakly supervised techniques to learn to generate equations directly from the problem description and answer. Through extensive experiments, we demonstrate that without using equations for supervision, our approach achieves accuracy gains of 4.5% and 32% over the current state-of-the-art weakly-supervised approach, on the standard Math23K and AllArith datasets respectively. Additionally, we curate and release new datasets of roughly 10k MWPs each in English and in Hindi (a low-resource language). These datasets are suitable for training weakly supervised models. We also present an extension of our model to semi-supervised learning and present further improvements on results, along with insights.
Oishik Chatterjee, Isha Pandey, Aashish Waikar, Vishwajeet Kumar, Ganesh Ramakrishnan
COLING1
2022 Auto-Query - A simple natural language to SQL query generator for an e-learning platform
abstract
Despite its difficulties, SQL is an essential tool for the users in an educational organisation who need quick and easy access to data to gauge the reception of their learning content by their students and potentially improve their content depending on the insights. To get these insights, the course instructors need real-time access to the database and also need to have relevant SQL knowledge to operate the database to retrieve the required data. The study explores ways to mitigate the difficulties of SQL by developing an application that takes natural language questions that the course instructors have and convert them into SQL queries using a sequence-to-sequence model that show them the data they asked for on a dashboard. The study found that there was a drastic reduction in the time it took for the users of the e-learning platform to get the data from the database without waiting for support from the database administrators. This in turn empowered the educators to study the data and get insights into the reception and working of the course and make suitable changes if necessary which might enhance the user experience for their students.
Parth Parikh, Oishik Chatterjee, Muskan Jain, Aman Harsh, Gaurav Shahani, Rathin Biswas, Kavi Arya
EDUCON2
2020 Robust Data Programming with Precision-guided Labeling Functions
abstract
Scarcity of labeled data is a bottleneck for supervised learning models. A paradigm that has evolved for dealing with this problem is data programming. An existing data programming paradigm allows human supervision to be provided as a set of discrete labeling functions (LF) that output possibly noisy labels to input instances and a generative model for consolidating the weak labels. We enhance and generalize this paradigm by supporting functions that output a continuous score (instead of a hard label) that noisily correlates with labels. We show across five applications that continuous LFs are more natural to program and lead to improved recall. We also show that accuracy of existing generative models is unstable with respect to initialization, training epochs, and learning rates. We give control to the data programmer to guide the training process by providing intuitive quality guides with each LF. We propose an elegant method of incorporating these guides into the generative model. Our overall method, called CAGE, makes the data programming paradigm more reliable than other tricks based on initialization, sign-penalties, or soft-accuracy constraints.
Oishik Chatterjee, Ganesh Ramakrishnan, Sunita Sarawagi
AAAI1
2016 Stability of Consensus Node Orderings Under Imperfect Network Data
abstract
In complex network analysis, the problem of ranking individual nodes based on their importance has attracted increasing attention from the scientific community due to its vast application, such as identification of influential spreaders for viral marketing or epidemic control, bottlenecks for traffic congestion control, and so on. The growing literature proposes a number of measures to determine the rank order of the network entities where complete information about the nodes and their interaction is available. Degree centrality, PageRank, eigenvector centrality, closeness centrality are few such popular measures. In most real-life scenarios, however, the information about the underlying network is incomplete or affected due to noise. The few works that study the effects of incomplete information on the rank orders show the vulnerability of the rank orders in various topologies. In this paper, we investigate the effects of noise, both random and nonrandom, on the aggregated rank orders determined from the degree, PageRank, eigenvector centrality, and closeness centrality-based rankings. This paper reveals an important insight that even the simple Borda Count ranking has the potential to improve on the accuracy of rank orders in networks with uncertainty. This paper shows the existence of stable nodes in various networks and indicates that the design of the consensus approach based on the properties of the stable nodes can further improve the stability of the rank orders.
Srinka Basu, Ujjwal Maulik, Oishik Chatterjee
IEEE Trans. Comput. Soc. Syst.3