Pooja Aggarwal

dblp:133/3638 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 5 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 ScriptSmith: A Unified LLM Framework for Enhancing IT Operations via Automated Bash Script Generation, Assessment, and Refinement
abstract
In the rapidly evolving landscape of site reliability engineering (SRE), the demand for efficient and effective solutions to manage and resolve issues in site and cloud applications is paramount. This paper presents an innovative approach to action automation using large language models (LLMs) for script generation, assessment, and refinement. By leveraging the capabilities of LLMs, we aim to significantly reduce the human effort involved in writing and debugging scripts, thereby enhancing the productivity of SRE teams. Our experiments focus on Bash scripts, a commonly used tool in SRE, and involve the CodeSift dataset of 100 tasks and the InterCode dataset of 153 tasks. The results show that LLMs can automatically assess and refine scripts efficiently, reducing the need for script validation in an execution environment. Results demonstrate that the framework shows an overall improvement of 7-10% in script generation.
Pooja Aggarwal, Oishik Chatterjee, Suranjana Samanta, Prateeti Mohapatra, Debanjana Kar, Ruchi Mahindru, Steve Barbieri, Eugen Postea, Brad Blancett, Arthur De Magalhaes
AAAI1
2025 ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
abstract
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench.
Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri
ICML23
2024 CodeSift: An LLM-Based Reference-Less Framework for Automatic Code Validation
abstract
The advent of large language models (LLMs) has greatly facilitated code generation, but ensuring the functional correctness of generated code remains a challenge. Traditional validation methods are often time-consuming, error-prone, and impractical for large volumes of code. We introduce CodeSift, a novel framework that leverages LLMs as the first-line filter of code validation without the need for execution, reference code, or human feedback, thereby reducing the validation effort. We assess the effectiveness of our method across three diverse datasets encompassing two programming languages. Our results indicate that CodeSift outperforms state-of-the-art code evaluation methods. Internal testing conducted with subject matter experts reveals that the output generated by CodeSift is in line with human preference, reinforcing its effectiveness as a dependable automated code validation tool.
Pooja Aggarwal, Oishik Chatterjee, Prateeti Mohapatra, Brent Paulovicks, Brad Blancett, Arthur De Magalhaes
CLOUD1
2024 CLA-RA: Collaborative Active Learning Amidst Relabeling Ambiguity
abstract
Obtaining diverse and high-quality labeled data for training efficient classifiers remains a practical challenge. Crowdsourcing, which involves employing multiple weak labelers, is a popular approach to address this issue. However, crowd labelers often introduce noise, inaccuracies, and possess limited domain knowledge. In this paper, we propose a novel framework CLA-RA to optimize the labeling process by determining what to label next and assigning tasks to the most suitable annotators. Our technique aims to optimize classifier efficiency by utilizing the collective wisdom of various annotators while limiting the influence of error-prone annotations. Key contributions include an instance selection mechanism based on annotator disagreement and an instance-dependent annotator confidence model. Experimental results over 9 datasets demonstrate significant improvements over state-of-the-art multi-annotator active learning methods, highlighting the effectiveness of our approach in obtaining high-quality labeled data for training classifiers with minimal labeling costs and errors.
Oishik Chatterjee, Kaizer Rahaman, Pooja Aggarwal
SSE3
2023 Learning Representations on Logs for AIOps
abstract
AI for IT Operations (AIOps) is a powerful platform that Site Reliability Engineers (SREs) use to automate and streamline operational workflows with minimal human intervention. Automated log analysis is a critical task in AIOps as it provides key insights for SREs to identify and address ongoing faults. Tasks such as log format detection, log classification, and log parsing are key components of automated log analysis. Most of these tasks require supervised learning; however, there are multiple challenges due to limited labeled log data and the diverse nature of log data. Large Language Models (LLMs) such as BERT and GPT3 are trained using self-supervision on a vast amount of unlabeled data. These models provide generalized representations that can be effectively used for various downstream tasks with limited labeled data. Motivated by the success of LLMs in specific domains like science and biology, this paper introduces a LLM for log data which is trained on public and proprietary log data. Results of our experiments demonstrate that the proposed LLM outperforms existing models on multiple downstream tasks. In summary, AIOps powered by LLMs offers an efficient and effective solution for automating log analysis tasks and enabling SREs to focus on higher-level tasks. Our proposed LLM, trained on public and proprietary log data, offers superior performance on multiple downstream tasks, making it a valuable addition to the AIOps platform.
Pranjal Gupta, Debanjana Kar, Karan Bhukar, Pooja Aggarwal, Prateeti Mohapatra
CLOUD5
2022 Towards More Effective and Explainable Fault Management Using Cross-Layer Service Topology
abstract
As microservice architecture becomes prominent, existing fault management techniques to deal with service disruption become limiting mainly due to the amount of data needed to be analyzed. This paper emphasizes the need to consider the cross-layer topology of the cloud service to intelligently identify and correlate the observability data and assist in implementing efficient and more accurate fault management techniques that can provide better explainability. Towards this goal, the paper presents a tool that discovers the cross-layer topology for a cloud microservice application and discusses the benefits of using cross-layer service topology to implement effective fault management.
Dhanya R. Mathews, Mudit Verma, J. Lakshmi, Pooja Aggarwal
CLOUD4
2021 Causal Modeling based Fault Localization in Cloud Systems using Golden Signals
abstract
In cloud-native applications, a large fraction of operational failures, known as outages, result in violations of Service Level Objectives (SLOs). SLOs are defined around specific measurable characteristics: availability, throughput, frequency, response time, and quality. Four metrics, latency, traffic, errors, and saturation, ensure coverage for most outages of an application. These are often called golden signals. The dynamicity and complexity of cloud-native applications complicate Site Reliability Engineers’ (SREs) efforts in problem determination, in particular in its fault localization. The fault localization is often a try-and-error process in which SREs rely on their domain knowledge and experience. It is laborious and frequently results in long Mean Time To Resolution (MTTR) for outages. This paper describes a lightweight fault localization system, that establishes causal relationships among the golden signal service errors and error logs, and further leverages PageRank centrality of the derived causal graph for generating a ranked list of faulty microservices.
Pooja Aggarwal, Seema Nagar, Larisa Shwartz, Prateeti Mohapatra, Qing Wang 0016, Amit M. Paradkar, Atri Mandal
CLOUD1
2021 Rightsizing Clusters for Time-Limited Tasks
abstract
Cluster rightsizing facilitates cost-performance trade-off in resource-constrained clouds. Multidimensional bin-packing algorithms can address this rightsizing problem, but these assume that every task on the cluster is always active. In contrast, real-world tasks may be active only during specific time-periods, which allows reusing resources via time sharing and optimal packing. This motivates our generalized problem of rightsizing for time-limited tasks: given a timeline, time-periods and resource demands for tasks, the objective is to place the tasks on a minimum cost cluster of nodes without violating node capacities at any time instance. We design a baseline two-phase algorithm that performs penalty-based mapping of task to node-type and then, solves each node-type independently. We prove that the algorithm has an approximation ratio of O(D. min(m, T)), where D, m and$T$are the number of resources, node-types and timeslots, respectively, We then present an improved linear programming based mapping strategy, enhanced further with a cross-node-type filling mechanism. Our experiments on synthetic and real-world cluster traces show significant cost reduction by LP-based mapping compared to the baseline, and the filling mechanism improves further to produce solutions within 20% of (a lower-bound to) the optimal solution.
Venkatesan T. Chakaravarthy, Padmanabha Venkatagiri Seshadri, Pooja Aggarwal, Anamitra R. Choudhury, Ashok Pon Kumar, Yogish Sabharwal, Amith Singhee
CLOUD3
2021 Insights into Multi-Layered Fault Propagation and Analysis in a Cloud Stack
abstract
Emerging application modernisation efforts are pushing new application services to be built and existing monoliths to be refactored as loosely coupled distributed components (e.g, microservices) for independent scaling and management in cloud. With dynamic operating conditions, component failures, complex component interconnections across the cloud stack, etc., it becomes a challenge to develop effective fault management techniques at the granularity of a multi-layered cloud application service. This paper emphasises on considering faults, errors and failure across the components in different layers of a cloud stack for effective fault management.
Dhanya R. Mathews, Mudit Verma, Pooja Aggarwal, J. Lakshmi
CLOUD3
2020 Question Quality Improvement: Deep Question Understanding for Incident Management in Technical Support Domain
abstract
Technical support domain involves solving problems from user queries through various channels: voice, web and chat, and is both time-consuming and labour intensive. The textual queries in web or chat mode are unstructured and often incomplete. This affects information retrieval and increases the difficulty level for agents to solve it. Such cases require multiple rounds of interaction between user and agent/chatbot in order to better understand the user query. This paper presents a deployed system called Question Quality Improvement (QQI), that aims to improve the quality of user utterance by understanding and extracting important parts of an utterance and gamifying the user interface, prompting them to enter the remaining relevant information. QQI is guided by an ontology designed for the technical support domain and uses co-reference resolution and deep parsing to understand the sentences. Using the syntactics and semantics in the deep parse tree structure various attributes in the ontology are extracted. The system has been in production for over two years supporting around 800 products resulting in a reduction in the time-to-resolve cases by around 29%, leading to huge cost savings. QQI being a core natural language understanding and metadata extraction technology, directly affects more than 8K tickets everyday. These cases are submitted after 50K edits done on the case based on QQI feedback. QQI outputs are used by other technologies such as search and retrieval, case routing for automated dispatch, case-difficulty-prediction, and by the chatbots supported in each product page.
Anupama Ray, Pooja Aggarwal, Csaba Hadhazi, Gargi Dasgupta, Amit M. Paradkar
AAAI2
2019 An attempt at using mass media data to analyze the political economy around some key ICTD policies in India
abstract
Policy making is influenced by a number of factors, including electoral politics, ideological biases of actors involved in the policy making process, and the interlocks between corporate and government entities. This influence is also exercised by shaping public opinion through mass media. In this paper, we study four ICTD policies in India, and explore the political economy around them by using data about how these policies are covered in the mass media. We study which actors are covered more in media, how they speak on the policy issues, and which aspects are given more coverage for these policies. We find that politicians get the highest coverage in mass media regarding discussions on policies, and that the politicians and business-persons often express similar ideologies related to these policies. We also observe that mass media is often biased towards issues related to its middle class reader base with a strong sense of technology driven high-modernism, and negative aspects of these policies and issues faced by the poor due to improper policy implementation are often not given significant coverage. Our key contribution is a methodology of using automated analysis of mass media data to reveal the factors that might be shaping the political economy behind policy making.
Anirban Sen, Priya Chhillar, Pooja Aggarwal, Sravan Verma, Debanjan Ghatak, Priya Kumari, Manpreet Singh Agandh, Aditya Guru, Aaditeshwar Seth
ICTD3
2019 Liquid Cloud Storage
abstract
A liquid system provides durable object storage based on spreading redundantly generated data across a network of hundreds to thousands of potentially unreliable storage nodes. A liquid system uses a combination of a large code , lazy repair , and flow storage organization . We show that a liquid system can be operated to enable flexible and essentially optimal combinations of storage durability, storage overhead, repair bandwidth usage, and access performance.
Michael Luby, Roberto Padovani, Tom Richardson 0001, Lorenz Minder, Pooja Aggarwal
ACM Trans. Storage5
2018 Semantic Parsing for Technical Support Questions
abstract
Technical support problems are very complex. In contrast to regular web queries (that contain few keywords) or factoid questions (which are a few sentences), these problems usually include attributes like a detailed description of what is failing (symptom), steps taken in an effort to remediate the failure (activity), and sometimes a specific request or ask (intent). Automating support is the task of automatically providing answers to these problems given a corpus of solution documents. Traditional approaches to this task rely on information retrieval and are keyword based; looking for keyword overlap between the question and solution documents and ignoring these attributes. We present an approach for semantic parsing of technical questions that uses grammatical structure to extract these attributes as a baseline, and a CRF based model that can improve performance considerably in the presence of annotated data for training. We also demonstrate that combined with reasoning, these attributes help outperform retrieval baselines.
Abhirut Gupta, Anupama Ray, Gargi Dasgupta, Gautam Singh, Pooja Aggarwal, Prateeti Mohapatra
COLING5
2018 Empirical Analysis of the Presence of Power Elite in Media
abstract
Politicians, politically connected business persons, bureaucrats, celebrities, and highly placed government officials (collectively termed as the power elite in sociology literature) can influence national and regional policy for personal and organizational benefit, which may not always be in the best interests of the people. Media is a crucial tool to shape public opinion, and is used heavily by the power elite to bring legitimacy to their policy decisions. In this paper, we empirically analyze the coverage given to the power elite in mainstream media on Demonetization, a significant, recent policy event in India. We compare the extent of coverage given to the elite and non-elite, the policy slant expressed by them, and differences in coverage between seven of the largest news media organizations in India. We find that among the power elite, powerful politicians and political parties are given the maximum coverage in the media, with conspicuous negligence in coverage given to expert opinions. Sentiment analysis clearly reveals that opposing political factions express opposing views towards the policy, and there is variation across different news sources as well. We are applying our methods on other contentious policy events to be able to do a more systematic analysis of how media can aid the power elite to shape public opinion by giving them disproportionately large coverage and visibility.
Anirban Sen, Priya, Pooja Aggarwal, Aditya Guru, Deepak Bansal, I. Mohammed, J. Goyal, K. Mittal, M. Goel, Varuni Madapur, Vipul Khatana, Aaditeshwar Seth
COMPASS3
2018 Domain Knowledge Driven Key Term Extraction for IT Services
Prateeti Mohapatra, Yu Deng 0004, Abhirut Gupta, Gargi Dasgupta, Amit M. Paradkar, Ruchi Mahindru, Daniela Rosu 0001, Shu Tao, Pooja Aggarwal
ICSOC9
2018 Citicafe: An Interactive Interface for Citizen Engagement
abstract
Community engagement is a new and emerging trend in urban cities driven by the mission of developing responsible citizenship. The platform ingests data from different sources, which is exploited by a virtual agent to enable informed interactions. It can help citizens to (a) report problems and (b) gather information related to civic issues for different locations and their neighborhoods. We report the results of a user study carried out to establish the effectiveness of our interface and draw a comparison with an existing platform. A detailed qualitative and quantitative analysis of the survey results shows a definite and statistically significant (p < 0.05) preference for our interface over the existing platform.
Shubham Atreja, Pooja Aggarwal, Prateeti Mohapatra, Amol Dumrewal, Anwesh Basu, Gargi Dasgupta
IUI2
2017 Expander: Lock-Free Cache for a Concurrent Data Structure
abstract
Parallel programming models and paradigms are increasingly becoming more expressive with a steady increase in the number of cores that can be placed on a single chip. Concurrent data structures for shared memory parallel pro- grams are now being used in operating systems, middle-ware, and device drivers. In such a shared memory model, processes communicate and synchronize by applying primitive operations on memory words. To implement concurrent data structures that are linearizable and possibly lock-free or wait-free, it is often necessary to add additional information to memory words in a data structure. This additional information can range from a single bit to multiple bits that typically represent thread ids, request ids, timestamps, and other application dependent fields. Since most processors can perform compare-And-Set (CAS) or load-link/store-conditional (LL/SC) operations on only 64 bits at a time, current approaches either use some bits in a memory word to pack additional information (packing), or use the bits to store a pointer to an object that contains additional information (redirection), and the original data item. The former approach restricts the number of bits for each additional field and this reduces the range of the field, and the latter approach is wasteful in terms of space. We propose a novel and universal method called a memory word expander in this paper. It caches information for a set of memory locations that need to be augmented with additional information. It supports traditional atomic get, set, and CAS operations, and tries to maintain state for a minimum number of entries. We experimentally demonstrate that it is possible to reduce the runtime memory footprint by 20-35% for algorithms that use redirection. For algorithms that use packing, the use of the EXPANDER can make them feasible. The performance overhead is within 2-13% for 32 threads. When we compare the performance of the EXPANDER based non-blocking algorithms with the version that uses locks, we have a performance gain of at least 10-100X.
Pooja Aggarwal, Smruti R. Sarangi
HiPC1
2016 Lock-Free and Wait-Free Slot Scheduling Algorithms
abstract
In this paper, we consider the design space of parallel non-blocking slot scheduling algorithms. Slot schedulers divide time into discrete quanta called slots, and schedule resources at the granularity of slots. They are typically used in high throughput I/O systems, data centers, video servers, and network drivers. We propose a family of parallel slot scheduling problems of increasing complexity, and then propose parallel lock-free and wait-free algorithms to solve them. In specific, we propose problems that can reserve, as well as free a set of contiguous slots in a non-blocking manner. We show that in a system with 64 threads, it is possible to get speedups of 10X by using lock-free algorithms as compared to a baseline implementation that uses locks. We additionally propose wait-free algorithms, whose mean performance is roughly the same as the version with locks. However, they suffer from significantly lower jitter and ensure a high degree of fairness among threads.
Pooja Aggarwal, Smruti R. Sarangi
IEEE Trans. Parallel Distributed Syst.1
2014 RADIR: Lock-free and wait-free bandwidth allocation models for solid state drives
abstract
Novel applications such as micro-blogging and algorithmic trading typically place a very high load on the underlying storage system. They are characterized by a stream of very short requests, and thus they require a very high I/O throughput. The traditional solution for supporting such applications is to use an array of hard disks. With the advent of solid state drives (SSDs), storage vendors are increasingly preferring them because their I/O throughput can scale up to a million IOPS (I/O operations per second). In this paper, we design a family of algorithms, RADIR, to schedule requests for such systems. Our algorithms are lock-free/wait-free, lineariz-able, and take the characteristics of requests into account such as the deadlines, request sizes, dependences, and the amount of available redundancy in RAID configurations. We perform simulations with workloads derived from traces provided by Microsoft and demonstrate a scheduling throughput of 900K IOPS on a 64 thread Intel server. Our algorithms are 2-3 orders of magnitude faster than the versions that use locks. We show detailed results for the effect of deadlines, request sizes, and the effect of RAID levels on the quality of the schedule.
Pooja Aggarwal, Giridhar Yasa, Smruti R. Sarangi
HiPC1
2014 ParTejas: A parallel simulator for multicore processors
abstract
In this paper, we present the design of a novel multicore simulator called ParTejas. It is a fast shared memory based parallel simulator written in Java. Unlike recently released parallel simulators that mainly rely on sampling, high level models, and highly relaxed synchronization, we primarily rely on novel concurrent data structures. In specific, we use a lock free parallel slot scheduler for synchronizing the accesses of multiple threads at a shared resource, and we use flexible barriers known as phasers to relax synchronization within bounds. We leverage additional language specific features of Java, and demonstrate a mean speedup of 11.8X (simulation speed of 4-8 MIPS) with 64 threads for a suite of Splash2 and Parsec benchmarks.
Geetika Malhotra, Pooja Aggarwal, Abhishek Sagar, Smruti R. Sarangi
ISPASS2
2013 Lock-Free and Wait-Free Slot Scheduling Algorithms
abstract
Scalable scheduling is being increasingly regarded as an important requirement in high performance systems. There is a demand for high throughput schedulers in servers, data-centers, networking hardware, large storage systems, and in multi-cores of the future. In this paper, we consider an important subset of schedulers namely slot schedulers that discretize time into quanta called slots. Slot schedulers are commonly used for scheduling jobs in a large number of applications. Current implementations of slot schedulers are either sequential, or use locks. Sadly, lock based synchronization can lead to blocking, and deadlocks, and effectively reduces concurrency. To mitigate these problems, we propose a set of parallel lock-free and wait-free slot scheduling algorithms. Our algorithms are immune to operating system jitter, and guarantee forward progress. Additionally, all our algorithms are linearizable and expose the scheduler's interface as a shared data structure with standard semantics. We empirically demonstrate the scalability of our algorithms for a setup with thousands of requests per second on a 24 thread server. The wait free algorithms are most of the time as fast as the lock-free versions (3X-8X slower in the worst case).
Pooja Aggarwal, Smruti R. Sarangi
IPDPS1