Khushbu Agarwal

dblp:72/8323 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-1892-2439ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Graph learning · 44% Language models and text generation · 24% Planning, search and constraint satisfaction · 12%
Databases, data mining, and information retrieval
3 papers
Data mining · 51% Knowledge graphs · 39% Web and social media mining · 10%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 62% Bioinformatics and computational biology · 38%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search
0.812024
CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback · ICML 2024
Natural language and speech › Language models and text generation
language-model-guided search
0.812024
CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback · ICML 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback · ICML 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
scientific discovery
0.812024
CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback · ICML 2024
Machine learning › Graph learning
graph neural network
0.712023
A Unification Framework for Euclidean and Hyperbolic Graph Neural Networks · IJCAI 2023
Machine learning › Graph learning › graph neural network › geometric graph neural network
hyperbolic graph neural network
0.712023
A Unification Framework for Euclidean and Hyperbolic Graph Neural Networks · IJCAI 2023
Machine learning › Graph learning
graph self-supervised learning
0.512021
Self-Supervised Learning of Contextual Embeddings for Link Prediction in Heterogeneous Networks · WWW 2021
Machine learning › Graph learning › network embedding
heterogeneous graph embedding
0.512021
Self-Supervised Learning of Contextual Embeddings for Link Prediction in Heterogeneous Networks · WWW 2021
Machine learning › Graph learning
link prediction
0.512021
Self-Supervised Learning of Contextual Embeddings for Link Prediction in Heterogeneous Networks · WWW 2021
Data mining › structured data mining › graph mining › dynamic network analysis
dynamic graph mining
0.312018
Percolator: Scalable Pattern Discovery in Dynamic Graphs · WSDM 2018
Data mining › pattern mining
graph pattern mining
0.312018
Percolator: Scalable Pattern Discovery in Dynamic Graphs · WSDM 2018
Computer vision › 3D vision › geometric estimation › geometric model fitting
hypothesis generation
0.312017
NOUS: Construction and Querying of Dynamic Knowledge Graphs · ICDE 2017
Knowledge graphs
knowledge graph construction
0.312017
NOUS: Construction and Querying of Dynamic Knowledge Graphs · ICDE 2017
Knowledge graphs
temporal knowledge graph
0.312017
NOUS: Construction and Querying of Dynamic Knowledge Graphs · ICDE 2017
Computational science and engineering
computational chemistry
0.212024
CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback · ICML 2024
Machine learning › Deep learning architectures and training › normalization
normalization layers
0.212023
A Unification Framework for Euclidean and Hyperbolic Graph Neural Networks · IJCAI 2023
Web and social media mining › web mining
web graph analysis
0.112021
Self-Supervised Learning of Contextual Embeddings for Link Prediction in Heterogeneous Networks · WWW 2021
Bioinformatics and computational biology
proteomics
0.112010
Machine learning based prediction for peptide drift times in ion mobility spectrometry · Bioinform. 2010
Data mining
incremental mining
0.112018
Percolator: Scalable Pattern Discovery in Dynamic Graphs · WSDM 2018
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.012010
Machine learning based prediction for peptide drift times in ion mobility spectrometry · Bioinform. 2010

Methods — techniques the papers use, named apart from their topics

quantum-chemistry simulation · 1.5large language model · 1.5graph neural network · 1.5self-supervised pretraining · 1.0meta-path aggregation · 1.0attention mechanism · 1.0poincaré disk model · 0.7hyperbolic normalization · 0.7curated knowledge graph integration · 0.6parallel pattern mining · 0.3incremental pattern mining · 0.3support vector regression · 0.1partial least squares regression · 0.1machine learning · 0.1
YearPublicationVenuePosition
2026 Increasing value in the Veterans Affairs Healthcare System (VA) with precision health: a continuing landmark collaboration with the Department of Energy
abstract
OBJECTIVE: Phase II of MVP-CHAMPION, a federal collaboration between the Veterans Affairs Healthcare System (VA) and the Department of Energy (DoE), leveraged large-scale clinical, geo-spatial, and genetic data with state-of-the-art artificial intelligence (AI), and high-performance computing (HPC) to improve value in healthcare. MATERIALS AND METHODS: Eight clinical priority projects for which AI was a critical missing capability were initiated to address: lung cancer screening (MVP 061), suicide risk screening (MVP 062), cardiovascular risk in obstructive sleep apnea (MVP 063), checkpoint inhibitor toxicity (MVP 064), heart failure (MVP 065), renal complications in diabetes (MVP 066), post COVID-19 sequelae (MVP 067), and antipsychotic medication toxicity (MVP 068). RESULTS: Building on a strong regulatory and administrative foundation, we developed multimorbidity-aware analytic frameworks, reusable computational tools, and analytic pipelines. These greatly facilitated identification of novel risk factors including genetic variants and specification of more discriminating prediction models. Novel genetic risk factors are informing development and repurposing of medications and discriminating prediction models promise to improve healthcare value. DISCUSSION: The research foundation developed in Phase I and extended in Phase II of MVP CHAMPION has supported an unprecedented federal collaboration and yielded significant scientific advances. Our clinical findings are poised for near-term application, while advances in machine learning and high-performance computing may accelerate the broader adoption of artificial intelligence in healthcare. CONCLUSION: This maturing VA-DoE federal collaboration is poised to transform the future of Veterans' healthcare and the broader national landscape of precision health.
Amy Justice, Benjamin H. McMahon, Daniel A. Jacobson, Kelly Cho, Anuj J. Kapadia, Samuel M. Aguayo, Zeynep H. Gümüs, Ioana Danciu, Jean C. Beckham, Nathan A. Kimbrel, Silvia Crivelli, Eilis A. Boudreau, Patrick D. Finley, Alex K. Bryant, Shinjae Yoo, Jacob Joseph, Peter Reaven, Shiuh-Wen Luoh, Ravi K. Madduri, Ayman Fanous, Khushbu Agarwal, Harshini Mukundan, Sumitra Muralidhar
J. Am. Medical Informatics Assoc.23
2024 CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback
abstract
The discovery of new catalysts is essential for the design of new and more efficient chemical processes in order to transition to a sustainable future. We introduce an AI-guided computational screening framework unifying linguistic reasoning with quantum-chemistry based feedback from 3D atomistic representations. Our approach formulates catalyst discovery as an uncertain environment where an agent actively searches for highly effective catalysts via the iterative combination of large language model (LLM)-derived hypotheses and atomistic graph neural network (GNN)-derived feedback. Identified catalysts in intermediate search steps undergo structural evaluation based on spatial orientation, reaction pathways, and stability. Scoring functions based on adsorption energies and reaction energy barriers steer the exploration in the LLM’s knowledge space toward energetically favorable, high-efficiency catalysts. We introduce planning methods that automatically guide the exploration without human input, providing competitive performance against expert-enumerated chemical descriptor-based implementations. By integrating language-guided reasoning with computational chemistry feedback, our work pioneers AI-accelerated, trustworthy catalyst discovery.
Henry Sprueill, Carl Edwards, Khushbu Agarwal, Mariefel V. Olarte, Udishnu Sanyal, Conrad Johnston, Heng Ji 0001, Sutanay Choudhury
ICML3
2023 A Unification Framework for Euclidean and Hyperbolic Graph Neural Networks
abstract
Hyperbolic neural networks can effectively capture the inherent hierarchy of graph datasets, and consequently a powerful choice of GNNs. However, they entangle multiple incongruent (gyro-)vector spaces within a layer, which makes them limited in terms of generalization and scalability. In this work, we propose the Poincaré disk model as our search space, and apply all approximations on the disk (as if the disk is a tangent space derived from the origin), thus getting rid of all inter-space transformations. Such an approach enables us to propose a hyperbolic normalization layer and to further simplify the entire hyperbolic model to a Euclidean model cascaded with our hyperbolic normalization layer. We applied our proposed nonlinear hyperbolic normalization to the current state-of-the-art homogeneous and multi-relational graph networks. We demonstrate that our model not only leverages the power of Euclidean networks such as interpretability and efficient execution of various model components, but also outperforms both Euclidean and hyperbolic counterparts on various benchmarks. Our code is made publicly available at https://github.com/oom-debugger/ijcai23.
Mehrdad Khatir, Nurendra Choudhary, Sutanay Choudhury, Khushbu Agarwal, Chandan K. Reddy
IJCAI4
2021 Tracking the Evolution of COVID-19 via Temporal Comorbidity Analysis from Multi-Modal Data
Sutanay Choudhury, Khushbu Agarwal, Colby Ham, Pritam Mukherjee, Siyi Tang, Sindhu Tipirneni, Veysel Kocaman, Suzanne Tamang, Robert Rallo, Chandan K. Reddy
AMIA2
2021 Self-Supervised Learning of Contextual Embeddings for Link Prediction in Heterogeneous Networks
abstract
Representation learning methods for heterogeneous networks produce a low-dimensional vector embedding (that is typically fixed for all tasks) for each node. Many of the existing methods focus on obtaining a static vector representation for a node in a way that is agnostic to the downstream application where it is being used. In practice, however, downstream tasks such as link prediction require specific contextual information that can be extracted from the subgraphs related to the nodes provided as input to the task. To tackle this challenge, we develop , a framework for bridging static representation learning methods using global information from the entire graph with localized attention driven mechanisms to learn contextual node representations. We first pre-train our model in a self-supervised manner by introducing higher-order semantic associations and masking nodes, and then fine-tune our model for a specific link prediction task. Instead of training node representations by aggregating information from all semantic neighbors connected via metapaths, we automatically learn the composition of different metapaths that characterize the context for a specific task without the need for any pre-defined metapaths. significantly outperforms both static and contextual embedding learning methods on several publicly available benchmark network datasets. We also demonstrate the interpretability, effectiveness of contextual learning, and the scalability of through extensive evaluation.
Ping Wang 0024, Khushbu Agarwal, Colby Ham, Sutanay Choudhury, Chandan K. Reddy
WWW2
2018 Percolator: Scalable Pattern Discovery in Dynamic Graphs
abstract
We demonstrate \perco, a distributed system for graph pattern discovery in dynamic graphs. In contrast to conventional mining systems, Percolator advocates efficient pattern mining schemes that (1) support pattern detection with keywords; (2) integrate incremental and parallel pattern mining; and (3) support analytical queries such as trend analysis. The core idea of \perco is to dynamically decide and verify a small fraction of patterns and their instances that must be inspected in response to buffered updates in dynamic graphs, with a total mining cost independent of graph size. We demonstrate a( the feasibility of incremental pattern mining by walking through each component of \perco, b) the efficiency and scalability of \perco over the sheer size of real-world dynamic graphs, and c) how the user-friendly \gui of \perco interacts with users to support keyword-based queries that detect, browse and inspect trending patterns. We demonstrate how \perco effectively supports event and trend analysis in social media streams and research publication, respectively.
Sutanay Choudhury, Sumit Purohit, Yinghui Wu 0001, Lawrence B. Holder, Khushbu Agarwal
WSDM6
2017 NOUS: Construction and Querying of Dynamic Knowledge Graphs
abstract
The ability to construct domain specific knowledge graphs (KG) and perform question-answering or hypothesis generation is a transformative capability. Despite their value, automated construction of knowledge graphs remains an expensive technical challenge that is beyond the reach for most enterprises and academic institutions. We propose an end-toend framework for developing custom knowledge graph driven analytics for arbitrary application domains. The uniqueness of our system lies A) in its combination of curated KGs along with knowledge extracted from unstructured text, B) support for advanced trending and explanatory questions on a dynamic KG, and C) the ability to answer queries where the answer is embedded across multiple data sources.
Sutanay Choudhury, Khushbu Agarwal, Sumit Purohit, Baichuan Zhang, Meg Pirrung, William P. Smith 0001, Mathew Thomas
ICDE2
2015 Large Scale Frequent Pattern Mining Using MPI One-Sided Model
abstract
In this paper, we propose a work-stealing runtime -- Library for Work Stealing LibWS -- using MPI one-sided model for designing scalable FP-Growth -- de facto frequent pattern mining algorithm -- on large scale systems. LibWS provides locality efficient and highly scalable work-stealing techniques for load balancing on a variety of data distributions. We also propose a novel communication algorithm for FP-growth data exchange phase, which reduces the communication complexity from state-of-the-art Θ(p) to Θ(f + p/f), for p processes and f frequent attributed-ids. FP-Growth is implemented using LibWS and evaluated on several work distributions and support counts. An experimental evaluation of the FP-Growth on LibWS using 4096 processes on an InfiniBand Cluster demonstrates excellent efficiency for several work distributions (91% efficiency for Power-law and 93% for Poisson). The proposed distributed FP-Tree merging algorithm provides 38x communication speedup on 4096 cores.
Abhinav Vishnu, Khushbu Agarwal
CLUSTER2
2015 A Selectivity based approach to Continuous Pattern Detection in Streaming Graphs
Sutanay Choudhury, Lawrence B. Holder, George Chin, Khushbu Agarwal, John Feo
EDBT4
2014 Synchronization Algorithms for Co-simulation of Power Grid and Communication Networks
abstract
The ongoing modernization of power grids consists of integrating them with communication networks in order to achieve robust and resilient control of grid operations. To understand the operation of the new smart grid, one approach is to use simulation software. Unfortunately, current power grid simulators at best utilize inadequate approximations to simulate communication networks, if at all. Cooperative simulation of specialized power grid and communication network simulators promises to more accurately reproduce the interactions of real smart grid deployments. However, co-simulation is a challenging problem. A co-simulation must manage the exchange of information, including the synchronization of simulator clocks, between all simulators while maintaining adequate computational performance. This paper describes two new conservative algorithms for reducing the overhead of time synchronization, namely Active Set Conservative and Reactive Conservative. We provide a detailed analysis of their performance characteristics with respect to the current state of the art including both conservative and optimistic synchronization algorithms. In addition, we provide guidelines for selecting the appropriate synchronization algorithm based on the requirements of the co-simulation. The newly proposed algorithms are shown to achieve as much as 14% and 63% improvement in performance, respectively, over the existing conservative algorithm.
Selim Ciraci, Jeff Daily, Khushbu Agarwal, Jason C. Fuller, Laurentiu Marinovici, Andrew Fisher 0003
MASCOTS3
2013 Scalable PGAS Metadata Management on Extreme Scale Systems
abstract
Programming models intended to run on exascale systems have a number of challenges to overcome, specially the sheer size of the system as measured by the number of concurrent software entities created and managed by the underlying runtime. It is clear from the size of these systems that any state maintained by the programming model has to be strictly sub-linear in size, in order not to overwhelm memory usage with pure overhead. A principal feature of Partitioned Global Address Space (PGAS) models is providing easy access to global-view distributed data structures. In order to provide efficient access to these distributed data structures, PGAS models must keep track of metadata such as where array sections are located with respect to processes/threads running on the HPC system. As PGAS models and applications become ubiquitous on very large trans-pet scale systems, a key component to their performance and scalability will be efficient and judicious use of memory for model overhead (metadata) compared to application data. We present an evaluation of several strategies to manage PGAS metadata that exhibit different space/time tradeoffs. We use two real-world PGAS applications to capture metadata usage patterns and gain insight into their communication behavior.
Daniel G. Chavarría-Miranda, Khushbu Agarwal, Tjerk P. Straatsma
CCGRID2
2011 Implementing High Performance Remote Method Invocation in CCA
abstract
We report our effort in engineering a high performance remote method invocation (RMI) mechanism for the Common Component Architecture (CCA). This mechanism provides a highly efficient and easy-to-use mechanism for distributed computing in CCA, enabling CCA applications to effectively leverage parallel systems to accelerate computations. This work is built on the previous work of Babel RMI. Babel is a high performance language interoperability tool that is used in CCA for scientific application writers to share, reuse, and compose applications from software components written in different programming languages. Babel provides a transparent and flexible RMI framework for distributed computing. However, the existing Babel RMI implementation is built on top of TCP and does not provide the level of performance required to distribute fine-grained tasks. We observed that the main reason the TCP based RMI does not perform well is because it does not utilize the high performance interconnect hardware on a cluster efficiently. We have implemented a high performance RMI protocol, HPCRMI. HPCRMI achieves low latency by building on top of a low-level portable communication library, Aggregated Remote Message Copy Interface (ARMCI), and minimizing communication for each RMI call. Our design allows a RMI operation to be completed by only two RDMA operations. We also aggressively optimize our system to reduce copying. In this paper, we discuss the design and our experimental evaluation of this protocol. Our experimental results show that our protocol can improve RMI performance by an order of magnitude.
Jian Yin 0002, Khushbu Agarwal, Manoj Krishnan, Daniel G. Chavarría-Miranda, Ian Gorton, Tom Epperly
CLUSTER2
2010 Scalable Communication Trace Compression
abstract
Characterizing the communication behavior of parallel programs through tracing can help understand an application's characteristics, model its performance, and predict behavior on future systems. However, lossless communication traces can get prohibitively large, causing programmers to resort to variety of other techniques. In this paper, we present a novel approach to lossless communication trace compression. We augment the sequitur compression algorithm to employ it in communication trace compression of parallel programs. We present optimizations to reduce the memory overhead, reduce size of the trace files generated, and enable compression across multiple processes in a parallel program. The evaluation shows improved compression and reduced overhead over other approaches, with up to 3 orders of magnitude improvement for the NAS MG benchmark. We also observe that, unlike existing schemes, the trace files sizes and the memory overhead incurred are less sensitive to, if not independent of, the problem size for the NAS benchmarks.
Sriram Krishnamoorthy, Khushbu Agarwal
CCGRID2
2010 Machine learning based prediction for peptide drift times in ion mobility spectrometry
abstract
MOTIVATION: Ion mobility spectrometry (IMS) has gained significant traction over the past few years for rapid, high-resolution separations of analytes based upon gas-phase ion structure, with significant potential impacts in the field of proteomic analysis. IMS coupled with mass spectrometry (MS) affords multiple improvements over traditional proteomics techniques, such as in the elucidation of secondary structure information, identification of post-translational modifications, as well as higher identification rates with reduced experiment times. The high throughput nature of this technique benefits from accurate calculation of cross sections, mobilities and associated drift times of peptides, thereby enhancing downstream data analysis. Here, we present a model that uses physicochemical properties of peptides to accurately predict a peptide's drift time directly from its amino acid sequence. This model is used in conjunction with two mathematical techniques, a partial least squares regression and a support vector regression setting. RESULTS: When tested on an experimentally created high confidence database of 8675 peptide sequences with measured drift times, both techniques statistically significantly outperform the intrinsic size parameters-based calculations, the currently held practice in the field, on all charge states (+2, +3 and +4). AVAILABILITY: The software executable, imPredict, is available for download from http:/omics.pnl.gov/software/imPredict.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anuj R. Shah, Khushbu Agarwal, Erin S. Baker, Mudita Singhal, Anoop M. Mayampurath, Yehia M. Ibrahim, Lars J. Kangas, Matthew E. Monroe, Mikhail E. Belov, Gordon A. Anderson, Richard D. Smith
Bioinform.2