Arvind Ramanathan

dblp:51/5675 · DBLP profile ↗
← Back
30ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-1622-5488ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 11 · 5 since 2021Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Attr-RAG: Attribution-Guided Retrieval-Augmented Generation for Scientific Experiment Design
abstract
Evidence-based science depends on the iterative integration of experimentation, a process traditionally driven by slow and error-prone human effort. This has inspired the vision of an automated "robot scientist" capable of conducting end-to-end experimentation. While Large Language Models (LLMs) can generate procedural instructions, they often struggle to accurately describe scientific experiments due to the limited availability of high-quality, domain-specific examples in their training data. Retrieval-Augmented Generation (RAG) helps bridge this gap by allowing LLMs to access up-to-date external information. However, despite being effective for short questions, RAG struggles with long-form scientific experimental queries due to information loss from chunk fragmentation and retrieval of irrelevant information. In this paper, we propose Attr-RAG, an attribution-guided RAG framework to remove irrelevant or misleading context and retaining only complete, relevant information. Unlike traditional RAG methods that rely solely on vector similarity, Attr-RAG introduces a refinement stage using occlusion-based attribution to identify which retrieved chunks truly influence the LLM’s response. This attribution-guided filtering ensures that only contextually coherent chunks are used for accurate and grounded final answer generation. Attr-RAG demonstrated superior performance in 9 out of 10 chemistry lab experiment tasks of the ChemEx dataset and outperformed baselines across most quantitative evaluation metrics. In qualitative evaluations conducted by state-of-the-art LLM judges (GPT-4o, Gemini 2.5, and Grok 3), the top mean scores of 27.8, 27.1, and 22.9, respectively, were achieved across six key evaluation criteria.
Fazle Rahat, M. Shifat Hossain, Arvind Ramanathan, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz
ICMLA3
2024 Equivariant Graph Neural Operator for Modeling 3D Dynamics
abstract
Modeling the complex three-dimensional (3D) dynamics of relational systems is an important problem in the natural sciences, with applications ranging from molecular simulations to particle mechanics. Machine learning methods have achieved good success by learning graph neural networks to model spatial interactions. However, these approaches do not faithfully capture temporal correlations since they only model next-step predictions. In this work, we propose Equivariant Graph Neural Operator (EGNO), a novel and principled method that directly models dynamics as trajectories instead of just next-step prediction. Different from existing methods, EGNO explicitly learns the temporal evolution of 3D dynamics where we formulate the dynamics as a function over time and learn neural operators to approximate it. To capture the temporal correlations while keeping the intrinsic SE(3)-equivariance, we develop equivariant temporal convolutions parameterized in the Fourier space and build EGNO by stacking the Fourier layers over equivariant networks. EGNO is the first operator learning framework that is capable of modeling solution dynamics functions over time while retaining 3D equivariance. Comprehensive experiments in multiple domains, including particle simulations, human motion capture, and molecular dynamics, demonstrate the significantly superior performance of EGNO against existing methods, thanks to the equivariant temporal modeling. Our code is available at https://github.com/MinkaiXu/egno.
Minkai Xu, Jiaqi Han 0001, Aaron Lou, Jean Kossaifi, Arvind Ramanathan, Kamyar Azizzadenesheli, Jure Leskovec, Stefano Ermon, Anima Anandkumar
ICML5
2024 MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
abstract
We present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows.
Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan
SC30
2023 Transferable Graph Neural Fingerprint Models for Quick Response to Future Bio-Threats
abstract
Fast screening of drug molecules based on the ligand binding affinity is an important step in the drug discovery pipeline. Graph neural fingerprint is a promising method for developing molecular docking surrogates with high throughput and great fidelity. In this study, we built a COVID-19 drug docking dataset of about 300,000 drug candidates on 23 coronavirus protein targets. With this dataset, we trained graph neural fin-gerprint docking models for high-throughput virtual COVID-19 drug screening. The graph neural fingerprint models yield high prediction accuracy on docking scores with the mean squared error lower than 0.21 kcal/mol for most of the docking targets, showing significant improvement over conventional circular fin-gerprint methods. To make the neural fingerprints transferable for unknown targets, we also propose a transferable graph neural fingerprint method trained on multiple targets. With comparable accuracy to target-specific graph neural fingerprint models, the training and data efficiency of the transferable model is several times higher. We highlight that the impact of this study extends beyond COVID-19 dataset, as our approach for fast virtual ligand screening can be easily adapted and integrated into a general machine learning-accelerated pipeline to battle future bio-threats.
Wei Chen 0043, Yihui Ren 0001, Ai Kagawa, Matthew R. Carbone, Samuel Yen-Chi Chen, Xiaohui Qu, Shinjae Yoo, Austin Clyde, Arvind Ramanathan, Rick L. Stevens, Huub J. J. Van Dam, Deyu Lu
ICMLA9
2023 ChemoGraph: Interactive Visual Exploration of the Chemical Space
abstract
Abstract Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.
Bharat Kale, Austin Clyde, Maoyuan Sun, Arvind Ramanathan, Rick L. Stevens, Michael E. Papka
Comput. Graph. Forum4
2023 Model certainty in cellular network-driven processes with missing data
abstract
Mathematical models are often used to explore network-driven cellular processes from a systems perspective. However, a dearth of quantitative data suitable for model calibration leads to models with parameter unidentifiability and questionable predictive power. Here we introduce a combined Bayesian and Machine Learning Measurement Model approach to explore how quantitative and non-quantitative data constrain models of apoptosis execution within a missing data context. We find model prediction accuracy and certainty strongly depend on rigorous data-driven formulations of the measurement, and the size and make-up of the datasets. For instance, two orders of magnitude more ordinal (e.g., immunoblot) data are necessary to achieve accuracy comparable to quantitative (e.g., fluorescence) data for calibration of an apoptosis execution model. Notably, ordinal and nominal (e.g., cell fate observations) non-quantitative data synergize to reduce model uncertainty and improve accuracy. Finally, we demonstrate the potential of a data-driven Measurement Model approach to identify model features that could lead to informative experimental measurements and improve model predictive power.
Michael W. Irvin, Arvind Ramanathan, Carlos F. Lopez
PLoS Comput. Biol.2
2022 Shaping Noise for Robust Attributions in Neural Stochastic Differential Equations
abstract
Neural SDEs with Brownian motion as noise lead to smoother attributions than traditional ResNets. Various attribution methods such as saliency maps, integrated gradients, DeepSHAP and DeepLIFT have been shown to be more robust for neural SDEs than for ResNets using the recently proposed sensitivity metric. In this paper, we show that neural SDEs with adaptive attribution-driven noise lead to even more robust attributions and smaller sensitivity metrics than traditional neural SDEs with Brownian motion as noise. In particular, attribution-driven shaping of noise leads to 6.7%, 6.9% and 19.4% smaller sensitivity metric for integrated gradients computed on three discrete approximations of neural SDEs with standard Brownian motion noise: stochastic ResNet-50, WideResNet-101 and ResNeXt-101 models respectively. The neural SDE model with adaptive attribution-driven noise leads to 25.7% and 4.8% improvement in the SIC metric over traditional ResNets and Neural SDEs with Brownian motion as noise. To the best of our knowledge, we are the first to propose the use of attributions for shaping the noise injected in neural SDEs, and demonstrate that this process leads to more robust attributions than traditional neural SDEs with standard Brownian motion as noise.
Sumit Kumar Jha 0001, Rickard Ewetz, Alvaro Velasquez, Arvind Ramanathan, Susmit Jha
AAAI4
2022 Coupling streaming AI and HPC ensembles to achieve 100-1000× faster biomolecular simulations
abstract
Machine learning (ML)-based steering can improve the performance of ensemble-based simulations by allowing for online selection of more scientifically meaningful computations. We present DeepDriveMD, a framework for ML-driven steering of scientific simulations that we have used to achieve orders-of-magnitude improvements in molecular dynamics (MD) performance via effective coupling of ML and HPC on large parallel computers. We discuss the design of DeepDriveMD and characterize its performance. We demonstrate that DeepDriveMD can achieve between 100-1000× acceleration for protein folding simulations relative to other methods, as measured by the amount of simulated time performed, while covering the same conformational landscape as quantified by the states sampled during a simulation. Experiments are performed on leadership-class platforms on up to 1020 nodes. The results establish DeepDriveMD as a high-performance framework for ML-driven HPC simulation scenarios, that supports diverse MD simulation and ML back-ends, and which enables new scientific insights by improving the length and time scales accessible with current computing capacity.
Alex Brace, Igor Yakushin, Anda Trifan, Todd S. Munson, Ian T. Foster, Arvind Ramanathan, Hyungro Lee, Matteo Turilli, Shantenu Jha
IPDPS7
2021 IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads
abstract
The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2–3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100 × to 1000 × improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ∼ 1011 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine.
Aymen Alsaadi, Dario Alfè, Yadu N. Babuji, Agastya Bhati, Ben Blaiszik, Alex Brace, Thomas S. Brettin, Kyle Chard, Ryan Chard, Austin Clyde, Peter V. Coveney, Ian T. Foster, Tom Gibbs, Shantenu Jha, Kristopher Keipert, Dieter Kranzlmüller, Thorsten Kurth, Hyungro Lee, Zhuozhao Li, Gerald Mathias, André Merzky, Alexander Partin, Arvind Ramanathan, Ashka Shah, Abraham C. Stern, Rick L. Stevens, Mikhail Titov, Anda Trifan, Aristeidis Tsaris, Matteo Turilli, Huub J. J. Van Dam, Shunzhou Wan, David Wifling, Junqi Yin
ICPP24
2021 The 4th International Workshop on Epidemiology meets Data Mining and Knowledge Discovery (epiDAMIK 4.0 @ KDD2021)
abstract
The 4th [email protected] workshop is a forum to discuss new insights into how data mining can play a bigger role in epidemiology and public health research. While the integration of data science methods into epidemiology has significant potential, it remains under studied. We aim to raise the profile of this emerging research area of data-driven and computational epidemiology, and create a venue for presenting state-of-the-art and in-progress results-in particular, results that would otherwise be difficult to present at a major data mining conference, including lessons learnt in the 'trenches'. The current COVID-19 pandemic has only showcased the urgency and importance of this area. Our target audience consists of data mining and machine learning researchers from both academia and industry who are interested in epidemiological and public-health applications of their work, and practitioners from the areas of mathematical epidemiology and public health.
Bijaya Adhikari, Ajitesh Srivastava, Sen Pei, Sarah Kefayati, Rose Yu, Amulya Yadav, Alexander Rodríguez, Arvind Ramanathan, Anil Vullikanti, B. Aditya Prakash
KDD8
2021 Prototypical Models for Classifying High-Risk Atypical Breast Lesions
Akash Parvatikar, Om Choudhary, Arvind Ramanathan, Rebekah Jenkins, Olga Navolotskaia, Gloria Carter, Akif Burak Tosun, Jeffrey L. Fine, S. Chakra Chennubhotla
MICCAI (8)3
2020 Modeling Histological Patterns for Differential Diagnosis of Atypical Breast Lesions
Akash Parvatikar, Om Choudhary, Arvind Ramanathan, Olga Navolotskaia, Gloria Carter, Akif Burak Tosun, Jeffrey L. Fine, S. Chakra Chennubhotla
MICCAI (5)3
2020 Distributed Bayesian optimization of deep reinforcement learning algorithms
abstract
Significant strides have been made in supervised learning settings thanks to the successful application of deep learning. Now, recent work has brought the techniques of deep learning to bear on sequential decision processes in the area of deep reinforcement learning (DRL). Currently, little is known regarding hyperparameter optimization for DRL algorithms. Given that DRL algorithms are computationally intensive to train, and are known to be sample inefficient, optimizing model hyperparameters for DRL presents significant challenges to established techniques. We provide an open source, distributed Bayesian model-based optimization algorithm, HyperSpace, and show that it consistently outperforms standard hyperparameter optimization techniques across three DRL algorithms.
M. Todd Young, Jacob D. Hinkle, Ramakrishnan Kannan, Arvind Ramanathan
J. Parallel Distributed Comput.4
2020 Attacking NIST biometric image software using nonlinear optimization
Sunny Raj, Jodh S. Pannu, Steven Lawrence Fernandes, Arvind Ramanathan, Laura L. Pullum, Sumit Kumar Jha 0001
Pattern Recognit. Lett.4
2019 Visual Analytics for Deep Embeddings of Large Scale Molecular Dynamics Simulations
abstract
Molecular Dynamics (MD) simulation have been emerging as an excellent candidate for understanding complex atomic and molecular scale mechanism of bio-molecules that control essential bio-physical phenomenon in a living organism. But this MD technique produces large-size and long-timescale data that are inherently high-dimensional and occupies many terabytes of data. Processing this immense amount of data in a meaningful way is becoming increasingly difficult. Therefore, specific dimensionality reduction algorithm using deep learning technique has been employed here to embed the high-dimensional data in a lower-dimension latent space that still preserves the inherent molecular characteristics i.e. retains biologically meaningful information. Subsequently, the results of the embedding models are visualized for model evaluation and analysis of the extracted underlying features. However, most of the existing visualizations for embeddings have limitations in evaluating the embedding models and understanding the complex simulation data. We propose an interactive visual analytics system for embeddings of MD simulations to not only evaluate and explain an embedding model but also analyze various characteristics of the simulations. Our system enables exploration and discovery of meaningful and semantic embedding results and supports the understanding and evaluation of results by the quantitatively described features of the MD simulations (even without specific labels).
Junghoon Chae, Debsindhu Bhowmik, Arvind Ramanathan, Chad A. Steed
IEEE BigData4
2019 Classifying cancer pathology reports with hierarchical self-attention networks
abstract
We introduce a deep learning architecture, hierarchical self-attention networks (HiSANs), designed for classifying pathology reports and show how its unique architecture leads to a new state-of-the-art in accuracy, faster training, and clear interpretability. We evaluate performance on a corpus of 374,899 pathology reports obtained from the National Cancer Institute's (NCI) Surveillance, Epidemiology, and End Results (SEER) program. Each pathology report is associated with five clinical classification tasks - site, laterality, behavior, histology, and grade. We compare the performance of the HiSAN against other machine learning and deep learning approaches commonly used on medical text data - Naive Bayes, logistic regression, convolutional neural networks, and hierarchical attention networks (the previous state-of-the-art). We show that HiSANs are superior to other machine learning and deep learning text classifiers in both accuracy and macro F-score across all five classification tasks. Compared to the previous state-of-the-art, hierarchical attention networks, HiSANs not only are an order of magnitude faster to train, but also achieve about 1% better relative accuracy and 5% better relative macro F-score.
Shang Gao 0008, John X. Qiu, Mohammed M. Alawad, Jacob D. Hinkle, Noah Schaefferkoetter, Hong-Jun Yoon, James Blair Christian, Paul A. Fearn, Lynne Penberthy, Xiao-Cheng Wu, Linda Coyle, Georgia D. Tourassi, Arvind Ramanathan
Artif. Intell. Medicine13
2019 Data-driven efficient network and surveillance-based immunization
Yao Zhang 0003, Arvind Ramanathan, Anil Vullikanti, Laura L. Pullum, B. Aditya Prakash
Knowl. Inf. Syst.2
2018 HyperSpace: Distributed Bayesian Hyperparameter Optimization
abstract
As machine learning models continue to increase in complexity, so does the potential number of free model parameters commonly known as hyperparameters. While there has been considerable progress toward finding optimal configurations of these hyperparameters, many optimization procedures are treated as black boxes. We believe optimization methods should not only return a set of optimized hyperparameters, but also give insight into the effects of model hyperparameter settings. To this end, we present HyperSpace, a parallel implementation of Bayesian sequential model-based optimization. HyperSpace leverages high performance computing (HPC) resources to better understand unknown, potentially non-convex hyperparameter search spaces. We show that it is possible to learn the dependencies between model hyperparameters through the optimization process. By partitioning large search spaces and running many optimization procedures in parallel, we also show that it is possible to discover families of good hyperparameter settings over a variety of models including unsupervised clustering, regression, and classification tasks.
M. Todd Young, Jacob D. Hinkle, Arvind Ramanathan, Ramakrishnan Kannan
SBAC-PAD3
2018 Deep clustering of protein folding simulations
abstract
BACKGROUND: We examine the problem of clustering biomolecular simulations using deep learning techniques. Since biomolecular simulation datasets are inherently high dimensional, it is often necessary to build low dimensional representations that can be used to extract quantitative insights into the atomistic mechanisms that underlie complex biological processes. RESULTS: We use a convolutional variational autoencoder (CVAE) to learn low dimensional, biophysically relevant latent features from long time-scale protein folding simulations in an unsupervised manner. We demonstrate our approach on three model protein folding systems, namely Fs-peptide (14 μs aggregate sampling), villin head piece (single trajectory of 125 μs) and β- β- α (BBA) protein (223 + 102 μs sampling across two independent trajectories). In these systems, we show that the CVAE latent features learned correspond to distinct conformational substates along the protein folding pathways. The CVAE model predicts, on average, nearly 89% of all contacts within the folding trajectories correctly, while being able to extract folded, unfolded and potentially misfolded states in an unsupervised manner. Further, the CVAE model can be used to learn latent features of protein folding that can be applied to other independent trajectories, making it particularly attractive for identifying intrinsic features that correspond to conformational substates that share similar structural features. CONCLUSIONS: Together, we show that the CVAE model can quantitatively describe complex biophysical processes such as protein folding.
Debsindhu Bhowmik, Shang Gao 0008, Michael T. Young, Arvind Ramanathan
BMC Bioinform.4
2018 Scalable deep text comprehension for Cancer surveillance on high-performance computing
abstract
BACKGROUND: Deep Learning (DL) has advanced the state-of-the-art capabilities in bioinformatics applications which has resulted in trends of increasingly sophisticated and computationally demanding models trained by larger and larger data sets. This vastly increased computational demand challenges the feasibility of conducting cutting-edge research. One solution is to distribute the vast computational workload across multiple computing cluster nodes with data parallelism algorithms. In this study, we used a High-Performance Computing environment and implemented the Downpour Stochastic Gradient Descent algorithm for data parallelism to train a Convolutional Neural Network (CNN) for the natural language processing task of information extraction from a massive dataset of cancer pathology reports. We evaluated the scalability improvements using data parallelism training and the Titan supercomputer at Oak Ridge Leadership Computing Facility. To evaluate scalability, we used different numbers of worker nodes and performed a set of experiments comparing the effects of different training batch sizes and optimizer functions. RESULTS: We found that Adadelta would consistently converge at a lower validation loss, though requiring over twice as many training epochs as the fastest converging optimizer, RMSProp. The Adam optimizer consistently achieved a close 2nd place minimum validation loss significantly faster; using a batch size of 16 and 32 allowed the network to converge in only 4.5 training epochs. CONCLUSIONS: We demonstrated that the networked training process is scalable across multiple compute nodes communicating with message passing interface while achieving higher classification accuracy compared to a traditional machine learning algorithm.
John X. Qiu, Hong-Jun Yoon, Kshitij Srivastava, Thomas P. Watson, James Blair Christian, Arvind Ramanathan, Xiao-Cheng Wu, Paul A. Fearn, Georgia D. Tourassi
BMC Bioinform.6
2018 Hierarchical attention networks for information extraction from cancer pathology reports
abstract
OBJECTIVE: We explored how a deep learning (DL) approach based on hierarchical attention networks (HANs) can improve model performance for multiple information extraction tasks from unstructured cancer pathology reports compared to conventional methods that do not sufficiently capture syntactic and semantic contexts from free-text documents. MATERIALS AND METHODS: Data for our analyses were obtained from 942 deidentified pathology reports collected by the National Cancer Institute Surveillance, Epidemiology, and End Results program. The HAN was implemented for 2 information extraction tasks: (1) primary site, matched to 12 International Classification of Diseases for Oncology topography codes (7 breast, 5 lung primary sites), and (2) histological grade classification, matched to G1-G4. Model performance metrics were compared to conventional machine learning (ML) approaches including naive Bayes, logistic regression, support vector machine, random forest, and extreme gradient boosting, and other DL models, including a recurrent neural network (RNN), a recurrent neural network with attention (RNN w/A), and a convolutional neural network. RESULTS: Our results demonstrate that for both information tasks, HAN performed significantly better compared to the conventional ML and DL techniques. In particular, across the 2 tasks, the mean micro and macro F-scores for the HAN with pretraining were (0.852,0.708), compared to naive Bayes (0.518, 0.213), logistic regression (0.682, 0.453), support vector machine (0.634, 0.434), random forest (0.698, 0.508), extreme gradient boosting (0.696, 0.522), RNN (0.505, 0.301), RNN w/A (0.637, 0.471), and convolutional neural network (0.714, 0.460). CONCLUSIONS: HAN-based DL models show promise in information abstraction tasks within unstructured clinical pathology reports.
Shang Gao 0008, Michael T. Young, John X. Qiu, Hong-Jun Yoon, James Blair Christian, Paul A. Fearn, Georgia D. Tourassi, Arvind Ramanathan
J. Am. Medical Informatics Assoc.8
2017 Data-Driven Immunization
abstract
Given a contact network and coarse-grained diagnostic information like electronic Healthcare Reimbursement Claims (eHRC) data, can we develop efficient intervention policies to control an epidemic? Immunization is an important problem in multiple areas especially epidemiology and public health. However, most existing studies focus on developing pre-emptive strategies assuming prior epidemiological models. In practice, disease spread is usually complicated, hence assuming an underlying model may deviate from true spreading patterns, leading to possibly inaccurate interventions. Additionally, the abundance of health care surveillance data (like eHRC) makes it possible to study data-driven strategies without too many restrictive assumptions. Hence, such an approach can help public-health experts take more practical decisions. In this paper, we take into account propagation log and contact networks for controlling propagation. We formulate the novel and challenging Data-Driven Immunization problem without assuming classical epidemiological models. To solve it, we first propose an efficient sampling approach to align surveillance data with contact networks, then develop an efficient algorithm with the provably approximate guarantee for immunization. Finally, we show the effectiveness and scalability of our methods via extensive experiments on multiple datasets, and conduct case studies on nation-wide real medical surveillance data.
Yao Zhang 0003, Arvind Ramanathan, Anil Vullikanti, Laura L. Pullum, B. Aditya Prakash
ICDM2
2016 Constellation: A science graph network for scalable data and knowledge discovery in extreme-scale scientific collaborations
abstract
Constellation's overarching goal is the federation of information from resources within an extreme-scale scientific collaboration to enable the scalable discovery of data and new knowledge pathways. The resource fabric is comprised of petascale supercomputers and storage systems, users, jobs, datasets and lifecycle artifacts. For an extreme-scale supercomputing center, normal operations can generate hundreds of millions of data products and metadata entries describing the resource fabric. Constellation federates the information extracted from the resources using a custom, transformative science graph network; constructs rich metadata indexes and higher-order derived metadata from the extracted information; and conducts scalable graph analytics to unravel hidden data pathways. Our implementation and deployment for a production, supercomputing facility shows that the graph can scale to more than 750 million vertices, its domain agnostic indexing can answer interesting science queries, and its analytics can aid in structural, topological and temporal analysis to identify usage hotspots.
Sudharshan S. Vazhkudai, John Harney, Raghul Gunasekaran, Dale Stansberry, Seung-Hwan Lim, Tom Barron, Andrew Nash, Arvind Ramanathan
IEEE BigData8
2016 Integrating symbolic and statistical methods for testing intelligent systems: Applications to machine learning and computer vision
Arvind Ramanathan, Laura L. Pullum, Faraz Hussain 0001, Dwaipayan Chakrabarty, Sumit Kumar Jha 0001
DATE1
2015 Sequential pattern mining of electronic healthcare reimbursement claims: Experiences and challenges in uncovering how patients are treated by physicians
abstract
We examine the use of electronic healthcare reimbursement claims (EHRC) for analyzing healthcare delivery and practice patterns across the United States (US). We show that EHRCs are correlated with disease incidence estimates published by the Centers for Disease Control. Further, by analyzing over 1 billion EHRCs, we track patterns of clinical procedures administered to patients with autism spectrum disorder (ASD), heart disease (HD) and breast cancer (BC) using sequential pattern mining algorithms. Our analyses reveal that in contrast to treating HD and BC, clinical procedures for ASD diagnoses are highly varied leading up to and after the ASD diagnoses. The discovered clinical procedure sequences also reveal significant differences in the overall costs incurred across different parts of the US, indicating a lack of consensus amongst practitioners in treating ASD patients. We show that a data-driven approach to understand clinical trajectories using EHRC can provide quantitative insights into how to better manage and treat patients. Based on our experience, we also discuss emerging challenges in using EHRC datasets for gaining insights into the state of contemporary healthcare delivery and practice in the US.
Kunal Malhotra, Tanner C. Hobson, Silvia Valkova, Laura L. Pullum, Arvind Ramanathan
IEEE BigData5
2015 NoC Architectures as Enablers of Biological Discovery for Personalized and Precision Medicine
abstract
This paper overviews the main computational issues in personalized and precision medicine (PPM), and present a cogent case for network-on-chip (NoC)-based multicore platforms as enablers in the process. We identify a series of challenges for the design and optimization of NoC-based solutions for PPM. To capture the characteristics of the cyber-physical sensing and processing, we propose a new computational model built on a dynamical heterogeneous hyper-graph description of application-to-architecture interactions. Starting from these premises, we summarize a few implications on NoC design methodologies, present some NoC-based solutions that deal with some of the challenges, and outline a few open problems.
Paul Bogdan, Turbo Majumder, Arvind Ramanathan, Yuankun Xue
NOCS3
2015 ORBiT: Oak Ridge biosurveillance toolkit for public health dynamics
abstract
BACKGROUND: The digitization of health-related information through electronic health records (EHR) and electronic healthcare reimbursement claims and the continued growth of self-reported health information through social media provides both tremendous opportunities and challenges in developing effective biosurveillance tools. With novel emerging infectious diseases being reported across different parts of the world, there is a need to build systems that can track, monitor and report such events in a timely manner. Further, it is also important to identify susceptible geographic regions and populations where emerging diseases may have a significant impact. METHODS: In this paper, we present an overview of Oak Ridge Biosurveillance Toolkit (ORBiT), which we have developed specifically to address data analytic challenges in the realm of public health surveillance. In particular, ORBiT provides an extensible environment to pull together diverse, large-scale datasets and analyze them to identify spatial and temporal patterns for various biosurveillance-related tasks. RESULTS: We demonstrate the utility of ORBiT in automatically extracting a small number of spatial and temporal patterns during the 2009-2010 pandemic H1N1 flu season using claims data. These patterns provide quantitative insights into the dynamics of how the pandemic flu spread across different parts of the country. We discovered that the claims data exhibits multi-scale patterns from which we could identify a small number of states in the United States (US) that act as "bridge regions" contributing to one or more specific influenza spread patterns. Similar to previous studies, the patterns show that the south-eastern regions of the US were widely affected by the H1N1 flu pandemic. Several of these south-eastern states act as bridge regions, which connect the north-east and central US in terms of flu occurrences. CONCLUSIONS: These quantitative insights show how the claims data combined with novel analytical techniques can provide important information to decision makers when an epidemic spreads throughout the country. Taken together ORBiT provides a scalable and extensible platform for public health surveillance.
Arvind Ramanathan, Laura L. Pullum, Tanner C. Hobson, Chad A. Steed, Shannon Quinn, S. Chakra Chennubhotla, Silvia Valkova
BMC Bioinform.1
2013 Performance modeling of microsecond scale biological molecular dynamics simulations on heterogeneous architectures
abstract
SUMMARY Performance improvements in biomolecular simulations based on molecular dynamics (MD) codes are widely desired. Unfortunately, the factors, which allowed past performance improvements, particularly the microprocessor clock frequencies, are no longer increasing. Hence, novel software and hardware solutions are being explored for accelerating performance of widely used MD codes. In this paper, we describe our efforts on porting, optimizing and tuning of Large‐scale Atomic/Molecular Massively Parallel Simulator, a popular MD framework, on heterogeneous architectures: multi‐core processors with graphical processing unit (GPU) accelerators. Our implementation is based on accelerating the most computationally expensive non‐bonded interaction terms on the GPUs and overlapping the computation on the CPU and GPUs. This functionality is built on top of message passing interface that allows multi‐level parallelism to be extracted even at the workstation level with the multi‐core CPUs and allows extension of the implementation on GPU‐enabled clusters. We hypothesize that the optimal benefit of heterogeneous architectures for applications will come by utilizing all possible resources (for example, CPU‐cores and GPU devices on GPU‐enabled clusters). Benchmarks for a range of biomolecular system sizes are provided, and an analysis is performed on four generations of NVIDIA's GPU devices. On GPU‐enabled Linux clusters, by overlapping and pipelining computation and communication, we observe up to 10‐folds application acceleration in multi‐core and multi‐GPU environments illustrating significant performance improvements. Detailed analysis of the implementation is presented that allows identification of bottlenecks in algorithm, indicating that code optimization and improvements on GPUs could allow microsecond scale simulation throughput on workstations and inexpensive GPU clusters, putting widely desired biologically relevant simulation time‐scales within reach of a large user community. In order to systematically optimize simulation throughput and to enable performance prediction, we have developed a parameterized performance model that will allow developers and users to explore the performance potential of future heterogeneous systems for biological simulations. Copyright © 2012 John Wiley & Sons, Ltd.
Pratul K. Agarwal, Scott S. Hampton, Jeffrey D. Poznanovic, Arvind Ramanathan, Sadaf R. Alam, Paul S. Crozier
Concurr. Comput. Pract. Exp.4
2011 QAARM: quasi-anharmonic autoregressive model reveals molecular recognition pathways in ubiquitin
abstract
MOTIVATION: Molecular dynamics (MD) simulations have dramatically improved the atomistic understanding of protein motions, energetics and function. These growing datasets have necessitated a corresponding emphasis on trajectory analysis methods for characterizing simulation data, particularly since functional protein motions and transitions are often rare and/or intricate events. Observing that such events give rise to long-tailed spatial distributions, we recently developed a higher-order statistics based dimensionality reduction method, called quasi-anharmonic analysis (QAA), for identifying biophysically-relevant reaction coordinates and substates within MD simulations. Further characterization of conformation space should consider the temporal dynamics specific to each identified substate. RESULTS: Our model uses hierarchical clustering to learn energetically coherent substates and dynamic modes of motion from a 0.5 μs ubiqutin simulation. Autoregressive (AR) modeling within and between states enables a compact and generative description of the conformational landscape as it relates to functional transitions between binding poses. Lacking a predictive component, QAA is extended here within a general AR model appreciative of the trajectory's temporal dependencies and the specific, local dynamics accessible to a protein within identified energy wells. These metastable states and their transition rates are extracted within a QAA-derived subspace using hierarchical Markov clustering to provide parameter sets for the second-order AR model. We show the learned model can be extrapolated to synthesize trajectories of arbitrary length. CONTACT: [email protected]; [email protected].
Andrej J. Savol, Virginia M. Burger, Pratul K. Agarwal, Arvind Ramanathan, S. Chakra Chennubhotla
Bioinform.4
2009 An Online Approach for Mining Collective Behaviors from Molecular Dynamics Simulations
Arvind Ramanathan, Pratul K. Agarwal, Maria G. Kurnikova, Christopher J. Langmead
RECOMB1