VLDB 2026 Research / reviewers in the wild / expert
Ankit Agrawal 0001
dblp:55/5176-1
· DBLP profile ↗
84ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-5519-0302ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 37 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 7 first-author · 3 since 2021Systems, architecture and hardware · 21 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REN: Anatomically-Informed Mixture-of-Experts for Interstitial Lung Disease DiagnosisabstractMixture-of-Experts (MoE) architectures achieve scalable learning by routing inputs to specialized subnetworks through conditional computation. However, conventional MoE designs assume homogeneous expert capability and domain-agnostic routing-assumptions that are fundamentally misaligned with medical imaging, where anatomical structure and regional disease heterogeneity govern pathological patterns. We introduce Regional Expert Networks (REN), the first anatomically-informed MoE framework for medical image classification. REN encodes anatomical priors by training seven specialized experts, each dedicated to a distinct lung lobe or bilateral lung combination, enabling precise modeling of region-specific pathological variation. Multi-modal gating mechanisms dynamically integrate radiomics biomarkers with deep learning (DL) features extracted by convolutional (CNN), Transformer (ViT), and state-space (Mamba) architectures to weight expert contributions at inference. Applied to interstitial lung disease (ILD) classification on a 597-patient, 1,898-scan longitudinal cohort, REN achieves consistently superior performance: the radiomics-guided ensemble attains an average AUC of $0.8646~\pm ~0.0467$ , a +12.5% improvement over the SwinUNETR single-model baseline (AUC 0.7685, ${p}={0}.{031}$ ). Lower-lobe experts reach AUCs of 0.88-0.90, outperforming DL baselines (CNN: 0.76-0.79) and mirroring known patterns of basal ILD progression. Evaluated under rigorous patient-level cross-validation, REN demonstrates strong generalizability and clinical interpretability, establishing a scalable, anatomically-guided framework potentially extensible to other structured medical imaging tasks. Code is available on our GitHub https://github.com/NUBagciLab/MoE-REN. Alec Peltekian, Halil Ertugrul Aktas, Gorkem Durak, Kevin M. Grudzinski, Bradford C. Bemiss, Carrie Richardson, Jane E. Dematte, G. R. Scott Budinger, Anthony J. Esposito, Alexander V. Misharin, Alok N. Choudhary, Ankit Agrawal 0001, Ulas Bagci |
IEEE Trans. Medical Imaging | 12 |
| 2025 | AI-Driven Prediction of Material Deformation: Stress-Strain Curves Faster Than Crystal Plasticity Finite Element SimulationabstractStress–strain curves capture the mechanical behavior of materials but are computationally expensive to generate using crystal-plasticity finite-element (CPFE) models, due to the profound nonlinearity of the response, and its relationship to crystallographic orientation. We propose an AI-driven framework for predicting bilinear approximate stress–strain curves of metallic alloys using supervised machine learning models. We evaluate its performance on three representative materials: aluminum (Al), nickel (Ni), and copper (Cu), commonly used in aerospace engineering. Trained on just 100 fully resolved CPFE curves per material, our model accurately reconstructs entire curves using features extracted after only a single CPFE step, effectively leading to orders of magnitude speedup with respect to CPFE simulation for predicting stress-strain curves of new orientations unseen by the AI model. The resulting predictions achieve a mean absolute error fraction of around 1.53% for nickel, 1.43% for aluminum, and 2.86% for copper, while producing orientation-specific stress–strain curves several times faster than conventional CPFE simulations. Sayak Chakrabarty, Shahriyar Keshavarz, Yuwei Mao, Andrew C. E. Reid, Alok N. Choudhary, Ankit Agrawal 0001 |
ICMLA | 6 |
| 2025 | Large language models accurately identify immunosuppression in intensive care unit patientsabstractOBJECTIVE: Rule-based structured data algorithms and natural language processing (NLP) approaches applied to unstructured clinical notes have limited accuracy and poor generalizability for identifying immunosuppression. Large language models (LLMs) may effectively identify patients with heterogenous types of immunosuppression from unstructured clinical notes. We compared the performance of LLMs applied to unstructured notes for identifying patients with immunosuppressive conditions or immunosuppressive medication use against 2 baselines: (1) structured data algorithms using diagnosis codes and medication orders and (2) NLP approaches applied to unstructured notes. MATERIALS AND METHODS: We used hospital admission notes from a primary cohort of 827 intensive care unit (ICU) patients at Northwestern Memorial Hospital and a validation cohort of 200 ICU patients at Beth Israel Deaconess Medical Center, along with diagnosis codes and medication orders from the primary cohort. We evaluated the performance of structured data algorithms, NLP approaches, and LLMs in identifying 7 immunosuppressive conditions and 6 immunosuppressive medications. RESULTS: In the primary cohort, structured data algorithms achieved peak F1 scores ranging from 0.30 to 0.97 for identifying immunosuppressive conditions and medications. NLP approaches achieved peak F1 scores ranging from 0 to 1. GPT-4o outperformed or matched structured data algorithms and NLP approaches across all conditions and medications, with F1 scores ranging from 0.51 to 1. GPT-4o also performed impressively in our validation cohort (F1 = 1 for 8/13 variables). DISCUSSION: LLMs, particularly GPT-4o, outperformed structured data algorithms and NLP approaches in identifying immunosuppressive conditions and medications with robust external validation. CONCLUSION: LLMs can be applied for improved cohort identification for research purposes. Vijeeth Guggilla, Mengjia Kang, Melissa J. Bak, Steven D. Tran, Anna Pawlowski, Prasanth Nannapaneni, Luke V. Rasmussen, Helen K. Donnelly, Ankit Agrawal 0001, David M. Liebovitz, Alexander V. Misharin, G. R. Scott Budinger, Richard G. Wunderink, Theresa Walunas, Catherine A. Gao, Alan R. Hauser, Alec Peltekian, Alexis Rose Wolfe, Alison L. Szabo, Alok N. Choudhary, Amy Ludwig, Anahid Amani Moghadam, Anjana V. Yeldandi, Ankit Bharat, Anna E. Pawlowski, Anthony M. Joudi, Arjun Prakash Tambe, Ashley J. Smith-Nunez, Benjamin D. Singer, Benjamin J. Ulrich, Betty Tran, Cara J. Gottardi, Chiagozie O. Pickens, Clara J. Schroedl, Daniel Meza, Dulce Sarai Garcia, Egon A. Ozer, Elen Gusman, Elisheva D. Shanes, Emily Mower Provost, Emily M. Olson, Erica Marie Hartmann, Erin A. Korth, Estefani Diaz, Estefany R. Guzman, Francisco J. Martinez, Gabrielle Matias, Hiam Abdala-Valencia, Jack T. Sumner, Jacob I Sznajder, Jacqueline M. Kruser, Jakub Glowala, James M. Walter, Jamie H. Rowell, Jason M. Arnold, John Coleman, Jon W. Lomasney, Joseph Isaac Bailey, Judd F. Hultquist, Justin A. Fiala, Justin Starren, Karen M. Ridge, Karolina Senkow, Kathryn A. Helmin, Khalilah L. Gates, Lacy Simmons, Lesley Pinzon, Lindsey D. Gradone, Lisa F. Wolfe, Lucy Luo, Luisa Morales-Nebreda, Manu Jain, Marc Sala, Maxwell Schleck, Melissa H. Ross, Melissa Querrey, Michael J. Cuttica, Michelle Hinsch Prickett, Nandita R. Nadig, Nathaniel Rhodes, Navdeep S. Chandel, Nikolay S. Markov, Peter H. S. Sporn, Qianli Liu, Rachel B. Kadar, Rachel L. Medernach, Ramon Lorenzo-Redondo, Ravi Kalhan, Rebecca K. Clepp, Richard I. Morimoto, Rogan A. Grant, Ruben J. Mylvaganam, Samuel Fenske, Scott A. Laurenzo, Seung Hye Han, Sophia Nozick, Srinivas Panchamukhi, Stephanie C. Eisenbarth, Suchitra Swaminathan, Susan R. Russell, Taylor A. Poor, Thaddeus Cybulski, Theresa A. Lombardo, Thomas Bolig, Thomas Stoeger, Tien Doan, Timothy Rowe, Wan-Ting Liao, Yuan Luo 0001, Yuliana Sokolenko, Ziyan Lu |
J. Am. Medical Informatics Assoc. | 10 |
| 2024 | Automated Nanoparticle Image Processing Pipeline for AI-Driven Materials CharacterizationabstractRecent innovations have made it possible to produce millions of distinct nanoparticles on a chip. These vast volumes of data are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image preprocessing step before model training can produce significantly higher-performing models in a fraction of the time and make the model more robust to different image noise levels and microscope acquisition settings. The proposed image processing pipeline effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures, resulting in higher performance and significant cost savings. Experiments demonstrate superior performance relative to our baseline, including a 15% improvement in recall and more than a 10% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives in our application, and our best-performing model obtains a precision of 97.3% and weighted F-score of 95.9% on an unseen test set. Additionally, model training time is reduced from 15.5 hours to 32 seconds. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will offer a considerable speedup in the laboratory, allowing researchers to rapidly and accurately analyze much greater volumes of data and accelerate materials discovery. Alexandra L. Day, Carolin B. Wahl, Roberto dos Reis, Wei-keng Liao, Vinayak P. Dravid, Alok N. Choudhary, Ankit Agrawal 0001 |
CIKM | 7 |
| 2024 | Combining Transfer Learning and Representation Learning to Improve Predictive Analytics on Small Materials DataabstractModern data mining methods have seen a widespread and growing application in the field of materials science for regression-based predictive modeling due to their effectiveness in extracting and utilizing the hidden information from the materials datasets. However, due to the costly and time-consuming nature of the methods involved in obtaining the experimental and computational data, the majority of the materials datasets are small in size. Moreover, limited hand-engineered representations available from the raw materials data make it harder to improve the accuracy of predictive models on such small and specialized training datasets. In this paper, we introduce a novel technique that combines transfer learning (TL) and representation learning (RL) using a pre-trained deep neural network to maximize accuracy without additional computational costs on inorganic material properties. The performance of the proposed method is compared against traditional machine learning (ML), and deep neural network models trained from scratch (SC) with elemental fraction (EF) as input, more informative physical attributes (PA) as input (for a stringent comparison), as well as conventional TL and RL techniques using deep neural networks. The results demonstrate that the proposed method can improve the accuracy as compared to SC models and conventional TL and RL techniques. Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
ICMLA | 4 |
| 2024 | Enhancing Deep Neural Network Classification Performance Through Novel Weight Initialization: t-SNE Supported Walsh Matrix ApproachabstractDeep Neural Networks, as a subset of AI, outper-form in understanding complex relationships. The key to this success lies in the network's ability to adapt to problem-specific nuances. During model training, the network dynamically optimizes its weights by updating them during backpropagation while trying to minimize the value of the loss function. Throughout this process, the shaping of model weights is crucially linked to how they were initialized. In this study, we introduce the auxiliary network model, called Sup-Walsh (Support Walsh), which reorganizes weights to enhance class boundaries. We tested our approach on three publicly available datasets using popular classification models. For instance, when using AlexNet [1] on the MNIST dataset [2], integrating Sup-Walsh led to a significant increase in accuracy after first epoch from 14.61% to 78.99%. Similarly, GoogleNet [3] on the FashionMNIST dataset [4] showed a notable 31.61% accuracy difference between configurations without and with Sup-Walsh after first epoch. Across nearly all experiments, our proposed method consistently outperformed existing approaches, demonstrating its potential to improve classification accuracy. Code availability: Code is available at Efficient-Weight-Initializer. Muhammed Nur Talha Kilic, Vishu Gupta, Yuwei Mao, Kewei Wang 0002, Alec Peltekian, Alok N. Choudhary, Wei-keng Liao, Ankit Agrawal 0001 |
ICMLA | 8 |
| 2024 | Tackling the Nonlinearity Problem in Inverse Modeling: Mixture Density Network-Backed Quantized AutoEncoderabstractGenerative models have been widely used in the field of computer vision due to their ability to produce unseen data points. Its application has proven to be useful in various scientific domains such as materials science for generating new microstructure images that require learning nonlinear and one-to-many property-microstructure relationships. However, existing simulation-based solutions for this application are inefficient and time-consuming. Moreover, nonlinearity from the lower to higher dimensions poses considerable challenges. In this work, we propose a novel Mixture Density Network (MDN) based Quan-tized Autoencoder Network designed to generate microstructure images from only a single data point by establishing one-to-many nonlinear relationships from property to microstructure. Once the autoencoder effectively compresses spatial information into the property domain, we use the combination of produced latent vectors and property values to create a supplementary dataset for MDN. Upon completion of the model training, generative structures are extracted and merged to create a framework that generates images based on target property values (i.e., absorption values in this study). The trained MDN demonstrates proficiency in generating latent vectors within distribution, while the proposed Vector Quantized Variational Autoencoder (VQ-VAE) efficiently maps the embedding table to the latent space, generating images from property values within the range of properties observed during training. We demonstrate that our proposed model consistently outperforms the baselines with respect to generating new microstructure images having target properties and overcoming the above-mentioned challenges. Muhammed Nur Talha Kilic, Yuwei Mao, Vishu Gupta, Alok N. Choudhary, Wei-keng Liao, Ankit Agrawal 0001 |
ICMLA | 6 |
| 2024 | Deep Learning Based Inverse Modeling for Materials Design: From Microstructure and Property to ProcessingabstractPolycrystalline materials are crucial in various industries, necessitating a comprehensive understanding of the processing-structure-property-performance (PSPP) relationships. Traditional experimental methods are laborious and slow, while computational approaches predominantly address forward problems, deriving structures and properties from processing conditions. Conversely, inferring processing parameters from desired microstructures and properties remains a crucial yet challenging inverse problem due to the complex and nonlinear mappings involved. In this work, we propose a deep learning-based framework exploring non-sequential and sequential models to address two key inverse problems: predicting processing parameters from microstructures and from properties. Focusing on microstructural texture defined by the orientation distribution function (ODF), we apply our framework to copper, generating a dataset of 31,588 unique processing strain rates ($s$-1) in [0, 1] with corresponding ODFs and homogenized properties through simulations. Our inverse prediction results on processing parameters demonstrate high accuracy, with average test RMSEs of 0.0152 from microstructures and 0.0295 from properties. These findings validate the framework's efficacy as a tool for polycrystalline materials process design, enabling the precise determination of processing methods to achieve desired microstructures and properties. Kewei Wang 0002, Yuwei Mao, Mahmudul Hasan 0016, Md Maruf Billah, Muhammed Nur Talha Kilic, Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Pinar Acar, Ankit Agrawal 0001 |
ICMLA | 10 |
| 2023 | A Case Study of Data Management Challenges Presented in Large-Scale Machine Learning WorkflowsabstractRunning scientific workflow applications on high-performance computing systems provides promising results in terms of accuracy and scalability. An example is the particle track reconstruction research in high-energy physics that consists of multiple machine-learning tasks. However, as the modern HPC system scales up, researchers spend more effort on coordinating the individual workflow tasks due to their increasing demands on computational power, large memory footprint, and data movement among various storage devices. These issues are further exacerbated when intermediate result data must be shared among different tasks and each is optimized to fulfill its own design goals, such as the shortest time or minimal memory footprint. In this paper, we investigate the data management challenges presented in scientific workflows. We observe that individual tasks, such as data generation, data curation, model training, and inference, often use data layouts only best for one's I/O performance but orthogonal to its successive tasks. We propose various solutions by employing alternative data structures and layouts in consideration of two tasks running consecutively in the workflow. Our experimental results show up to a 16.46x and 3.42x speedup for initialization time and I/O time respectively, compared to previous approaches. Claire Songhyun Lee, V. Hewes, Giuseppe Cerati, Jim Kowalkowski, Adam Aurisano, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
CCGrid | 6 |
| 2023 | A Deep Learning Framework for Time-Series Processing-Microstructure-Property PredictionabstractA process simulator provides valuable insights into the evolution of microstructures under various elementary processes, employing the Orientation Distribution Function (ODF) as a representation of the microstructure's texture. However, such simulations often involve complex physical computations, making them time-consuming. To address this, our study introduces an artificial intelligence (AI)-based framework to predict the microstructural texture of polycrystalline materials using a specified deformation process. As a case study, we apply our framework to copper. The dataset includes 3,125 unique processing parameter combinations and their corresponding ODF vectors generated using a process simulator. The resulting predictions enable the calculation of homogenized properties. As opposed to traditional material processing simulations, our AI-driven framework offers faster results with minimal error rates (less than 0.5%). This indicates that our approach is a promising tool for rapidly predicting processing-specific microstructures and properties, thereby offering significant improvements over conventional simulation techniques. Yuwei Mao, Mahmudul Hasan 0016, Claire Songhyun Lee, Muhammed Nur Talha Kilic, Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Pinar Acar, Ankit Agrawal 0001 |
ICMLA | 9 |
| 2023 | Pre-Activation based Representation Learning to Enhance Predictive Analytics on Small Materials DataabstractArtificial intelligence based predictive modeling has become increasingly sought-after in the field of materials science for training property prediction models due to their promising ability to extract and utilize data-driven information from materials data. However, current methods typically use limited hand-engineered fixed-length representations obtained from available composition-based information only, making model inputs a stumbling block when handling small and specialized training datasets. In this paper, we study and propose a method to perform representation learning (RL) that is both applicable and adaptive for generalized use across various domains. We introduce a RL technique that utilizes pre-activation based representations extracted from a model pre-trained using a deep neural network to maximize the accuracy. We perform model training for inorganic material properties using composition-based numerical vectors representing the elemental fractions (EF) of the materials by leveraging source models trained on large datasets to build target models on small datasets and then compare its performance against traditional machine learning (ML), deep neural network and RL-based graph neural network (GNN) models trained from scratch (SC) with EF as input, more informative physical attributes (PA) as input, as well as conventional TL/RL techniques. Using large$(\sim 345K)$datasets for source model training and small computational$(\sim 28K)$and experimental$(\sim 2K)$datasets for target model training and testing, we show that the proposed RL methods help significantly improve the accuracy of the model as compared to the SC models and conventional TL/RL techniques for all data sizes and properties by using only EF as input. We also perform a statistical significance analysis by calculating the p-value to find that the observed improvement in the accuracy of proposed RL model over SC, RL-based GNN, and conventional TL/RL models is indeed significant. Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
IJCNN | 4 |
| 2023 | AI for Learning Deformation Behavior of a Material: Predicting Stress-Strain Curves 4000x Faster Than SimulationsabstractStress-strain curves are important representations of a given material's mechanical properties, which depend primarily on the orientation of the individual crystals in the microstructure. Generating stress-strain curves from numerical methods such as the crystal plasticity finite element (CPFE) simulations is computationally intensive. As a result, it is difficult to generate complete stress-strain curves for all possible orientations of a material. In this work, we propose a bilinear stress-strain curve prediction framework for metallic alloys by integrating supervised and unsupervised deep learning methods via transfer learning principles. As a specific case-study, we focus on predicting stress-strain curves of Nickel (Ni)-based superalloys that have important applications in aerospace industry. Using a small training set of just 100 complete stress-strain curves (4,000 strain steps each) of different orientations generated by CPFE simulation code, we were able to build a model that could accurately predict stress-strain curves (<2 % error) using simple features that could be obtained by running the CPFE simulation for just a single strain step. The proposed model can thus predict the complete stress-strain curve for a given orientation of Ni-based superalloys in a fraction of a second, which amounts to a speedup of over 4000x as compared to the simulation. Yuwei Mao, Shahriyar Keshavarz, Vishu Gupta, Andrew C. E. Reid, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
IJCNN | 7 |
| 2023 | I/O in WRF: A Case Study in Modern Parallel I/O TechniquesabstractLarge-scale parallel applications can face significant I/O performance bottlenecks, making efficient I/O crucial. This work presents a comparative study of several parallel I/O implementations in the Weather Research and Forecasting model, including PnetCDF blocking and non-blocking I/O options, netCDF4, HDF5 Log VOL, and ADIOS. For I/O methods creating files in a canonical data layout, PnetCDF's non-blocking option offers up to 2x improvement over its blocking option and up to 4.5x over HDF5 via netCDF4, demonstrating the effectiveness of the write request aggregation technique. The HDF5 Log VOL outperforms ADIOS with a 4x improvement in write performance when creating files in the log layout, although both require non-negligible time to convert the file back to canonical order for post-run analysis. From these results we extract some observations that can guide I/O strategies for modern parallel codes. Zanhua Huang, Kaiyuan Hou, Ankit Agrawal 0001, Alok N. Choudhary, Robert B. Ross, Wei-keng Liao |
SC | 3 |
| 2022 | Using Multi-Resolution Data to Accelerate Neural Network Training in Scientific ApplicationsabstractNeural networks are powerful solutions to many scientific applications; however, they usually require long model training time due to large training data sets or large model size. Research has been focused on developing numerical optimization algorithms and parallel processing to reduce the training time. In this work, we propose a multi-resolution strategy that can reduce the training time by training the model with the reduced-resolution data samples at the beginning and later switching to the original resolution data samples. This strategy is motivated by the observation that coarser versions of many applications can be solved faster than their denser counterparts, and the solution to a coarser problem could be used to initialize the solution to the denser problem. When applying the idea to neural network training, coarse data can have a similar effect on the learning curves at the early stage as the dense data but requires less time. Once the curves no longer improve significantly, our strategy switches to using the data in original resolution. The key in this process is the ability to generate multiple resolutions of a problem automatically, which could usually be done with scientific applications with spatial and temporal continuity. We use two real-world scientific applications, CosmoFlow and DeepCAM, to evaluate the proposed mixed-resolution training strategy. Our experiment results demonstrate that the proposed training strategy effectively reduces the end-to-end training time while achieving a comparable accuracy to that of the training only with the original data. While maintaining the same model accuracy, our multi-resolution training strategy reduces the end-to-end training time up to 30% and 23% for CosmoFlow and DeepCAM, respectively. Kewei Wang 0002, Sunwoo Lee 0001, Jan Balewski, Alex Sim, Peter Nugent, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu, Wei-keng Liao |
CCGRID | 6 |
| 2022 | Machine Learning for Materials Science (MLMS)abstractArtificial intelligence and machine learning are being increasingly used in scientific domains such as computational fluid dynamics and chemistry. Particularly notable is a recently renewed interest in solving partial differential equations using machine learning models, especially deep neural networks, as partial differential equations arise in many scientific problems of interest. Within materials science literature, there has been a surge in publications on AI-enabled materials discovery, in the last five years. Despite this, the interaction between machine learning researchers and materials scientists (especially, scientists working on structural materials, their microstructures, textures and so on) has been very sparse. On the other hand, AI/ML techniques can clearly be integrated into materials design frameworks (e.g., MGI efforts) to support accelerated materials development, novel simulation methodologies and advanced data analytics. Hence there is an immediate need for exchange of ideas and collaborations between machine learning and materials science communities. We believe a workshop dedicated to this theme would be well- suited to foster such collaborations. The aim of this workshop is to bring together the computer science and materials science communities and foster deeper collaborations between the two to accelerate the adoption of AI/ML in materials science. We hope and envision this workshop to facilitate in building a community of researchers in this interdisciplinar area in the years ahead. Avadhut Sardeshmukh, Sreedhar Reddy, Gautham B. P., Ankit Agrawal 0001 |
KDD | 4 |
| 2022 | BRNet: Branched Residual Network for Fast and Accurate Predictive Modeling of Materials PropertiesabstractMachine Learning (ML) and Deep Learning (DL) have become increasingly popular in the field of materials science for building property prediction models owing to their ability to efficiently extract and understand data-driven relationships between materials composition, structure, and properties. In general, materials property prediction are regression problems with a vector-based input material representation. While fully connected layers have been widely used in deep neural networks to predict materials properties, simply adding more and more layers to create a deep model often degrades their performance due to the vanishing gradient problem, thereby limiting usage. In this paper, we study and propose architectural principles for building deep regression neural networks comprising fully connected layers with numerical vectors that bypass manual feature engineering. We introduce a novel deep regression neural network with branched residual learning, BRNet, consisting of branching of layers to maximize variation of features learned from the input or previous layer and places skip connections after each layer to minimize the information loss due to vanishing gradient. We perform BRNet model training for inorganic material properties using numerical vectors representing the elemental fractions of the compositions of the respective materials and compare its performance against other traditional ML and DL techniques, including ElemNet and IRNet. Using multiple datasets (such as OQMD, MP, JARVIS) for training and testing, we show that BRNet models are significantly more accurate than the state-of-the-art ML methods and DL models for all data sizes by using only raw elemental fractions as input. We also show that BRNet's branched residual learning requires fewer parameters and leads to better convergence during the training phase than other neural networks, thus resulting in faster model training. Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
SDM | 4 |
| 2022 | Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency
Sunwoo Lee 0001, Qiao Kang, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
J. Parallel Distributed Comput. | 4 |
| 2022 | A case study on parallel HDF5 dataset concatenation for high energy physics data analysis
Sunwoo Lee 0001, Kaiyuan Hou, Kewei Wang 0002, Saba Sehrish, Marc F. Paterno, Jim Kowalkowski, Quincey Koziol, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
Parallel Comput. | 9 |
| 2021 | Supporting Data Compression in PnetCDFabstractRecently, the dramatic increase of the data amounts drives up the demand for data compression among HPC applications. Although many file systems and I/O middlewares have incorporated compression features, few high-level parallel I/O libraries support data compression due to the challenges of achieving scalable performance on HPC systems. This paper presents the design and implementation of the variable compression feature in the Parallel NetCDF library. Our design employs the same concept of chunking used by the HDF5 library, but we focus on enabling I/O aggregation across multiple requests to address the challenges on performance and scalability. We evaluate our solution using the I/O kernel of real-world scientific applications and analyze the impacts of data compression on parallel I/O performance. Our result suggests that handling multiple requests at once can significantly improve the parallel I/O performance on chunked and compressed data. Kaiyuan Hou, Qiao Kang, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
IEEE BigData | 4 |
| 2021 | Asynchronous I/O Strategy for Large-Scale Deep Learning ApplicationsabstractMany scientific applications have started using deep learning methods for their classification or regression problems. However, for data-intensive scientific applications, I/O performance can be the major performance bottleneck. In order to effectively solve important real-world problems using deep learning methods on High-Performance Computing (HPC) systems, it is essential to address the poor I/O performance issue in large-scale neural network training. In this paper, we propose an asynchronous I/O strategy that can be generally applied to deep learning applications. Our I/O strategy employs an I/O -dedicated thread per process, that performs I/O operations independently of the training progress. The I/O thread reads many training samples at once to reduce the total number of I/O operations per epoch. Given the fixed amount of training data, the fewer the I/O operations per epoch, the shorter the overall I/O time. The I/O operations are also overlapped with the computations using the double-buffering method. We evaluate our I/O strategy using two real-world scientific applications, CosmoFlow and Neuron-Inverter. Our experimental results demonstrate that the proposed I/O strategy significantly improves the scaling performance without affecting the regression performance. Sunwoo Lee 0001, Qiao Kang, Kewei Wang 0002, Jan Balewski, Alex Sim, Ankit Agrawal 0001, Alok N. Choudhary, Peter Nugent, Kesheng Wu, Wei-keng Liao |
HiPC | 6 |
| 2021 | SIGRNN: Synthetic Minority Instances Generation in Imbalanced Datasets using a Recurrent Neural Network
Reda Al-Bahrani, Dipendra Jha, Qiao Kang, Sunwoo Lee 0001, Zijiang Yang 0008, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary |
ICPRAM | 7 |
| 2021 | Enhancing Phase Mapping for High-throughput X-ray Diffraction Experiments using Fuzzy Clustering
Dipendra Jha, K. V. L. V. Narayanachari, Denis T. Keane, Wei-keng Liao, Alok N. Choudhary, Yip-Wah Chung, Michael J. Bedzyk, Ankit Agrawal 0001 |
ICPRAM | 9 |
| 2020 | Communication-Efficient Local Stochastic Gradient Descent for Scalable Deep LearningabstractSynchronous Stochastic Gradient Descent (SGD) with data parallelism, the most popular parallel training strategy for deep learning, suffers from expensive gradient communications. Local SGD with periodic model averaging is a promising alternative to synchronous SGD. The algorithm allows each worker to locally update its own model, and periodically averages the model parameters across all the workers. While this algorithm enjoys less frequent communications, the convergence rate is strongly affected by the number of workers. In order to scale up the local SGD training without losing accuracy, the number of workers should be sufficiently small so that the model converges reasonably fast. In this paper, we discuss how to exploit the degree of parallelism in local SGD while maintaining model accuracy. Our training strategy employs multiple groups of processes and each group trains a local model based on data parallelism. The local models are periodically averaged across all the groups. Based on this hierarchical parallelism, we design a model averaging algorithm that has a cheaper communication cost than allreduce-based approach. We also propose a practical metric for finding the maximum number of workers that does not cause a significant accuracy loss. Our experimental results demonstrate that our proposed training strategy provides a significantly improved scalability while achieving a comparable model accuracy to synchronous SGD. Sunwoo Lee 0001, Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
IEEE BigData | 3 |
| 2020 | Predicting Resource Requirement in Intermediate Palomar Transient Factory WorkflowabstractQuickly identifying astronomical transients from synoptic surveys is critical to many recent astrophysical discoveries. However, each of the data processing pipelines in these surveys contains dozens of stages with highly varying time and space requirements. Properly predicting the resources required to run these pipelines is critical for the allocation of computing resources and reducing the discovery response time. We propose a machine learning strategy for this prediction task and demonstrate its effectiveness using a set of timing measurements from the intermediate Palomar Transient Factory (iPTF) workflow. The proposed model utilizes the spatiotemporal correlation of astronomical images, where nearby patches of the sky (space) are likely to have a similar number of objects of interest and workflows executed in the recent past (time) are likely to use a similar amount of time because the machines and data storage systems are likely to be in similar states. We capture the relationship among these spatial and temporal features in a Bayesian network and study how they impact the prediction accuracy. This Bayesian network helps us to identify the most influential features for predictions. With proper features, our models achieve errors close to the random variance boundary within batches of images taken at the same time, which can be regarded as the intrinsic limit of prediction accuracy. Qiao Kang, Alex Sim, Peter Nugent, Sunwoo Lee 0001, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu |
CCGRID | 6 |
| 2020 | Improving all-to-many personalized communication in two-phase I/OabstractAs modern parallel computers enter the exascale era, the communication cost for redistributing requests becomes a significant bottleneck in MPIIO routines. The communication kernel for request redistribution, which has an all-to-many personalized communication pattern for application programs with a large number of noncontiguous requests, plays an essential role in the overall performance. This paper explores the available communication kernels for two-phase I/O communication. We generalize the spread-out algorithm to adapt to the all-to-many communication pattern of two-phase I/O by reducing the communication straggler effect. Communication throttling methods that reduce communication contention for asynchronous MPI implementation are adopted to improve communication performance further. Experimental results are presented using different communication kernels running on Cray XC40 Cori and IBM AC922 Summit supercomputers with different I/O patterns. Our study shows that adjusting communication kernel algorithms for different I/O patterns can improve the end-to-end performance up to 10 times compared with default MPI-IO implementations. Qiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 0001, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
SC | 5 |
| 2020 | Improving MPI Collective I/O for High Volume Non-Contiguous Requests With Intra-Node AggregationabstractTwo-phase I/O is a well-known strategy for implementing collective MPI-IO functions. It redistributes I/O requests among the calling processes into a form that minimizes the file access costs. As modern parallel computers continue to grow into the exascale era, the communication cost of such request redistribution can quickly overwhelm collective I/O performance. This effect has been observed from parallel jobs that run on multiple compute nodes with a high count of MPI processes on each node. To reduce the communication cost, we present a new design for collective I/O by adding an extra communication layer that performs request aggregation among processes within the same compute nodes. This approach can significantly reduce inter-node communication contention when redistributing the I/O requests. We evaluate the performance and compare it with the original two-phase I/O on Cray XC40 parallel computers (Theta and Cori) with Intel KNL and Haswell processors. Using I/O patterns from two large-scale production applications and an I/O benchmark, we show our proposed method effectively reduces the communication cost and hence maintains the scalability for a large number of processes. Qiao Kang, Sunwoo Lee 0001, Kaiyuan Hou, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Spatiotemporal Real-Time Anomaly Detection for Supercomputing SystemsabstractThe demands of increasingly large scientific application workflows lead to the need for more powerful supercomputers. As the scale of supercomputing systems have grown, the prediction of fault tolerance has become an increasingly critical area of study, since the prediction of system failures can improve performance by saving checkpoints in advance. We propose a real-time failure detection algorithm that adopts an event-based prediction model. The prediction model is a convolutional neural network that utilizes both traditional event attributes and additional spatio-temporal features. We present a case study using our proposed method with six years of reliability, availability, and serviceability event logs recorded by Mira, a Blue Gene/Q supercomputer at Argonne National Laboratory. In the case study, we have shown that our failure prediction model is not limited to predict the occurrence of failures in general. It is capable of accurately detecting specific types of critical failures such as coolant and power problems within reasonable lead time ranges. Our case study shows that the proposed method can achieve a F1score of 0.56 for general failures, 0.97 for coolant failures, and 0.86 for power failures. Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Alex Sim, Kesheng Wu, Rajkumar Kettimuthu, Pete Beckman, Zhengchun Liu, Wei-keng Liao |
IEEE BigData | 2 |
| 2019 | Improving Scalability of Parallel CNN Training by Adjusting Mini-Batch Size at Run-TimeabstractTraining Convolutional Neural Network (CNN) is a computationally intensive task, requiring efficient parallelization to shorten the execution time. Considering the ever-increasing size of available training data, the parallelization of CNN training becomes more important. Data-parallelism, a popular parallelization strategy that distributes the input data among compute processes, requires the mini-batch size to be sufficiently large to achieve a high degree of parallelism. However, training with large batch size is known to produce a low convergence accuracy. In image restoration problems, for example, the batch size is typically tuned to a small value between 16 ~ 64, making it challenging to scale up the training. In this paper, we propose a parallel CNN training strategy that gradually increases the mini-batch size and learning rate at run-time. While improving the scalability, this strategy also maintains the accuracy close to that of the training with a fixed small batch size. We evaluate the performance of the proposed parallel CNN training algorithm with image regression and classification applications using various models and datasets. Sunwoo Lee 0001, Qiao Kang, Sandeep Madireddy, Prasanna Balaprakash, Ankit Agrawal 0001, Alok N. Choudhary, Rick Archibald, Wei-keng Liao |
IEEE BigData | 5 |
| 2019 | Martensite Start Temperature Predictor for Steels Using Ensemble Data MiningabstractMartensite start temperature (MsT) is an important characteristic of steels, knowledge of which is vital for materials engineers to guide the structural design process of steels. It is defined as the highest temperature at which the austenite phase in steel begins to transform to martensite phase during rapid cooling. Here we describe the development and deployment of predictive models for MsT, given the chemical composition of the material. The data-driven models described here are built on a dataset of about 1000 experimental observations reported in published literature, and the best model developed was found to significantly outperform several existing MsT prediction methods. The data-driven analyses also revealed several interesting insights about the relationship between MsT and the constituent alloying elements of steels. The most accurate predictive model resulting from this work has been deployed in an online web-tool that takes as input the elemental alloying composition of a given steel and predicts its MsT. The online MsT predictor is available at http://info.eecs.northwestern.edu/MsTpredictor. Ankit Agrawal 0001, Abhinav Saboo, Gregory B. Olson, Alok N. Choudhary |
DSAA | 1 |
| 2019 | A Real-Time Iterative Machine Learning Approach for Temperature Profile Prediction in Additive Manufacturing ProcessesabstractAdditive Manufacturing (AM) is a manufacturing paradigm that builds three-dimensional objects from a computer-aided design model by successively adding material layer by layer. AM has become very popular in the past decade due to its utility for fast prototyping such as 3D printing as well as manufacturing functional parts with complex geometries using processes such as laser metal deposition that would be difficult to create using traditional machining. As the process for creating an intricate part for an expensive metal such as Titanium is prohibitive with respect to cost, computational models are used to simulate the behavior of AM processes before the experimental run. However, as the simulations are computationally costly and time-consuming for predicting multiscale multi-physics phenomena in AM, physics-informed data-driven machine-learning systems for predicting the behavior of AM processes are immensely beneficial. Such models accelerate not only multiscale simulation tools but also empower real-time control systems using in-situ data. In this paper, we design and develop essential components of a scientific framework for developing a data-driven model-based real-time control system. Finite element methods are employed for solving time-dependent heat equations and developing the database. The proposed framework uses extremely randomized trees - an ensemble of bagged decision trees as the regression algorithm iteratively using temperatures of prior voxels and laser information as inputs to predict temperatures of subsequent voxels. The models achieve mean absolute percentage errors below 1% for predicting temperature profiles for AM processes. The code is made available for the research community at https://github.com/paularindam/ml-iter-additive. Arindam Paul, Mojtaba Mozaffar, Zijiang Yang 0008, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
DSAA | 7 |
| 2019 | Peak Area Detection Network for Directly Learning Phase Regions from Raw X-ray Diffraction PatternsabstractX-ray diffraction (XRD) is a well-known technique used by scientists and engineers to determine the atomic-scale structures as a basis for understanding the composition-structure-property relationship of materials. The current approach for the analysis of XRD data is a multi-stage process requiring several intensive computations such as integration along 2θ for conversion to 1D patterns (intensity-2θ), background removal by polynomial fitting, and indexing against a large database of reference peaks. It impacts the decisions about the subsequent experiments of the materials under investigation and delays the overall process. In this paper, we focus on eliminating such multi-stage XRD analysis by directly learning the phase regions from the raw (2D) XRD image. We introduce a peak area detection network (PADNet) that directly learns to predict the phase regions using the raw XRD patterns without any need for explicit preprocessing and background removal. PADNet contains specially designed large symmetrical convolutional filters at the first layer to capture the peaks and automatically remove the background by computing the difference in intensity counts across different symmetries. We evaluate PADNet using two sets of XRD patterns collected from SLAC and Bruker D-8 for the Sn-Ti-Zn-O composition space; each set contains 177 experimental XRD patterns with their phase regions. We find that PADNet can successfully classify the XRD patterns independent of the presence of background noise and perform better than the current approach of extrapolating phase region labels based on 1D XRD patterns. Dipendra Jha, Aaron Gilad Kusne, Reda Al-Bahrani, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
IJCNN | 7 |
| 2019 | Transfer Learning Using Ensemble Neural Networks for Organic Solar Cell ScreeningabstractOrganic Solar Cells are a promising technology for solving the clean energy crisis in the world. However, generating candidate chemical compounds for solar cells is a time-consuming process requiring thousands of hours of laboratory analysis. For a solar cell, the most important property is the power conversion efficiency which is dependent on the highest occupied molecular orbitals (HOMO) values of the donor molecules. Recently, machine learning techniques have proved to be very useful in building predictive models for HOMO values of donor structures of Organic Photovoltaic Cells (OPVs). Since experimental datasets are limited in size, current machine learning models are trained on data derived from calculations based on density functional theory (DFT). Molecular line notations such as SMILES or InChI are popular input representations for describing the molecular structure of donor molecules. The two types of line representations encode different information, such as SMILES defines the bond types while InChi defines protonation. In this work, we present an ensemble deep neural network architecture, called SINet, which harnesses both the SMILES and InChI molecular representations to predict HOMO values and leverage the potential of transfer learning from a sizeable DFT-computed dataset- Harvard CEP to build more robust predictive models for relatively smaller HOPV datasets. Harvard CEP dataset contains molecular structures and properties for 2.3 million candidate donor structures for OPV while HOPV contains DFT-computed and experimental values of 350 and 243 molecules respectively. Our results demonstrate significant performance improvement from the use of transfer learning and leveraging both molecular representations. Arindam Paul, Dipendra Jha, Reda Al-Bahrani, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
IJCNN | 6 |
| 2019 | Deep learning based domain knowledge integration for small datasets: Illustrative applications in materials informaticsabstractDeep learning has shown its superiority to traditional machine learning methods in various fields, and in general, its success depends on the availability of large amounts of reliable data. However, in some scientific fields such as materials science, such big data is often expensive or even impossible to collect. Thus given relatively small datasets, most of data-driven methods are based on traditional machine learning methods, and it is challenging to apply deep learning for many tasks in these fields. In order to take the advantage of deep learning even for small datasets, a domain knowledge integration approach is proposed in this work. The efficacy of the proposed approach is tested on two materials science datasets with different types of inputs and outputs, for which domain knowledge-aware convolutional neural networks (CNNs) are developed and evaluated against traditional machine learning methods and standard CNN-based approaches. Experiment results demonstrate that integrating domain knowledge into deep learning can not only improve the model's performance for small datasets, but also make the prediction results more explainable based on domain knowledge. Zijiang Yang 0008, Reda Al-Bahrani, Andrew C. E. Reid, Stefanos Papanikolaou, Surya R. Kalidindi, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
IJCNN | 8 |
| 2019 | IRNet: A General Purpose Deep Residual Regression Framework for Materials DiscoveryabstractMaterials discovery is crucial for making scientific advances in many domains. Collections of data from experiments and first-principle computations have spurred interest in applying machine learning methods to create predictive models capable of mapping from composition and crystal structures to materials properties. Generally, these are regression problems with the input being a 1D vector composed of numerical attributes representing the material composition and/or crystal structure. While neural networks consisting of fully connected layers have been applied to such problems, their performance often suffers from the vanishing gradient problem when network depth is increased. Hence, predictive modeling for such tasks has been mainly limited to traditional machine learning techniques such as Random Forest. In this paper, we study and propose design principles for building deep regression networks composed of fully connected layers with numerical vectors as input. We introduce a novel deep regression network with individual residual learning, IRNet, that places shortcut connections after each layer so that each layer learns the residual mapping between its output and input. We use the problem of learning properties of inorganic materials from numerical attributes derived from material composition and/or crystal structure to compare IRNet's performance against that of other machine learning techniques. Using multiple datasets from the Open Quantum Materials Database (OQMD) and Materials Project for training and evaluation, we show that IRNet provides significantly better prediction performance than the state-of-the-art machine learning approaches currently used by domain scientists. We also show that IRNet's use of individual residual learning leads to better convergence during the training phase than when shortcut connections are between multi-layer stacks while maintaining the same number of parameters. Dipendra Jha, Logan T. Ward, Zijiang Yang 0008, Christopher Wolverton, Ian T. Foster, Wei-keng Liao, Alok N. Choudhary, Ankit Agrawal 0001 |
KDD | 8 |
| 2019 | Scalable Algorithms for MPI Intergroup Allgather and Allgatherv
Qiao Kang, Jesper Larsson Träff, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
Parallel Comput. | 4 |
| 2018 | Parallel DBSCAN Algorithm Using a Data Partitioning Strategy with Spark ImplementationabstractDBSCAN is a well-known clustering algorithm which is based on density and is able to identify arbitrary shaped clusters and eliminate noise data. However, existing parallel implementation strategies based on MPI lack fault tolerance and there is no guarantee that their workload is balanced. Although some of Hadoop-based approaches have been proposed, they do not perform well in terms of scalability since the merge process is not efficient.We propose a scalable parallel DBSCAN algorithm by applying a partitioning strategy. It is implemented in Apache Spark. In order to reduce search time, kdtree is used in our algorithm. To achieve better performance and scalability based on kdtree, we adopt an effective partitioning technique aimed at producing balanced sub-domains which can be computed within Spark executors. Moreover, we came up with a new merging technique: through mapping the relationship between the local points and their bordering neighbors, all the partial clusters which are generated in executors are merged to form the final complete clusters. We have observed and verified (through experiments) that this merging approach is very effective in reducing the time taken for the merge phase and very scalable with increasing the number of processing cores and the generated partial clusters.We implemented the algorithm in Java, evaluated its scalability by using different number of processing cores, and using real and synthetic datasets containing up to several hundred million high-dimensional points. We used three scales of datasets to evaluate our implementation. For small scale, we use 50k, 100k, and 500k data points, obtaining up to a factor of 14.9 speedup when using 16 cores. For medium scale, we use 1.0m, 1.5m, and 1.9m data points, obtaining a factor of 109.2 speedup when using 128 cores. For large scale, we use 61.0m, 91.5m, and 115.9m data points, obtaining a factor of 8344.5 speedup when using 16384 cores. Dianwei Han, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
IEEE BigData | 2 |
| 2018 | Full-Duplex Inter-Group All-to-All Broadcast Algorithms with Optimal BandwidthabstractMPI inter-group collective communication patterns can be viewed as bipartite graphs that divide processes into two disjoint groups in which messages are transferred between but not within the groups. Such communication patterns can serve as basic operations for scientific application workflows. In this paper, we present parallel algorithms for inter-group all-to-all broadcast (Allgather) communication with optimal bandwidth for any message size and process number under single-port communication constraints. We implement the algorithms using MPI point-to-point and intra-group collective communication functions and evaluate their performance on the Cori supercomputer at NERSC. Using message sizes ranging from 256B to 64MB, the experiments show a significant performance improvement achieved by our algorithm, which is up to 9.27 times faster than production MPI libraries that adopt the so called root-gathering algorithm. Qiao Kang, Jesper Larsson Träff, Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
EuroMPI | 4 |
| 2017 | Distinguish Polarity in Bag-of-Words Visualization
Yusheng Xie, Zhengzhang Chen, Ankit Agrawal 0001, Alok N. Choudhary |
AAAI | 3 |
| 2017 | Parallel Deep Convolutional Neural Network Training by Exploiting the Overlapping of Computation and CommunicationabstractTraining Convolutional Neural Network (CNN) is a computationally intensive task whose parallelization has become critical in order to complete the training in an acceptable time. However, there are two obstacles to developing a scalable parallel CNN in a distributed-memory computing environment. One is the high degree of data dependency exhibited in the model parameters across every two adjacent minibatches and the other is the large amount of data to be transferred across the communication channel. In this paper, we present a parallelization strategy that maximizes the overlap of inter-process communication with the computation. The overlapping is achieved by using a thread per compute node to initiate communication after the gradients are available. The output data of backpropagation stage is generated at each model layer, and the communication for the data can run concurrently with the computation of other layers. To study the effectiveness of the overlapping and its impact on the scalability, we evaluated various model architectures and hyperparameter settings. When training VGG-A model using ImageNet data sets, we achieve speedups of 62.97× and 77.97× on 128 compute nodes using mini-batch sizes of 256 and 512, respectively. Sunwoo Lee 0001, Dipendra Jha, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao |
HiPC | 3 |
| 2017 | Building Halo Merger Trees from the Q Continuum SimulationabstractCosmological N-body simulations rank among the most computationally intensive efforts today. A key challenge is the analysis of structure, substructure, and the merger history for many billions of compact particle clusters, called halos. Effectively representing the merging history of halos is essential for many galaxy formation models used to generate synthetic sky catalogs, an important application of modern cosmological simulations. Generating realistic mock catalogs requires computing the halo formation history from simulations with large volumes and billions of halos over many time steps, taking hundreds of terabytes of analysis data. We present fast parallel algorithms for producing halo merger trees and tracking halo substructure from a single-level, density-based clustering algorithm. Merger trees are created from analyzing the halo-particle membership function in adjacent snapshots, and substructure is identified by tracking the "cores" of merging halos – sets of particles near the halo center. Core tracking is performed after creating merger trees and uses the relationships found during tree construction to associate substructures with hosts. The algorithms are implemented with MPI and evaluated on a Cray XK7 supercomputer using up to 16,384 processes on data from HACC, a modern cosmological simulation framework. We present results for creating merger trees from 101 analysis snapshots taken from the Q Continuum, a large volume, high mass resolution, cosmological simulation evolving half a trillion particles. Esteban Rangel, Nicholas Frontiere, Salman Habib 0002, Katrin Heitmann, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary |
HiPC | 6 |
| 2017 | SILVERBACK+: scalable association mining via fast list intersection for columnar social data
Yusheng Xie, Zhengzhang Chen, Diana Palsetia, Goce Trajcevski, Ankit Agrawal 0001, Alok N. Choudhary |
Knowl. Inf. Syst. | 5 |
| 2016 | Evaluation of K-means data clustering algorithm on Intel Xeon PhiabstractIntel Xeon Phi is a processor based on MIC architecture that contains a large number of compute cores with a high local memory bandwidth and 512-bit vector processing units. To achieve high performance on Xeon Phi, it is important for programmers to explore all the software features provided by the Intel compiler and libraries to fully utilize the new hardware resources. In this paper, we use the K-Means algorithm to study the performance of various Intel software settings available for Xeon Phi and their impacts to the performance of K-means. At first we examine different memory layouts for storing data points using Intel compiler-intrinsic functions. During distance calculation, the computational kernel of K-means, when the size of individual input data points is not vector-friendly, we pad the data points to align with the VPU width. At last, we implement a parallel reduction to increase memory access parallelism and cache hits. These techniques enable us to successfully take advantage of thread-level parallelism and data-level parallelism on Xeon Phi. Experimental results demonstrate large performance gains over the default auto-vectorization approach. The K-Means implemented with the proposed techniques achieves up to 68.65% and 56.14% performance improvements for aligned datasets and unaligned datasets, respectively. For high-dimensional aligned datasets, we achieved up to 53.49% performance improvement on a large-scale parallel computer. Sunwoo Lee 0001, Wei-keng Liao, Ankit Agrawal 0001, Nikos Hardavellas, Alok N. Choudhary |
IEEE BigData | 3 |
| 2016 | Materials discovery: Understanding polycrystals from large-scale electron patternsabstractThis paper explores the idea of modeling a large image data collection of polycrystal electron patterns, in order to detect insights in understanding materials discovery. There is an emerging interest in applying big data processing, management and modeling methods to scientific images, which often come in a form and with patterns only interpretable to domain experts. While large-scale machine learning approaches have demonstrated certain superiority in analyzing, summarizing, and providing an understandable route to data types like natural images, speeches and texts, scientific images is still a relatively unexplored area. Deep convolutional neural networks, despite their recent triumph in natural image understanding, are still rarely seen adapted to experimental microscopic images, especially in a large scale. To the best of our knowledge, we present the first deep learning solution towards a scientific image indexing problem using a collection of over 300K microscopic images. The result obtained is 54% better than a dictionary lookup method which is state-of-the-art in the materials science society. Rosanne Liu, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary, Marc De Graef |
IEEE BigData | 2 |
| 2016 | PinterNet: A thematic label curation tool for large image datasetsabstractRecent progress in big data and computer vision with deep learning models has gained a lot of attention. Deep learning has been performed on tasks such as image classification, object detection, image segmentation, image captioning, visual question and answering, using large collections of annotated images. This calls for more curated large image datasets with clearer descriptions, cleaner contents, and diversified usability. However, the curation and labeling of such datasets can be labor-intensive. In this paper, we present PinterNet, an algorithm for automatic curation and label generation from noisy textual descriptions, and also publish a big image dataset containing over 110K images automatically labeled with their themes. Our dataset is hierarchical in nature, it has high level category information which we refer as verticals with fine-grained thematic labels at lower level. This advocates a new type of hierarchical theme classification problem closer to human cognition and of business value. We provide benchmark performances using deep learning models based on AlexNet architecture with different pre-training schemes for this novel task and new data. Rosanne Liu, Diana Palsetia, Arindam Paul, Reda Al-Bahrani, Dipendra Jha, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary |
IEEE BigData | 7 |
| 2016 | A Fatigue Strength Predictor for Steels Using Ensemble Data Mining: Steel Fatigue Strength PredictorabstractFatigue strength is one of the most important mechanical properties of steel. High cost and time for fatigue testing, and potentially disastrous consequences of fatigue failures motivates the development of predictive models for this property. We have developed advanced data-driven ensemble predictive models for this purpose with an extremely high cross-validated accuracy of >98\%, and have deployed these models in a user-friendly online web-tool, which can make very fast predictions of fatigue strength for a given steel represented by its composition and processing information. Such a tool with fast and accurate models is expected to be a very useful resource for the materials science researchers and practitioners to assist in their search for new and improved quality steels. The web-tool is available at http://info.eecs.northwestern.edu/SteelFatigueStrengthPredictor Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 1 |
| 2016 | A Filtering-based Clustering Algorithm for Improving Spatio-temporal Kriging Interpolation AccuracyabstractGeostatistical interpolation is the process that uses existing data and statistical models as inputs to predict data in unobserved spatio-temporal contexts as output. Kriging is a well-known geostatistical interpolation method that minimizes mean square error of prediction. The result interpolated by Kriging is accurate when consistency of statistical properties in data is assumed. However, without this assumption, Kriging interpolation has poor accuracy. To address this problem, this paper presents a new filtering-based clustering algorithm that partitions data into clusters such that the interpolation error within each cluster is significantly reduced, which in turn improves the overall accuracy. Comparisons to traditional Kriging are made with two real-world datasets using two error criteria: normalized mean square error(NMSE) and χ2 test statistics for normalized deviation measurement. Our method has reduced NMSE by more than 50% for both datasets over traditional Kriging. Moreover, χ2 tests have also shown significant improvements of our approach over traditional Kriging. Qiao Kang, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 3 |
| 2016 | Parallel DTFE Surface Density Field ReconstructionabstractWe improve the interpolation accuracy and efficiency of the Delaunay tessellation field estimator (DTFE) for surface density field reconstruction by proposing an algorithm that takes advantage of the adaptive triangular mesh for line-of-sight integration. The costly computation of an intermediate 3D grid is completely avoided by our method and only optimally chosen interpolation points are computed, thus, the overall computational cost is significantly reduced. The algorithm is implemented as a parallel shared-memory kernel for large-scale grid rendered field reconstructions in our distributed-memory framework designed for N-body gravitational lensing simulations in large volumes. We also introduce a load balancing scheme to optimize the efficiency of processing a large number of field reconstructions. Our results show our kernel outperforms existing software packages for volume weighted density field reconstruction, achieving~10x speedup, and our load balancing algorithm gains an additional~3.6x speedup at scales with~16k processes. Esteban Rangel, Nan Li 0022, Salman Habib 0002, Tom Peterka, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
CLUSTER | 5 |
| 2016 | Parallel Implementation of Lossy Data Compression for Temporal Data SetsabstractMany scientific data sets contain temporal dimensions. These are the data storing information at the same spatial location but different time stamps. Some of the biggest temporal datasets are produced by parallel computing applications such as simulations of climate change and fluid dynamics. Temporal datasets can be very large and cost a huge amount of time to transfer among storage locations. Using data compression techniques, files can be transferred faster and save storage space. NUMARCK is a lossy data compression algorithm for temporal data sets that can learn emerging distributions of element-wise change ratios along the temporal dimension and encodes them into an index table to be concisely represented. This paper presents a parallel implementation of NUMARCK. Evaluated with six data sets obtained from climate and astrophysics simulations, parallel NUMARCK achieved scalable speedups of up to 8788 when running 12800 MPI processes on a parallel computer. We also compare the compression ratios against two lossy data compression algorithms, ISABELA and ZFP. The results show that NUMARCK achieved higher compression ratio than ISABELA and ZFP. William Hendrix, Seung Woo Son 0001, Christoph Federrath, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
HiPC | 5 |
| 2015 | Mining Social Media Streams to Improve Public Health Allergy SurveillanceabstractAllergies are one of the most common chronic diseases worldwide. One in five Americans suffer from either allergy or asthma symptoms. With the prevalence of social media, people sharing experiences and opinions on personal health symptoms and concerns on social media are increasing. Mining those publicly available health related data potentially provides valuable healthcare insights. In this paper, we propose a real-time allergy surveillance system that first classifies tweets to identify those that mention actual allergy incidents using bag-of-words model and NaiveBayesMultinomial classifier and applies in-depth text and spatiotemporal analysis. Our experimental results show that the proposed system can detect predominant allergy types with high precision and that allergy-related tweet volume is highly correlated to the weather data (daily maximum temperature). We believe that this is the first study that examines a large-scale social media stream for in-depth analysis of allergy activities. Kathy Lee, Ankit Agrawal 0001, Alok N. Choudhary |
ASONAM | 2 |
| 2015 | Running MAP Inference on Million Node Graphical Models: A High Performance Computing PerspectiveabstractAn important problem in discrete graphical models is the maximum a posterior (MAP) inference problem. Recent research has been focusing on the development of parallel MAP inference algorithm, which scales to graphical models of millions of nodes. In this paper, we introduce a parallel implementation of the recently proposed Bethe-ADMM algorithm using Message Passing Interface (MPI), which allows us to fully utilize the computing power provided by the modern supercomputers with thousands of cores. Experimental results demonstrate that for a broad class of problems, our parallel implementation of Bethe-ADMM scales almost linearly even with thousands of cores. Qiang Fu 0021, Huahua Wang, William Hendrix, Zhengzhang Chen, Ankit Agrawal 0001, Arindam Banerjee 0001, Alok N. Choudhary |
CCGRID | 6 |
| 2015 | Legislative Prediction with Dual Uncertainty Minimization from Heterogeneous InformationabstractVoting on legislative bills to form new laws serves as a key function of most legislature. Predicting the votes of such deliberative bodies leads to better understanding of government policies and generates actionable strategies for social good. In this paper, we present a novel prediction model that maximizes the usage of publicly accessible heterogeneous data, i.e., bill text and lawmakers' profile data, to carry out effective legislative prediction. In particular, we propose to design a probabilistic prediction model which achieves high consistency with past vote records while ensuring the minimum uncertainty of the vote prediction reflecting the firm legal ground often held by the lawmakers. In addition, the proposed legislative prediction model enjoys the following properties: inductive and analytical solution, abilities to deal with the prediction on new bills and new legislators, and robustness to the missing vote issue. We conduct extensive empirical study using the real legislative data and compare with other representative methods in both quantitative political science and data mining communities. The experimental results clearly corroborate that the proposed method provides superior prediction accuracy with visible performance gain. Yu Cheng 0001, Ankit Agrawal 0001, Huan Liu 0001, Alok N. Choudhary |
SDM | 2 |
| 2014 | Indexing bipartite memberships in web graphsabstractMassive bipartite graphs are ubiquitous in real world and have important applications in social networks, biological mechanisms, etc. Consider one billion plus people on Facebook making trillions of connections with millions of organizations. Such big social bipartite graphs are often very skewed and unbalanced, on which traditional indexing algorithms do not perform optimally. In this paper, we propose Arowana, a data-driven algorithm for indexing large unbalanced bipartite graphs. Arowana achieves a high-performance efficiency by building an index tree that incorporates the semantic affinity among unbalanced graphs. Arowana uses probabilistic data structures to minimize space overhead and optimize search. In the experiments, we show that Arowana exhibits significant performance improvements and reduces space overhead over traditional indexing techniques. Yusheng Xie, Zhengzhang Chen, Diana Palsetia, Ankit Agrawal 0001, Alok N. Choudhary |
ASONAM | 4 |
| 2014 | Clique guided community detectionabstractDiscovering communities to understand and model network structures has been a fundamental problem in several fields including social networks, physics, and biology. Many algorithms have been developed for finding the communities. Modularity based technique is fairly new relative to clustering, though it is very popular currently. Although some fast modularity based algorithms exist for detecting communities, the quality of these solutions is limited. At the other extreme, a clique embodies a basic community as it has the greatest possible edge density. However, the requirement that each pair of vertices be connected is too strict. Therefore, techniques to merge partitioned cliques using a hill-climbing greedy algorithm have been studied to form communities. However, the task of finding cliques is computationally expensive. In this paper, we present a new approach for fast and efficient community detection. We propose a clique guided community detection framework that consists of two phases. In the first phase, the framework finds disjoint cliques. In the second phase, the cliques from the first phase are used to guide the merging of individual vertices until a good quality solution is obtained. For the first phase, we develop an algorithm named MaCH (Maximum Clique Heuristic), which is a new approach to compute disjoint cliques using a heuristic-based branch-and-bound technique. We provide experimental results to demonstrate the efficiency of the new algorithm and compare our approach with other previously proposed algorithms. Diana Palsetia, Md. Mostofa Ali Patwary, William Hendrix, Ankit Agrawal 0001, Alok N. Choudhary |
IEEE BigData | 4 |
| 2014 | SILVERBACK: Scalable association mining for temporal data in columnar probabilistic databasesabstractWe address the problem of large scale probabilistic association rule mining and consider the trade-offs between accuracy of the mining results and quest of scalability on modest hardware infrastructure. We demonstrate how extensions and adaptations of research findings can be integrated in an industrial application, and we present the commercially deployed SILVERBACK framework, developed at Voxsup Inc. SILVERBACK tackles the storage efficiency problem by proposing a probabilistic columnar infrastructure and using Bloom filters and reservoir sampling techniques. In addition, a probabilistic pruning technique has been introduced based on Apriori for mining frequent item-sets. The proposed target-driven technique yields a significant reduction on the size of the frequent item-set candidates. We present extensive experimental evaluations which demonstrate the benefits of a context-aware incorporation of infrastructure limitations into corresponding research techniques. The experiments indicate that, when compared to the traditional Hadoop-based approach for improving scalability by adding more hosts, SILVERBACK - which has been commercially deployed and developed at Voxsup Inc. since May 2011 - has much better run-time performance with negligible accuracy sacrifices. Yusheng Xie, Diana Palsetia, Goce Trajcevski, Ankit Agrawal 0001, Alok N. Choudhary |
ICDE | 4 |
| 2014 | Social Role Identification via Dual Uncertainty Minimization RegularizationabstractIn this paper, we study a challenging problem of inferring individuals' role and statuses in a professional social network, which is of central importance in workforce optimization and human capital management. Realizing the natural setting of social nodes associated with dual view information, i.e., The local node characteristics and the global network influence, we present a novel model that explores graph regularization techniques and integrates such information to achieve improved prediction performance. In particular, our prediction model is built upon the graph transductive learning framework that encodes an uncertainty regularization term in the conventional empirical risk minimization principle. Through taking advantage of the information from both the local profile and the global network characteristics, the final inference of the role or statues achieves minimum an empirical loss on the labeled set, as well as a minimum uncertainty on the unlabeled social nodes. We perform extensive empirical study using real-world data and compare with representative peer approaches. The experimental results on three real social network data sets show that the proposed model greatly outperforms a number of baseline models and is able to effectively infer in a wide range of scenarios. Yu Cheng 0001, Ankit Agrawal 0001, Alok N. Choudhary, Huan Liu 0001, Tao Zhang 0006 |
ICDM | 2 |
| 2014 | NUMARCK: Machine Learning Algorithm for Resiliency and CheckpointingabstractData check pointing is an important fault tolerance technique in High Performance Computing (HPC) systems. As the HPC systems move towards exascale, the storage space and time costs of check pointing threaten to overwhelm not only the simulation but also the post-simulation data analysis. One common practice to address this problem is to apply compression algorithms to reduce the data size. However, traditional lossless compression techniques that look for repeated patterns are ineffective for scientific data in which high-precision data is used and hence common patterns are rare to find. This paper exploits the fact that in many scientific applications, the relative changes in data values from one simulation iteration to the next are not very significantly different from each other. Thus, capturing the distribution of relative changes in data instead of storing the data itself allows us to incorporate the temporal dimension of the data and learn the evolving distribution of the changes. We show that an order of magnitude data reduction becomes achievable within guaranteed user-defined error bounds for each data point. We propose NUMARCK, North western University Machine learning Algorithm for Resiliency and Check pointing, that makes use of the emerging distributions of data changes between consecutive simulation iterations and encodes them into an indexing space that can be concisely represented. We evaluate NUMARCK using two production scientific simulations, FLASH and CMIP5, and demonstrate a superior performance in terms of compression ratio and compression accuracy. More importantly, our algorithm allows users to specify the maximum tolerable error on a per point basis, while compressing the data by an order of magnitude. Zhengzhang Chen, Seung Woo Son 0001, William Hendrix, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
SC | 4 |
| 2013 | A probabilistic graphical model for brand reputation assessment in social networksabstractSocial media has become a popular platform that connects people who share information, in particular personal opinions. Through such a fast information exchange mechanism, reputation of individuals, consumer products, or business companies can be quickly built up within a social network. Recently, applications mining social network data start emerging to find the communities sharing the same interests for marketing purposes. Knowing the reputation of social network entities, such as celebrities or business companies, can help develop better strategies for election campaigns or new product advertisements. In this paper, we propose a probabilistic graphical model to collectively measure reputations of entities in social networks. By collecting and analyzing large amount of user activities on Facebook, our model can effectively and efficiently rank entities, such as presidential candidates, professional sport teams, musician bands, and companies, based on their social reputation. The proposed model produces results largely consistent with the two publicly available systems - movie ranking in Internet Movie Database and business school ranking by the US news & World Report - with the correlation coefficients of 0.75 and -0.71, respectively. Kunpeng Zhang 0001, Doug Downey, Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
ASONAM | 6 |
| 2013 | Lung transplant outcome prediction using UNOS dataabstractWe analyze lung transplant data from the United Network for Organ Sharing (UNOS) program with the aim of developing accurate risk prediction models for mortality within 1 year of lung transplant using data mining techniques. The data used in this study is de-identified and consists of 62 predictor attributes, and 1-year posttranplant survial outcome for patients who underwent lung transplant between the years 2005 and 2009. Our dataset had 5,319 such patient instances. Several data mining classification techniques were used on this data along with various data mining optimizations and validations to build predictive models for the abovementioned outcome. Prediction results were evaluated using c-statistic metric, and the highest c-statistic obtained was 0.68. Further, we also applied feature selection techniques to reduce the number of attributes in the model from 50 to 8, without any degradation in c-statistic. The final model was also found to outperform logistic regression, which is the most commonly used technique in predictive healthcare informatics. We believe that the resulting predictive model on the reduced dataset can be quite useful to integrate in a risk calculator to aid both physicians and patients in risk assessment. Ankit Agrawal 0001, Reda Al-Bahrani, Mark J. Russo, Jaishankar Raman, Alok N. Choudhary |
IEEE BigData | 1 |
| 2013 | Colon cancer survival prediction using ensemble data mining on SEER dataabstractWe analyze the colon cancer data available from the SEER program with the aim of developing accurate survival prediction models for colon cancer. Carefully designed preprocessing steps resulted in removal of several attributes and applying several supervised classification methods. We also adopt synthetic minority over-sampling technique (SMOTE) to balance the survival and non-survival classes we have. In our experiments, ensemble voting of the three of the top performing classifiers was found to result in the best prediction performance in terms of prediction accuracy and area under the ROC curve. We evaluated multiple classification schemes to estimate the risk of mortality after 1 year, 2 years and 5 years of diagnosis, on a subset of 65 attributes after the data clean up process, 13 attribute carefully selected using attribute selection techniques, and SMOTE balanced set of the same 13 attributes, while trying to retain the predictive power of the original set of attributes. Moreover, we demonstrate the importance of balancing the classes of the data set to yield better results. Reda Al-Bahrani, Ankit Agrawal 0001, Alok N. Choudhary |
IEEE BigData | 2 |
| 2013 | Elver: Recommending Facebook pages in cold start situation without content featuresabstractRecommender systems are vital to the success of online retailers and content providers. One particular challenge in recommender systems is the “cold start” problem. The word “cold” refers to the items that are not yet rated by any user or the users who have not yet rated any items. We propose Elver to recommend and optimize page-interest targeting on Facebook. Existing techniques for cold recommendation mostly rely on content features in the event of lacking user ratings. Since it is very hard to construct universally meaningful features for the millions of Facebook pages, Elver makes minimal assumption of content features. Elver employs iterative matrix completion technology and nonnegative factorization procedure to work with meagre content inklings. Experiments on Facebook data shows the effectiveness of Elver at different levels of sparsity. Yusheng Xie, Zhengzhang Chen, Kunpeng Zhang 0001, Yu Cheng 0001, Ankit Agrawal 0001, Alok N. Choudhary |
IEEE BigData | 6 |
| 2013 | Feedback-driven multiclass active learning for data streamsabstractActive learning is a promising way to efficiently build up training sets with minimal supervision. Most existing methods consider the learning problem in a pool-based setting. However, in a lot of real-world learning tasks, such as crowdsourcing, the unlabeled samples, arrive sequentially in the form of continuous rapid streams. Thus, preparing a pool of unlabeled data for active learning is impractical. Moreover, performing exhaustive search in a data pool is expensive, and therefore unsuitable for supporting on-the-fly interactive learning in large scale data. In this paper, we present a systematic framework for stream-based multi-class active learning. Following the reinforcement learning framework, we propose a feedback-driven active learning approach by adaptively combining different criteria in a time-varying manner. Our method is able to balance exploration and exploitation during the learning process. Extensive evaluation on various benchmark and real-world datasets demonstrates the superiority of our framework over existing methods. Yu Cheng 0001, Zhengzhang Chen, Lu Liu 0005, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 5 |
| 2013 | Bootstrapping active name disambiguation with crowdsourcingabstractName disambiguation is a challenging and important problem in many domains, such as digital libraries, social media management and people search systems. Traditional methods, based on direct assignment using supervised machine learning techniques, seem to be the most effective, but their performances are highly dependent on the amount of training data, while large data annotation can be expensive and time-consuming requiring hours of manual inspection by a domain expert. To efficiently acquire labeled data, we propose a bootstrapping algorithm for the name disambiguation task based on active learning and crowdsourced labeling. We show that the proposed method can leverage the advantages of exploration and exploitation by combining two strategies, thereby improving the overall quality of the training data at minimal expense. The experimental results on two datasets DBLP and ArnetMiner demonstrate the superiority of our framework over existing methods. Yu Cheng 0001, Zhengzhang Chen, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 4 |
| 2013 | Mining diabetes complication and treatment patterns for clinical decision supportabstractThe fast development of hospital information systems (HIS) produces a large volume of electronic medical records, which provides a comprehensive source for exploratory analysis and statistics to support clinical decision-making. In this paper, we investigate how to utilize the heterogeneous medical records to aid the clinical treatments of diabetes mellitus. Diabetes mellitus, simply diabetes, is a group of metabolic diseases, which is often accompanied with many complications. We propose a Symptom-Diagnosis-Treatment model to mine the diabetes complication patterns and to unveil the latent association mechanism between treatments and symptoms from large volume of electronic medical records. Furthermore, we study the demographic statistics of patient population w.r.t. complication patterns in real data and observe several interesting phenomena. The discovered complication and treatment patterns can help physicians better understand their specialty and learn previous experiences. Our experiments on a collection of one-year diabetes clinical records from a famous geriatric hospital demonstrate the effectiveness of our approaches. Lu Liu 0005, Jie Tang 0001, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
CIKM | 4 |
| 2013 | Random walk-based graphical sampling in unbalanced heterogeneous bipartite social graphsabstractWe investigate sampling techniques in unbalanced heterogeneous bipartite graphs (UHBGs), which have wide applications in real world web-scale social networks. We propose random walked-based link sampling and stratified sampling for UHBGs and show that they have advantages over generic random walk samplers. In addition, each sampler's node degree distribution parameter estimator statistic is analytically derived to be used as a quality indicator. In the experiments, we apply the two sampling techniques, with a baseline node sampling method, to both synthetic and real Facebook data. The experimental results show that random walk-based stratified sampler has significant advantage over node sampler and link sampler on UHBGs. Yusheng Xie, Zhengzhang Chen, Ankit Agrawal 0001, Alok N. Choudhary, Lu Liu 0005 |
CIKM | 3 |
| 2013 | Forecast Oriented Classification of Spatio-Temporal Extreme Events
Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Kunpeng Zhang 0001, Ankit Agrawal 0001, Wei-keng Liao, Nagiza F. Samatova, Alok N. Choudhary |
IJCAI | 5 |
| 2013 | JobMiner: a real-time system for mining job-related patterns from social mediaabstractThe various kinds of booming social media not only provide a platform where people can communicate with each other, but also spread useful domain information, such as career and job market information. For example, LinkedIn publishes a large amount of messages either about people who want to seek jobs or companies who want to recruit new members. By collecting information, we can have a better understanding of the job market and provide insights to job-seekers, companies and even decision makers. In this paper, we analyze the job information from the social network point of view. We first collect the job-related information from various social media sources. Then we construct an inter-company job-hopping network, with the vertices denoting companies and the edges denoting flow of personnel between companies. We subsequently employ graphmining techniques to mine influential companies and related company groups based on the job-hopping network model. Demonstration on LinkedIn data shows that our system JobMiner can provide a better understanding of the dynamic processes and a more accurate identification of important entities in the job market. Yu Cheng 0001, Yusheng Xie, Zhengzhang Chen, Ankit Agrawal 0001, Alok N. Choudhary, Songtao Guo |
KDD | 4 |
| 2013 | Real-time disease surveillance using Twitter data: demonstration on flu and cancerabstractSocial media is producing massive amounts of data on an unprecedented scale. Here people share their experiences and opinions on various topics, including personal health issues, symptoms, treatments, side-effects, and so on. This makes publicly available social media data an invaluable resource for mining interesting and actionable healthcare insights. In this paper, we describe a novel real-time flu and cancer surveillance system that uses spatial, temporal, and text mining on Twitter data. The real-time analysis results are reported visually in terms of US disease surveillance maps, distribution and timelines of disease types, symptoms, and treatments, in addition to overall disease activity timelines on our project website. Our surveillance system can be very useful not only for early prediction of seasonal disease outbreaks such as flu, but also for monitoring distribution of cancer patients with different cancer types and symptoms in each state and the popularity of treatments used. The resulting insights are expected to help facilitate faster response to and preparation for epidemics and also be very useful for both patients and doctors to make more informed decisions. Kathy Lee, Ankit Agrawal 0001, Alok N. Choudhary |
KDD | 2 |
| 2013 | Scalable parallel OPTICS data clustering using graph algorithmic techniquesabstractOPTICS is a hierarchical density-based data clustering algorithm that discovers arbitrary-shaped clusters and eliminates noise using adjustable reachability distance thresholds. Parallelizing OPTICS is considered challenging as the algorithm exhibits a strongly sequential data access order. We present a scalable parallel OPTICS algorithm (Poptics) designed using graph algorithmic concepts. To break the data access sequentiality, POPTICS exploits the similarities between the OPTICS algorithm and Prim's Minimum Spanning Tree algorithm. Additionally, we use the disjoint-set data structure to achieve a high parallelism for distributed cluster extraction. Using high dimensional datasets containing up to a billion floating point numbers, we show scalable speedups of up to 27.5 for our OpenMP implementation on a 40-core shared-memory machine, and up to 3,008 for our MPI implementation on a 4,096-core distributed-memory machine. We also show that the quality of the results given by POPTICS is comparable to those given by the classical OPTICS algorithm. Md. Mostofa Ali Patwary, Diana Palsetia, Ankit Agrawal 0001, Wei-keng Liao, Fredrik Manne, Alok N. Choudhary |
SC | 3 |
| 2013 | Graphical Modeling of Macro Behavioral Targeting in Social NetworksabstractWe investigate a class of emerging online marketing challenges in social networks; macro behavioral targeting (MBT) is introduced as non-personalized broadcasting efforts to massive populations. We propose a new probabilistic graphical model for MBT. Further, a linear-time approximation method is proposed to circumvent an intractable parametric representation of user behaviors. We compare the proposed model with the existing state-of-the-art method on real datasets from social networks. Our model outperforms in all categories by comfortable margins. Ankit Agrawal 0001, Zhengzhang Chen, Yu Cheng 0001, Alok N. Choudhary, Md. Mostofa Ali Patwary, Yusheng Xie, Kunpeng Zhang 0001 |
SDM | 1 |
| 2012 | On active learning in hierarchical classificationabstractMost of the existing active learning algorithms assume all the category labels as independent or consider them in a "flat" structure. However, in reality, there are many applications in which the set of possible labels are often organized in a hierarchical structure. In this paper, we consider the problem of active learning when the categories are represented as a tree. Our goal is to exploit the structure information of the label tree in active learning to select the most informative samples to be labeled. We propose an algorithm that estimates the semantic space, embedding the category hierarchy. In this space, each category label is represented as a prototype and the uncertainty is measured using a variance-based fashion. We also demonstrate notable performance improvement with the proposed approach on synthetic and real datasets. Yu Cheng 0001, Kunpeng Zhang 0001, Yusheng Xie, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 4 |
| 2012 | Parallel hierarchical clustering on shared memory platformsabstractHierarchical clustering has many advantages over traditional clustering algorithms like k-means, but it suffers from higher computational costs and a less obvious parallel structure. Thus, in order to scale this technique up to larger datasets, we present SHRINK, a novel shared-memory algorithm for single-linkage hierarchical clustering based on merging the solutions from overlapping sub-problems. In our experiments, we find that SHRINK provides a speedup of 18–20 on 36 cores on both real and synthetic datasets of up to 250,000 points. Source code for SHRINK is available for download on our website, http://cucis.ece.northwestern.edu. William Hendrix, Md. Mostofa Ali Patwary, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
HiPC | 3 |
| 2012 | VOXSUP: a social engagement frameworkabstractSocial media websites are currently central hubs on the Internet. Major online social media platforms are not only places for individual users to socialize but are increasingly more important as channels for companies to advertise, public figures to engage, etc. In order to optimize such advertising and engaging efforts, there is an emerging challenge for knowledge discovery on today's Internet. The goal of knowledge discovery is to understand the entire online social landscape instead of merely summarizing the statistics. To answer this challenge, we have created VOXSUP as a unified social engagement framework. Unlike most existing tools, VOXSUP not only aggregates and filters social data from the Internet, but also provides what we call Voxsupian Knowledge Discovery (VKD). VKD consists of an almost human-level understanding of social conversations at any level of granularity from a single comment sentiment to multi-lingual inter-platform user demographics. Here we describe the technologies that are crucial to VKD, and subsequently go beyond experimental verification and present case studies from our live VOXSUP system. Yusheng Xie, Daniel Honbo, Alok N. Choudhary, Kunpeng Zhang 0001, Yu Cheng 0001, Ankit Agrawal 0001 |
KDD | 6 |
| 2012 | A new scalable parallel DBSCAN algorithm using the disjoint-set data structureabstractDBSCAN is a well-known density based clustering algorithm capable of discovering arbitrary shaped clusters and eliminating noise data. However, parallelization of DBSCAN is challenging as it exhibits an inherent sequential data access order. Moreover, existing parallel implementations adopt a master-slave strategy which can easily cause an unbalanced workload and hence result in low parallel efficiency. We present a new parallel DBSCAN algorithm (PDSDBSCAN) using graph algorithmic concepts. More specifically, we employ the disjoint-set data structure to break the access sequentiality of DBSCAN. In addition, we use a tree-based bottom-up approach to construct the clusters. This yields a better-balanced workload distribution. We implement the algorithm both for shared and for distributed memory. Using data sets containing up to several hundred million high-dimensional points, we show that PDSDBSCAN significantly outperforms the master-slave approach, achieving speedups up to 25.97 using 40 cores on shared memory architecture, and speedups up to 5,765 using 8,192 cores on distributed memory architecture. Md. Mostofa Ali Patwary, Diana Palsetia, Ankit Agrawal 0001, Wei-keng Liao, Fredrik Manne, Alok N. Choudhary |
SC | 3 |
| 2012 | Sentiment identification by incorporating syntax, semantics and context informationabstractThis paper proposes a method based on conditional random fields to incorporate sentence structure (syntax and semantics) and context information to identify sentiments of sentences within a document. It also proposes and evaluates two different active learning strategies for labeling sentiment data. The experiments with the proposed approach demonstrate a 5-15% improvement in accuracy on Amazon customer reviews compared to existing supervised learning and rule-based methods. Kunpeng Zhang 0001, Yusheng Xie, Yu Cheng 0001, Daniel Honbo, Doug Downey, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
SIGIR | 6 |
| 2012 | Accelerating pairwise statistical significance estimation for local alignment by harvesting GPU's powerabstractBACKGROUND: Pairwise statistical significance has been recognized to be able to accurately identify related sequences, which is a very important cornerstone procedure in numerous bioinformatics applications. However, it is both computationally and data intensive, which poses a big challenge in terms of performance and scalability. RESULTS: We present a GPU implementation to accelerate pairwise statistical significance estimation of local sequence alignment using standard substitution matrices. By carefully studying the algorithm's data access characteristics, we developed a tile-based scheme that can produce a contiguous data access in the GPU global memory and sustain a large number of threads to achieve a high GPU occupancy. We further extend the parallelization technique to estimate pairwise statistical significance using position-specific substitution matrices, which has earlier demonstrated significantly better sequence comparison accuracy than using standard substitution matrices. The implementation is also extended to take advantage of dual-GPUs. We observe end-to-end speedups of nearly 250 (370) × using single-GPU Tesla C2050 GPU (dual-Tesla C2050) over the CPU implementation using Intel Corei7 CPU 920 processor. CONCLUSIONS: Harvesting the high performance of modern GPUs is a promising approach to accelerate pairwise statistical significance estimation for local sequence alignment. Sanchit Misra, Ankit Agrawal 0001, Md. Mostofa Ali Patwary, Wei-keng Liao, Zhiguang Qin, Alok N. Choudhary |
BMC Bioinform. | 3 |
| 2011 | Anatomy of a hash-based long read sequence mapping algorithm for next generation DNA sequencingabstractMOTIVATION: Recently, a number of programs have been proposed for mapping short reads to a reference genome. Many of them are heavily optimized for short-read mapping and hence are very efficient for shorter queries, but that makes them inefficient or not applicable for reads longer than 200 bp. However, many sequencers are already generating longer reads and more are expected to follow. For long read sequence mapping, there are limited options; BLAT, SSAHA2, FANGS and BWA-SW are among the popular ones. However, resequencing and personalized medicine need much faster software to map these long sequencing reads to a reference genome to identify SNPs or rare transcripts. RESULTS: We present AGILE (AliGnIng Long rEads), a hash table based high-throughput sequence mapping algorithm for longer 454 reads that uses diagonal multiple seed-match criteria, customized q-gram filtering and a dynamic incremental search approach among other heuristics to optimize every step of the mapping process. In our experiments, we observe that AGILE is more accurate than BLAT, and comparable to BWA-SW and SSAHA2. For practical error rates (< 5%) and read lengths (200-1000 bp), AGILE is significantly faster than BLAT, SSAHA2 and BWA-SW. Even for the other cases, AGILE is comparable to BWA-SW and several times faster than BLAT and SSAHA2. AVAILABILITY: http://www.ece.northwestern.edu/~smi539/agile.html. Sanchit Misra, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
Bioinform. | 2 |
| 2011 | Parallel pairwise statistical significance estimation of local sequence alignment using Message Passing Interface libraryabstractSUMMARY Homology detection is a fundamental step in sequence analysis. In the recent years, pairwise statistical significance has emerged as a promising alternative to database statistical significance for homology detection. Although more accurate, currently it is much time consuming because it involves generating tens of hundreds of alignment scores to construct the empirical score distribution. This paper presents a parallel algorithm for pairwise statistical significance estimation, called MPIPairwiseStatSig, implemented in C using MPI library. We further apply the parallelization technique to estimate non‐conservative pairwise statistical significance using standard, sequence‐specific, and position‐specific substitution matrices, which has earlier demonstrated superior sequence comparison accuracy than original pairwise statistical significance. Distributing the most compute‐intensive portions of the pairwise statistical significance estimation procedure across multiple processors has been shown to result in near‐linear speed‐ups for the application. The MPIPairwiseStatSig program for pairwise statistical significance estimation is available for free academic use at www.cs.iastate.edu~ankitag/MPIPairwiseStatSig.html . Copyright © 2011 John Wiley & Sons, Ltd. Ankit Agrawal 0001, Sanchit Misra, Daniel Honbo, Alok N. Choudhary |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | Pairwise Statistical Significance of Local Sequence Alignment Using Sequence-Specific and Position-Specific Substitution MatricesabstractPairwise sequence alignment is a central problem in bioinformatics, which forms the basis of various other applications. Two related sequences are expected to have a high alignment score, but relatedness is usually judged by statistical significance rather than by alignment score. Recently, it was shown that pairwise statistical significance gives promising results as an alternative to database statistical significance for getting individual significance estimates of pairwise alignment scores. The improvement was mainly attributed to making the statistical significance estimation process more sequence-specific and database-independent. In this paper, we use sequence-specific and position-specific substitution matrices to derive the estimates of pairwise statistical significance, which is expected to use more sequence-specific information in estimating pairwise statistical significance. Experiments on a benchmark database with sequence-specific substitution matrices at different levels of sequence-specific contribution were conducted, and results confirm that using sequence-specific substitution matrices for estimating pairwise statistical significance is significantly better than using a standard matrix like BLOSUM62, and than database statistical significance estimates reported by popular database search programs like BLAST, PSI-BLAST (without pretrained PSSMs), and SSEARCH on a benchmark database, but with pretrained PSSMs, PSI-BLAST results are significantly better. Further, using position-specific substitution matrices for estimating pairwise statistical significance gives significantly better results even than PSI-BLAST using pretrained PSSMs. Ankit Agrawal 0001, Xiaoqiu Huang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | MPIPairwiseStatSig: parallel pairwise statistical significance estimation of local sequence alignmentabstractSequence comparison is considered as a cornerstone application in bioinformatics, which forms the basis of many other applications. In particular, pairwise sequence alignment is a fundamental step in numerous sequence comparison based applications, where the typical purpose of pairwise sequence alignment step is homology detection, i.e., identifying related sequences. Estimation of statistical significance of a pairwise sequence alignment is crucial in homology detection. A recent development in the field is the use of pairwise statistical significance as an alternative to database statistical significance. Although pairwise statistical significance has been shown to be potentially superior than database statistical significance for homology detection (evaluated in terms of retrieval accuracy), currently it is much time consuming since it involves generating an empirical score distribution by aligning one sequence of the sequence-pair with N random shuffles of the other sequence. In this paper, we present a parallel algorithm for pairwise statistical significance estimation, called MPIPairwiseStatSig, implemented in C using MPI. Distributing the most compute-intensive portions of the pairwise statistical significance estimation procedure across multiple processors has been shown to result in near-linear speed-ups for the application. Ankit Agrawal 0001, Sanchit Misra, Daniel Honbo, Alok N. Choudhary |
HPDC | 1 |
| 2009 | PSIBLAST_PairwiseStatSig: reordering PSI-BLAST hits using pairwise statistical significanceabstractAbstract Summary: We present an add-on to BLAST and PSI-BLAST programs to reorder their hits using pairwise statistical significance. Using position-specific substitution matrices to estimate pairwise statistical significance has been recently shown to give promising results in terms of retrieval accuracy, which motivates its use to refine PSI-BLAST results, since PSI-BLAST also constructs a position-specific substitution matrix for the query sequence during the search. The obvious advantage of the approach is more accurate estimates of statistical significance because of pairwise statistical significance, along with the advantage of BLAST/PSI-BLAST in terms of speed. Availability: The implementation as a C library is freely available at www.cs.iastate.edu/∼ankitag/PSIBLAST_PairwiseStatSig.html Contact: [email protected] Supplementary information: Supplementary data are available at Bionformatics online. Ankit Agrawal 0001, Xiaoqiu Huang 0001 |
Bioinform. | 1 |
| 2009 | Pairwise statistical significance of local sequence alignment using multiple parameter sets and empirical justification of parameter set change penaltyabstractBACKGROUND: Accurate estimation of statistical significance of a pairwise alignment is an important problem in sequence comparison. Recently, a comparative study of pairwise statistical significance with database statistical significance was conducted. In this paper, we extend the earlier work on pairwise statistical significance by incorporating with it the use of multiple parameter sets. RESULTS: Results for a knowledge discovery application of homology detection reveal that using multiple parameter sets for pairwise statistical significance estimates gives better coverage than using a single parameter set, at least at some error levels. Further, the results of pairwise statistical significance using multiple parameter sets are shown to be significantly better than database statistical significance estimates reported by BLAST and PSI-BLAST, and comparable and at times significantly better than SSEARCH. Using non-zero parameter set change penalty values give better performance than zero penalty. CONCLUSION: The fact that the homology detection performance does not degrade when using multiple parameter sets is a strong evidence for the validity of the assumption that the alignment score distribution follows an extreme value distribution even when using multiple parameter sets. Parameter set change penalty is a useful parameter for alignment using multiple parameter sets. Pairwise statistical significance using multiple parameter sets can be effectively used to determine the relatedness of a (or a few) pair(s) of sequences without performing a time-consuming database search. Ankit Agrawal 0001, Xiaoqiu Huang 0001 |
BMC Bioinform. | 1 |
| 2008 | Conservative, Non-conservative and Average Pairwise Statistical Significance of Local Sequence AlignmentabstractEstimation of statistical significance of a pairwise alignment is an important problem in sequence comparison. Recently, it was shown that pairwise statistical significance does better in practice than database statistical significance in terms of retrieval accuracy of homologs. In this paper, we introduce the concept of conservative, non-conservative, and average pairwise statistical significance which can be easily derived from original pairwise statistical significance estimates and use more information specific to the sequence pair under consideration using multiple shuffle spaces. Experimental results for homology detection reveal that the proposed measures give at least comparable or significantly better retrieval accuracy than original pairwise statistical significance and database statistical significance reported by BLAST, PSI-BLAST, and SSEARCH. The use of the proposed measures is further shown to be extremely useful when using sequence-specific substitution matrices. Ankit Agrawal 0001, Xiaoqiu Huang 0001 |
BIBM | 1 |
| 2008 | Pairwise Statistical Significance Versus Database Statistical Significance for Local Alignment of Protein Sequences
Ankit Agrawal 0001, Volker Brendel, Xiaoqiu Huang 0001 |
ISBRA | 1 |
| 2008 | Estimating Pairwise Statistical Significance of Protein Local Alignments Using a Clustering-Classification Approach Based on Amino Acid Composition
Ankit Agrawal 0001, Arka P. Ghosh, Xiaoqiu Huang 0001 |
ISBRA | 1 |