VLDB 2026 Research / reviewers in the wild / expert
Ramakrishnan Kannan
dblp:98/5996
· DBLP profile ↗
28ranked-venue papers in the field
5as first author
13since 2021 · last 2025
0000-0002-5852-4806ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 15 (1 first)Data Mining & Knowledge Discovery · 12 (3 first)Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast Active-Set Thresholding Method for Nonnegative Least Squares
Benjamin Cobb, Ramakrishnan Kannan, Konstantin Pieper, Piyush Sao, Yongseok Soh, Jee W. Choi, Richard W. Vuduc, Haesun Park |
IEEE Big Data | 2 |
| 2025 | Scalable and Efficient Tensor Message-Passing Hypergraph Neural Networks
Syed Ahmed Taimoor, Shruti Shivakumar, Ramakrishnan Kannan, Jiajia Li 0001 |
IEEE Big Data | 4 |
| 2025 | Fragile Earth: Innovative AI For Climate Risk MitigationabstractThe Fragile Earth Workshop is a recurring event in ACM's KDD Conference on research in knowledge discovery and data mining that gathers the research community to find and explore how data science can measure and progress climate and social issues, following the United Nations Sustainable Development Goals (SDGs) framework. Emre Eftelioglu, Naoki Abe, Ramakrishnan Kannan, Kathleen Buckingham, Auroop R. Ganguly, James Hodson 0003 |
KDD (2) | 3 |
| 2024 | On Rank Selection for Nonnegative Matrix FactorizationabstractRank selection, i.e. the choice of factorization rank, is the first step in constructing Nonnegative Matrix Factorization (NMF) models. It is a long-standing problem which is not unique to NMF, but arises in most models which attempt to decompose data into its underlying components. Since these models are often used in the unsupervised setting, the rank selection problem is further complicated by the lack of ground truth labels. In this paper, we review and empirically evaluate the most commonly used schemes for NMF rank selection. Srinivas Eswar, Koby Hayashi, Benjamin Cobb, Ramakrishnan Kannan, Grey Ballard, Richard W. Vuduc, Haesun Park |
IEEE Big Data | 4 |
| 2024 | Counter Data Paucity through Adversarial Invariance Encoding: A Case Study on Modeling Battery Thermal RunawayabstractLithium-ion batteries, widely used for their durability and high energy storage, face the risk of internal short circuits leading to catastrophic thermal runaway events. These events, triggered by external stimuli like mechanical loads, pose safety concerns in applications such as electric vehicles. Detecting and understanding thermal runaway events is crucial, but physics-driven models struggle to explain the non-linear evolution of battery temperature during these events, considering factors like material composition and state-of-charge. Due to the rarity of these events and the cost of data collection, we propose a deep learning (DL) model to predict battery temperature responses during thermal runaway. The challenge lies in the scarcity of data, making traditional DL models prone to overfitting and learning low-quality representations of the complex process.Our approach introduces a novel few-shot architecture that incorporates an adversarially governed invariant encoding process. This architecture aims to distill "invariant" relationships by addressing distributional shifts in data across various battery properties, facilitating the detection of thermal runaway events. Specifically, our results demonstrate that deep learning models conditioned on these "invariant" representations outperform state-of-the-art baselines, achieving a remarkable 96.8% performance improvement in terms of the popular metric MAPE. This framework presents a promising direction for enhancing battery safety modeling, particularly in the context of rare and complex events like thermal runaway. Our code and code and dataset used for the paper are public1. Anika Tabassum, Srikanth Allu, Ramakrishnan Kannan, Nikhil Muralidhar |
IEEE Big Data | 3 |
| 2024 | Fragile Earth: Generative and Foundational Models for Sustainable DevelopmentabstractThe Fragile Earth Workshop is a recurring event in ACM's KDD Conference on research in knowledge discovery and data mining that gathers the research community to find and explore how data science can measure and progress climate and social issues, following the United Nations Sustainable Development Goals (SDGs) framework. Emre Eftelioglu, Bistra Dilkina, Naoki Abe, Ramakrishnan Kannan, Yulia R. Gel, Kathleen Buckingham, Auroop R. Ganguly, James Hodson 0003, Jiafu Mao |
KDD | 4 |
| 2023 | A Deep Learning Pipeline for Optimizing Large-scale Phase Field SimulationsabstractPhase field (PF) simulations are computationally expensive but remain a key analysis tool to understand the complex mechanisms of additive manufacturing (AM) processes. Each PF simulation-aided analysis requires thousands of node hours on leadership-class supercomputers. One of the main goals of these analyses is the study of microstructure evolution during the build process which begins with the onset of nucleation. Nucleation occurs under certain thermomechanical conditions which are not known a priori and many PF simulations are required to identify ranges of input thermo-mechanical parameters that can result in the onset of nucleation. Since many of the simulations do not result in nucleation, an analysis campaign often ends up wasting tremendous amounts of precious computing resources executing nucleation-absent simulations. The goal of this work is to design and train deep learning models to inform a PF simulation about the likelihood of the occurrence of nucleation in a future simulation time-step based on the state summary over a finite number of past time-steps of a running simulation. If the prediction determines that the running simulation is unlikely to reach nucleation in the allotted time, then its execution is stopped immediately ultimately resulting in vast reduction in wasted computations when accrued over all the PF simulations typically performed in a single or multiple analysis campaign(s). The paper presents the performance of a machine learning pipeline that uses a convolutional neural network (CNN) model to learn an embedding which is then used with a self-attention network to build a multi-task deep learning model to predict the likelihood of nucleation. The model also predicts the input parameters used in a simulation. Performance is compared with a baseline pipeline that uses an off-the-shelf LeNet-5 model to learn the initial embedding. Despite their smaller size, performance results indicate significant improvement in accuracy of the proposed models compared to the larger baseline models. Ramakrishnan Kannan, Cristina Garcia-Cardona, Balasubramaniam Radhakrishnan, Sudip K. Seal |
IEEE Big Data | 1 |
| 2023 | Fragile Earth: AI for Climate Sustainability - From Wildfire Disaster Management to Public Health and BeyondabstractThe Fragile Earth Workshop is a recurring event in ACM's KDD Conference on research in knowledge discovery and data mining that gathers the research community to find and explore how data science can measure and progress climate and social issues, fol- lowing the United Nations Sustainable Development Goals (SDGs) framework. Naoki Abe, Kathleen Buckingham, Bistra Dilkina, Emre Eftelioglu, Auroop R. Ganguly, Yulia R. Gel, James Hodson 0003, Ramakrishnan Kannan, Huikyo Lee, Jiafu Mao, Rose Yu |
KDD | 9 |
| 2022 | MatPhase: Material phase prediction for Li-ion Battery Reconstruction using Hierarchical Curriculum LearningabstractLi-ion Batteries (LIB), one of the most efficient energy storage devices, are used extensively in many industrial applications. These batteries consist of electrodes that are put together with heterogeneous material compositions. Imaging data of these battery electrodes obtained from X-ray tomography can explain the distribution of material constituents and allow reconstructions to study electron transport pathways. Such reconstructions of material constituents help quantify various associated properties of electrodes (e.g., volume-specific surface area, porosity) which determine the performance of batteries. These images often suffer from low image contrast between multiple material constituents, hence making it difficult for humans to distinguish and characterize these constituents through visual inspection. A minor error in detecting distributions of the material constituents can lead to magnified errors in the calculated parameters of material properties (e.g., porosity). We present MatPhase, a novel hierarchical curriculum learning technique to address the complex task of estimating material constituent distribution in battery electrodes. MatPhase comprises three modules: (i) an uncertainty-aware global model trained to yield inferences conditioned upon global knowledge of material distribution, (ii) a local model to capture relatively more fine-grained (local) distributional signals, (iii) an aggregator model to appropriately fuse the local and global effects towards obtaining the final distribution. On average, MatPhase improves prediction up to 8.5% relative to other sophisticated modeling pipelines and state-of-the-arts (SOTA) object detection models employed in the performance comparison. Anika Tabassum, Nikhil Muralidhar, Ramakrishnan Kannan, Srikanth Allu |
IEEE Big Data | 3 |
| 2022 | Fragile Earth: AI for Climate Mitigation, Adaptation, and Environmental JusticeabstractThe Fragile EarthWorkshop is a recurring event that gathers the research community to find and explore howdata science can measure and progress climate and social issues, following the framework of the United Nations Sustainable Development Goals (SDGs). Naoki Abe, Kathleen Buckingham, Bistra Dilkina, Emre Eftelioglu, Auroop R. Ganguly, James Hodson 0003, Ramakrishnan Kannan, Rose Yu |
KDD | 7 |
| 2021 | Science-Guided Machine Learning for Wall-Modeled Large Eddy SimulationabstractA deep learning algorithm is trained to predict wall-shear stress based on flow in the outer region of transitional and turbulent boundary layer flow. Flow variables sampled from the boundary layer region of eight high quality wall-resolved Large Eddy Simulations of transonic compressor cascades in the presence of shock-boundary layer interaction effects are used to train the model. About 1.1 TB of data stored in compressed text format generated from the LES calculations is used to perform the model training. The model is shown to be generic and able to predict complex boundary layer physics including laminar to turbulent transition caused by the shock-boundary layer interaction. Rathakrishnan Bhaskaran, Ramakrishnan Kannan, Brian Barr, Stephan Priebe |
IEEE BigData | 2 |
| 2021 | Visual Understanding of COVID-19 Knowledge Graph for Predictive AnalysisabstractThis study aims to effectively analyze and visualize the concept to concept network derived from the COVID-19 Open Research Dataset (CORD-19) dataset, where we have more than 48,000 concepts with more than 300,000 relationships between concepts. In analyzing networks, we focus on finding relationship patterns between the coronavirus disease 2019 (COVID-19) concepts and other concepts. Given the node and edge datasets, we construct directional graphs and calculate all pair shortest paths based on multiple edge weight schemes. However, statistical metrics are not sufficient to identify specific relationships represented in the network. Therefore, we also propose a visual analytics approach to effectively understand the knowledge graph. Our highly interactive visual analytics allows users to effectively analyze the evolving graphs and (COVID-19) concept nodes and other nodes related to the COVID-19 nodes. We envision that this study will pave the path to develop strategies to provide more accurate and scalable predictive analysis on knowledge graphs related to CORD19 and other biomedical knowledge graphs. Seung-Hwan Lim, Junghoon Chae, Guojing Cong, Drahomira Herrmannova, Robert M. Patton, Ramakrishnan Kannan, Thomas E. Potok |
IEEE BigData | 6 |
| 2021 | Fragile Earth: Accelerating Progress towards Equitable SustainabilityabstractFragile Earth 2021, our annual workshop is taking place as part of the Earth Day events at ACM's KDD 2021 Conference on research in Machine Learning and its applications. The 5th edition of Fragile Earth will bring together the research community, industry, and policymakers to develop radically new technological foundations for advancing and meeting the Sustainable Development Goals in a way that ensures equitable and inclusive progress. Naoki Abe, Kathleen Buckingham, Bistra Dilkina, Emre Eftelioglu, Auroop R. Ganguly, James Hodson 0003, Ramakrishnan Kannan |
KDD | 7 |
| 2020 | Structure Prediction from Neutron Scattering Profiles: A Data Sciences ApproachabstractOne of the main goals of neutron data analysis is to determine the internal structure of materials from their neutron scattering profiles. These structures are defined by a crystallographic class label and a set of real-valued parameters specific to that class. Existing structure analysis approaches use computationally expensive loop refinements methods that routinely take days, and even weeks, to complete. Additionally, the outcomes often rely on the fidelity of physical models that are computed during the refinement process. Here, we evaluate the feasibffity of using trained data-driven machine learning models as fast and accurate substitutes for these expensive methods. We report on the efficacies of a variety of ML models, including convolutional neural networks, auto-encoders, random forests and combinations thereof, in addition to techniques such as transfer learning in predicting these structural parameters. Specifically, we evaluate two categories of models which we call class-conditional and integrated. The first relies on a two-stage inference pipeline in which a crystallographic class label is first predicted followed by regression to predict the length/angle parameters. In the second category, the classification and regression tasks are performed as a single learning task. We train these models on synthetically generated data, validate them against experimental observa-tions and show that integrated models outperform their class-conditional counterparts opening up the possibffity of deep learning models as a viable alternative to existing resource-intensive loop refinement methods in neutron data analysis. Cristina Garcia-Cardona, Ramakrishnan Kannan, J. Travis Johnston, Thomas Proffen, Sudip K. Seal |
IEEE BigData | 2 |
| 2019 | Learning to Predict Material Structure from Neutron Scattering DataabstractUnderstanding structural properties of materials and how they relate to its atomic structure, while extremely challenging, is a key scientific quest that has dominated the landscape of materials research for decades. Neutron and X-ray scattering is a state-of-the-art method to investigate material structure on the atomic scale. Traditional methods of processing neutron scattering data to decipher the structure of target materials have relied on computing scattering patterns using physics-based forward models and comparing them with experimentally gathered scattering profiles within a computationally expensive optimization loop. Here, we report an initial design of a data-driven machine learning pipeline for material structure prediction that is computationally faster (once trained) and potentially more accurate. We describe the architecture of the ML pipeline and a preliminary benchmarking study of shallow machine learning models in terms of their prediction accuracy and limitations. We show that material structure prediction from neutron scattering data using shallow learning models is feasible to within 90% prediction accuracy for certain classes of materials but deeper models are required for more general material structure predictions. Cristina Garcia-Cardona, Ramakrishnan Kannan, J. Travis Johnston, Thomas Proffen, Katharine Page, Sudip K. Seal |
IEEE BigData | 2 |
| 2019 | A Scalable Graph Analytics Framework for Programming with Big Data in R (pbdR)abstractMany disciplines such as biology, economics, engineering, physics, and the social sciences represent their data as graphs to capture patterns, trends, and associations. There are are many commercially available graph libraries in different programming languages to analyze these complex graphs. But there is no distributed graph library package in R - the popular statistical programming language to analyze graphs that bigger than a single machine's memory. Many domain experts prefer R over the numerous other alternatives. Towards this, we present a distributed graph analytics framework for R called programming with big graph using R (pBGR.) Our proposed framework leverages the Programming with Big Data in R (pbdR) ecosystem that provides scalable R packages for distributed computing in data science. We present an early prototype implementation of this framework using the distributed-memory parallel graph library CombBLAS and evaluate the framework's performance on leadership class computing platforms. Our experimental results demonstrate that the proposed framework is capable of performing large-scale parallel graph mining through the easyto-use R language. This enhanced graph processing capability coupled with other statistical tools already available in R, should be valuable to many domain experts. S. M. Shamimul Hasan, Drew Schmidt, Ramakrishnan Kannan, Neena Imam |
IEEE BigData | 3 |
| 2018 | A Flexible-blocking Based Approach for Performance Tuning of Matrix Multiplication Routines for Large Matrices with Edge CasesabstractEfficient and scalable matrix operations are being highly demanding in the recent era of Machine Learning, Deep Learning, and Big Data Analytics. The two commonly used matrix-matrix operations in the Basic Linear Algebra Subprograms (BLAS) specification are General Matrix-Matrix multiplication (GEMM) and Symmetric Rank-k update (SYRK). The SYRK routine is a specialization of the GEMM routine, where half of the multiplications are skipped as the resultant matrix is known to be symmetric. Fortunately, several linear algebra libraries implement these BLAS routines quite efficiently. The libraries usually partition the input matrices into blocks and place them in processor caches, thus improving performance by leveraging the caches. However, the contemporary libraries are highly optimized for squarish matrices, but the performance degrades significantly for the matrices with edge case (strictly thin or strictly fat shapes) in the multicore machine. The primary reason is that the current state-of-the-art libraries make fixed block shapes based on a processor architecture, and do not consider the shape of the input matrices. In this paper, we propose a new blocking approach, we name it Flexible-blocking, to mitigate the scalability issues. In contrast to the contemporary libraries, our approach formulates the blocks of the input matrices based on the shapes of the matrices as well as the number of threads used in the implementation. Our proposed technique shows noticeable performance improvement on multicore shared-memory machines for the edge case matrices. Md Mosharaf Hossain, Thomas M. Hines, Sheikh K. Ghafoor, Sheikh Rabiul Islam, Ramakrishnan Kannan, Sreenivas R. Sukumar 0001 |
IEEE BigData | 5 |
| 2018 | VisIRR: A Visual Analytics System for Information Retrieval and Recommendation for Large-Scale Document DataabstractIn this article, we present an interactive visual information retrieval and recommendation system, called VisIRR, for large-scale document discovery. VisIRR effectively combines the paradigms of (1) a passive pull through query processes for retrieval and (2) an active push that recommends items of potential interest to users based on their preferences. Equipped with an efficient dynamic query interface against a large-scale corpus, VisIRR organizes the retrieved documents into high-level topics and visualizes them in a 2D space, representing the relationships among the topics along with their keyword summary. In addition, based on interactive personalized preference feedback with regard to documents, VisIRR provides document recommendations from the entire corpus, which are beyond the retrieved sets. Such recommended documents are visualized in the same space as the retrieved documents, so that users can seamlessly analyze both existing and newly recommended ones. This article presents novel computational methods, which make these integrated representations and fast interactions possible for a large-scale document corpus. We illustrate how the system works by providing detailed usage scenarios. Additionally, we present preliminary user study results for evaluating the effectiveness of the system. Jaegul Choo, Hannah Kim 0001, Edward Clarkson, Zhicheng Liu 0001, Fuxin Li, Hanseung Lee, Ramakrishnan Kannan, Charles D. Stolper, John T. Stasko, Haesun Park |
ACM Trans. Knowl. Discov. Data | 8 |
| 2018 | MPI-FAUN: An MPI-Based Framework for Alternating-Updating Nonnegative Matrix FactorizationabstractNon-negative matrix factorization (NMF) is the problem of determining two non-negative low rank factors Wand H, for the given input matrix A, such that A WH. NMF is a useful tool for many applications in different domains such as topic modeling in text mining, background separation in video analysis, and community detection in social networks. Despite its popularity in the data mining community, there is a lack of efficient parallel algorithms to solve the problem for big data sets. The main contribution of this work is a new, high-performance parallel computational framework for a broad class of NMF algorithms that iteratively solves alternating non-negative least squares (NLS) subproblems for W and H. It maintains the data and factor matrices in memory (distributed across processors), uses MPI for interprocessor communication, and, in the dense case, provably minimizes communication costs (under mild assumptions). The framework is flexible and able to leverage a variety of NMF and NLS algorithms, including Multiplicative Update, Hierarchical Alternating Least Squares, and Block Principal Pivoting. Our implementation allows us to benchmark and compare different algorithms on massive dense and sparse data matrices of size that spans from few hundreds of millions to billions. We demonstrate the scalability of our algorithm and compare it with baseline implementations, showing significant performance improvements. The code and the datasets used for conducting the experiments are available online. Ramakrishnan Kannan, Grey Ballard, Haesun Park |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Scaling up data-parallel analytics platforms: Linear algebraic operation casesabstractLinear algebraic operations such as matrix manipulations form the kernel of many machine learning and other crucial algorithms. Scaling up as well as scaling out such algorithms are key to supporting large scale data analysis that require efficient processing over millions of data samples. To this end, we present, ARION, a hardware acceleration based approach for scaling-up individual tasks of Spark, a popular data-parallel analytics platform. We support both linear algebraic operations of between two dense matrices, and between sparse and dense matrices in distributed environments. ARION provides a flexible control of acceleration according to matrix density, along with efficient scheduling based on runtime resource utilization. We demonstrate the benefit of our approach for general matrix multiplication operations over large matrices with up to four billion elements by using Gramian matrix computation that is commonly used in machine learning. Experiments show that our approach achieves more than 2× and 1.5× end-to-end performance speedups for dense and sparse matrices, respectively, and up to 57.04× faster computation compared to MLlib, a state of the art Spark-based implementation. This work is sponsored in part by the NSF under the grants: CNS-1565314, CNS-1405697, and CNS-1615411. The manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a non-exclusive, paid-up, irrevocable, world-wide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan (http://energy.gov/downloads/doe-public-access-plan). This research used resources of the Oak Ridge Leadership Computing Facility at the Oak Ridge National Laboratory, which is supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC05-00OR22725. Luna Xu, Seung-Hwan Lim, Ali Raza Butt, Ramakrishnan Kannan |
IEEE BigData | 5 |
| 2017 | STExNMF: Spatio-Temporally Exclusive Topic Discovery for Anomalous Event DetectionabstractUnderstanding newly emerging events or topics associated with a particular region of a given day can provide deep insight on the critical events occurring in highly evolving metropolitan cities. We propose herein a novel topic modeling approach on text documents with spatio-temporal information (e.g., when and where a document was published) such as location-based social media data to discover prevalent topics or newly emerging events with respect to an area and a time point. We consider a map view composed of regular grids or tiles with each showing topic keywords from documents of the corresponding region. To this end, we present a tilebased spatio-temporally exclusive topic modeling approach called STExNMF, based on a novel nonnegative matrix factorization (NMF) technique. STExNMF mainly works based on the two following stages: (1) first running a standard NMF of each tile to obtain general topics of the tile and (2) running a spatiotemporally exclusive NMF on a weighted residual matrix. These topics likely reveal information on newly emerging events or topics of interest within a region. We demonstrate the advantages of our approach using the geo-tagged Twitter data of New York City. We also provide quantitative comparisons in terms of the topic quality, spatio-temporal exclusiveness, topic variation, and qualitative evaluations of our method using several usage scenarios. In addition, we present a fast topic modeling technique of our model by leveraging parallel computing. Dear Sungbok Shin, Minsuk Choi, Jinho Choi 0005, Scott Langevin, Christopher Bethune, Philippe Horne, Nathan Kronenfeld, Ramakrishnan Kannan, Barry L. Drake, Haesun Park, Jaegul Choo |
ICDM | 8 |
| 2017 | Outlier Detection for Text DataabstractThe problem of outlier detection is extremely challenging in many domains such as text, in which the attribute values are typically non-negative, and most values are zero. In such cases, it often becomes difficult to separate the outliers from the natural variations in the patterns in the underlying data. In this paper, we present a matrix factorization method, which is naturally able to distinguish the anomalies with the use of low rank approximations of the underlying data. Our iterative algorithm TONMF is based on Block Coordinate Descent (BCD) framework. Our approach has significant advantages over traditional methods for text outlier detection. Finally, we present experimental results illustrating the effectiveness of our method over competing methods. Ramakrishnan Kannan, Hyenkyun Woo, Charu C. Aggarwal, Haesun Park |
SDM | 1 |
| 2016 | Kernels for scalable data analysis in science: Towards an architecture-portable futureabstractIn this paper, we pose and address some of the unique challenges in the analysis of scientific Big Data on supercomputing platforms. Our approach identifies, implements and scales numerical kernels that are critical to the instantiation of theory-inspired analytic workflows on modern computing architectures. We present the benefits of scalable kernels towards constructing algorithms such as principal component analysis and non-negative matrix factorization on an image-analysis use case at the Oak Ridge Leadership Computing Facility (OLCF). Based on experience with the use-case, we conclude that piecing scalable analytic kernels into user-defined analytic workflows are a flexible, modular and agile way to enable architecture-portable productivity for the data-intensive sciences. Sreenivas R. Sukumar 0001, Ramakrishnan Kannan, Seung-Hwan Lim, Michael A. Matheson |
IEEE BigData | 2 |
| 2016 | Mini-apps for high performance data analysisabstractScaling-up scientific data analysis and machine learning algorithms for data-driven discovery is a grand challenge that we face today. Despite the growing need for analysis from science domains that are generating ‘Big Data’ from instruments and simulations, building high-performance analytical workflows of data-intensive algorithms have been daunting because: (i) the ‘Big Data’ hardware and software architecture landscape is constantly evolving, (ii) newer architectures impose new programming models, and (iii) data-parallel kernels of analysis algorithms and their performance facets on different architectures are poorly understood. To address these problems, we have: (i) identified scalable data-parallel kernels of popular data analysis algorithms, (ii) implemented ‘Mini-Apps’ of those kernels using different programming models (e.g. Map Reduce, MPI, etc.), (iii) benchmarked and validated the performance of the kernels in diverse architectures. In this paper, we discuss two of those Mini-Apps and show the execution of principal component analysis built as a workflow of the Mini-Apps. We show that Mini-Apps enable scientists to (i) write domain-specific data analysis code that scales on most HPC hardware and (ii) and offers the ability (most times with over a 10x speed-up) to analyze data sizes 100 times the size of what off-the-shelf desktop/workstations of today can handle. Sreenivas R. Sukumar 0001, Michael A. Matheson, Ramakrishnan Kannan, Seung-Hwan Lim |
IEEE BigData | 3 |
| 2014 | Bounded matrix factorization for recommender system
Ramakrishnan Kannan, Mariya Ishteva, Haesun Park |
Knowl. Inf. Syst. | 1 |
| 2012 | Bounded Matrix Low Rank ApproximationabstractMatrix lower rank approximations such as non-negative matrix factorization (NMF) have been successfully used to solve many data mining tasks. In this paper, we propose a new matrix lower rank approximation called Bounded Matrix Low Rank Approximation (BMA) which imposes a lower and an upper bound on every element of a lower rank matrix that best approximates a given matrix with missing elements. This new approximation models many real world problems, such as recommender systems, and performs better than other methods, such as singular value decompositions (SVD) or NMF. We present an efficient algorithm to solve BMA based on coordinate descent method. BMA is different from NMF as it imposes bounds on the approximation itself rather than on each of the low rank factors. We show that our algorithm is scalable for large matrices with missing elements on multi core systems with low memory. We present substantial experimental results illustrating that the proposed method outperforms the state of the art algorithms for recommender systems such as Stochastic Gradient Descent, Alternating Least Squares with regularization, SVD++, Bias-SVD on real world data sets such as Jester, Movie lens, Book crossing, Online dating and Netflix. Ramakrishnan Kannan, Mariya Ishteva, Haesun Park |
ICDM | 1 |
| 2011 | NIMBLE: a toolkit for the implementation of parallel data mining and machine learning algorithms on mapreduceabstractIn the last decade, advances in data collection and storage technologies have led to an increased interest in designing and implementing large-scale parallel algorithms for machine learning and data mining (ML-DM). Existing programming paradigms for expressing large-scale parallelism such as MapReduce (MR) and the Message Passing Interface (MPI) have been the de facto choices for implementing these ML-DM algorithms. The MR programming paradigm has been of particular interest as it gracefully handles large datasets and has built-in resilience against failures. However, the existing parallel programming paradigms are too low-level and ill-suited for implementing ML-DM algorithms. To address this deficiency, we present NIMBLE, a portable infrastructure that has been specifically designed to enable the rapid implementation of parallel ML-DM algorithms. The infrastructure allows one to compose parallel ML-DM algorithms using reusable (serial and parallel) building blocks that can be efficiently executed using MR and other parallel programming models; it currently runs on top of Hadoop, which is an open-source MR implementation. We show how NIMBLE can be used to realize scalable implementations of ML-DM algorithms and present a performance evaluation. Amol Ghoting, Prabhanjan Kambadur, Edwin P. D. Pednault, Ramakrishnan Kannan |
KDD | 4 |
| 2011 | Fast Rule Mining Over Multi-Dimensional WindowsabstractAssociation rule mining is an indispensable tool for discovering insights from large databases and data warehouses.The data in a warehouse being multi-dimensional, it is often useful to mine rules over subsets of data defined by selections over the dimensions.Such interactive rule mining over multi-dimensional query windows is difficult since rule mining is computationally expensive.Current methods using pre-computation of frequent itemsets require counting of some itemsets by revisiting the transaction database at query time, which is very expensive.We develop a method (RMW) that identifies the minimal set of itemsets to compute and store for each cell, so that rule mining over any query window may be performed without going back to the transaction database.We give formal proofs that the set of itemsets chosen by RMW is sufficient to answer any query and also prove that it is the optimal set to be computed for 1 dimensional queries.We demonstrate through an extensive empirical evaluation that RMW achieves extremely fast query response time compared to existing methods, with only moderate overhead in pre-computation and storage. Mahashweta Das, Deepak P 0001, Prasad Deshpande, Ramakrishnan Kannan |
SDM | 4 |