EDBT 2026 Demo / reviewers in the wild / expert
B. Aditya Prakash
dblp:06/3956
· DBLP profile ↗
80ranked-venue papers in the field
10as first author
20since 2021 · last 2026
0000-0002-3252-455XORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 58 (8 first)Information Retrieval & Web Search · 10 (1 first)Database Systems & Data Management · 9 (1 first)Big Data, Cloud & Distributed Data Systems · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Industrial Demand Forecasting with Temporal and Uncertainty ExplanationsabstractHierarchical time-series forecasting is essential for demand prediction across various industries. While machine learning models have obtained significant accuracy and scalability on such forecasting tasks, the interpretability of their predictions, informed by application, is still largely unexplored. To bridge this gap, we introduce a novel interpretability method for large hierarchical probabilistic time-series forecasting, adapting generic interpretability techniques while addressing challenges associated with hierarchical structures and uncertainty. Our approach offers valuable interpretative insights in response to real-world industrial supply chain scenarios, including 1) the significance of various time-series within the hierarchy and external variables at specific time points, 2) the impact of different variables on forecast uncertainty, and 3) explanations for forecast changes in response to modifications in the training dataset. To evaluate the explainability method, we generate semi-synthetic datasets based on real-world scenarios of explaining hierarchical demands for over ten thousand products at a large chemical company. The experiments showed that our explainability method successfully explained state-of-the-art industrial forecasting methods with significantly higher explainability accuracy. Furthermore, we provide multiple real-world case studies that show the efficacy of our approach in identifying important patterns and explanations that help stakeholders better understand the forecasts. Additionally, our method facilitates the identification of key drivers behind forecasted demand, enabling more informed decision-making and strategic planning. Our approach helps build trust and confidence among users, ultimately leading to better adoption and utilization of hierarchical forecasting models in practice. Harshavardhan Kamarthi, Shangqing Xu, Xinjie Tong, James Peters, Joseph Czyzyk, B. Aditya Prakash |
ICDE | 7 |
| 2026 | TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand ForecastingabstractMulti-horizon time series forecasting has many practical applications such as demand forecasting. Accurate demand prediction is critical to help make buying and inventory decisions for supply chain management of e-commerce and physical retailers, and such predictions are typically required for future horizons extending tens of weeks. This is especially challenging during high-stake sales events when demand peaks are particularly difficult to predict accurately. However, these events are important not only for managing supply chain operations but also for ensuring a seamless shopping experience for customers. To address this challenge, we propose Temporal-Aligned Transformer (TAT), a multi-horizon forecaster leveraging apriori-known context variables such as holiday and promotion events information for improving predictive performance. Our model consists of an encoder and decoder, both embedded with a novel Temporal Alignment Attention (TAA), designed to learn context-dependent alignment for peak demand forecasting. We conduct extensive empirical analysis on two large-scale proprietary datasets from a large e-commerce retailer. We demonstrate that TAT brings up to 30% accuracy improvement on peak demand forecasting while maintaining competitive overall performance compared to other state-of-the-art methods. Zhiyuan Zhao 0002, Sitan Yang, Kin G. Olivares, Boris N. Oreshkin, Stan Vitebsky, Michael W. Mahoney, B. Aditya Prakash, Dmitry Efimov |
ICDE | 7 |
| 2026 | AHA: Scalable Alternative History Analysis for Operational Timeseries Applications
Harshavardhan Kamarthi, Harshil Shah, Henry Milner, Sayan Sinha, B. Aditya Prakash, Vyas Sekar |
KDD (1) | 6 |
| 2025 | In-context Pre-trained Time-Series Foundation Models adapt to Unseen TasksabstractTime-series foundation models (TSFMs) have demonstrated strong generalization capabilities across diverse datasets and tasks. However, existing foundation models are typically pre-trained to enhance performance on specific tasks and often struggle to generalize to unseen tasks without fine-tuning. To address this limitation, we propose augmenting TSFMs with In-Context Learning (ICL) capabilities, enabling them to perform test-time inference by dynamically adapting to input-output relationships provided within the context. Our framework, In-Context Time-series Pre-training (ICTP), restructures the original pre-training data to equip the backbone TSFM with ICL capabilities, enabling adaptation to unseen tasks. Experiments demonstrate that ICT improves the performance of state-of-the-art TSFMs by approximately 11.4% on unseen tasks without requiring fine-tuning. Shangqing Xu, Harshavardhan Kamarthi, Haoxin Liu 0001, B. Aditya Prakash |
CIKM | 4 |
| 2025 | SciSoc LLM Workshop: Large Language Models for Scientific and Societal AdvancesabstractThe proposed ''SciSoc LLM Workshop: Large Language Models for Scientific and Societal Advances'' aims to explore the profound implications and potential of Large Language Models (LLMs) in driving forward scientific inquiry and addressing critical societal challenges. As LLMs such as GPT-4 continue to redefine boundaries in both complexity and capability, their integration into the scientific and societal domains is not just beneficial but essential. In particular, LLMs have demonstrated substantial value in improving our understanding of complex datasets and generating insights across various fields such as healthcare, environmental science, education, and public policy. By bringing together experts and enthusiasts from diverse fields, the workshop aims to foster a comprehensive understanding of how LLMs can redefine traditional research methodologies. Participants will explore innovative ways to harness the power of LLMs for greater efficiency and innovation in their respective fields, potentially catalyzing a new era of scientific and societal advancement. Wei Jin 0009, Lu Cheng 0001, Wenpeng Yin 0001, Xianfeng Tang, Qingsong Wen, Danai Koutra, B. Aditya Prakash, Yan Liu 0002 |
KDD (2) | 7 |
| 2025 | Performative Time-Series ForecastingabstractTime-series forecasting is a critical challenge in various domains and has witnessed substantial progress in recent years. Many real-life scenarios, such as public health, economics, and social applications, involve feedback loops where predictive models can trigger actions that influence the outcome they aim to predict, subsequently altering the target variable's distribution. This phenomenon, known as performativity, introduces the potential for 'self-negating' or 'self-fulfilling' predictions. Despite extensive studies on performativity in classification problems across domains, this phenomenon remains largely unexplored in the context of time-series forecasting from a machine-learning perspective. Zhiyuan Zhao 0002, Haoxin Liu 0001, Alexander Rodríguez, B. Aditya Prakash |
KDD (2) | 4 |
| 2025 | Accurately Estimating Unreported Infections using Information TheoryabstractOne of the most significant challenges in combating against the spread of infectious diseases was the difficulty in estimating the true magnitude of infections. Unreported infections could drive up disease spread, making it very hard to accurately estimate the infectivity of the pathogen, therewith hampering our ability to react effectively. Despite the use of surveillance-based methods such as serological studies, identifying the true magnitude is still challenging. This paper proposes an information theoretic approach for accurately estimating the number of total infections. Our approach is built on top of Ordinary Differential Equations (ODE) based models, which are commonly used in epidemiology and for estimating such infections. We show how we can help such models to better compute the number of total infections and identify the parametrization by which we need the fewest bits to describe the observed dynamics of reported infections. Our experiments on COVID-19 spread show that our approach leads to not only substantially better estimates of the number of total infections but also better forecasts of infections than standard model calibration based methods. We additionally show how our learned parametrization helps in modeling more accurate what-if scenarios with non-pharmaceutical interventions. Our approach provides a general method for improving epidemic modeling which is applicable broadly. Jiaming Cui, Bijaya Adhikari, Arash Haddadan, A. S. M. Ahsan-Ul-Haque, Jilles Vreeken, Anil Vullikanti, B. Aditya Prakash |
SDM | 7 |
| 2024 | A Review of Graph Neural Networks in Epidemic ModelingabstractSince the onset of the COVID-19 pandemic, there has been a growing interest in studying epidemiological models. Traditional mechanistic models mathematically describe the transmission mechanisms of infectious diseases. However, they often fall short when confronted with the growing challenges of today. Consequently, Graph Neural Networks (GNNs) have emerged as a progressively popular tool in epidemic research. In this paper, we endeavor to furnish a comprehensive review of GNNs in epidemic tasks and highlight potential future directions. To accomplish this objective, we introduce hierarchical taxonomies for both epidemic tasks and methodologies, offering a trajectory of development within this domain. For epidemic tasks, we establish a taxonomy akin to those typically employed within the epidemic domain. For methodology, we categorize existing work into Neural Models and Hybrid Models. Following this, we perform an exhaustive and systematic examination of the methodologies, encompassing both the tasks and their technical details. Furthermore, we discuss the limitations of existing methods from diverse perspectives and systematically propose future research directions. This survey aims to bridge literature gaps and promote the progression of this promising field. We hope that it will facilitate synergies between the communities of GNNs and epidemiology, and contribute to their collective progress. Zewen Liu 0005, Guancheng Wan, B. Aditya Prakash, Max S. Y. Lau, Wei Jin 0009 |
KDD | 3 |
| 2024 | Large Scale Hierarchical Industrial Demand Time-Series Forecasting incorporating SparsityabstractHierarchical time-series forecasting (HTSF) is an important problem for many real-world business applications where the goal is to simultaneously forecast multiple time-series that are related to each other via a hierarchical relation. Recent works, however, do not address two important challenges that are typically observed in many demand forecasting applications at large companies. First, many time-series at lower levels of the hierarchy have high sparsity i.e., they have a significant number of zeros. Most HTSF methods do not address this varying sparsity across the hierarchy. Further, they do not scale well to the large size of the real-world hierarchy typically unseen in benchmarks used in literature. We resolve both these challenges by proposing HAILS, a novel probabilistic hierarchical model that enables accurate and calibrated probabilistic forecasts across the hierarchy by adaptively modeling sparse and dense time-series with different distributional assumptions and reconciling them to adhere to hierarchical constraints. We show the scalability and effectiveness of our methods by evaluating them against real-world demand forecasting datasets. We deploy HAILS at a large chemical manufacturing company for a product demand forecasting application with over ten thousand products and observe a significant 8.5% improvement in forecast accuracy and 23% better improvement for sparse time-series. The enhanced accuracy and scalability make HAILS a valuable tool for improved business planning and customer experience. Harshavardhan Kamarthi, Aditya B. Sasanur, Xinjie Tong, James Peters, Joe Czyzyk, B. Aditya Prakash |
KDD | 7 |
| 2024 | epiDAMIK 2024: The 7th International Workshop on Epidemiology meets Data Mining and Knowledge DiscoveryabstractWhile the worst of COVID-19 pandemic has most likely passed us, an occurrence of equally devastating global pandemic or regional epidemic cannot be ruled out in future. H1N1, Zika, SARS, MERS, and Ebola outbreaks over the past few decades have sharply illustrated our enormous vulnerability to emerging infectious diseases. While the data mining research community has demonstrated increased interest in epidemiological applications, much is still left to be desired. For example, there is an urgent need to develop sound theoretical principles and transformative computational approaches that will allow us to address the escalating threat of current and future pandemics. Data mining and knowledge discovery have an important role to play in this regard. Different aspects of infectious disease modeling, analysis, and control have traditionally been studied within the confines of individual disciplines, such as mathematical epidemiology and public health, and data mining and machine learning. Coupled with increasing data generation across multiple domains/sources (e.g., wastewater surveillance, electronic medical records, and social media), there is a clear need for analyzing them to inform public health policies and outcomes timely. Recent advances in disease surveillance and forecasting, and initiatives such as the CDC Flu Challenge, CDC COVID-19 Forecasting Hub etc., have brought these disciplines closer together. On the one hand, public health practitioners seek to use novel datasets, such as Safegraph, Unacast, and Google mobility data, and techniques like Graph Neural Networks. On the other hand, researchers from data mining and machine learning develop novel tools for solving many fundamental problems in the public health policy planning and decision-making process, leveraging novel datasets (e.g., COVID-19 behavioral health surveys, contact tracing trees, and satellite images of urban streets) and combining them with more traditional time series information (e.g., surveillance, hospitalization, and death records). We believe the next stage of advances will result from closer collaborations between these two groups, which is the main objective of epiDAMIK. Alexander Rodríguez, Bijaya Adhikari, Ajitesh Srivastava, Sen Pei, Marie-Laure Charpignon, Kai Wang 0040, Serina Chang, Anil Vullikanti, B. Aditya Prakash |
KDD | 9 |
| 2024 | H2ABM: Heterogeneous Agent-based Model on Hypergraphs to Capture Group InteractionsabstractHeterogeneous agent-based models (HABMs) can simulate the dynamics of multiple types of entities and their interactions on contact networks. In recent years, they have gathered great interest and are widely applied in multiple fields, such as personalized recommendations, publication ranking, and epidemic modeling. Nevertheless, conventional HABMs on graphs can only capture pair-wise interactions between agents but fail to capture the more complex dynamics of group interactions (e.g., multiple people in the same location simultaneously), consequently leading to suboptimal performance. To address this, we propose using hypergraphs to capture such group interactions better and extend the current graph-based HABMs to hypergraphs. Specifically, we use MRSA (Methicillin-resistant Staphylococcus aureus, a kind of infectious disease acquired by patients during treatment at healthcare facilities) spread in the University of Virginia hospital as an example to showcase how we extend an existing graph-based HABM, Graph-HeterSIS, to a hypergraph-based HABM (H2ABM), Hypergraph-HeterSIS. We show how the hyper-graphs can capture the structural difference between contacts before and during the first wave of COVID-19 outbreak in Virginia better than graphs. Our experiments show that H2ABM better captures the underlying group interactions and better fits and forecasts MRSA cases. Vivek Anand, Jiaming Cui, Jack Heavey, Anil Vullikanti, B. Aditya Prakash |
SDM | 5 |
| 2023 | epiDAMIK 6.0: The 6th International Workshop on Epidemiology meets Data Mining and Knowledge DiscoveryabstractThe epiDAMIK workshop serves as a platform for advancing the utilization of data-driven methods in the fields of epidemiology and public health research. These fields have seen relatively limited exploration of data-driven approaches compared to other disciplines. Therefore, our primary objective is to foster the growth and recognition of the emerging discipline of data-driven and computational epidemiology, providing a valuable avenue for sharing state-of-the-art research and ongoing projects. The workshop also seeks to showcase results that are not typically presented at major computing conferences, including valuable insights gained from practical experiences. Our target audience encompasses researchers in AI, machine learning, and data science from both academia and industry, who have a keen interest in applying their work to epidemiological and public health contexts. Additionally, we welcome practitioners from mathematical epidemiology and public health, as their expertise and contributions greatly enrich the discussions. Homepage: https://epidamik.github.io/ Bijaya Adhikari, Alexander Rodríguez, Amulya Yadav, Sen Pei, Ajitesh Srivastava, Marie-Laure Charpignon, Anil Vullikanti, B. Aditya Prakash |
KDD | 8 |
| 2023 | When Rigidity Hurts: Soft Consistency Regularization for Probabilistic Hierarchical Time Series ForecastingabstractProbabilistic hierarchical time-series forecasting is an important variant of time-series forecasting, where the goal is to model and forecast multivariate time-series that have hierarchical relations. Previous works assume rigid consistency over the given hierarchies and do not adapt well to real-world data that show deviation from this assumption. Moreover, recent state-of-art neural probabilistic methods also impose hierarchical relations on point predictions and samples of the predictive distribution. This does not account for full forecast distributions being consistent with the hierarchy and leading to poorly calibrated forecasts. We close both these gaps and propose PROFHiT, a probabilistic hierarchical forecasting model that jointly models forecast distributions over the entire hierarchy. PROFHiT (1) uses a flexible probabilistic Bayesian approach and (2) introduces soft distributional consistency regularization that enables end-to-end learning of the entire forecast distribution leveraging information from the underlying hierarchy. This enables calibrated forecasts as well as adaptation to real-life data with varied hierarchical consistency. PROFHiT provides 41-88% better performance in accuracy and significantly better calibration over a wide range of dataset consistency. Furthermore, PROFHiT adapts to missing data and can provide reliable forecasts even if up to 10% of input time-series data is missing, whereas other methods' performance severely degrades by over 70% Harshavardhan Kamarthi, Alexander Rodríguez, Chao Zhang 0014, B. Aditya Prakash |
KDD | 5 |
| 2023 | Uncertainty Quantification in Deep LearningabstractDeep neural networks (DNNs) have achieved enormous success in a wide range of domains, such as computer vision, natural language processing and scientific areas. However, one key bottleneck of DNNs is that they are ignorant about the uncertainties in their predictions. They can produce wildly wrong predictions without realizing, and can even be confident about their mistakes. Such mistakes can cause misguided decisions-sometimes catastrophic in critical applications, ranging from self-driving cars to cyber security to automatic medical diagnosis. In this tutorial, we present recent advancements in uncertainty quantification for DNNs and their applications across various domains. We first provide an overview of the motivation behind uncertainty quantification, different sources of uncertainty, and evaluation metrics. Then, we delve into several representative uncertainty quantification methods for predictive models, including ensembles, Bayesian neural networks, conformal prediction, and others. We go on to discuss how uncertainty can be utilized for label-efficient learning, continual learning, robust decision-making, and experimental design. Furthermore, we showcase examples of uncertainty-aware DNNs in various domains, such as health, robotics, and scientific machine learning. Finally, we summarize open challenges and future directions in this area. Harshavardhan Kamarthi, Peng Chen 0024, B. Aditya Prakash, Chao Zhang 0014 |
KDD | 4 |
| 2022 | epiDAMIK 5.0: The 5th International Workshop on Epidemiology meets Data Mining and Knowledge DiscoveryabstractSimilar to previous iterations, the epiDAMIK @ KDD workshop is a forum to promote data driven approaches in epidemiology and public health research. Even after the devastating impact of COVID-19 pandemic, data driven approaches are not as widely studied in epidemiology, as they are in other spaces. We aim to promote and raise the profile of the emerging research area of data-driven and computational epidemiology, and create a venue for presenting state-of-the-art and in-progress results-in particular, results that would otherwise be difficult to present at a major data mining conference, including lessons learnt in the 'trenches'. The current COVID-19 pandemic has only showcased the urgency and importance of this area. Our target audience consists of data mining and machine learning researchers from both academia and industry who are interested in epidemiological and public-health applications of their work, and practitioners from the areas of mathematical epidemiology and public health. Homepage: https://epidamik.github.io/. Bijaya Adhikari, Amulya Yadav, Sen Pei, Ajitesh Srivastava, Sarah Kefayati, Alexander Rodríguez, Marie-Laure Charpignon, Anil Vullikanti, B. Aditya Prakash |
KDD | 9 |
| 2022 | Epidemic Forecasting with a Data-Centric LensabstractThe recent COVID-19 pandemic has reinforced the importance of epidemic forecasting to equip decision makers in multiple domains, ranging from public health to economics. However, forecasting the epidemic progression remains a non-trivial task as the spread of diseases is subject to multiple confounding factors spanning human behavior, pathogen dynamics and environmental conditions, etc. Research interest has been fueled by the increased availability of rich data sources capturing previously unseen facets of the epidemic spread and initiatives from government public health and funding agencies like forecasting challenges and funding calls. This has resulted in recent works covering many aspects of epidemic forecasting. Data-centered solutions have specifically shown potential by leveraging non-traditional data sources as well as recent innovations in AI and machine learning. This tutorial will explore various data-driven methodological and practical advancements. First, we will enumerate epidemiological datasets and novel data streams capturing various factors like symptomatic online surveys, retail and commerce, mobility and genomics data. Next, we discuss methods and modeling paradigms with a focus on the recent data-driven statistical and deep-learning based methods as well as novel class of hybrid models that combine domain knowledge of mechanistic models with the effectiveness and flexibility of statistical approaches. We also discuss experiences and challenges that arise in real-world deployment of these forecasting systems including decision-making informed by forecasts. Finally, we highlight some open problems found across the forecasting pipeline. Alexander Rodríguez, Harshavardhan Kamarthi, B. Aditya Prakash |
KDD | 3 |
| 2022 | CAMul: Calibrated and Accurate Multi-view Time-Series ForecastingabstractProbabilistic time-series forecasting enables reliable decision making across many domains. Most forecasting problems have diverse sources of data containing multiple modalities and structures. Leveraging information from these data sources for accurate and well-calibrated forecasts is an important but challenging problem. Most previous works on multi-view time-series forecasting aggregate features from each data view by simple summation or concatenation and do not explicitly model uncertainty for each data view. We propose a general probabilistic multi-view forecasting framework CAMul, which can learn representations and uncertainty from diverse data sources. It integrates the information and uncertainty from each data view in a dynamic context-specific manner, assigning more importance to useful views to model a well-calibrated forecast distribution. We use CAMul for multiple domains with varied sources and modalities and show that CAMul outperforms other state-of-art probabilistic forecasting models by over 25% in accuracy and calibration. Harshavardhan Kamarthi, Alexander Rodríguez, Chao Zhang 0014, B. Aditya Prakash |
WWW | 5 |
| 2021 | Efficient Contingency Analysis in Power Systems via Network Trigger NodesabstractModeling failure dynamics within a power system is a complex and challenging process due to multiple inter-dependencies and convoluted inter-domain relationships. Subject matter experts (SMEs) are interested in understanding these failure dynamics for reducing the impact from future disasters (i.e., losses or failures of power system components, such as transmission lines). Contingency analysis (CA) tools enable such ’what-if’ scenario analyses to evaluate the impacts on the power system. Analyzing all possible contingencies among N system components can be computationally expensive. An important step for performing CA is identifying a set of k ‘trigger’ components, which when failed initially can significantly impact the overall system by causing multiple failures. Currently SMEs focus on identifying these trigger components by running expensive simulations on all possible subsets, which quickly becomes infeasible. Hence finding a relevant set of trigger components (contingencies) rapidly to enable efficient and useful CA is crucial.In a collaboration between computer scientists and power system experts, we propose an efficient method for performing CA by exploiting network inter-dependencies in power system components. First, we construct a network with multiple electric grid infrastructure components and dependencies as connections among them. We reformulate the problem of finding a set of trigger components as a problem of identifying critical nodes in the network, which can cascade power failures through connected nodes and cause significant damage to the network. To guide the practical CA tools, we develop a network-based model with a probabilistic edge-weights setup using intricate domain rules. Then we conduct an empirical study on real power system data in the US for both regional and national levels. Firstly, we use power system datasets for the US to create a national-scale domain-driven model. Secondly, we demonstrate that network-based model outperforms the outputs from a real CA tool and show on average 25 × improved selection of contingencies, thereby showcasing practical benefits to the power experts. Anika Tabassum, Supriya Chinthavali, Sangkeun Matt Lee, Nils M. Stenvig, Bill Kay, P. Teja Kuruganti, B. Aditya Prakash |
IEEE BigData | 7 |
| 2021 | Actionable Insights in Urban Multivariate Time-seriesabstractMultivariate time-series data are gaining popularity in various urban applications, such as emergency management, public health, etc. Segmentation algorithms mostly focus on identifying discrete events with changing phases in such data. For example, consider a power outage scenario during a hurricane. Each time-series can represent the number of power failures in a county for a time period. Segments in such time-series are found in terms of different phases, such as, when a hurricane starts, counties face severe damage, and hurricane ends. Disaster management domain experts typically want to identify the most affected counties (time-series of interests) during these phases. These can be effective for retrospective analysis and decision-making for resource allocation to those regions to lessen the damage. However, getting these actionable counties directly (either by simple visualization or looking into the segmentation algorithm) is typically hard. Hence we introduce and formalize a novel problem RaTSS (Rationalization for time-series segmentation) that aims to find such time-series (rationalizations), which are actionable for the segmentation. We also propose an algorithm Find-RaTSS to find them for any black-box segmentation. We show Find-RaTSS outperforms non-trivial baselines on generalized synthetic and real data, also provides actionable insights in multiple urban domains, especially disasters and public health. Anika Tabassum, Supriya Chinthavali, Varisara Tansakul, B. Aditya Prakash |
CIKM | 4 |
| 2021 | The 4th International Workshop on Epidemiology meets Data Mining and Knowledge Discovery (epiDAMIK 4.0 @ KDD2021)abstractThe 4th [email protected] workshop is a forum to discuss new insights into how data mining can play a bigger role in epidemiology and public health research. While the integration of data science methods into epidemiology has significant potential, it remains under studied. We aim to raise the profile of this emerging research area of data-driven and computational epidemiology, and create a venue for presenting state-of-the-art and in-progress results-in particular, results that would otherwise be difficult to present at a major data mining conference, including lessons learnt in the 'trenches'. The current COVID-19 pandemic has only showcased the urgency and importance of this area. Our target audience consists of data mining and machine learning researchers from both academia and industry who are interested in epidemiological and public-health applications of their work, and practitioners from the areas of mathematical epidemiology and public health. Bijaya Adhikari, Ajitesh Srivastava, Sen Pei, Sarah Kefayati, Rose Yu, Amulya Yadav, Alexander Rodríguez, Arvind Ramanathan, Anil Vullikanti, B. Aditya Prakash |
KDD | 10 |
| 2020 | Mapping Network States using Connectivity QueriesabstractCan we infer all the failed components of an infrastructure network, given a sample of reachable nodes from supply nodes? One of the most critical post-disruption processes after a natural disaster is to quickly determine the damage or failure states of critical infrastructure components. However, this is nontrivial, considering that often only a fraction of components may be accessible or observable after a disruptive event. Past work has looked into inferring failed components given point probes, i.e. with a direct sample of failed components. In contrast, we study the harder problem of inferring failed components given partial information of some `serviceable' reachable nodes and a small sample of point probes, being the first often more practical to obtain. We formulate this novel problem using the Minimum Description Length (MDL) principle, and then present a greedy algorithm that minimizes MDL cost effectively. We evaluate our algorithm on domain-expert simulations of real networks in the aftermath of an earthquake. Our algorithm successfully identifies failed components, especially the critical ones affecting the overall system performance. Alexander Rodríguez, Bijaya Adhikari, Andrés D. González, Charles D. Nicholson, Anil Vullikanti, B. Aditya Prakash |
IEEE BigData | 6 |
| 2020 | Cut-n-Reveal: Time Series Segmentations with ExplanationsabstractRecent hurricane events have caused unprecedented amounts of damage on critical infrastructure systems and have severely threatened our public safety and economic health. The most observable (and severe) impact of these hurricanes is the loss of electric power in many regions, which causes breakdowns in essential public services. Understanding power outages and how they evolve during a hurricane provides insights on how to reduce outages in the future, and how to improve the robustness of the underlying critical infrastructure systems. In this article, we propose a novel scalable segmentation with explanations framework to help experts understand such datasets. Our method, CnR (Cut-n-Reveal), first finds a segmentation of the outage sequences based on the temporal variations of the power outage failure process so as to capture major pattern changes. This temporal segmentation procedure is capable of accounting for both the spatial and temporal correlations of the underlying power outage process. We then propose a novel explanation optimization formulation to find an intuitive explanation of the segmentation such that the explanation highlights theculprittime series of the change in each segment. Through extensive experiments, we show that our method consistently outperforms competitors in multiple real datasets with ground truth. We further study real county-level power outage data from several recent hurricanes (Matthew, Harvey, Irma) and show that CnR recovers important, non-trivial, and actionable patterns for domain experts, whereas baselines typically do not give meaningful results. Nikhil Muralidhar, Anika Tabassum, Liangzhe Chen, Supriya Chinthavali, Naren Ramakrishnan, B. Aditya Prakash |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2019 | EpiDeep: Exploiting Embeddings for Epidemic ForecastingabstractInfluenza leads to regular losses of lives annually and requires careful monitoring and control by health organizations. Annual influenza forecasts help policymakers implement effective countermeasures to control both seasonal and pandemic outbreaks. Existing forecasting techniques suffer from problems such as poor forecasting performance, lack of modeling flexibility, data sparsity, and/or lack of intepretability. We propose EpiDeep, a novel deep neural network approach for epidemic forecasting which tackles all of these issues by learning meaningful representations of incidence curves in a continuous feature space and accurately predicting future incidences, peak intensity, peak time, and onset of the upcoming season. We present extensive experiments on forecasting ILI (influenza-like illnesses) in the United States, leveraging multiple metrics to quantify success. Our results demonstrate that EpiDeep is successful at learning meaningful embeddings and, more importantly, that these embeddings evolve as the season progresses. Furthermore, our approach outperforms non-trivial baselines by up to 40%. Bijaya Adhikari, Xinfeng Xu, Naren Ramakrishnan, B. Aditya Prakash |
KDD | 4 |
| 2019 | Joint Post and Link-level Influence Modeling on Social MediaabstractMicroblogging websites, like Twitter and Weibo, are used by billions of people to create and spread information. This activity depends on various factors such as the friendship links between users, their topic interests and social influence between them. Social influence can be thought of as a latent factor, that may alter users posting and linking behaviors. Making sense of these behaviors is very important for fully understanding and utilizing these platforms. Most prior work in this space either ignores the effect of social influence, or considers its effect only on link formation or post generation. In contrast, we propose PoLIM, leveraging simple weak supervision, a novel model which jointly models the effect of influence on both link and post generation. We also give PoLIM-FIT, an efficient parallel inference algorithm which scales to large datasets. In our experiments on a large tweets corpus, we detect meaningful topical communities, celebrities, as well as the influence strengths patterns among them. Further, we find that there are significant portions of posts and links that are caused by influence, and this portion increases when the data focuses on a specific event. We also show that differentiating and identifying these influenced content benefits other specific quantitative downstream tasks as well, like predicting future tweets and link formation, where we significantly outperform state-of-the-art. Liangzhe Chen, B. Aditya Prakash |
SDM | 2 |
| 2019 | Data-driven efficient network and surveillance-based immunization
Yao Zhang 0003, Arvind Ramanathan, Anil Vullikanti, Laura L. Pullum, B. Aditya Prakash |
Knowl. Inf. Syst. | 5 |
| 2018 | NetGist: Learning to Generate Task-Based Network SummariesabstractGiven a network, can we visualize it for any given task, highlighting the important characteristics? Networks are widespread, and hence summarizing and visualizing them is of primary interest for many applications such as viral marketing, extracting communities and immunization. Summaries can help in solving new problems in visualization, sense-making, and in many other goals. However, most prior work focuses on generic structural summarization techniques or on developing specific algorithms for specific tasks. This is both tedious and challenging. As a result, for several popular tasks, there do not exist readymade summarization methods. In this paper, we explore a promising alternative approach instead. We propose NetGist, a framework which automatically learns how to generate a summary for a given task on a given network. In addition to generating the required summary, this also allows us to reuse the learned process on other similar networks. We formulate a novel task-based graph summarization problem and leverage reinforcement learning to design a flexible framework for our solution. Via extensive experiments, we show that NetGist robustly and effectively learns meaningful summaries, and helps solve challenging problems, and aids in complex task-based sense-making of networks. Sorour E. Amiri, Bijaya Adhikari, Aditya Bharadwaj, B. Aditya Prakash |
ICDM | 4 |
| 2018 | DeepDiffuse: Predicting the 'Who' and 'When' in CascadesabstractCascades are an accepted model to capturing how information diffuses across social network platforms. A large body of research has been focused on dissecting the anatomy of such cascades and forecasting their progression. One recurring theme involves predicting the next stage(s) of cascades utilizing pertinent information such as the underlying social network, structural properties of nodes (e.g., degree) and (partial) histories of cascade propagation. However, such type of granular information is rarely available in practice. We study in this paper the problem of cascade prediction utilizing only two types of (coarse) information, viz. which node is infected and its corresponding infection time. We first construct several simple baselines to solve this cascade prediction problem. Then we describe the shortcomings of these methods and propose a new solution leveraging recent progress in embeddings and attention models from representation learning. We also perform an exhaustive analysis of our methods on several real world datasets. Our proposed model outperforms the baselines and several other state-of-the-art methods. Mohammad Raihanul Islam, Sathappan Muthiah, Bijaya Adhikari, B. Aditya Prakash, Naren Ramakrishnan |
ICDM | 4 |
| 2018 | Sub2Vec: Feature Learning for Subgraphs
Bijaya Adhikari, Yao Zhang 0003, Naren Ramakrishnan, B. Aditya Prakash |
PAKDD (2) | 4 |
| 2018 | SIGNet: Scalable Embeddings for Signed Networks
Mohammad Raihanul Islam, B. Aditya Prakash, Naren Ramakrishnan |
PAKDD (2) | 2 |
| 2018 | Near-Optimal Mapping of Network States using ProbesabstractIn many applications, such as the Internet and infrastructure networks, nodes fail or get congested dynamically. We study the problem of inferring all the failed nodes, when only a sample of the failures is known, and there exist correlations between node failures/congestion in networks. We formalize this as the GraphStateInf problem, using the Minimum Description Length (MDL) principle. We propose the GraphMap algorithm for minimizing the MDL cost, and show that it gives an additive approximation, relative to the optimal. We evaluate our methods on synthetic and real datasets, which includes one from WAZE which gives traffic incident reports for the city of Boston. We find that our method gives promising results in recovering the missing failures. Bijaya Adhikari, Pavan Rangudu, B. Aditya Prakash, Anil Vullikanti |
SDM | 3 |
| 2018 | Mining E-Commerce Query Relations using Customer Interaction NetworksabstractCustomer Interaction Networks (CINs) are a natural framework for representing and mining customer interactions with E-Commerce search engines. Customer interactions begin with the submission of a query formulated based on an initial product intent, followed by a sequence of product engagement and query reformulation actions. Engagement with a product (e.g. clicks) indicates its relevance to the customer»s product intent. Reformulation to a new query indicates either dissatisfaction with current results, or an evolution in the customer»s product intent. Analyzing such interactions within and across sessions, enables us to discover various query-query and query-product relationships. In this work, we begin by studying the properties of CINs developed using Walmart.com»s product search logs. We observe that the properties exhibited by CINs make it possible to mine intent relationships between queries based purely on their structural information. We show how these relations can be exploited for a) clustering queries based on intents, b) significantly improve search quality for poorly performing queries, and c) identify the most influential (aka. »critical») queries whose performance have the highest impact on performance of other queries. Bijaya Adhikari, Parikshit Sondhi, Wenke Zhang, Mohit Sharma 0002, B. Aditya Prakash |
WWW | 5 |
| 2018 | Efficiently summarizing attributed diffusion networks
Sorour E. Amiri, Liangzhe Chen, B. Aditya Prakash |
Data Min. Knowl. Discov. | 3 |
| 2018 | Propagation-Based Temporal Network SummarizationabstractModern networks are very large in size and also evolve with time. As their sizes grow, the complexity of performing network analysis grows as well. Getting a smaller representation of a temporal network with similar properties will help in various data mining tasks. In this paper, we study the novel problem of getting a smaller diffusion-equivalent representation of a set of time-evolving networks. We first formulate a well-founded and general temporal-network condensation problem based on the so-called systemmatrix of the network. We then propose NETCONDENSE, a scalable and effective algorithm which solves this problem using careful transformations in sub-quadratic running time, and linear space complexities. Our extensive experiments show that we can reduce the size of large real temporal networks (from multiple domains such as social, co-authorship, and email) significantly without much loss of information. We also show the wide-applicability of NETCONDENSE by leveraging it for several tasks: for example, we use it to understand, explore, and visualize the original datasets and to also speed-up algorithms for the influence-maximization and event detection problems on temporal networks. Bijaya Adhikari, Yao Zhang 0003, Sorour E. Amiri, Aditya Bharadwaj, B. Aditya Prakash |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Automatic Segmentation of Dynamic Network Sequences with Node LabelsabstractGiven a sequence of snapshots of flu propagating over a population network, can we find a segmentation when the patterns of the disease spread change, possibly due to interventions? In this paper, we study the problem of segmenting graph sequences with labeled nodes. Memes on the Twitter network, diseases over a contact network, movie-cascades over a social network, etc. are all graph sequences with labeled nodes. Most related work on this subject is on plain graphs and hence ignores the label dynamics. Others require fix parameters or feature engineering. We propose SNAPNETS, to automatically find segmentations of such graph sequences, with different characteristics of nodes of each label in adjacent segments. It satisfies all the desired properties (being parameter free, comprehensive and scalable) by leveraging a principled, multi-level, flexible framework which maps the problem to a path optimization problem over a weighted DAG. Also, we develop the parallel framework of SNAPNETS which speeds up its running time. Finally, we propose an extension of SNAPNETS to handle the dynamic graph structures and use it to detect anomalies (and events) in network sequences. Extensive experiments on several diverse real datasets show that it finds cut points matching ground-truth or meaningful external signals and detects anomalies outperforming non-trivial baselines. We also show that the segmentations are easily interpretable, and that SNAPNETS scales near-linearly with the size of the input. Finally, we show how to use SNAPNETS to detect anomaly in a sequence of dynamic networks. Sorour E. Amiri, Liangzhe Chen, B. Aditya Prakash |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | HotSpots: Failure Cascades on Heterogeneous Critical Infrastructure NetworksabstractCritical Infrastructure Systems such as transportation, water and power grid systems are vital to our national security, economy, and public safety. Recent events, like the 2012 hurricane Sandy, show how the interdependencies among different CI networks lead to catastrophic failures among the whole system. Hence, analyzing these CI networks, and modeling failure cascades on them becomes a very important problem. However, traditional models either do not take multiple CIs or the dynamics of the system into account, or model it simplistically. In this paper, we study this problem using a heterogeneous network viewpoint. We first construct heterogeneous CI networks with multiple components using national-level datasets. Then we study novel failure maximization problems on these networks, to compute critical nodes in such systems. We then provide HotSpots, a scalable and effective algorithm for these problems, based on careful transformations. Finally, we conduct extensive experiments on real CIS data from multiple US states, and show that our method HotSpots outperforms non-trivial baselines, gives meaningful results and that our approach gives immediate benefits in providing situational-awareness during large-scale failures. Liangzhe Chen, Xinfeng Xu, Sangkeun Matt Lee, Sisi Duan, Alfonso G. Tarditi, Supriya Chinthavali, B. Aditya Prakash |
CIKM | 7 |
| 2017 | Data-Driven ImmunizationabstractGiven a contact network and coarse-grained diagnostic information like electronic Healthcare Reimbursement Claims (eHRC) data, can we develop efficient intervention policies to control an epidemic? Immunization is an important problem in multiple areas especially epidemiology and public health. However, most existing studies focus on developing pre-emptive strategies assuming prior epidemiological models. In practice, disease spread is usually complicated, hence assuming an underlying model may deviate from true spreading patterns, leading to possibly inaccurate interventions. Additionally, the abundance of health care surveillance data (like eHRC) makes it possible to study data-driven strategies without too many restrictive assumptions. Hence, such an approach can help public-health experts take more practical decisions. In this paper, we take into account propagation log and contact networks for controlling propagation. We formulate the novel and challenging Data-Driven Immunization problem without assuming classical epidemiological models. To solve it, we first propose an efficient sampling approach to align surveillance data with contact networks, then develop an efficient algorithm with the provably approximate guarantee for immunization. Finally, we show the effectiveness and scalability of our methods via extensive experiments on multiple datasets, and conduct case studies on nation-wide real medical surveillance data. Yao Zhang 0003, Arvind Ramanathan, Anil Vullikanti, Laura L. Pullum, B. Aditya Prakash |
ICDM | 5 |
| 2017 | Condensing Temporal Networks using PropagationabstractModern networks are very large in size and also evolve with time. As their size grows, the complexity of performing network analysis grows as well. Getting a smaller representation of a temporal network with similar properties will help in various data mining tasks. In this paper, we study the novel problem of getting a smaller diffusion-equivalent representation of a set of time-evolving networks. We first formulate a well-founded and general temporal-network condensation problem based on the so-called system-matrix of the network. We then propose NetCon-dense, a scalable and effective algorithm which solves this problem using careful transformations in sub-quadratic running time, and linear space complexities. Our extensive experiments show that we can reduce the size of large real temporal networks (from multiple domains such as social, co-authorship and email) significantly without much loss of information. We also show the wide-applicability of Net-Condense by leveraging it for several tasks: for example, we use it to understand, explore and visualize the original datasets and to also speed-up algorithms for the influence-maximization problem on temporal networks. Bijaya Adhikari, Yao Zhang 0003, Aditya Bharadwaj, B. Aditya Prakash |
SDM | 4 |
| 2017 | MeiKe: Influence-based Communities in NetworksabstractGiven a social network, how to find communities of nodes based on their diffusive characteristics? There exist two important types of nodes, for information propagation: nodes that are influential (“kernel nodes”), and nodes that serve as “bridges” to boost the diffusion (“media nodes”). How to find these nodes and uncover connections between them? In addition, it is also important to discover the hidden community structure of these nodes, which can help study their interactions, predict links and also understand the information flow in such networks. In this paper, we give an intuitive and novel optimization-based formulation for this task, which aims to discover media nodes as well as community structures of kernel nodes. We prove our task is computationally challenging, and develop an effective and practical algorithm MeiKe (pronounced as ‘Mike’). It first obtains media nodes via a new successive summarization based approach, and then finds kernel nodes including their community structures. Experimental results show that MeiKe finds high-quality media and kernel communities which match our expectations and ground-truth (outperforming non-trivial baselines by 40% in F1-score). Our case studies also demonstrate the applicability of MeiKe on a variety of datasets. Yao Zhang 0003, Bijaya Adhikari, Steve T. K. Jan, B. Aditya Prakash |
SDM | 4 |
| 2017 | Detecting Large Reshare Cascades in Social NetworksabstractDetecting large reshare cascades is an important problem in online social networks. There are a variety of attempts to model this problem, from using time series analysis methods to stochastic processes. Most of these approaches heavily depend on the underlying network features and use network information to detect the virality of cascades. In most cases, however, getting such detailed network information can be hard or even impossible. Karthik Subbian, B. Aditya Prakash, Lada A. Adamic |
WWW | 2 |
| 2017 | Understanding the Relationship between Human Behavior and Susceptibility to Cyber Attacks: A Data-Driven ApproachabstractDespite growing speculation about the role of human behavior in cyber-security of machines, concrete data-driven analysis and evidence have been lacking. Using Symantec’s WINE platform, we conduct a detailed study of 1.6 million machines over an 8-month period in order to learn the relationship between user behavior and cyber attacks against their personal computers. We classify users into 4 categories (gamers, professionals, software developers, and others, plus a fifth category comprising everyone) and identify a total of 7 features that act as proxies for human behavior. For each of the 35 possible combinations (5 categories times 7 features), we studied the relationship between each of these seven features and one dependent variable, namely the number of attempted malware attacks detected by Symantec on the machine. Our results show that there is a strong relationship between several features and the number of attempted malware attacks. Had these hosts not been protected by Symantec’s anti-virus product or a similar product, they would likely have been infected. Surprisingly, our results show that software developers are more at risk of engaging in risky cyber-behavior than other categories. Michael Ovelgönne, Tudor Dumitras, B. Aditya Prakash, V. S. Subrahmanian, Benjamin Wang |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | Nonlinear Dynamics of Information Diffusion in Social NetworksabstractThe recent explosion in the adoption of search engines and new media such as blogs and Twitter have facilitated the faster propagation of news and rumors. How quickly does a piece of news spread over these media? How does its popularity diminish over time? Does the rising and falling pattern follow a simple universal law? In this article, we propose SpikeM, a concise yet flexible analytical model of the rise and fall patterns of information diffusion. Our model has the following advantages. First, unification power: it explains earlier empirical observations and generalizes theoretical models including the SI and SIR models. We provide the threshold of the take-off versus die-out conditions for SpikeM and discuss the generality of our model by applying it to an arbitrary graph topology. Second, practicality: it matches the observed behavior of diverse sets of real data. Third, parsimony: it requires only a handful of parameters. Fourth, usefulness: it makes it possible to perform analytic tasks such as forecasting, spotting anomalies, and interpretation by reverse engineering the system parameters of interest (quality of news, number of interested bloggers, etc.). We also introduce an efficient and effective algorithm for the real-time monitoring of information diffusion, namely SpikeStream, which identifies multiple diffusion patterns in a large collection of online event streams. Extensive experiments on real datasets demonstrate that SpikeM accurately and succinctly describes all patterns of the rise and fall spikes in social networks. Yasuko Matsubara, Yasushi Sakurai, B. Aditya Prakash, Lei Li 0005, Christos Faloutsos |
ACM Trans. Web | 3 |
| 2016 | URBAN-NET: A network-based infrastructure monitoring and analysis system for emergency management and public safetyabstractCritical Infrastructures (CIs) such as energy, water, and transportation are complex networks that are crucial for sustaining day-to-day commodity flows vital to national security, economic stability, and public safety. The nature of these CIs is such that failures caused by an extreme weather event or a man-made incident can trigger widespread cascading failures, sending ripple effects at regional or even national scales. To minimize such effects, it is critical for emergency responders to identify existing or potential vulnerabilities within CIs during such stressor events in a systematic and quantifiable manner and take appropriate mitigating actions. We present here a novel critical infrastructure monitoring and analysis system named URBAN-NET. The system includes a software stack and tools for monitoring CIs, pre-processing data, interconnecting multiple CI datasets as a heterogeneous network, identifying vulnerabilities through graph-based topological analysis, and predicting consequences based on “what-if” simulations along with visualization. As a proof-of-concept, we present several case studies to show the capabilities of our system. We also discuss remaining challenges and future work. Sangkeun Matt Lee, Liangzhe Chen, Sisi Duan, Supriya Chinthavali, Mallikarjun Shankar, B. Aditya Prakash |
IEEE BigData | 6 |
| 2016 | Leveraging Propagation for Data Mining: Models, Algorithms and ApplicationsabstractCan we infer if a user is sick from her tweet? How do opinions get formed in online forums? Which people should we immunize to prevent an epidemic as fast as possible? How do we quickly zoom out of a graph? Graphs---also known as networks---are powerful tools for modeling processes and situations of interest in real life domains of social systems, cyber-security, epidemiology, and biology. They are ubiquitous, from online social networks, gene-regulatory networks, to router graphs. B. Aditya Prakash, Naren Ramakrishnan |
KDD | 1 |
| 2016 | Reconstructing an Epidemic Over TimeabstractWe consider the problem of reconstructing an epidemic over time, or, more general, reconstructing the propagation of an activity in a network. Our input consists of a temporal network, which contains information about when two nodes interacted, and a sample of nodes that have been reported as infected. The goal is to recover the flow of the spread, including discovering the starting nodes, and identifying other likely-infected nodes that are not reported. The problem we consider has multiple applications, from public health to social media and viral marketing purposes. Polina Rozenshtein, Aristides Gionis, B. Aditya Prakash, Jilles Vreeken |
KDD | 3 |
| 2016 | Unstable Communities in Network EnsemblesabstractEnsembles of graphs arise in several natural applications. Many techniques exist to compute frequent, dense subgraphs in these ensembles. In contrast, in this paper, we propose to discover maximally variable regions of the graphs, i.e., sets of nodes that induce very different subgraphs across the ensemble. We first develop two intuitive and novel definitions of such node sets, which we then show can be efficiently enumerated using a level-wise algorithm. Finally, using extensive experiments on multiple real datasets, we show how these sets capture the main structural variations of the given set of networks and also provide us with interesting and relevant insights about these datasets. Ahsanur Rahman 0001, Steve T. K. Jan, B. Aditya Prakash, T. M. Murali 0001 |
SDM | 4 |
| 2016 | Ensemble Models for Data-driven Prediction of Malware InfectionsabstractGiven a history of detected malware attacks, can we predict the number of malware infections in a country? Can we do this for different malware and countries? This is an important question which has numerous implications for cyber security, right from designing better anti-virus software, to designing and implementing targeted patches to more accurately measuring the economic impact of breaches. This problem is compounded by the fact that, as externals, we can only detect a fraction of actual malware infections. In this paper we address this problem using data from Symantec covering more than 1.4 million hosts and 50 malware spread across 2 years and multiple countries. We first carefully design domain-based features from both malware and machine-hosts perspectives. Secondly, inspired by epidemiological and information diffusion models, we design a novel temporal non-linear model for malware spread and detection. Finally we present ESM, an ensemble-based approach which combines both these methods to construct a more accurate algorithm. Using extensive experiments spanning multiple malware and countries, we show that ESM can effectively predict malware infection ratios over time (both the actual number and trend) upto 4 times better compared to several baselines on various metrics. Furthermore, ESM's performance is stable and robust even when the number of detected infections is low. Chanhyun Kang, Noseong Park, B. Aditya Prakash, Edoardo Serra, V. S. Subrahmanian |
WSDM | 3 |
| 2016 | Syndromic surveillance of Flu on Twitter using weakly supervised temporal topic models
Liangzhe Chen, K. S. M. Tozammel Hossain, Patrick Butler, Naren Ramakrishnan, B. Aditya Prakash |
Data Min. Knowl. Discov. | 5 |
| 2016 | Eigen-Optimization on Large Graphs by Edge ManipulationabstractLarge graphs are prevalent in many applications and enable a variety of information dissemination processes, e.g., meme, virus, and influence propagation. How can we optimize the underlying graph structure to affect the outcome of such dissemination processes in a desired way (e.g., stop a virus propagation, facilitate the propagation of a piece of good idea, etc)? Existing research suggests that the leading eigenvalue of the underlying graph is the key metric in determining the so-called epidemic threshold for a variety of dissemination models. In this paper, we study the problem of how to optimally place a set of edges (e.g., edge deletion and edge addition) to optimize the leading eigenvalue of the underlying graph, so that we can guide the dissemination process in a desired way. We propose effective, scalable algorithms for edge deletion and edge addition, respectively. In addition, we reveal the intrinsic relationship between edge deletion and node deletion problems. Experimental results validate the effectiveness and efficiency of the proposed algorithms. Chen Chen 0022, Hanghang Tong, B. Aditya Prakash, Tina Eliassi-Rad, Michalis Faloutsos, Christos Faloutsos |
ACM Trans. Knowl. Discov. Data | 3 |
| 2016 | Node Immunization on Large Graphs: Theory and AlgorithmsabstractGiven a large graph, like a computer communication network, which k nodes should we immunize (or monitor, or remove), to make it as robust as possible against a computer virus attack? This problem, referred to as the node immunization problem, is the core building block in many high-impact applications, ranging from public health, cybersecurity to viral marketing. A central component in node immunization is to find the best k bridges of a given graph. In this setting, we typically want to determine the relative importance of a node (or a set of nodes) within the graph, for example, how valuable (as a bridge) a person or a group of persons is in a social network. First of all, we propose a novel `bridging' score Dλ, inspired by immunology, and we show that its results agree with intuition for several realistic settings. Since the straightforward way to compute Dλ is computationally intractable, we then focus on the computational issues and propose a surprisingly efficient way (O(nk2+ m)) to estimate it. Experimental results on real graphs show that (1) the proposed `bridging' score gives mining results consistent with intuition; and (2) the proposed fast solution is up to seven orders of magnitude faster than straightforward alternatives. Chen Chen 0022, Hanghang Tong, B. Aditya Prakash, Charalampos E. Tsourakakis, Tina Eliassi-Rad, Christos Faloutsos, Polo Chau |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Near-Optimal Algorithms for Controlling Propagation at Group Scale on NetworksabstractGiven a network with groups, such as a contact-network grouped by ages, which are the best groups to immunize to control the epidemic? Equivalently, how to choose best communities in social media like Facebook to stop rumors from spreading? Immunization is an important problem in multiple different domains like epidemiology, public health, cyber security, and social media. Additionally, clearly immunization at group scale (like schools and communities) is more realistic due to constraints in implementations and compliance (e.g., it is hard to ensure specific individuals take the adequate vaccine). Hence, efficient algorithms for such a “group-based” problem can help public-health experts take more practical decisions. However, most prior work has looked into individual-scale immunization. In this paper, we study the problem of controlling propagation at group scale. We formulate a set of novel Group Immunization problems for multiple natural settings (for both threshold and cascade-based contagion models under both node-level and edge-level interventions) and develop multiple efficient algorithms, including provably approximate solutions. Finally, we show the effectiveness of our methods via extensive experiments on real and synthetic datasets. Yao Zhang 0003, Abhijin Adiga, Sudip Saha, Anil Vullikanti, B. Aditya Prakash |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2015 | Controlling Propagation at Group Scale on NetworksabstractGiven a network with groups, such as a contact-network grouped by ages, which are the best groups to immunize to control the epidemic? Equivalently, how to best choose communities in social networks like Facebook to stop rumors from spreading? Immunization is an important problem in multiple different domains like epidemiology, public health, cyber security and social media. Additionally, clearly immunization at group scale (like schools and communities) is more realistic due to constraints in implementations and compliance (e.g., it is hard to ensure specific individuals take the adequate vaccine). Hence efficient algorithms for such a "group-based" problem can help public-health experts take more practical decisions. However most prior work has looked into individual-scale immunization. In this paper, we study the problem of controlling propagation at group scale. We formulate novel so-called Group Immunization problems for multiple natural settings (for both threshold and cascade-based contagion models under both node-level and edge-level interventions) and develop multiple efficient algorithms, including provably approximate solutions. Finally, we show the effectiveness of our methods via extensive experiments on real and synthetic datasets. Yao Zhang 0003, Abhijin Adiga, Anil Vullikanti, B. Aditya Prakash |
ICDM | 4 |
| 2015 | Approximation Algorithms for Reducing the Spectral Radius to Control Epidemic SpreadabstractThe largest eigenvalue of the adjacency matrix of a network (referred to as the spectral radius) is an important metric in its own right. Further, for several models of epidemic spread on networks (e.g., the ‘flu-like’ SIS model), it has been shown that an epidemic dies out quickly if the spectral radius of the graph is below a certain threshold that depends on the model parameters. This motivates a strategy to control epidemic spread by reducing the spectral radius of the underlying network. In this paper, we develop a suite of provable approximation algorithms for reducing the spectral radius by removing the minimum cost set of edges (modeling quarantining) or nodes (modeling vaccinations), with different time and quality tradeoffs. Our main algorithm, GREEDYWALK, is based on the idea of hitting closed walks of a given length, and gives an O(log2 n)-approximation, where n denotes the number of nodes; it also performs much better in practice compared to all prior heuristics proposed for this problem. We further present a novel sparsification method to improve its running time. In addition, we give a new primal-dual based algorithm with an even better approximation guarantee (O(log n)), albeit with slower running time. We also give lower bounds on the worst-case performance of some of the popular heuristics. Finally we demonstrate the applicability of our algorithms and the properties of our solutions via extensive experiments on multiple synthetic and real networks. Sudip Saha, Abhijin Adiga, B. Aditya Prakash, Anil Vullikanti |
SDM | 3 |
| 2015 | Hidden Hazards: Finding Missing Nodes in Large Graph EpidemicsabstractGiven a noisy or sampled snapshot of an infection in a large graph, can we automatically and reliably recover the truly infected yet somehow missed nodes? And, what about the seeds, the nodes from which the infection started to spread? These are important questions in diverse contexts, ranging from epidemiology to social media. In this paper, we address the problem of simultaneously recovering the missing infections and the source nodes of the epidemic given noisy data. We formulate the problem by the Minimum Description Length principle, and propose NETFILL, an efficient algorithm that automatically and highly accurately identifies the number and identities of both missing nodes and the infection seed nodes. Experimental evaluation on synthetic and real datasets, including using data from information cascades over 96 million blog posts and news articles, shows that our method outperforms other baselines, scales near-linearly, and is highly effective in recovering missing nodes and sources. Shashidhar Sundareisan, Jilles Vreeken, B. Aditya Prakash |
SDM | 3 |
| 2015 | Data-Aware Vaccine Allocation Over Large NetworksabstractGiven a graph, like a social/computer network or the blogosphere, in which an infection (or meme or virus) has been spreading for some time, how to select thekbest nodes for immunization/quarantining immediately? Most previous works for controlling propagation (say via immunization) have concentrated on developing strategies for vaccinationpreemptivelybefore the start of the epidemic. While very useful to provide insights in to which baseline policies can best control an infection, they may not be ideal to makereal-timedecisions as the infection is progressing. In this paper, we study how to immunize healthy nodes, in the presence of already infected nodes. Efficient algorithms for such a problem can help public-health experts make more informed choices, tailoring their decisions to the actual distribution of the epidemic on the ground. First we formulate theData-Aware Vaccinationproblem, and prove it is NP-hard and also that it is hard to approximate. Secondly, we propose three effective polynomial-time heuristics DAVA, DAVA-prune and DAVA-fast, of varying degrees of efficiency and performance. Finally, we also demonstrate the scalability and effectiveness of our algorithms through extensive experiments on multiple real networks including large epidemiology datasets (containing millions of interactions). Our algorithms show substantial gains of up toten timesmore healthy nodes at the end against many other intuitive and nontrivial competitors. Yao Zhang 0003, B. Aditya Prakash |
ACM Trans. Knowl. Discov. Data | 2 |
| 2014 | SansText: Classifying temporal topic dynamics of Twitter cascades without tweet textabstractUnderstanding the dynamics of cascades in Twitter is an important modeling problem with multiple applications like viral marketing and the detection and forecasting of emerging events. Key hashtags rise in popularity to a peak and fall, with profiles characteristic to the specific topical area of the hashtag. Traditional text-based classification approaches are inadequate as new hashtags get created dynamically and because social media vocabulary evolves. We demonstrate a text-free approach SansText to classify emerging cascades by modeling the phenomenological patterns of rise and fall. We illustrate the utility of this approach over several specific event classes as well as more general topics in a collection of more than 2 million tweets from multiple countries of Latin America. Shashidhar Sundareisan, Abhay Bhadriraju, M. Saquib Khan, Naren Ramakrishnan, B. Aditya Prakash |
ASONAM | 5 |
| 2014 | Scalable Vaccine Distribution in Large Graphs given Uncertain DataabstractGiven an noisy or sampled snapshot of a network, like a contact-network or the blogosphere, in which an infection (or meme/virus) has been spreading for some time, what are the best nodes to immunize (vaccinate)? Manipulating graphs via node removal by itself is an important problem in multiple different domains like epidemiology, public health and social media. Moreover, it is important to account for uncertainty as typically surveillance data on who is infected is limited or the data is sampled. Efficient algorithms for such a problem can help public-health experts take more informed decisions. Yao Zhang 0003, B. Aditya Prakash |
CIKM | 2 |
| 2014 | Flu Gone Viral: Syndromic Surveillance of Flu on Twitter Using Temporal Topic ModelsabstractSurveillance of epidemic outbreaks and spread from social media is an important tool for governments and public health authorities. Machine learning techniques for now casting the flu have made significant inroads into correlating social media trends to case counts and prevalence of epidemics in a population. There is a disconnect between data-driven methods for forecasting flu incidence and epidemiological models that adopt a state based understanding of transitions, that can lead to sub-optimal predictions. Furthermore, models for epidemiological activity and social activity like on Twitter predict different shapes and have important differences. We propose a temporal topic model to capture hidden states of a user from his tweets and aggregate states in a geographical region for better estimation of trends. We show that our approach helps fill the gap between phenomenological methods for disease surveillance and epidemiological models. We validate this approach by modeling the flu using Twitter in multiple countries of South America. We demonstrate that our model can consistently outperform plain vocabulary assessment in flu case-count predictions, and at the same time get better flu-peak predictions than competitors. We also show that our fine-grained modeling can reconcile some contrasting behaviors between epidemiological and social models. Liangzhe Chen, K. S. M. Tozammel Hossain, Patrick Butler, Naren Ramakrishnan, B. Aditya Prakash |
ICDM | 5 |
| 2014 | Modeling mass protest adoption in social network communities using geometric brownian motionabstractModeling the movement of information within social media outlets, like Twitter, is key to understanding to how ideas spread but quantifying such movement runs into several difficulties. Two specific areas that elude a clear characterization are (i) the intrinsic random nature of individuals to potentially adopt and subsequently broadcast a Twitter topic, and (ii) the dissemination of information via non-Twitter sources, such as news outlets and word of mouth, and its impact on Twitter propagation. These distinct yet inter-connected areas must be incorporated to generate a comprehensive model of information diffusion. We propose a bispace model to capture propagation in the union of (exclusively) Twitter and non-Twitter environments. To quantify the stochastic nature of Twitter topic propagation, we combine principles of geometric Brownian motion and traditional network graph theory. We apply Poisson process functions to model information diffusion outside of the Twitter mentions network. We discuss techniques to unify the two sub-models to accurately model information dissemination. We demonstrate the novel application of these techniques on real Twitter datasets related to mass protest adoption in social communities. Fang Jin, Rupinder Paul Khandpur, Nathan Self, Edward R. Dougherty, Sheng Guo 0002, Feng Chen 0001, B. Aditya Prakash, Naren Ramakrishnan |
KDD | 7 |
| 2014 | Fast influence-based coarsening for large networksabstractGiven a social network, can we quickly 'zoom-out' of the graph? Is there a smaller equivalent representation of the graph that preserves its propagation characteristics? Can we group nodes together based on their influence properties? These are important problems with applications to influence analysis, epidemiology and viral marketing applications. Manish Purohit, B. Aditya Prakash, Chanhyun Kang, Yao Zhang 0003, V. S. Subrahmanian |
KDD | 2 |
| 2014 | DAVA: Distributing Vaccines over Networks under Prior InformationabstractGiven a graph, like a social/computer network or the blogosphere, in which an infection (or meme or virus) has been spreading for some time, how to select the k best nodes for immunization/quarantining immediately? Most previous works for controlling propagation (say via immunization) have concentrated on developing strategies for vaccination pre-emptively before the start of the epidemic. While very useful to provide insights in to which baseline policies can best control an infection, they may not be ideal to make real-time decisions as the infection is progressing. In this paper, we study how to immunize healthy nodes, in presence of already infected nodes. Efficient algorithms for such a problem can help public-health experts make more informed choices. First we formulate the Data-Aware Vaccination problem, and prove it is NP-hard and also that it is hard to approximate. Secondly, we propose two effective polynomial-time heuristics DAVA and DAVA-fast. Finally, we also demonstrate the scalability and effectiveness of our algorithms through extensive experiments on multiple real networks including epidemiology datasets, which show substantial gains of up to 10 times more healthy nodes at the end. Yao Zhang 0003, B. Aditya Prakash |
SDM | 2 |
| 2014 | Efficiently spotting the starting points of an epidemic in a large graph
B. Aditya Prakash, Jilles Vreeken, Christos Faloutsos |
Knowl. Inf. Syst. | 1 |
| 2013 | Spatio-temporal mining of software adoption & penetrationabstractHow does malware propagate? Does it form spikes over time? Does it resemble the propagation pattern of benign files, such as software patches? Does it spread uniformly over countries? How long does it take for a URL that distributes malware to be detected and shut down? Evangelos E. Papalexakis, Tudor Dumitras, Polo Chau, B. Aditya Prakash, Christos Faloutsos |
ASONAM | 4 |
| 2013 | Patterns amongst Competing Task Frequencies: Super-Linearities, and the Almond-DG Model
Danai Koutra, Vasileios Koutras, B. Aditya Prakash, Christos Faloutsos |
PAKDD (1) | 3 |
| 2013 | Fractional Immunization in NetworksabstractPreventing contagion in networks is an important problem in public health and other domains. Targeting nodes to immunize based on their network interactions has been shown to be far more effective at stemming infection spread than immunizing random subsets of nodes. However, the assumption that selected nodes can be rendered completely immune does not hold for infections for which there is no vaccination or effective treatment. Instead, one can confer fractional immunity to some nodes by allocating variable amounts of infection-prevention resource to them. We formulate the problem to distribute a fixed amount of resource across nodes in a network such that the infection rate is minimized, prove that it is NP-complete and derive a highly effective and efficient linear-time algorithm. We demonstrate the efficiency and accuracy of our algorithm compared to several other methods using simulation on real-world network datasets including US-MEDICARE and state-level interhospital patient transfer data. We find that concentrating resources at a small subset of nodes using our algorithm is up to 6 times more effective than distributing them uniformly (as is current practice) or using network-based heuristics. To the best of our knowledge, we are the first to formulate the problem, use truly nation-scale network data and propose effective algorithms. Lada A. Adamic, Christos Faloutsos, Theodore J. Iwashyna, B. Aditya Prakash, Hanghang Tong |
SDM | 4 |
| 2012 | Gelling, and melting, large graphs by edge manipulationabstractControlling the dissemination of an entity (e.g., meme, virus, etc) on a large graph is an interesting problem in many disciplines. Examples include epidemiology, computer security, marketing, etc. So far, previous studies have mostly focused on removing or inoculating nodes to achieve the desired outcome. Hanghang Tong, B. Aditya Prakash, Tina Eliassi-Rad, Michalis Faloutsos, Christos Faloutsos |
CIKM | 2 |
| 2012 | Spotting Culprits in Epidemics: How Many and Which Ones?abstractGiven a snapshot of a large graph, in which an infection has been spreading for some time, can we identify those nodes from which the infection started to spread? In other words, can we reliably tell who the culprits are? In this paper we answer this question affirmatively, and give an efficient method called NETSLEUTH for the well-known Susceptible-Infected virus propagation model. Essentially, we are after that set of seed nodes that best explain the given snapshot. We propose to employ the Minimum Description Length principle to identify the best set of seed nodes and virus propagation ripple, as the one by which we can most succinctly describe the infected graph. We give an highly efficient algorithm to identify likely sets of seed nodes given a snapshot. Then, given these seed nodes, we show we can optimize the virus propagation ripple in a principled way by maximizing likelihood. With all three combined, NETSLEUTH can automatically identify the correct number of seed nodes, as well as which nodes are the culprits. Experimentation on our method shows high accuracy in the detection of seed nodes, in addition to the correct automatic identification of their number. Moreover, we show NETSLEUTH scales linearly in the number of nodes of the graph. B. Aditya Prakash, Jilles Vreeken, Christos Faloutsos |
ICDM | 1 |
| 2012 | Interacting viruses in networks: can both survive?abstractSuppose we have two competing ideas/products/viruses, that propagate over a social or other network. Suppose that they are strong/virulent enough, so that each, if left alone, could lead to an epidemic. What will happen when both operate on the network? Earlier models assume that there is perfect competition: if a user buys product 'A' (or gets infected with virus 'X'), she will never buy product 'B' (or virus 'Y'). This is not always true: for example, a user could install and use both Firefox and Google Chrome as browsers. Similarly, one type of flu may give partial immunity against some other similar disease. Alex Beutel, B. Aditya Prakash, Ronald Rosenfeld, Christos Faloutsos |
KDD | 2 |
| 2012 | Rise and fall patterns of information diffusion: model and implicationsabstractThe recent explosion in the adoption of search engines and new media such as blogs and Twitter have facilitated faster propagation of news and rumors. How quickly does a piece of news spread over these media? How does its popularity diminish over time? Does the rising and falling pattern follow a simple universal law? Yasuko Matsubara, Yasushi Sakurai, B. Aditya Prakash, Lei Li 0005, Christos Faloutsos |
KDD | 3 |
| 2012 | Winner takes all: competing viruses or ideas on fair-play networksabstractGiven two competing products (or memes, or viruses etc.) spreading over a given network, can we predict what will happen at the end, that is, which product will 'win', in terms of highest market share? One may naively expect that the better product (stronger virus) will just have a larger footprint, proportional to the quality ratio of the products (or strength ratio of the viruses). However, we prove the surprising result that, under realistic conditions, for any graph topology, the stronger virus completely wipes-out the weaker one, thus not merely 'winning' but 'taking it all'. In addition to the proofs, we also demonstrate our result with simulations over diverse, real graph topologies, including the social-contact graph of the city of Portland OR (about 31 million edges and 1 million nodes) and internet AS router graphs. Finally, we also provide real data about competing products from Google-Insights, like Facebook-Myspace, and we show again that they agree with our analysis. B. Aditya Prakash, Alex Beutel, Ronald Rosenfeld, Christos Faloutsos |
WWW | 1 |
| 2012 | Threshold conditions for arbitrary cascade models on arbitrary networks
B. Aditya Prakash, Deepayan Chakrabarti, Nicholas Valler, Michalis Faloutsos, Christos Faloutsos |
Knowl. Inf. Syst. | 1 |
| 2012 | Understanding and Managing Cascades on Large GraphsabstractHow do contagions spread in population networks? Which group should we market to, for maximizing product penetration? Will a given YouTube video go viral? Who are the best people to vaccinate? What happens when two products compete? The objective of this tutorial is to provide an intuitive and concise overview of most important theoretical results and algorithms to help us understand and manipulate such propagation-style processes on large networks. The tutorial contains three parts: (a) Theoretical results on the behavior of fundamental models; (b) Scalable Algorithms for changing the behavior of these processes e.g., for immunization, marketing etc.; and (c) Empirical Studies of diffusion on blogs and on-line websites like Twitter. The problems we focus on are central in surprisingly diverse areas: from computer science and engineering, epidemiology and public health, product marketing to information dissemination. Our emphasis is on intuition behind each topic, and guidelines for the practitioner. B. Aditya Prakash, Christos Faloutsos |
Proc. VLDB Endow. | 1 |
| 2011 | Threshold Conditions for Arbitrary Cascade Models on Arbitrary NetworksabstractGiven a network of who-contacts-whom or who links-to-whom, will a contagious virus/product/meme spread and 'take-over' (cause an epidemic) or die-out quickly? What will change if nodes have partial, temporary or permanent immunity? The epidemic threshold is the minimum level of virulence to prevent a viral contagion from dying out quickly and determining it is a fundamental question in epidemiology and related areas. Most earlier work focuses either on special types of graphs or on specific epidemiological/cascade models. We are the first to show the G2-threshold (twice generalized) theorem, which nicely de-couples the effect of the topology and the virus model. Our result unifies and includes as special case older results and shows that the threshold depends on the first eigenvalue of the connectivity matrix, (a) for any graph and (b) for all propagation models in standard literature (more than 25, including H.I.V.) [20], [12]. Our discovery has broad implications for the vulnerability of real, complex networks, and numerous applications, including viral marketing, blog dynamics, influence propagation, easy answers to 'what-if' questions, and simplified design and evaluation of immunization policies. We also demonstrate our result using extensive simulations on one of the biggest available social contact graphs containing more than 31 million interactions among more than 1 million people representing the city of Portland, Oregon, USA. B. Aditya Prakash, Deepayan Chakrabarti, Michalis Faloutsos, Nicholas Valler, Christos Faloutsos |
ICDM | 1 |
| 2010 | On the Vulnerability of Large GraphsabstractGiven a large graph, like a computer network, which k nodes should we immunize (or monitor, or remove), to make it as robust as possible against a computer virus attack? We need (a) a measure of the 'Vulnerability' of a given network, (b) a measure of the 'Shield-value' of a specific set of k nodes and (c) a fast algorithm to choose the best such k nodes. We answer all these three questions: we give the justification behind our choices, we show that they agree with intuition as well as recent results in immunology. Moreover, we propose NetShield a fast and scalable algorithm. Finally, we give experiments on large real graphs, where NetShield achieves tremendous speed savings exceeding 7 orders of magnitude, against straightforward competitors. Hanghang Tong, B. Aditya Prakash, Charalampos E. Tsourakakis, Tina Eliassi-Rad, Christos Faloutsos, Polo Chau |
ICDM | 2 |
| 2010 | Metric forensics: a multi-level approach for mining volatile graphsabstractAdvances in data collection and storage capacity have made it increasingly possible to collect highly volatile graph data for analysis. Existing graph analysis techniques are not appropriate for such data, especially in cases where streaming or near-real-time results are required. An example that has drawn significant research interest is the cyber-security domain, where internet communication traces are collected and real-time discovery of events, behaviors, patterns, and anomalies is desired. We propose MetricForensics, a scalable framework for analysis of volatile graphs. MetricForensics combines a multi-level "drill down" approach, a collection of user-selected graph metrics, and a collection of analysis techniques. At each successive level, more sophisticated metrics are computed and the graph is viewed at finer temporal resolutions. In this way, MetricForensics scales to highly volatile graphs by only allocating resources for computationally expensive analysis when an interesting event is discovered at a coarser resolution first. We test MetricForensics on three real-world graphs: an enterprise IP trace, a trace of legitimate and malicious network traffic from a research institution, and the MIT Reality Mining proximity sensor data. Our largest graph has 3M vertices and 32M edges, spanning 4.5 days. The results demonstrate the scalability and capability of MetricForensics in analyzing volatile graphs; and highlight four novel phenomena in such graphs: elbows, broken correlations, prolonged spikes, and lightweight stars. Keith Henderson, Tina Eliassi-Rad, Christos Faloutsos, Leman Akoglu, Lei Li 0005, Koji Maruhashi, B. Aditya Prakash, Hanghang Tong |
KDD | 7 |
| 2010 | EigenSpokes: Surprising Patterns and Scalable Community Chipping in Large Graphs
B. Aditya Prakash, Ashwin Sridharan, Mukund Seshadri, Sridhar Machiraju, Christos Faloutsos |
PAKDD (2) | 1 |
| 2010 | Virus Propagation on Time-Varying Networks: Theory and Immunization Algorithms
B. Aditya Prakash, Hanghang Tong, Nicholas Valler, Michalis Faloutsos, Christos Faloutsos |
ECML/PKDD (3) | 1 |
| 2010 | Parsimonious Linear Fingerprinting for Time SeriesabstractWe study the problem of mining and summarizing multiple time series effectively and efficiently. We propose PLiF, a novel method to discover essential characteristics ("fingerprints"), by exploiting the joint dynamics in numerical sequences. Our fingerprinting method has the following benefits: (a) it leads to interpretable features; (b) it is versatile: PLiF enables numerous mining tasks, including clustering, compression, visualization, forecasting, and segmentation, matching top competitors in each task; and (c) it is fast and scalable , with linear complexity on the length of the sequences. We did experiments on both synthetic and real datasets, including human motion capture data (17MB of human motions), sensor data (166 sensors), and network router traffic data (18 million raw updates over 2 years). Despite its generality, PLiF outperforms the top clustering methods on clustering; the top compression methods on compression (3 times better reconstruction error, for the same compression ratio); it gives meaningful visualization and at the same time, enjoys a linear scale-up. Lei Li 0005, B. Aditya Prakash, Christos Faloutsos |
Proc. VLDB Endow. | 2 |
| 2009 | BGP-lens: patterns and anomalies in internet routing updatesabstractThe Border Gateway Protocol (BGP) is one of the fundamental computer communication protocols. Monitoring and mining BGP update messages can directly reveal the health and stability of Internet routing. Here we make two contributions: firstly we find patterns in BGP updates, like self-similarity, power-law and lognormal marginals; secondly using these patterns, we find anomalies. Specifically, we develop BGP-lens, an automated BGP updates analysis tool, that has three desirable properties: (a) It is effective, able to identify phenomena that would otherwise go unnoticed, such as a peculiar 'clothesline' behavior or prolonged 'spikes' that last as long as 8 hours; (b) It is scalable, using algorithms are all linear on the number of time-ticks; and (c) It is admin-friendly, giving useful leads for phenomenon of interest. B. Aditya Prakash, Nicholas Valler, David G. Andersen, Michalis Faloutsos, Christos Faloutsos |
KDD | 1 |
| 2009 | FRAPP: a framework for high-accuracy privacy-preserving mining
Shipra Agrawal 0001, Jayant R. Haritsa, B. Aditya Prakash |
Data Min. Knowl. Discov. | 3 |
| 2007 | Complex Group-By Queries for XMLabstractThe popularity of XML as a data exchange standard has led to the emergence of powerful XML query languages like XQuery and studies on XML query optimization. Of late, there is considerable interest in analytical processing of XML data. As pointed out by Borkar and Carey, even for data integration, there is a compelling need for performing various group-by style aggregate operations. A core operator needed for analytics is the group-by operator, which is widely used in relational as well as OLAP database applications. XQuery requires group-by operations to be simulated using nesting. Chaitanya Gokhale, Nitin Gupta 0003, Pranav Kumar, Laks V. S. Lakshmanan, Raymond T. Ng, B. Aditya Prakash |
ICDE | 6 |