VLDB 2026 Research / reviewers in the wild / expert
Kamalika Das
dblp:31/2815
· DBLP profile ↗
30ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 14 · 6 first-author · 2 since 2021Computer networks · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule LearningabstractXiang Gao, Yuguang Yao, Qi Zhang, Kaiwen Dong, Avinash Baidya, Ruocheng Guo, Hilaf Hasson, Kamalika Das. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiang Gao 0011, Yuguang Yao, Kaiwen Dong, Avinash Baidya, Ruocheng Guo, Hilaf Hasson, Kamalika Das |
ACL (1) | 8 |
| 2025 | SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt OptimizationabstractWendi Cui, Jiaxin Zhang, Zhuohang Li, Hao Sun, Damien Lopez, Kamalika Das, Bradley A. Malin, Sricharan Kumar. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wendi Cui, Jiaxin Zhang 0005, Damien Lopez, Kamalika Das, Bradley A. Malin, Kumar Sricharan |
ACL (1) | 6 |
| 2025 | XOOD: A Self-supervised Algorithm for Detecting Out-of-Distribution Data for Image Classification
Frej Berglind, Magesh Rajasekaran, Md Saiful Islam Sajol, Haron Temam, Supratik Mukhopadhyay, Kamalika Das, Kumar Sricharan, Kumar Kallurupalli |
ICANN (1) | 6 |
| 2025 | Transaction Categorization with Relational Deep Learning in QuickBooks
Kaiwen Dong, Padmaja Jonnalagedda, Xiang Gao 0011, Ayan Acharya, Maria Kissa, Mauricio Flores, Nitesh V. Chawla, Kamalika Das |
ECML/PKDD (9) | 8 |
| 2024 | Customizing Language Model Responses with Contrastive In-Context LearningabstractLarge language models (LLMs) are becoming increasingly important for machine learning applications. However, it can be challenging to align LLMs with our intent, particularly when we want to generate content that is preferable over others or when we want the LLM to respond in a certain style or tone that is hard to describe. To address this challenge, we propose an approach that uses contrastive examples to better describe our intent. This involves providing positive examples that illustrate the true intent, along with negative examples that show what characteristics we want LLMs to avoid. The negative examples can be retrieved from labeled data, written by a human, or generated by the LLM itself. Before generating an answer, we ask the model to analyze the examples to teach itself what to avoid. This reasoning step provides the model with the appropriate articulation of the user's need and guides it towards generting a better answer. We tested our approach on both synthesized and real-world datasets, including StackExchange and Reddit, and found that it significantly improves performance compared to standard few-shot prompting. Xiang Gao 0011, Kamalika Das |
AAAI | 2 |
| 2024 | Discriminant Distance-Aware Representation on Deterministic Uncertainty Quantification MethodsabstractUncertainty estimation is a crucial aspect of deploying dependable deep learning models in safety-critical systems. In this study, we introduce a novel and efficient method for deterministic uncertainty estimation called Discriminant Distance-Awareness Representation (DDAR). Our approach involves constructing a DNN model that incorporates a set of prototypes in its latent representations, enabling us to analyze valuable feature information from the input data. By leveraging a distinction maximization layer over optimal trainable prototypes, DDAR can learn a discriminant distance-awareness representation. We demonstrate that DDAR overcomes feature collapse by relaxing the Lipschitz constraint that hinders the practicality of deterministic uncertainty methods (DUMs) architectures. Our experiments show that DDAR is a flexible and architecture-agnostic method that can be easily integrated as a pluggable layer with distance-sensitive metrics, outperforming state-of-the-art uncertainty estimation methods on multiple benchmark problems. Jiaxin Zhang 0005, Kamalika Das, Kumar Sricharan |
AISTATS | 2 |
| 2024 | SPUQ: Perturbation-Based Uncertainty Quantification for Large Language ModelsabstractIn recent years, large language models (LLMs) have become increasingly prevalent, offering remarkable text generation capabilities.However, a pressing challenge is their tendency to make confidently wrong predictions, highlighting the critical need for uncertainty quantification (UQ) in LLMs.While previous works have mainly focused on addressing aleatoric uncertainty, the full spectrum of uncertainties, including epistemic, remains inadequately explored.Motivated by this gap, we introduce a novel UQ method, sampling with perturbation for UQ (SPUQ), designed to tackle both aleatoric and epistemic uncertainties.The method entails generating a set of perturbations for LLM inputs, sampling outputs for each perturbation, and incorporating an aggregation module that generalizes the sampling uncertainty approach for text generation tasks.Through extensive experiments on various datasets, we investigated different perturbation and aggregation techniques.Our findings show a substantial improvement in model uncertainty calibration, with a reduction in Expected Calibration Error (ECE) by 50% on average.Our findings suggest that our proposed UQ method offers promising steps toward enhancing the reliability and trustworthiness of LLMs 1 . Xiang Gao 0011, Jiaxin Zhang 0005, Lalla Mouatadid, Kamalika Das |
EACL (1) | 4 |
| 2024 | Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented GenerationabstractLanguage models (LMs) are known to suffer from hallucinations and misinformation.Retrieval augmented generation (RAG) that retrieves verifiable information from an external knowledge corpus to complement the parametric knowledge in LMs provides a tangible solution to these problems.However, the generation quality of RAG is highly dependent on the relevance between a user's query and the retrieved documents.Inaccurate responses may be generated when the query is outside of the scope of knowledge represented in the external knowledge corpus or if the information in the corpus is out-of-date.In this work, we establish a statistical framework that assesses how well a query can be answered by an RAG system by capturing the relevance of knowledge.We introduce an online testing procedure that employs goodness-of-fit (GoF) tests to inspect the relevance of each user query to detect out-of-knowledge queries with low knowledge relevance.Additionally, we develop an offline testing framework that examines a collection of user queries, aiming to detect significant shifts in the query distribution which indicates the knowledge corpus is no longer sufficiently capable of supporting the interests of the users.We demonstrate the capabilities of these strategies through a systematic evaluation on eight question-answering (QA) datasets, the results of which indicate that the new testing framework is an efficient solution to enhance the reliability of existing RAG systems. Jiaxin Zhang 0005, Chao Yan 0004, Kamalika Das, Kumar Sricharan, Murat Kantarcioglu, Bradley A. Malin |
EMNLP | 4 |
| 2024 | Synthetic Knowledge Ingestion: Towards Knowledge Refinement and Injection for Enhancing Large Language ModelsabstractLarge language models (LLMs) are proficient in capturing factual knowledge across various domains.However, refining their capabilities on previously seen knowledge or integrating new knowledge from external sources remains a significant challenge.In this work, we propose a novel synthetic knowledge ingestion method called Ski, which leverages fine-grained synthesis, interleaved generation, and assemble augmentation strategies to construct high-quality data representations from raw knowledge sources.We then integrate Ski and its variations with three knowledge injection techniques: Retrieval Augmented Generation (RAG), Supervised Fine-tuning (SFT), and Continual Pre-training (CPT) to inject and refine knowledge in language models.Extensive empirical experiments are conducted on various question-answering tasks spanning finance, biomedicine, and open-generation domains to demonstrate that Ski significantly outperforms baseline methods by facilitating effective knowledge injection.We believe that our work is an important step towards enhancing the factual accuracy of LLM outputs by refining knowledge representation and injection capabilities.1Raw knowledge context: "The annual contribution limit for a health savings account (HSA) in 2024 is $4,150 for individuals with self-only coverage and $8,300 for individuals with family coverage.These limits are about a 7% increase from 2023.Individuals who are 55 or older can contribute an additional $1,000, bringing the total to $4,850 for individuals and $8,750 for families…" Generate hypothetical questions given ngram knowledge contexts -"Holistic Detail-Oriented"• 1-gram Q -"What is the HSA contribution limit for individuals with self-only coverage in 2024" • 2-gram Q -"What is the HSA contribution limit for 2023?" • 3-gram Q -"How much more can individuals contribute if aged over 55 than younger?" Fine-grained SynthesisGenerate question and answer simultaneously given knowledge -"Aligned Harmony"• Context: The annual contribution limit for a health savings account (HSA) in 2024 is $4,150 for individuals with selfonly coverage and $8,300 for … • Question: What is the HSA contribution limit for individuals with self-only coverage in 2024?• Answer: $4150 Interleaved GenerationAssemble n-gram synthesis and question /answer/context pairs -"Repetition with diversity"• Assembles of question, answer, and context pairs: Jiaxin Zhang 0005, Wendi Cui, Kamalika Das, Kumar Sricharan |
EMNLP | 4 |
| 2024 | COMBOOD: A Semiparametric Approach for Detecting Out-of-distribution Data for Image ClassificationabstractIdentifying out-of-distribution (OOD) data at inference time is crucial for many machine learning applications, especially for automation. We present a novel unsupervised semi-parametric framework COMBOOD for OOD detection with respect to image recognition. Our framework combines signals from two distance metrics, nearest-neighbor and Mahalanobis, to derive a confidence score for an inference point to be out-of-distribution. The former provides a non-parametric approach to OOD detection. The latter provides a parametric, simple, yet effective method for detecting OOD data points, especially, in the far OOD scenario, where the inference point is far apart from the training data set in the embedding space. However, its performance is not satisfactory in the near OOD scenarios that arise in practical situations. Our COMBOOD framework combines the two signals in a semi-parametric setting to provide a confidence score that is accurate both for the near-OOD and far-OOD scenarios. We show experimental results with the COMBOOD framework for different types of feature extraction strategies. We demonstrate experimentally that COMBOOD outperforms state-of-the-art OOD detection methods on the OpenOOD (both version 1 and most recent version 1.5) benchmark datasets (for both far-OOD and near-OOD) as well as on the documents dataset in terms of accuracy. Magesh Rajasekaran, Md Saiful Islam Sajol, Frej Berglind, Supratik Mukhopadhyay, Kamalika Das |
SDM | 5 |
| 2024 | DECDM: Document Enhancement using Cycle-Consistent Diffusion ModelsabstractThe performance of optical character recognition (OCR) heavily relies on document image quality, which is crucial for automatic document processing and document intelligence. However, most existing document enhancement methods require supervised data pairs, which raises concerns about data separation and privacy protection, and makes it challenging to adapt these methods to new domain pairs. To address these issues, we propose DECDM, an end-to-end document-level image translation method inspired by recent advances in diffusion models. Our method overcomes the limitations of paired training by independently training the source (noisy input) and target (clean output) models, making it possible to apply domain-specific diffusion models to other pairs. DECDM trains on one dataset at a time, eliminating the need to scan both datasets concurrently, and effectively preserving data privacy from the source or target domain. We also introduce simple data augmentation strategies to improve character-glyph conservation during translation. We compare DECDM with state-of-the-art methods on multiple synthetic data and benchmark datasets, such as document denoising and shadow removal, and demonstrate the superiority of performance quantitatively and qualitatively. Jiaxin Zhang 0005, Joy Rimchala, Lalla Mouatadid, Kamalika Das, Kumar Sricharan |
WACV | 4 |
| 2023 | Interactive Multi-fidelity Learning for Cost-effective Adaptation of Language Model with Sparse Human SupervisionabstractLarge language models (LLMs) have demonstrated remarkable capabilities in various tasks. However, their suitability for domain-specific tasks, is limited due to their immense scale at deployment, susceptibility to misinformation, and more importantly, high data annotation costs. We propose a novel Interactive Multi-Fidelity Learning (IMFL) framework for cost-effective development of small domain-specific LMs under limited annotation budgets. Our approach formulates the domain-specific fine-tuning process as a multi-fidelity learning problem, focusing on identifying the optimal acquisition strategy that balances between low-fidelity automatic LLM annotations and high-fidelity human annotations to maximize model performance. We further propose an exploration-exploitation query strategy that enhances annotation diversity and informativeness, incorporating two innovative designs: 1) prompt retrieval that selects in-context examples from human-annotated samples to improve LLM annotation, and 2) variable batch size that controls the order for choosing each fidelity to facilitate knowledge distillation, ultimately enhancing annotation quality. Extensive experiments on financial and medical tasks demonstrate that IMFL achieves superior performance compared with single fidelity annotations. Given a limited budget of human annotation, IMFL significantly outperforms the $\bf 3\times$ human annotation baselines in all four tasks and achieves very close performance as $\bf 5\times$ human annotation on two of the tasks. These promising results suggest that the high human annotation costs in domain-specific tasks can be significantly reduced by employing IMFL, which utilizes fewer human annotations, supplemented with cheaper and faster LLM (e.g., GPT-3.5) annotations to achieve comparable performance. Jiaxin Zhang 0005, Kamalika Das, Kumar Sricharan |
NeurIPS | 3 |
| 2020 | Learning Instrument Invariant Characteristics for Generating High-resolution Global Coral Reef MapsabstractCoral reefs are one of the most biologically complex and diverse ecosystems within the shallow marine environment. Unfortunately, these underwater ecosystems are threatened by a number of anthropogenic challenges, including ocean acidification and warming, overfishing, and the continued increase of marine debris in oceans. This requires a comprehensive assessment of the world's coastal environments, including a quantitative analysis on the health and extent of coral reefs and other associated marine species, as a vital Earth Science measurement. However, limitations in observational and technological capabilities inhibit global sustained imaging of the marine environment. Harmonizing multimodal data sets acquired using different remote sensing instruments presents additional challenges, thereby limiting the availability of good quality labeled data for analysis. In this work, we develop a deep learning model for extracting domain invariant features from multimodal remote sensing imagery and creating high-resolution global maps of coral reefs by combining various sources of imagery and limited hand-labeled data available for certain regions. This framework allows us to generate, for the first time, coral reef segmentation maps at 2-meter resolution, which is a significant improvement over the kilometer-scale state-of-the-art maps. Additionally, this framework doubles accuracy and IoU metrics over baselines that do not account for domain invariance. Ata Akbari Asanjan, Kamalika Das, Alan S. Li, Ved Chirayath, Juan Torres-Perez, Soroosh Sorooshian |
KDD | 2 |
| 2018 | Understanding Climate-Vegetation Interactions in Global Rainforests Through a GP-Tree Analysis
Anuradha Kodali, Marcin Szubert, Kamalika Das, Sangram Ganguly, Josh C. Bongard |
PPSN (1) | 3 |
| 2017 | ASK-the-Expert: Active Learning Based Knowledge Discovery Using the Expert
Kamalika Das, Ilya Avrekh, Bryan L. Matthews, Manali Sharma, Nikunj C. Oza |
ECML/PKDD (3) | 1 |
| 2016 | Reducing Antagonism between Behavioral Diversity and Fitness in Semantic Genetic ProgrammingabstractMaintaining population diversity has long been considered fundamental to the effectiveness of evolutionary algorithms. Recently, with the advent of novelty search, there has been an increasing interest in sustaining behavioral diversity by using both fitness and behavioral novelty as separate search objectives. However, since the novelty objective explicitly rewards diverging from other individuals, it can antagonize the original fitness objective that rewards convergence toward the solution(s). As a result, fostering behavioral diversity may prevent proper exploitation of the most interesting regions of the behavioral space, and thus adversely affect the overall search performance. In this paper, we argue that an antagonism between behavioral diversity and fitness can indeed exist in semantic genetic programming applied to symbolic regression. Minimizing error draws individuals toward the target semantics but promoting novelty, defined as a distance in the semantic space, scatters them away from it. We introduce a less conflicting novelty metric, defined as an angular distance between two program semantics with respect to the target semantics. The experimental results show that this metric, in contrast to the other considered diversity promoting objectives, allows to consistently improve the performance of genetic programming regardless of whether it employs a syntactic or a semantic search operator. Marcin Szubert, Anuradha Kodali, Sangram Ganguly, Kamalika Das, Josh C. Bongard |
GECCO | 4 |
| 2016 | Active Learning with Rationales for Identifying Operationally Significant Anomalies in Aviation
Manali Sharma, Kamalika Das, Mustafa Bilgic 0001, Bryan L. Matthews, David Nielsen, Nikunj C. Oza |
ECML/PKDD (3) | 2 |
| 2016 | Semantic Forward Propagation for Symbolic Regression
Marcin Szubert, Anuradha Kodali, Sangram Ganguly, Kamalika Das, Josh C. Bongard |
PPSN | 4 |
| 2015 | Large scale support vector regression for aviation safetyabstractRegression problems on massive data sets are ubiquitous in many application domains including the Internet, earth and space sciences, and aviation. Support vector regression (SVR) is a popular technique for modeling the input-output relations of a set of variables under the added constraint of maximizing the margin, thereby leading to a very generalizable and regularized model. However, for a dataset with m training points, it is challenging to build SVR models due to the O(m3) cost involved in building them. In this paper we propose ParitoSVR - a parallel iterated optimizer for Support Vector Regression in the primal that can be deployed over a network of machines, where each machine iteratively solves a small (sub-)problem based only on the data observed locally and these solutions are then combined to form the solution to the global problem. Our proposed method is based on the Alternating Direction Method of Multipliers (ADMM) optimization technique. Unlike many other existing techniques, ParitoSVR is provably convergent to the results obtained from the centralized algorithm, where the optimization has access to the entire data set. The experimental results show that the algorithm is scalable both with respect to accuracy and time to convergence. We use ParitoSVR to identify flights having anomalous fuel consumption from a large fleet-wide commercial aviation database containing thousands of flights. Along with the algorithmic contributions, this paper also describes the process of deployment of the ADMM-based SVR method on a multicore architecture, namely, the NASA Pleiades supercomputing infrastructure. We have been successful in running ParitoSVR on millions of training data points and hundreds of compute nodes. Kamalika Das, Kanishka Bhaduri, Bryan L. Matthews, Nikunj C. Oza |
IEEE BigData | 1 |
| 2015 | PerCCS: person-count from carbon dioxide using sparse non-negative matrix factorizationabstractOccupancy count in rooms is valuable for applications such as room utilization, opportunistic meeting support, and efficient heating-cooling operations. Few buildings, however, have the means of knowing occupancy beyond simple binary presence-absence. In this paper we present the PerCCS algorithm that explores the possibility of estimating person count from CO2 sensors already integrated in everyday room air-conditioning infrastructure. PerCSS uses task-driven Sparse Non-negative Matrix Factorization (SNMF) to learn a nonnegative low-dimensional representation of the CO2 data in the preprocessing stage. This denoised CO2 acts as the predictor variable for estimating occupancy count using Ensemble Least Square Regression. We tested the algorithm to estimate 15 minutes average occupancy count from a classroom of capacity 42 and compared its performance against existing methods from the literature. PerCSS estimates occupancy with a normalized mean squared error (NMSE) of 0.075 and outperformed our comparative methods in predicting occupancy count with 91 % and 15 % for exact occupancy estimation, when the room was unoccupied and occupied respectively, whereas the competing methods failed mostly. Chandrayee Basu, Christian Koehler 0002, Kamalika Das, Anind K. Dey |
UbiComp | 3 |
| 2014 | Localizing anomalous changes in time-evolving graphsabstractGiven a time-evolving sequence of undirected, weighted graphs, we address the problem of localizing anomalous changes in graph structure over time. In this paper, we use the term `localization' to refer to the problem of identifying abnormal changes in node relationships (edges) that cause anomalous changes in graph structure. While there already exist several methods that can detect whether a graph transition is anomalous or not, these methods are not well suited for localizing the edges which are responsible for a transition being marked as an anomaly. This is a limitation in applications such as insider threat detection, where identifying the anomalous graph transitions is not sufficient, but rather, identifying the anomalous node relationships and associated nodes is key. To this end, we propose a novel, fast method based on commute time distance called CAD (Commute-time based Anomaly detection in Dynamic graphs) that detects node relationships responsible for abnormal changes in graph structure. In particular, CAD localizes anomalous edges by tracking a measure that combines information regarding changes in graph structure (in terms of commute time distance) as well as changes in edge weights. For large, sparse graphs, CAD returns a list of these anomalous edges and associated nodes in O(n\log n) time per graph instance in the sequence, where $n$ is the number of nodes. We analyze the performance of CAD on several synthetic and real-world data sets such as the Enron email network, the DBLP co-authorship network and a worldwide precipitation network data. Based on experiments conducted, we conclude that CAD consistently and efficiently identifies anomalous changes in relationships between nodes over time. Kumar Sricharan, Kamalika Das |
SIGMOD Conference | 2 |
| 2013 | Anomaly Detection in Vertically Partitioned Data by Distributed Core Vector Machines
Marco Stolpe, Kanishka Bhaduri, Kamalika Das, Katharina Morik |
ECML/PKDD (3) | 3 |
| 2011 | Distributed Monitoring of the R2 Statistic for Linear RegressionabstractThe problem of monitoring a multivariate linear regression model is relevant in studying the evolving relationship between a set of input variables (features) and one or more dependent target variables. This problem becomes challenging for large scale data in a distributed computing environment when only a subset of instances is available at individual nodes and the local data changes frequently. Data centralization and periodic model recomputation can add high overhead to tasks like anomaly detection in such dynamic settings. Therefore, the goal is to develop techniques for monitoring and updating the model over the union of all nodes' data in a communication-efficient fashion. Correctness guarantees on such techniques are also often highly desirable, especially in safety-critical application scenarios. In this paper we develop DReMo—a distributed algorithm with very low resource overhead, for monitoring the quality of a regression model in terms of its coefficient of determination (R2 statistic). When the nodes collectively determine that R2 has dropped below a fixed threshold, the linear regression model is recomputed via a network-wide convergecast and the updated model is broadcast back to all nodes. We show empirically, using both synthetic and real data, that our proposed method is highly communication-efficient and scalable, and also provide theoretical guarantees on correctness. Kanishka Bhaduri, Kamalika Das, Chris Giannella |
SDM | 2 |
| 2011 | Multi-objective optimization based privacy preserving distributed data mining in Peer-to-Peer networks
Kamalika Das, Kanishka Bhaduri, Hillol Kargupta |
Peer-to-Peer Netw. Appl. | 1 |
| 2010 | Block-GP: Scalable Gaussian Process Regression for Multimodal DataabstractRegression problems on massive data sets are ubiquitous in many application domains including the Internet, earth and space sciences, and finances. In many cases, regression algorithms such as linear regression or neural networks attempt to fit the target variable as a function of the input variables without regard to the underlying joint distribution of the variables. As a result, these global models are not sensitive to variations in the local structure of the input space. Several algorithms, including the mixture of experts model, classification and regression trees (CART), and others have been developed, motivated by the fact that a variability in the local distribution of inputs may be reflective of a significant change in the target variable. While these methods can handle the non-stationarity in the relationships to varying degrees, they are often not scalable and, therefore, not used in large scale data mining applications. In this paper we develop Block-GP, a Gaussian Process regression framework for multimodal data, that can be an order of magnitude more scalable than existing state-of-the-art nonlinear regression algorithms. The framework builds local Gaussian Processes on semantically meaningful partitions of the data and provides higher prediction accuracy than a single global model with very high confidence. The method relies on approximating the covariance matrix of the entire input space by smaller covariance matrices that can be modeled independently, and can therefore be parallelized for faster execution. Theoretical analysis and empirical studies on various synthetic and real data sets show high accuracy and scalability of Block-GP compared to existing nonlinear regression techniques. Kamalika Das, Ashok N. Srivastava |
ICDM | 1 |
| 2010 | A local asynchronous distributed privacy preserving feature selection algorithm for large peer-to-peer networks
Kamalika Das, Kanishka Bhaduri, Hillol Kargupta |
Knowl. Inf. Syst. | 1 |
| 2009 | A Local Distributed Peer-to-Peer Algorithm Using Multi-Party Optimization Based Privacy Preservation for Data Mining Primitive ComputationabstractThis paper proposes a scalable, local privacy-preserving algorithm for distributed peer-to-peer (P2P) data aggregation useful for many advanced data mining/analysis tasks such as average/sum computation, decision tree induction, feature selection, and more. Unlike most multi-party privacy-preserving data mining algorithms, this approach works in an asynchronous manner through local interactions and therefore, is highly scalable. It particularly deals with the distributed computation of the sum of a set of numbers stored at different peers in a P2P network in the context of a P2P Web mining application. The proposed optimization-based privacy-preserving technique for computing the sum allows different peers to specify different privacy requirements without having to adhere to a global set of parameters for the chosen privacy model. Since distributed sum computation is a frequently used primitive, the proposed approach is likely to have significant impact on many data mining tasks such as multi-party privacy-preserving clustering, frequent itemset mining, and statistical aggregate computation. Kamalika Das, Kanishka Bhaduri, Hillol Kargupta |
Peer-to-Peer Computing | 1 |
| 2009 | Scalable Distributed Change Detection from Astronomy Data Streams Using Local, Asynchronous Eigen Monitoring AlgorithmsabstractThis paper considers the problem of change detection using local distributed eigen monitoring algorithms for next generation of astronomy petascale data pipelines such as the Large Synoptic Survey Telescopes (LSST). This telescope will take repeat images of the night sky every 20 seconds, thereby generating 30 terabytes of calibrated imagery every night that will need to be co-analyzed with other astronomical data stored at different locations around the world. Change point detection and event classification in such data sets may provide useful insights to unique astronomical phenomenon displaying astrophysically significant variations: quasars, supernovae, variable stars, and potentially hazardous asteroids. However, performing such data mining tasks is a challenging problem for such high-throughput distributed data streams. In this paper we propose a highly scalable and distributed asynchronous algorithm for monitoring the principal components (PC) of such dynamic data streams. We demonstrate the algorithm on a large set of distributed astronomical data to accomplish well-known astronomy tasks such as measuring variations in the fundamental plane of galaxy parameters. The proposed algorithm is provably correct (i.e. converges to the correct PCs without centralizing any data) and can seamlessly handle changes to the data or the network. Real experiments performed on Sloan Digital Sky Survey (SDSS) catalogue data show the effectiveness of the algorithm. Kamalika Das, Kanishka Bhaduri, Sugandha Arora, Wesley Griffin, Kirk D. Borne, Chris Giannella, Hillol Kargupta |
SDM | 1 |
| 2008 | Distributed Identification of Top-l Inner Product Elements and its Application in a Peer-to-Peer NetworkabstractThe inner product measures how closely two feature vectors are related. It is an important primitive for many popular data mining tasks, for example, clustering, classification, correlation computation, and decision tree construction. If the entire data set is available at a single site, then computing the inner product matrix and identifying the top (in terms of magnitude) entries is trivial. However, in many real-world scenarios, data is distributed across many locations and transmitting the data to a central server would be quite communication intensive and not scalable. This paper presents an approximate local algorithm for identifying top-l, inner products among pairs of feature vectors in a large asynchronous distributed environment such as a peer-to-peer (P2P) network. We develop a probabilistic algorithm for this purpose using order statistics and the Hoeffding bound. We present experimental results to show the effectiveness and scalability of the algorithm. Finally, we demonstrate an application of this technique for interest-based community formation in a P2P environment. Kamalika Das, Kanishka Bhaduri, Kun Liu 0001, Hillol Kargupta |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2007 | Multi-party, Privacy-Preserving Distributed Data Mining Using a Game Theoretic Framework
Hillol Kargupta, Kamalika Das, Kun Liu 0001 |
PKDD | 2 |