Sunil Gupta 0001

dblp:47/333-1 · also Sunil Kumar Gupta 0001 · DBLP profile ↗
← Back
51ranked-venue papers in the field
8as first author
15since 2021 · last 2025
ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 50 (8 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2025 Score-Based Integrated Gradient for Root Cause Explanations of Outliers
abstract
Identifying the root causes of outliers is a fundamental problem in causal inference and anomaly detection. Traditional approaches based on heuristics or counterfactual reasoning often struggle under uncertainty and high-dimensional dependencies. We introduce SIREN, a novel and scalable method that attributes the root causes of outliers by estimating the score functions of the data likelihood. Attribution is computed via integrated gradients that accumulate score contributions along paths from the outlier toward the normal data distribution. Our method satisfies three of the four classic Shapley value axioms-dummy, efficiency, and linearity-as well as an asymmetry axiom derived from the underlying causal structure. Unlike prior work, SIREN operates directly on the score function, enabling tractable and uncertainty-aware root cause attribution in nonlinear, high-dimensional, and heteroscedastic causal models. Extensive experiments on synthetic random graphs and real-world cloud service and supply chain datasets show that SIREN outperforms state-of-the-art baselines in both attribution accuracy and computational efficiency.
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Svetha Venkatesh
ICDM3
2025 Federated Domain Generalization with Latent Space Inversion
abstract
Federated domain generalization (FedDG) addresses distribution shifts among clients in a federated learning frame-work. FedDG methods aggregate the parameters of locally trained client models to form a global model that generalizes to unseen clients while preserving data privacy. While improving the generalization capability of the global model, many existing approaches in FedDG jeopardize privacy by sharing statistics of client data between themselves. Our solution addresses this problem by contributing new ways to perform local client training and model aggregation. To improve local client training, we enforce (domain) invariance across local models with the help of a novel technique, latent space inversion, which enables better client privacy. When clients are not i.i.d, aggregating their local models may discard certain local adaptations. To overcome this, we propose an important weight aggregation strategy to prioritize parameters that significantly influence predictions of local models during aggregation. Our extensive experiments show that our approach achieves superior results over state-of-the-art methods with less communication overhead. Our code is available here.
Ragja Palakkadavath, Hung Le 0002, Thanh Nguyen-Tang, Svetha Venkatesh, Sunil Gupta 0001
ICDM5
2025 Defense Against Multi-target Multi-trigger Backdoor Attacks
Haripriya Harikumar, Santu Rana, Kien Do, Sunil Gupta 0001, Wei Zong, Willy Susilo, Svetha Venkatesh
PAKDD (6)4
2025 Designing Search Space for Unbounded Bayesian Optimization via Transfer Learning
Quoc Anh Hoang Nguyen, Hung The Tran, Sunil Gupta 0001, Dung D. Le
ECML/PKDD (5)3
2025 Hybrid Cross-Domain Robust Reinforcement Learning
Linh Le Pham Van, Hung Le 0002, Hung The Tran, Sunil Gupta 0001
ECML/PKDD (6)5
2024 Generating Realistic Tabular Data with Large Language Models
abstract
While most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation.
Dang Nguyen 0002, Sunil Gupta 0001, Kien Do, Thin Nguyen, Svetha Venkatesh
ICDM2
2024 Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
A. V. Arun Kumar, Alistair Shilton, Sunil Gupta 0001, Santu Rana, Stewart Greenhill, Svetha Venkatesh
ECML/PKDD (6)3
2024 PINN-BO: A Black-Box Optimization Algorithm Using Physics-Informed Neural Networks
Dat Phan-Trong, Hung The Tran, Alistair Shilton, Sunil Gupta 0001
ECML/PKDD (2)4
2024 Improving Diversity in Black-Box Few-Shot Knowledge Distillation
Tri-Nhan Vo, Dang Nguyen 0002, Kien Do, Sunil Gupta 0001
ECML/PKDD (2)4
2022 Real-Time Skill Discovery in Intelligent Virtual Assistants
Preeti Gopal, Sunil Gupta 0001, Santu Rana, Vuong Le, Trong Nguyen, Svetha Venkatesh
PAKDD (1)2
2021 Sparse Spectrum Gaussian Process for Bayesian Optimization
Cheng Li 0003, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
PAKDD (2)4
2021 Knowledge Distillation with Distribution Mismatch
Dang Nguyen 0002, Sunil Gupta 0001, Trong Nguyen, Santu Rana, Phuoc Nguyen, Truyen Tran 0001, Ky Le, Shannon Ryan, Svetha Venkatesh
ECML/PKDD (2)2
2021 Variational Hyper-encoding Networks
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Santu Rana, Hieu-Chi Dam, Svetha Venkatesh
ECML/PKDD (2)3
2021 Fast Conditional Network Compression Using Bayesian HyperNetworks
Phuoc Nguyen, Truyen Tran 0001, Ky Le, Sunil Gupta 0001, Santu Rana, Dang Nguyen 0002, Trong Nguyen, Shannon Ryan, Svetha Venkatesh
ECML/PKDD (3)4
2021 Fairness improvement for black-box classifiers with Gaussian process
Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Alistair Shilton, Svetha Venkatesh
Inf. Sci.2
2020 Level Set Estimation with Search Space Warping
Manisha Senadeera, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
PAKDD (2)3
2020 Scalable Backdoor Detection in Neural Networks
Haripriya Harikumar, Vuong Le, Santu Rana, Sourangshu Bhattacharya, Sunil Gupta 0001, Svetha Venkatesh
ECML/PKDD (2)5
2020 Bayesian Optimization with Missing Inputs
Phuc Luong, Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
ECML/PKDD (2)3
2019 Efficient Bayesian Optimization for Uncertainty Reduction Over Perceived Optima Locations
abstract
Bayesian optimization (BO) is concerned with efficient optimization using probabilistic methods. Predictive entropy search (PES) is a popular and successful BO strategy to find a point that maximizes the information gained about the optima location of an unknown function. Since the PES analytical form is intractable, it requires approximations and is computationally expensive. These approximations may degrade PES performance in terms of accuracy and efficiency. In this paper, we propose an alternative scheme - predictive variance reduction search (PVRS) - to find a point that maximally reduces the uncertainty at the perceived optima locations. The optimization converges to the true optimum when the uncertainty at all perceived optima locations is vanished. Our novel modification is beneficial in two ways. First, PVRS can be computed in closed-form, unlike the approximations made in PES. Second, PVRS is simple and easy to implement. As a result, the proposed PVRS gains huge speed up for scalable BO whilst showing favorable optimization efficiency. Furthermore, we extend our PVRS framework for batch setting where we select multiple experiments for parallel evaluations at each iteration. Empirically, we demonstrate the effectiveness of the PVRS on both benchmark functions and real-world applications in standard and batch BO settings.
Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, My T. Thai, Cheng Li 0003, Svetha Venkatesh
ICDM2
2019 Incomplete Conditional Density Estimation for Fast Materials Discovery
abstract
Designing new physical products and processes requires enormous experimentation. The scientific simulators play a fundamental role for such design tasks. To design a new product with certain target characteristics, a search is performed in the design space by trying out a large number of design combinations through simulators before reaching to the target characteristics. However, searching for the target design using simulators is generally expensive and becomes prohibitive when the target is either revised or only partially specified. To address this problem, we use a machine learning model to predict the design in single step using the target product specifications as input. We overcome two technical challenges: the first caused due to one-to-many mapping when learning the inverse problem and the second caused due to a user specifying the target specifications only partially. We unify a conditional variational auto-encoder model (to address the partial target specification) with mixture density networks (to address the one-to-many mapping) and train an end-to-end model to predict the optimum design.
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Santu Rana, Matthew Barnett, Svetha Venkatesh
SDM3
2019 Filtering Bayesian optimization approach in weakly specified search space
Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, Cheng Li 0003, Svetha Venkatesh
Knowl. Inf. Syst.2
2018 Accelerating Experimental Design by Incorporating Experimenter Hunches
abstract
Experimental design is a process of obtaining a product with target property via experimentation. Bayesian optimization offers a sample-efficient tool for experimental design when experiments are expensive. Often, expert experimenters have 'hunches' about the behavior of the experimental system, offering potentials to further improve the efficiency. In this paper, we consider per-variable monotonic trend in the underlying property that results in a unimodal trend in those variables for a target value optimization. For example, sweetness of a candy is monotonic to the sugar content. However, to obtain a target sweetness, the utility of the sugar content becomes a unimodal function, which peaks at the value giving the target sweetness and falls off both ways. In this paper, we propose a novel method to solve such problems that achieves two main objectives: (a) the monotonicity information is used to the fullest extent possible, whilst ensuring that (b) the convergence guarantee remains intact. This is achieved by a two-stage Gaussian process modeling, where the first stage uses the monotonicity trend to model the underlying property, and the second stage uses 'virtual' samples, sampled from the first, to model the target value optimization function. The process is made theoretically consistent by adding appropriate adjustment factor in the posterior computation, necessitated because of using the 'virtual' samples. The proposed method is evaluated through both simulations and real world experimental design problems of (a) new short polymer fiber with the target length, and (b) designing of a new three dimensional porous scaffolding with a target porosity. In all scenarios our method demonstrates faster convergence than the basic Bayesian optimization approach not using such 'hunches'.
Cheng Li 0003, Santu Rana, Sunil Gupta 0001, Vu Nguyen 0001, Svetha Venkatesh, Alessandra Sutti, David Rubin de Celis Leal, Teo Slezak, Murray Height, Mazher Mohammed, Ian Gibson
ICDM3
2018 Differentially Private Prescriptive Analytics
abstract
Privacy preservation is important. Prescriptive analytics is a method to extract corrective actions to avoid undesirable outcomes. We propose a privacy preserving prescriptive analytics algorithm to protect the data used during the construction of the prescriptive analytics algorithm. We use differential privacy mechanism to achieve strong privacy guarantee. Differential privacy mechanism requires computation of sensitivity: maximum change in the output between two training datasets, which is differed by only one instance. The main challenge we addressed is the computation of sensitivity of the prescription vector. In absence of any analytical form, we construct a nested global optimization problem to compute the sensitivity. We solve the optimization problem using constrained Bayesian optimization, as the nested structure makes the objective function expensive. We demonstrate our algorithm on two real world datasets and observe that the prescription vectors remains useful even after making them private.
Haripriya Harikumar, Santu Rana, Sunil Gupta 0001, Thin Nguyen, M. R. Kaimal 0001, Svetha Venkatesh
ICDM3
2018 Prescriptive Analytics Through Constrained Bayesian Optimization
Haripriya Harikumar, Santu Rana, Sunil Gupta 0001, Thin Nguyen, M. R. Kaimal 0001, Svetha Venkatesh
PAKDD (1)3
2018 A Privacy Preserving Bayesian Optimization with High Efficiency
Thanh Dai Nguyen, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
PAKDD (3)2
2018 Exploration Enhanced Expected Improvement for Bayesian Optimization
Julian Berk, Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
ECML/PKDD (2)3
2018 Information-Theoretic Transfer Learning Framework for Bayesian Optimisation
Anil Ramachandran, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
ECML/PKDD (2)2
2017 Bayesian Optimization in Weakly Specified Search Space
abstract
Bayesian optimization (BO) has recently emerged as a powerful and flexible tool for hyper-parameter tuning and more generally for the efficient global optimization of expensive black-box functions. Systems implementing BO has successfully solved difficult problems in automatic design choices and machine learning hyper-parameters tunings. Many recent advances in the methodologies and theories underlying Bayesian optimization have extended the framework to new applications and provided greater insights into the behavior of these algorithms. Still, these established techniques always require a user-defined space to perform optimization. This pre-defined space specifies the ranges of hyper-parameter values. In many situations, however, it can be difficult to prescribe such spaces, as a prior knowledge is often unavailable. Setting these regions arbitrarily can lead to inefficient optimization - if a space is too large, we can miss the optimum with a limited budget, on the other hand, if a space is too small, it may not contain the optimum point that we want to get. The unknown search space problem is intractable to solve in practice. Therefore, in this paper, we narrow down to consider specifically the setting of "weakly specified" search space for Bayesian optimization. By weakly specified space, we mean that the pre-defined space is placed at a sufficiently good region so that the optimization can expand and reach to the optimum. However, this pre-defined space need not include the global optimum. We tackle this problem by proposing the filtering expansion strategy for Bayesian optimization. Our approach starts from the initial region and gradually expands the search space. Wedevelop an efficient algorithm for this strategy and derive its regret bound. These theoretical results are complemented by an extensive set of experiments on benchmark functions and tworeal-world applications which demonstrate the benefits of our proposed approach.
Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, Cheng Li 0003, Svetha Venkatesh
ICDM2
2017 Stable Bayesian Optimization
Thanh Dai Nguyen, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
PAKDD (2)2
2017 Effective sparse imputation of patient conditions in electronic medical records for emergency risk predictions
Budhaditya Saha, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh
Knowl. Inf. Syst.2
2017 Nonparametric discovery and analysis of learning patterns and autism subgroups from therapeutic data
Pratibha Vellanki, Thi V. Duong, Sunil Gupta 0001, Svetha Venkatesh, Dinh Q. Phung
Knowl. Inf. Syst.3
2016 Understanding Behavioral Differences Between Short and Long-Term Drinking Abstainers from Social Media
Haripriya Harikumar, Thin Nguyen, Sunil Gupta 0001, Santu Rana, M. R. Kaimal 0001, Svetha Venkatesh
ADMA3
2016 Extracting Key Challenges in Achieving Sobriety Through Shared Subspace Learning
Haripriya Harikumar, Thin Nguyen, Santu Rana, Sunil Gupta 0001, M. R. Kaimal 0001, Svetha Venkatesh
ADMA4
2016 Budgeted Batch Bayesian Optimization
abstract
Parameter settings profoundly impact the performance of machine learning algorithms and laboratory experiments. The classical trial-error methods are exponentially expensive in large parameter spaces, and Bayesian optimization (BO) offers an elegant alternative for global optimization of black box functions. In situations where the functions can be evaluated at multiple points simultaneously, batch Bayesian optimization is used. Current batch BO approaches are restrictive in fixing the number of evaluations per batch, and this can be wasteful when the number of specified evaluations is larger than the number of real maxima in the underlying acquisition function. We present the budgeted batch Bayesian optimization (B3O) for hyper-parameter tuning and experimental design - we identify the appropriate batch size for each iteration in an elegant way. In particular, we use the infinite Gaussian mixture model (IGMM) for automatically identifying the number of peaks in the underlying acquisition functions. We solve the intractability of estimating the IGMM directly from the acquisition function by formulating the batch generalized slice sampling to efficiently draw samples from the acquisition function. We perform extensive experiments for benchmark functions and two real world applications - machine learning hyper-parameter tuning and experimental design for alloy hardening. We show empirically that the proposed B3O outperforms the existing fixed batch BO approaches in finding the optimum whilst requiring a fewer number of evaluations, thus saving cost and time.
Vu Nguyen 0001, Santu Rana, Sunil Gupta 0001, Cheng Li 0003, Svetha Venkatesh
ICDM3
2016 Flexible Transfer Learning Framework for Bayesian Optimisation
Tinu Theckel Joy, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
PAKDD (1)3
2016 Toxicity Prediction in Cancer Using Multiple Instance Learning in a Multi-task Framework
Cheng Li 0003, Sunil Gupta 0001, Santu Rana, Wei Luo 0001, Svetha Venkatesh, David Ashely, Dinh Q. Phung
PAKDD (1)2
2016 Privacy Aware K-Means Clustering with High Utility
Thanh Dai Nguyen, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
PAKDD (2)2
2016 A new transfer learning framework with application to model-agnostic multi-task learning
Sunil Gupta 0001, Santu Rana, Budhaditya Saha, Dinh Q. Phung, Svetha Venkatesh
Knowl. Inf. Syst.1
2016 Multiple task transfer learning with small sample sizes
Budhaditya Saha, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh
Knowl. Inf. Syst.2
2015 Exploiting feature relationships towards stable feature selection
abstract
Feature selection is an important step in building predictive models for most real-world problems. One of the popular methods in feature selection is Lasso. However, it shows instability in selecting features when dealing with correlated features. In this work, we propose a new method that aims to increase the stability of Lasso by encouraging similarities between features based on their relatedness, which is captured via a feature covariance matrix. Besides modeling positive feature correlations, our method can also identify negative correlations between features. We propose a convex formulation for our model along with an alternating optimization algorithm that can learn the weights of the features as well as the relationship between them. Using both synthetic and real-world data, we show that the proposed method is more stable than Lasso and many state-of-the-art shrinkage and feature selection methods. Also, its predictive performance is comparable to other methods.
Iman Kamkar, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh
DSAA2
2015 Improved risk predictions via sparse imputation of patient conditions in electronic medical records
abstract
Electronic Medical Records (EMR) are increasingly used for risk prediction. EMR analysis is complicated by missing entries. There are two reasons - the “primary reason for admission” is included in EMR, but the co-morbidities (other chronic diseases) are left uncoded, and, many zero values in the data are accurate, reflecting that a patient has not accessed medical facilities. A key challenge is to deal with the peculiarities of this data - unlike many other datasets, EMR is sparse, reflecting the fact that patients have some, but not all diseases. We propose a novel model to fill-in these missing values, and use the new representation for prediction of key hospital events. To “fill-in” missing values, we represent the feature-patient matrix as a product of two low rank factors, preserving the sparsity property in the product. Intuitively, the product regularization allows sparse imputation of patient conditions reflecting common comorbidities across patients. We develop a scalable optimization algorithm based on Block coordinate descent method to find an optimal solution. We evaluate the proposed framework on two real world EMR cohorts: Cancer (7000 admissions) and Acute Myocardial Infarction (2652 admissions). Our result shows that the AUC for 3 months admission prediction is improved significantly from (0.741 to 0.786) for Cancer data and (0.678 to 0.724) for AMI data. We also extend the proposed method to a supervised model for predicting of multiple related risk outcomes (e.g. emergency presentations and admissions in hospital over 3, 6 and 12 months period) in an integrated framework. For this model, the AUC averaged over outcomes is improved significantly from (0.768 to 0.806) for Cancer data and (0.685 to 0.748) for AMI data.
Budhaditya Saha, Sunil Gupta 0001, Svetha Venkatesh
DSAA2
2015 Differentially Private Random Forest with High Utility
abstract
Privacy-preserving data mining has become an active focus of the research community in the domains where data are sensitive and personal in nature. For example, highly sensitive digital repositories of medical or financial records offer enormous values for risk prediction and decision making. However, prediction models derived from such repositories should maintain strict privacy of individuals. We propose a novel random forest algorithm under the framework of differential privacy. Unlike previous works that strictly follow differential privacy and keep the complete data distribution approximately invariant to change in one data instance, we only keep the necessary statistics (e.g. variance of the estimate) invariant. This relaxation results in significantly higher utility. To realize our approach, we propose a novel differentially private decision tree induction algorithm and use them to create an ensemble of decision trees. We also propose feasible adversary models to infer about the attribute and class label of unknown data in presence of the knowledge of all other data. Under these adversary models, we derive bounds on the maximum number of trees that are allowed in the ensemble while maintaining privacy. We focus on binary classification problem and demonstrate our approach on four real-world datasets. Compared to the existing privacy preserving approaches we achieve significantly higher utility.
Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
ICDM2
2015 Collaborating Differently on Different Topics: A Multi-Relational Approach to Multi-Task Learning
Sunil Gupta 0001, Santu Rana, Dinh Q. Phung, Svetha Venkatesh
PAKDD (1)1
2015 Prediciton of Emergency Events: A Multi-Task Multi-Label Learning Approach
Budhaditya Saha, Sunil Gupta 0001, Svetha Venkatesh
PAKDD (1)2
2015 What shall I share and with Whom? - A Multi-Task Learning Formulation using Multi-Faceted Task Relationships
abstract
Multi-task learning is a learning paradigm that improves the performance of “related” tasks through their joint learning. To do this each task answers the question “Which other task should I share with”? This task relatedness can be complex - a task may be related to one set of tasks based on one subset of features and to other tasks based on other subsets. Existing multi-task learning methods do not explicitly model this reality, learning a single-faceted task relationship over all the features. This degrades performance by forcing a task to become similar to other tasks even on their unrelated features. Addressing this gap, we propose a novel multi-task learning model that learns multi-faceted task relationship, allowing tasks to collaborate differentially on different feature subsets. This is achieved by simultaneously learning a low dimensional subspace for task parameters and inducing task groups over each latent subspace basis using a novel combination of L1 and pairwise L∞ norms. Further, our model can induce grouping across both positively and negatively related tasks, which helps towards exploiting knowledge from all types of related tasks. We validate our model on two synthetic and five real datasets, and show significant performance improvements over several state-of-the-art multi-task learning techniques. Thus our model effectively answers for each task: What shall I share and with whom?
Sunil Gupta 0001, Santu Rana, Dinh Q. Phung, Svetha Venkatesh
SDM1
2014 Intervention-Driven Predictive Framework for Modeling Healthcare Data
Santu Rana, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh
PAKDD (1)2
2014 Keeping up with Innovation: A Predictive Framework for Modeling Healthcare Data with Evolving Clinical Interventions
abstract
Medical outcomes are inexorably linked to patient illness and clinical interventions. Interventions change the course of disease, crucially determining outcome. Traditional outcome prediction models build a single classifier by augmenting interventions with disease information. Interventions, however, differentially affect prognosis, thus a single prediction rule may not suffice to capture variations. Interventions also evolve over time as more advanced interventions replace older ones. To this end, we propose a Bayesian nonparametric, supervised framework that models a set of intervention groups through a mixture distribution building a separate prediction rule for each group, and allows the mixture distribution to change with time. This is achieved by using a hierarchical Dirichlet process mixture model over the interventions. The outcome is then modeled as conditional on both the latent grouping and the disease information through a Bayesian logistic regression. Experiments on synthetic and medical cohorts for 30-day readmission prediction demonstrate the superiority of the proposed model over clinical and data mining baselines.
Sunil Gupta 0001, Santu Rana, Dinh Q. Phung, Svetha Venkatesh
SDM1
2013 Regularized nonnegative shared subspace learning
Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Svetha Venkatesh
Data Min. Knowl. Discov.1
2012 A Bayesian Nonparametric Joint Factor Model for Learning Shared and Individual Subspaces from Multiple Data Sources
abstract
Joint analysis of multiple data sources is becoming increasingly popular in transfer learning, multi-task learning and cross-domain data mining.One promising approach to model the data jointly is through learning the shared and individual factor subspaces.However, performance of this approach depends on the subspace dimensionalities and the level of sharing needs to be specified a priori.To this end, we propose a nonparametric joint factor analysis framework for modeling multiple related data sources.Our model utilizes the hierarchical beta process as a nonparametric prior to automatically infer the number of shared and individual factors.For posterior inference, we provide a Gibbs sampling scheme using auxiliary variables.The effectiveness of the proposed framework is validated through its application on two real world problemstransfer learning in text and image retrieval.
Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh
SDM1
2011 A Bayesian Framework for Learning Shared and Individual Subspaces from Multiple Data Sources
Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Svetha Venkatesh
PAKDD (1)1
2010 Nonnegative shared subspace learning and its application to social media retrieval
abstract
Although tagging has become increasingly popular in online image and video sharing systems, tags are known to be noisy, ambiguous, incomplete and subjective. These factors can seriously affect the precision of a social tag-based web retrieval system. Therefore improving the precision performance of these social tag-based web retrieval systems has become an increasingly important research topic. To this end, we propose a shared subspace learning framework to leverage a secondary source to improve retrieval performance from a primary dataset. This is achieved by learning a shared subspace between the two sources under a joint Nonnegative Matrix Factorization in which the level of subspace sharing can be explicitly controlled. We derive an efficient algorithm for learning the factorization, analyze its complexity, and provide proof of convergence. We validate the framework on image and video retrieval tasks in which tags from the LabelMe dataset are used to improve image retrieval performance from a Flickr dataset and video retrieval performance from a YouTube dataset. This has implications for how to exploit and transfer knowledge from readily available auxiliary tagging resources to improve another social web retrieval system. Our shared subspace learning framework is applicable to a range of problems where one needs to exploit the strengths existing among multiple and heterogeneous datasets.
Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Truyen Tran 0001, Svetha Venkatesh
KDD1