EDBT 2026 Demo / reviewers in the wild / expert
Svetha Venkatesh
dblp:81/1984
· DBLP profile ↗
113ranked-venue papers in the field
0as first author
12since 2021 · last 2025
0000-0001-8675-6631ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 91Information Retrieval & Web Search · 15Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 2Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Score-Based Integrated Gradient for Root Cause Explanations of OutliersabstractIdentifying the root causes of outliers is a fundamental problem in causal inference and anomaly detection. Traditional approaches based on heuristics or counterfactual reasoning often struggle under uncertainty and high-dimensional dependencies. We introduce SIREN, a novel and scalable method that attributes the root causes of outliers by estimating the score functions of the data likelihood. Attribution is computed via integrated gradients that accumulate score contributions along paths from the outlier toward the normal data distribution. Our method satisfies three of the four classic Shapley value axioms-dummy, efficiency, and linearity-as well as an asymmetry axiom derived from the underlying causal structure. Unlike prior work, SIREN operates directly on the score function, enabling tractable and uncertainty-aware root cause attribution in nonlinear, high-dimensional, and heteroscedastic causal models. Extensive experiments on synthetic random graphs and real-world cloud service and supply chain datasets show that SIREN outperforms state-of-the-art baselines in both attribution accuracy and computational efficiency. Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Svetha Venkatesh |
ICDM | 4 |
| 2025 | Federated Domain Generalization with Latent Space InversionabstractFederated domain generalization (FedDG) addresses distribution shifts among clients in a federated learning frame-work. FedDG methods aggregate the parameters of locally trained client models to form a global model that generalizes to unseen clients while preserving data privacy. While improving the generalization capability of the global model, many existing approaches in FedDG jeopardize privacy by sharing statistics of client data between themselves. Our solution addresses this problem by contributing new ways to perform local client training and model aggregation. To improve local client training, we enforce (domain) invariance across local models with the help of a novel technique, latent space inversion, which enables better client privacy. When clients are not i.i.d, aggregating their local models may discard certain local adaptations. To overcome this, we propose an important weight aggregation strategy to prioritize parameters that significantly influence predictions of local models during aggregation. Our extensive experiments show that our approach achieves superior results over state-of-the-art methods with less communication overhead. Our code is available here. Ragja Palakkadavath, Hung Le 0002, Thanh Nguyen-Tang, Svetha Venkatesh, Sunil Gupta 0001 |
ICDM | 4 |
| 2025 | Defense Against Multi-target Multi-trigger Backdoor Attacks
Haripriya Harikumar, Santu Rana, Kien Do, Sunil Gupta 0001, Wei Zong, Willy Susilo, Svetha Venkatesh |
PAKDD (6) | 7 |
| 2024 | Generating Realistic Tabular Data with Large Language ModelsabstractWhile most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation. Dang Nguyen 0002, Sunil Gupta 0001, Kien Do, Thin Nguyen, Svetha Venkatesh |
ICDM | 5 |
| 2024 | Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
A. V. Arun Kumar, Alistair Shilton, Sunil Gupta 0001, Santu Rana, Stewart Greenhill, Svetha Venkatesh |
ECML/PKDD (6) | 6 |
| 2024 | Variable-Agnostic Causal Exploration for Reinforcement Learning
Hung Le 0002, Svetha Venkatesh |
ECML/PKDD (2) | 3 |
| 2022 | Real-Time Skill Discovery in Intelligent Virtual Assistants
Preeti Gopal, Sunil Gupta 0001, Santu Rana, Vuong Le, Trong Nguyen, Svetha Venkatesh |
PAKDD (1) | 6 |
| 2021 | Sparse Spectrum Gaussian Process for Bayesian Optimization
Cheng Li 0003, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh |
PAKDD (2) | 5 |
| 2021 | Knowledge Distillation with Distribution Mismatch
Dang Nguyen 0002, Sunil Gupta 0001, Trong Nguyen, Santu Rana, Phuoc Nguyen, Truyen Tran 0001, Ky Le, Shannon Ryan, Svetha Venkatesh |
ECML/PKDD (2) | 9 |
| 2021 | Variational Hyper-encoding Networks
Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Santu Rana, Hieu-Chi Dam, Svetha Venkatesh |
ECML/PKDD (2) | 6 |
| 2021 | Fast Conditional Network Compression Using Bayesian HyperNetworks
Phuoc Nguyen, Truyen Tran 0001, Ky Le, Sunil Gupta 0001, Santu Rana, Dang Nguyen 0002, Trong Nguyen, Shannon Ryan, Svetha Venkatesh |
ECML/PKDD (3) | 9 |
| 2021 | Fairness improvement for black-box classifiers with Gaussian process
Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Alistair Shilton, Svetha Venkatesh |
Inf. Sci. | 5 |
| 2020 | Level Set Estimation with Search Space Warping
Manisha Senadeera, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2020 | Scalable Backdoor Detection in Neural Networks
Haripriya Harikumar, Vuong Le, Santu Rana, Sourangshu Bhattacharya, Sunil Gupta 0001, Svetha Venkatesh |
ECML/PKDD (2) | 6 |
| 2020 | Bayesian Optimization with Missing Inputs
Phuc Luong, Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
ECML/PKDD (2) | 5 |
| 2019 | Efficient Bayesian Optimization for Uncertainty Reduction Over Perceived Optima LocationsabstractBayesian optimization (BO) is concerned with efficient optimization using probabilistic methods. Predictive entropy search (PES) is a popular and successful BO strategy to find a point that maximizes the information gained about the optima location of an unknown function. Since the PES analytical form is intractable, it requires approximations and is computationally expensive. These approximations may degrade PES performance in terms of accuracy and efficiency. In this paper, we propose an alternative scheme - predictive variance reduction search (PVRS) - to find a point that maximally reduces the uncertainty at the perceived optima locations. The optimization converges to the true optimum when the uncertainty at all perceived optima locations is vanished. Our novel modification is beneficial in two ways. First, PVRS can be computed in closed-form, unlike the approximations made in PES. Second, PVRS is simple and easy to implement. As a result, the proposed PVRS gains huge speed up for scalable BO whilst showing favorable optimization efficiency. Furthermore, we extend our PVRS framework for batch setting where we select multiple experiments for parallel evaluations at each iteration. Empirically, we demonstrate the effectiveness of the PVRS on both benchmark functions and real-world applications in standard and batch BO settings. Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, My T. Thai, Cheng Li 0003, Svetha Venkatesh |
ICDM | 6 |
| 2019 | Graph Transformation Policy Network for Chemical Reaction PredictionabstractWe address a fundamental problem in chemistry known as chemical reaction product prediction. Our main insight is that the input reactant and reagent molecules can be jointly represented as a graph, and the process of generating product molecules from reactant molecules can be formulated as a sequence of graph transformations. To this end, we propose Graph Transformation Policy Network (GTPN) - a novel generic method that combines the strengths of graph neural networks and reinforcement learning to learn reactions directly from data with minimal chemical knowledge. Compared to previous methods, GTPN has some appealing properties such as: end-to-end learning, and making no assumption about the length or the order of graph transformations. In order to guide model search through the complex discrete space of sets of bond changes effectively, we extend the standard policy gradient loss by adding useful constraints. Evaluation results show that GTPN improves the top-1 accuracy over the current state-of-the-art method by about 3% on the large USPTO dataset. Kien Do, Truyen Tran 0001, Svetha Venkatesh |
KDD | 3 |
| 2019 | Incomplete Conditional Density Estimation for Fast Materials DiscoveryabstractDesigning new physical products and processes requires enormous experimentation. The scientific simulators play a fundamental role for such design tasks. To design a new product with certain target characteristics, a search is performed in the design space by trying out a large number of design combinations through simulators before reaching to the target characteristics. However, searching for the target design using simulators is generally expensive and becomes prohibitive when the target is either revised or only partially specified. To address this problem, we use a machine learning model to predict the design in single step using the target product specifications as input. We overcome two technical challenges: the first caused due to one-to-many mapping when learning the inverse problem and the second caused due to a user specifying the target specifications only partially. We unify a conditional variational auto-encoder model (to address the partial target specification) with mixture density networks (to address the one-to-many mapping) and train an end-to-end model to predict the optimum design. Phuoc Nguyen, Truyen Tran 0001, Sunil Gupta 0001, Santu Rana, Matthew Barnett, Svetha Venkatesh |
SDM | 6 |
| 2019 | Filtering Bayesian optimization approach in weakly specified search space
Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, Cheng Li 0003, Svetha Venkatesh |
Knowl. Inf. Syst. | 5 |
| 2018 | Accelerating Experimental Design by Incorporating Experimenter HunchesabstractExperimental design is a process of obtaining a product with target property via experimentation. Bayesian optimization offers a sample-efficient tool for experimental design when experiments are expensive. Often, expert experimenters have 'hunches' about the behavior of the experimental system, offering potentials to further improve the efficiency. In this paper, we consider per-variable monotonic trend in the underlying property that results in a unimodal trend in those variables for a target value optimization. For example, sweetness of a candy is monotonic to the sugar content. However, to obtain a target sweetness, the utility of the sugar content becomes a unimodal function, which peaks at the value giving the target sweetness and falls off both ways. In this paper, we propose a novel method to solve such problems that achieves two main objectives: (a) the monotonicity information is used to the fullest extent possible, whilst ensuring that (b) the convergence guarantee remains intact. This is achieved by a two-stage Gaussian process modeling, where the first stage uses the monotonicity trend to model the underlying property, and the second stage uses 'virtual' samples, sampled from the first, to model the target value optimization function. The process is made theoretically consistent by adding appropriate adjustment factor in the posterior computation, necessitated because of using the 'virtual' samples. The proposed method is evaluated through both simulations and real world experimental design problems of (a) new short polymer fiber with the target length, and (b) designing of a new three dimensional porous scaffolding with a target porosity. In all scenarios our method demonstrates faster convergence than the basic Bayesian optimization approach not using such 'hunches'. Cheng Li 0003, Santu Rana, Sunil Gupta 0001, Vu Nguyen 0001, Svetha Venkatesh, Alessandra Sutti, David Rubin de Celis Leal, Teo Slezak, Murray Height, Mazher Mohammed, Ian Gibson |
ICDM | 5 |
| 2018 | Differentially Private Prescriptive AnalyticsabstractPrivacy preservation is important. Prescriptive analytics is a method to extract corrective actions to avoid undesirable outcomes. We propose a privacy preserving prescriptive analytics algorithm to protect the data used during the construction of the prescriptive analytics algorithm. We use differential privacy mechanism to achieve strong privacy guarantee. Differential privacy mechanism requires computation of sensitivity: maximum change in the output between two training datasets, which is differed by only one instance. The main challenge we addressed is the computation of sensitivity of the prescription vector. In absence of any analytical form, we construct a nested global optimization problem to compute the sensitivity. We solve the optimization problem using constrained Bayesian optimization, as the nested structure makes the objective function expensive. We demonstrate our algorithm on two real world datasets and observe that the prescription vectors remains useful even after making them private. Haripriya Harikumar, Santu Rana, Sunil Gupta 0001, Thin Nguyen, M. R. Kaimal 0001, Svetha Venkatesh |
ICDM | 6 |
| 2018 | Dual Memory Neural Computer for Asynchronous Two-view Sequential LearningabstractOne of the core tasks in multi-view learning is to capture relations among views. For sequential data, the relations not only span across views, but also extend throughout the view length to form long-term intra-view and inter-view interactions. In this paper, we present a new memory augmented neural network that aims to model these complex interactions between two asynchronous sequential views. Our model uses two encoders for reading from and writing to two external memories for encoding input views. The intra-view interactions and the long-term dependencies are captured by the use of memories during this encoding process. There are two modes of memory accessing in our system: late-fusion and early-fusion, corresponding to late and early inter-view interactions. In the late-fusion mode, the two memories are separated, containing only view-specific contents. In the early-fusion mode, the two memories share the same addressing space, allowing cross-memory accessing. In both cases, the knowledge from the memories will be combined by a decoder to make predictions over the output space. The resulting dual memory neural computer is demonstrated on a comprehensive set of experiments, including a synthetic task of summing two sequences and the tasks of drug prescription and disease progression in healthcare. The results demonstrate competitive performance over both traditional algorithms and deep learning methods designed for multi-view problems. Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh |
KDD | 3 |
| 2018 | Trans2Vec: Learning Transaction Embedding via Items and Frequent Itemsets
Dang Nguyen 0002, Tu Dinh Nguyen, Wei Luo 0001, Svetha Venkatesh |
PAKDD (3) | 4 |
| 2018 | Prescriptive Analytics Through Constrained Bayesian Optimization
Haripriya Harikumar, Santu Rana, Sunil Gupta 0001, Thin Nguyen, M. R. Kaimal 0001, Svetha Venkatesh |
PAKDD (1) | 6 |
| 2018 | Dual Control Memory Augmented Neural Networks for Treatment Recommendations
Hung Le 0002, Truyen Tran 0001, Svetha Venkatesh |
PAKDD (3) | 3 |
| 2018 | A Privacy Preserving Bayesian Optimization with High Efficiency
Thanh Dai Nguyen, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
PAKDD (3) | 4 |
| 2018 | Sqn2Vec: Learning Sequence Representation via Sequential Patterns with a Gap Constraint
Dang Nguyen 0002, Wei Luo 0001, Tu Dinh Nguyen, Svetha Venkatesh, Dinh Q. Phung |
ECML/PKDD (2) | 4 |
| 2018 | Exploration Enhanced Expected Improvement for Bayesian Optimization
Julian Berk, Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
ECML/PKDD (2) | 5 |
| 2018 | Information-Theoretic Transfer Learning Framework for Bayesian Optimisation
Anil Ramachandran, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
ECML/PKDD (2) | 4 |
| 2018 | Learning Graph Representation via Frequent SubgraphsabstractWe propose a novel approach to learn distributed representation for graph data. Our idea is to combine a recently introduced neural document embedding model with a traditional pattern mining technique, by treating a graph as a document and frequent subgraphs as atomic units for the embedding process. Compared to the latest graph embedding methods, our proposed method offers three key advantages: fully unsupervised learning, entire-graph embedding, and edge label leveraging. We demonstrate our method on several datasets in comparison with a comprehensive list of up-to-date state-of-the-art baselines where we show its advantages for both classification and clustering tasks. Dang Nguyen 0002, Wei Luo 0001, Tu Dinh Nguyen, Svetha Venkatesh, Dinh Q. Phung |
SDM | 4 |
| 2018 | Jointly Predicting Affective and Mental Health Scores Using Deep Neural Networks of Visual Cues on the Web
Van Nguyen 0002, Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen |
WISE (2) | 9 |
| 2018 | Discovering topic structures of a temporally evolving document corpusabstractIn this paper we describe a novel framework for the discovery of the topical content of a data corpus, and the tracking of its complex structural changes across the temporal dimension. In contrast to previous work our model does not impose a prior on the rate at which documents are added to the corpus nor does it adopt the Markovian assumption which overly restricts the type of changes that the model can capture. Our key technical contribution is a framework based on (i) discretization of time into epochs, (ii) epoch-wise topic discovery using a hierarchical Dirichlet process-based model, and (iii) a temporal similarity graph which allows for the modelling of complex topic changes: emergence and disappearance, evolution, splitting, and merging. The power of the proposed framework is demonstrated on two medical literature corpora concerned with the autism spectrum disorder (ASD) and the metabolic syndrome (MetS)—both increasingly important research subjects with significant social and healthcare consequences. In addition to the collected ASD and metabolic syndrome literature corpora which we made freely available, our contribution also includes an extensive empirical analysis of the proposed framework. We describe a detailed and careful examination of the effects that our algorithms’s free parameters have on its output and discuss the significance of the findings both in the context of the practical application of our algorithm as well as in the context of the existing body of work on temporal topic analysis. Our quantitative analysis is followed by several qualitative case studies highly relevant to the current research on ASD and MetS, on which our algorithm is shown to capture well the actual developments in these fields. Adham Beykikhoshk, Ognjen Arandjelovic, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2018 | Energy-based anomaly detection for mixed data
Kien Do, Truyen Tran 0001, Svetha Venkatesh |
Knowl. Inf. Syst. | 3 |
| 2017 | Bayesian Optimization in Weakly Specified Search SpaceabstractBayesian optimization (BO) has recently emerged as a powerful and flexible tool for hyper-parameter tuning and more generally for the efficient global optimization of expensive black-box functions. Systems implementing BO has successfully solved difficult problems in automatic design choices and machine learning hyper-parameters tunings. Many recent advances in the methodologies and theories underlying Bayesian optimization have extended the framework to new applications and provided greater insights into the behavior of these algorithms. Still, these established techniques always require a user-defined space to perform optimization. This pre-defined space specifies the ranges of hyper-parameter values. In many situations, however, it can be difficult to prescribe such spaces, as a prior knowledge is often unavailable. Setting these regions arbitrarily can lead to inefficient optimization - if a space is too large, we can miss the optimum with a limited budget, on the other hand, if a space is too small, it may not contain the optimum point that we want to get. The unknown search space problem is intractable to solve in practice. Therefore, in this paper, we narrow down to consider specifically the setting of "weakly specified" search space for Bayesian optimization. By weakly specified space, we mean that the pre-defined space is placed at a sufficiently good region so that the optimization can expand and reach to the optimum. However, this pre-defined space need not include the global optimum. We tackle this problem by proposing the filtering expansion strategy for Bayesian optimization. Our approach starts from the initial region and gradually expands the search space. Wedevelop an efficient algorithm for this strategy and derive its regret bound. These theoretical results are complemented by an extensive set of experiments on benchmark functions and tworeal-world applications which demonstrate the benefits of our proposed approach. Vu Nguyen 0001, Sunil Gupta 0001, Santu Rana, Cheng Li 0003, Svetha Venkatesh |
ICDM | 5 |
| 2017 | Stable Bayesian Optimization
Thanh Dai Nguyen, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2017 | Energy-Based Localized Anomaly Detection in Video Surveillance
Hung Vu, Tu Dinh Nguyen, Anthony Travers, Svetha Venkatesh, Dinh Q. Phung |
PAKDD (1) | 4 |
| 2017 | Estimating Support Scores of Autism Communities in Large-Scale Web Information Systems
Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
WISE (1) | 3 |
| 2017 | Effective sparse imputation of patient conditions in electronic medical records for emergency risk predictions
Budhaditya Saha, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2017 | Nonparametric discovery and analysis of learning patterns and autism subgroups from therapeutic data
Pratibha Vellanki, Thi V. Duong, Sunil Gupta 0001, Svetha Venkatesh, Dinh Q. Phung |
Knowl. Inf. Syst. | 4 |
| 2016 | Outlier Detection on Mixed-Type Data: An Energy-Based Approach
Kien Do, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
ADMA | 4 |
| 2016 | Stabilizing Linear Prediction Models Using Autoencoder
Shivapratap Gopakumar, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
ADMA | 4 |
| 2016 | Understanding Behavioral Differences Between Short and Long-Term Drinking Abstainers from Social Media
Haripriya Harikumar, Thin Nguyen, Sunil Gupta 0001, Santu Rana, M. R. Kaimal 0001, Svetha Venkatesh |
ADMA | 6 |
| 2016 | Extracting Key Challenges in Achieving Sobriety Through Shared Subspace Learning
Haripriya Harikumar, Thin Nguyen, Santu Rana, Sunil Gupta 0001, M. R. Kaimal 0001, Svetha Venkatesh |
ADMA | 6 |
| 2016 | Textual Cues for Online Depression in Community and Personal Settings
Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
ADMA | 2 |
| 2016 | Analysing the History of Autism Spectrum Disorder Using Topic ModelsabstractWe describe a novel framework for the discovery of underlying topics of a longitudinal collection of scholarly data, and the tracking of their lifetime and popularity over time. Unlike the social media or news data where the underlying topics evolve over time, the topic nuances in science result in new scientific directions to emerge. Therefore, we model the longitudinal literature data with a new approach that uses topics which remain identifiable over the course of time. Current studies either disregard the time dimension or treat it as an exchangeable covariate when they fix the topics over time or do not share the topics over epochs when they model the time naturally. We address these issues by adopting a non-parametric Bayesian approach. We assume the data is partially exchangeable and divide it into consecutive epochs. Then, by fixing the topics in a recurrent Chinese restaurant franchise, we impose a static topical structure on the corpus such that the topics are shared across epochs and the documents within epochs. We demonstrate the effectiveness of the proposed framework on a collection of medical literature related to autism spectrum disorder. We collect a large corpus of publications and carefully examine two important research issues of the domain as case studies. Moreover, we make the results of our experiment and the source code of the model, freely available to the public. This aids other researchers to analyse our results or apply the model to their data collections. Adham Beykikhoshk, Dinh Q. Phung, Ognjen Arandjelovic, Svetha Venkatesh |
DSAA | 4 |
| 2016 | Learning Multifaceted Latent Activities from Heterogeneous Mobile DataabstractInferring abstract contexts and activities from heterogeneous data is vital to context-aware ubiquitous applications but still remains one of the most challenging problems. Recent advances in Bayesian nonparametric machine learning, in particular the theory of topic models based on Hierarchical Dirichlet Process (HDP), has provided an elegant solution towards these challenges. However, limited existing methods have addressed the problem of inferring latent multifaceted activities and contexts from heterogeneous data sources such as those collected from mobile devices. In this paper, we extend the original HDP to model heterogeneous data using a richer structure of the base measure being a product-space. The proposed model, called product-space HDP (PS-HDP), naturally handles the heterogeneous data from multiple sources and identify the unknown number of latent structures in a principle way. Although this framework is generic, our current work primarily focuses on inferring (latent) threefold activities of who-when-where simultaneously, which corresponds to inducing activities from data collected for identity, location and time. We demonstrate our model on synthetic data as well as on a real-world dataset – the StudentLife dataset. We report results and provide analysis on the discovered activities and patterns to demonstrate the merit of the model. We also quantitatively evaluate the performance of PS-HDP model using standard metrics including F1-score, NMI, RI, purity, and compare them with well-known existing baseline methods. Binh T. Nguyen 0001, Vu Nguyen 0001, Nguyen Cong Thuong, Svetha Venkatesh, Mohan Kumar, Dinh Q. Phung |
DSAA | 4 |
| 2016 | One-Pass Logistic Regression for Label-Drift and Large-Scale Classification on Distributed SystemsabstractLogistic regression (LR) for classification is the workhorse in industry, where a set of predefined classes is required. The model, however, fails to work in the case where the class labels are not known in advance, a problem we term label-drift classification. Label-drift classification problem naturally occurs in many applications, especially in the context of streaming settings where the incoming data may contain samples categorized with new classes that have not been previously seen. Additionally, in the wave of big data, traditional LR methods may fail due to their expense of running time. In this paper, we introduce a novel variant of LR, namely one-pass logistic regression (OLR) to offer a principled treatment for label-drift and large-scale classifications. To handle largescale classification for big data, we further extend our OLR to a distributed setting for parallelization, termed sparkling OLR (Spark-OLR). We demonstrate the scalability of our proposed methods on large-scale datasets with more than one hundred million data points. The experimental results show that the predictive performances of our methods are comparable orbetter than those of state-of-the-art baselines whilst the executiontime is much faster at an order of magnitude. In addition, the OLR and Spark-OLR are invariant to data shuffling and have no hyperparameter to tune that significantly benefits data practitioners and overcomes the curse of big data cross-validationto select optimal hyperparameters. Vu Nguyen 0001, Tu Dinh Nguyen, Trung Le 0001, Svetha Venkatesh, Dinh Q. Phung |
ICDM | 4 |
| 2016 | Budgeted Batch Bayesian OptimizationabstractParameter settings profoundly impact the performance of machine learning algorithms and laboratory experiments. The classical trial-error methods are exponentially expensive in large parameter spaces, and Bayesian optimization (BO) offers an elegant alternative for global optimization of black box functions. In situations where the functions can be evaluated at multiple points simultaneously, batch Bayesian optimization is used. Current batch BO approaches are restrictive in fixing the number of evaluations per batch, and this can be wasteful when the number of specified evaluations is larger than the number of real maxima in the underlying acquisition function. We present the budgeted batch Bayesian optimization (B3O) for hyper-parameter tuning and experimental design - we identify the appropriate batch size for each iteration in an elegant way. In particular, we use the infinite Gaussian mixture model (IGMM) for automatically identifying the number of peaks in the underlying acquisition functions. We solve the intractability of estimating the IGMM directly from the acquisition function by formulating the batch generalized slice sampling to efficiently draw samples from the acquisition function. We perform extensive experiments for benchmark functions and two real world applications - machine learning hyper-parameter tuning and experimental design for alloy hardening. We show empirically that the proposed B3O outperforms the existing fixed batch BO approaches in finding the optimum whilst requiring a fewer number of evaluations, thus saving cost and time. Vu Nguyen 0001, Santu Rana, Sunil Gupta 0001, Cheng Li 0003, Svetha Venkatesh |
ICDM | 5 |
| 2016 | Flexible Transfer Learning Framework for Bayesian Optimisation
Tinu Theckel Joy, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh |
PAKDD (1) | 4 |
| 2016 | Toxicity Prediction in Cancer Using Multiple Instance Learning in a Multi-task Framework
Cheng Li 0003, Sunil Gupta 0001, Santu Rana, Wei Luo 0001, Svetha Venkatesh, David Ashely, Dinh Q. Phung |
PAKDD (1) | 5 |
| 2016 | Privacy Aware K-Means Clustering with High Utility
Thanh Dai Nguyen, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2016 | DeepCare: A Deep Dynamic Memory Model for Predictive Medicine
Trang Pham, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2016 | Discriminative Cues for Different Stages of Smoking Cessation in Online Community
Thin Nguyen, Ron Borland, John Yearwood, Hua-Hie Yong, Svetha Venkatesh, Dinh Q. Phung |
WISE (2) | 5 |
| 2016 | Large-Scale Stylistic Analysis of Formality in Academia and Social Media
Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
WISE (2) | 2 |
| 2016 | Graph-induced restricted Boltzmann machines for document modeling
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
Inf. Sci. | 4 |
| 2016 | Collaborative filtering via sparse Markov random fields
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
Inf. Sci. | 3 |
| 2016 | A new transfer learning framework with application to model-agnostic multi-task learning
Sunil Gupta 0001, Santu Rana, Budhaditya Saha, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 5 |
| 2016 | Data clustering using side information dependent Chinese restaurant processes
Cheng Li 0003, Santu Rana, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2016 | Multiple task transfer learning with small sample sizes
Budhaditya Saha, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2016 | Modelling human preferences for ranking and collaborative filtering: a probabilistic ordered partition approach
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 3 |
| 2015 | Overcoming Data Scarcity of Twitter: Using Tweets as Bootstrap with Application to Autism-Related Topic Content AnalysisabstractNotwithstanding recent work which has demonstrated the potential of using Twitter messages for content-specific data mining and analysis, the depth of such analysis is inherently limited by the scarcity of data imposed by the 140 character tweet limit. In this paper we describe a novel approach for targeted knowledge exploration which uses tweet content analysis as a preliminary step. This step is used to bootstrap more sophisticated data collection from directly related but much richer content sources. In particular we demonstrate that valuable information can be collected by following URLs included in tweets. We automatically extract content from the corresponding web pages and treating each web page as a document linked to the original tweet show how a temporal topic model based on a hierarchical Dirichlet process can be used to track the evolution of a complex topic structure of a Twitter community. Using autism-related tweets we demonstrate that our method is capable of capturing a much more meaningful picture of information exchange than user-chosen hashtags. Adham Beykikhoshk, Ognjen Arandjelovic, Dinh Q. Phung, Svetha Venkatesh |
ASONAM | 4 |
| 2015 | Nonparametric discovery of online mental health-related communitiesabstractPeople are increasingly using social media, especially online communities, to discuss mental health issues and seek supports. Understanding topics, interaction, sentiment and clustering structures of these communities informs important aspects of mental health. It can potentially add knowledge to the underlying cognitive dynamics, mood swings patterns, shared interests, and interaction. There has been growing research interest in analyzing online mental health communities; however sentiment analysis of these communities has been largely under-explored. This study presents an analysis of online Live Journal communities with and without mental health-related conditions including depression and autism. Latent topics for mood tags, affective words, and generic words in the content of the posts made in these communities were learned using nonparametric topic modelling. These representations were then input into a nonparametric clustering to discover meta-groups among the communities. The best performance results can be achieved on clustering communities with latent mood-based representation for such communities. The study also found significant differences in usage latent topics for mood tags and affective features between online communities with and without affective disorders. The findings reveal useful insights into hyper-group detection of online mental health-related communities. Bo Dao, Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
DSAA | 3 |
| 2015 | Exploiting feature relationships towards stable feature selectionabstractFeature selection is an important step in building predictive models for most real-world problems. One of the popular methods in feature selection is Lasso. However, it shows instability in selecting features when dealing with correlated features. In this work, we propose a new method that aims to increase the stability of Lasso by encouraging similarities between features based on their relatedness, which is captured via a feature covariance matrix. Besides modeling positive feature correlations, our method can also identify negative correlations between features. We propose a convex formulation for our model along with an alternating optimization algorithm that can learn the weights of the features as well as the relationship between them. Using both synthetic and real-world data, we show that the proposed method is more stable than Lasso and many state-of-the-art shrinkage and feature selection methods. Also, its predictive performance is comparable to other methods. Iman Kamkar, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh |
DSAA | 4 |
| 2015 | Improved risk predictions via sparse imputation of patient conditions in electronic medical recordsabstractElectronic Medical Records (EMR) are increasingly used for risk prediction. EMR analysis is complicated by missing entries. There are two reasons - the “primary reason for admission” is included in EMR, but the co-morbidities (other chronic diseases) are left uncoded, and, many zero values in the data are accurate, reflecting that a patient has not accessed medical facilities. A key challenge is to deal with the peculiarities of this data - unlike many other datasets, EMR is sparse, reflecting the fact that patients have some, but not all diseases. We propose a novel model to fill-in these missing values, and use the new representation for prediction of key hospital events. To “fill-in” missing values, we represent the feature-patient matrix as a product of two low rank factors, preserving the sparsity property in the product. Intuitively, the product regularization allows sparse imputation of patient conditions reflecting common comorbidities across patients. We develop a scalable optimization algorithm based on Block coordinate descent method to find an optimal solution. We evaluate the proposed framework on two real world EMR cohorts: Cancer (7000 admissions) and Acute Myocardial Infarction (2652 admissions). Our result shows that the AUC for 3 months admission prediction is improved significantly from (0.741 to 0.786) for Cancer data and (0.678 to 0.724) for AMI data. We also extend the proposed method to a supervised model for predicting of multiple related risk outcomes (e.g. emergency presentations and admissions in hospital over 3, 6 and 12 months period) in an integrated framework. For this model, the AUC averaged over outcomes is improved significantly from (0.768 to 0.806) for Cancer data and (0.685 to 0.748) for AMI data. Budhaditya Saha, Sunil Gupta 0001, Svetha Venkatesh |
DSAA | 3 |
| 2015 | Differentially Private Random Forest with High UtilityabstractPrivacy-preserving data mining has become an active focus of the research community in the domains where data are sensitive and personal in nature. For example, highly sensitive digital repositories of medical or financial records offer enormous values for risk prediction and decision making. However, prediction models derived from such repositories should maintain strict privacy of individuals. We propose a novel random forest algorithm under the framework of differential privacy. Unlike previous works that strictly follow differential privacy and keep the complete data distribution approximately invariant to change in one data instance, we only keep the necessary statistics (e.g. variance of the estimate) invariant. This relaxation results in significantly higher utility. To realize our approach, we propose a novel differentially private decision tree induction algorithm and use them to create an ensemble of decision trees. We also propose feasible adversary models to infer about the attribute and class label of unknown data in presence of the knowledge of all other data. Under these adversary models, we derive bounds on the maximum number of trees that are allowed in the ensemble while maintaining privacy. We focus on binary classification problem and demonstrate our approach on four real-world datasets. Compared to the existing privacy preserving approaches we achieve significantly higher utility. Santu Rana, Sunil Gupta 0001, Svetha Venkatesh |
ICDM | 3 |
| 2015 | Hierarchical Dirichlet Process for Tracking Complex Topical Structure Evolution and Its Application to Autism Research Literature
Adham Beykikhoshk, Ognjen Arandjelovic, Svetha Venkatesh, Dinh Q. Phung |
PAKDD (1) | 3 |
| 2015 | Stabilizing Sparse Cox Model Using Statistic and Semantic Structures in Electronic Medical Records
Shivapratap Gopakumar, Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 5 |
| 2015 | Collaborating Differently on Different Topics: A Multi-Relational Approach to Multi-Task Learning
Sunil Gupta 0001, Santu Rana, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (1) | 4 |
| 2015 | Learning Conditional Latent Structures from Multiple Data Sources
Viet Huynh, Dinh Q. Phung, XuanLong Nguyen, Svetha Venkatesh, Hung Hai Bui |
PAKDD (1) | 4 |
| 2015 | Fast One-Class Support Vector Machine for Novelty Detection
Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2015 | Small-Variance Asymptotics for Bayesian Nonparametric Models with Constraints
Cheng Li 0003, Santu Rana, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2015 | A Bayesian Nonparametric Approach to Multilevel Regression
Vu Nguyen 0001, Dinh Q. Phung, Svetha Venkatesh, Hung Hai Bui |
PAKDD (1) | 3 |
| 2015 | Prediciton of Emergency Events: A Multi-Task Multi-Label Learning Approach
Budhaditya Saha, Sunil Gupta 0001, Svetha Venkatesh |
PAKDD (1) | 3 |
| 2015 | What shall I share and with Whom? - A Multi-Task Learning Formulation using Multi-Faceted Task RelationshipsabstractMulti-task learning is a learning paradigm that improves the performance of “related” tasks through their joint learning. To do this each task answers the question “Which other task should I share with”? This task relatedness can be complex - a task may be related to one set of tasks based on one subset of features and to other tasks based on other subsets. Existing multi-task learning methods do not explicitly model this reality, learning a single-faceted task relationship over all the features. This degrades performance by forcing a task to become similar to other tasks even on their unrelated features. Addressing this gap, we propose a novel multi-task learning model that learns multi-faceted task relationship, allowing tasks to collaborate differentially on different feature subsets. This is achieved by simultaneously learning a low dimensional subspace for task parameters and inducing task groups over each latent subspace basis using a novel combination of L1 and pairwise L∞ norms. Further, our model can induce grouping across both positively and negatively related tasks, which helps towards exploiting knowledge from all types of related tasks. We validate our model on two synthetic and five real datasets, and show significant performance improvements over several state-of-the-art multi-task learning techniques. Thus our model effectively answers for each task: What shall I share and with whom? Sunil Gupta 0001, Santu Rana, Dinh Q. Phung, Svetha Venkatesh |
SDM | 4 |
| 2015 | Differentiating Sub-groups of Online Depression-Related Communities Using Textual Cues
Thin Nguyen, Bridianne O'Dea, Mark E. Larsen, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen |
WISE (2) | 5 |
| 2015 | Stabilized sparse ordinal regression for medical risk stratification
Truyen Tran 0001, Dinh Q. Phung, Wei Luo 0001, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2014 | Data-mining twitter and the autism spectrum disorder: A Pilot studyabstractThe autism spectrum disorder (ASD) is increasingly being recognized as a major public health issue which affects approximately 0.5-0.6% of the population. Promoting the general awareness of the disorder, increasing the engagement with the affected individuals and their carers, and understanding the success of penetration of the current clinical recommendations in the target communities, is crucial in driving research as well as policy. The aim of the present work is to investigate if Twitter, as a highly popular platform for information exchange, can be used as a data-mining source which could aid in the aforementioned challenges. Specifically, using a large data set of harvested tweets, we present a series of experiments which examine a range of linguistic and semantic aspects of messages posted by individuals interested in ASD. Our findings, the first of their nature in the published scientific literature, strongly motivate additional research on this topic and present a methodological basis for further work. Adham Beykikhoshk, Ognjen Arandjelovic, Dinh Q. Phung, Svetha Venkatesh, Terry Caelli |
ASONAM | 4 |
| 2014 | Analysis of circadian rhythms from online communities of individuals with affective disordersabstractThe circadian system regulates 24 hour rhythms in biological creatures. It impacts mood regulation. The disruptions of circadian rhythms cause destabilization in individuals with affective disorders, such as depression and bipolar disorders. Previous work has examined the role of the circadian system on effects of light interactions on mood-related systems, the effects of light manipulation on brain, the impact of chronic stress on rhythms. However, such studies have been conducted in small, preselected populations. The deluge of data is now changing the landscape of research practice. The unprecedented growth of social media data allows one to study individual behavior across large and diverse populations. In particular, individuals with affective disorders from online communities have not been examined rigorously. In this paper, we aim to use social media as a sensor to identify circadian patterns for individuals with affective disorders in online communities.We use a large scale study cohort of data collecting from online affective disorder communities. We analyze changes in hourly, daily, weekly and seasonal affect of these clinical groups in contrast with control groups of general communities. By comparing the behaviors between the clinical groups and the control groups, our findings show that individuals with affective disorders show a significant distinction in their circadian rhythms across the online activity. The results shed light on the potential of using social media for identifying diurnal individual variation in affective state, providing key indicators and risk factors for noninvasive wellbeing monitoring and prediction. Bo Dao, Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
DSAA | 3 |
| 2014 | Individualized arrhythmia detection with ECG signals from wearable devicesabstractLow cost pervasive electrocardiogram (ECG) monitors is changing how sinus arrhythmia are diagnosed among patients with mild symptoms. With the large amount of data generated from long-term monitoring, come new data science and analytical challenges. Although traditional rule-based detection algorithms still work on relatively short clinical quality ECG, they are not optimal for pervasive signals collected from wearable devices—they don't adapt to individual difference and assume accurate identification of ECG fiducial points. To overcome these short-comings of the rule-based methods, this paper introduces an arrhythmia detection approach for low quality pervasive ECG signals. To achieve the robustness needed, two techniques were applied. First, a set of ECG features with minimal reliance on fiducial point identification were selected. Next, the features were normalized using robust statistics to factors out baseline individual differences and clinically irrelevant temporal drift that is common in pervasive ECG. The proposed method was evaluated using pervasive ECG signals we collected, in combination with clinician validated ECG signals from Physiobank. Empirical evaluation confirms accuracy improvements of the proposed approach over the traditional clinical rules. Binh T. Nguyen 0001, Wei Luo 0001, Terry Caelli, Svetha Venkatesh, Dinh Q. Phung |
DSAA | 4 |
| 2014 | Intervention-Driven Predictive Framework for Modeling Healthcare Data
Santu Rana, Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (1) | 4 |
| 2014 | Keeping up with Innovation: A Predictive Framework for Modeling Healthcare Data with Evolving Clinical InterventionsabstractMedical outcomes are inexorably linked to patient illness and clinical interventions. Interventions change the course of disease, crucially determining outcome. Traditional outcome prediction models build a single classifier by augmenting interventions with disease information. Interventions, however, differentially affect prognosis, thus a single prediction rule may not suffice to capture variations. Interventions also evolve over time as more advanced interventions replace older ones. To this end, we propose a Bayesian nonparametric, supervised framework that models a set of intervention groups through a mixture distribution building a separate prediction rule for each group, and allows the mixture distribution to change with time. This is achieved by using a hierarchical Dirichlet process mixture model over the interventions. The outcome is then modeled as conditional on both the latent grouping and the disease information through a Bayesian logistic regression. Experiments on synthetic and medical cohorts for 30-day readmission prediction demonstrate the superiority of the proposed model over clinical and data mining baselines. Sunil Gupta 0001, Santu Rana, Dinh Q. Phung, Svetha Venkatesh |
SDM | 4 |
| 2014 | Effect of Mood, Social Connectivity and Age in Online Depression Community via Topic and Linguistic Analysis
Bo Dao, Thin Nguyen, Dinh Q. Phung, Svetha Venkatesh |
WISE (1) | 4 |
| 2014 | Affective, Linguistic and Topic Patterns in Online Autism Communities
Thin Nguyen, Thi V. Duong, Dinh Q. Phung, Svetha Venkatesh |
WISE (2) | 4 |
| 2014 | iPoll: Automatic Polling Using Online Search
Thin Nguyen, Dinh Q. Phung, Wei Luo 0001, Truyen Tran 0001, Svetha Venkatesh |
WISE (1) | 5 |
| 2014 | Anomaly detection in large-scale data stream networks
Duc-Son Pham 0001, Svetha Venkatesh, Mihai M. Lazarescu, Budhaditya Saha |
Data Min. Knowl. Discov. | 2 |
| 2014 | Mood sensing from social media texts and its applications
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2013 | Online Social Capital: Mood, Topical and Psycholinguistic Analysis
Thin Nguyen, Bo Dao, Dinh Q. Phung, Svetha Venkatesh, Michael Berk |
ICWSM | 4 |
| 2013 | An integrated framework for suicide risk predictionabstractSuicide is a major concern in society. Despite of great attention paid by the community with very substantive medico-legal implications, there has been no satisfying method that can reliably predict the future attempted or completed suicide. We present an integrated machine learning framework to tackle this challenge. Our proposed framework consists of a novel feature extraction scheme, an embedded feature selection process, a set of risk classifiers and finally, a risk calibration procedure. For temporal feature extraction, we cast the patient's clinical history into a temporal image to which a bank of one-side filters are applied. The responses are then partly transformed into mid-level features and then selected in l1-norm framework under the extreme value theory. A set of probabilistic ordinal risk classifiers are then applied to compute the risk probabilities and further re-rank the features. Finally, the predicted risks are calibrated. Together with our Australian partner, we perform comprehensive study on data collected for the mental health cohort, and the experiments validate that our proposed framework outperforms risk assessment instruments by medical practitioners. Truyen Tran 0001, Dinh Q. Phung, Wei Luo 0001, Richard Harvey 0002, Michael Berk, Svetha Venkatesh |
KDD | 6 |
| 2013 | Latent Patient Profile Modelling and Applications with Mixed-Variate Restricted Boltzmann Machine
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (1) | 4 |
| 2013 | Split-Merge Augmented Gibbs Sampling for Hierarchical Dirichlet Processes
Santu Rana, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 3 |
| 2013 | Clustering Patient Medical Records via Sparse Subspace Representation
Budhaditya Saha, Duc-Son Pham 0001, Dinh Q. Phung, Svetha Venkatesh |
PAKDD (2) | 4 |
| 2013 | Sparse Subspace Clustering via Group Sparse CodingabstractWe propose in this paper a novel sparse subspace clustering method that regularizes sparse subspace representation by exploiting the structural sharing between tasks and data points via group sparse coding. We derive simple, provably convergent, and computationally efficient algorithms for solving the proposed group formulations. We demonstrate the advantage of the framework on three challenging benchmark datasets ranging from medical record data to image and text clustering and show that they consistently outperforms rival methods. Duc-Son Pham 0001, Dinh Q. Phung, Budhaditya Saha, Svetha Venkatesh |
SDM | 4 |
| 2013 | Regularized nonnegative shared subspace learning
Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
Data Min. Knowl. Discov. | 4 |
| 2013 | Event extraction using behaviors of sentiment signals and burst structure in social media
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2013 | Detection of cross-channel anomalies
Duc-Son Pham 0001, Budhaditya Saha, Dinh Q. Phung, Svetha Venkatesh |
Knowl. Inf. Syst. | 4 |
| 2012 | Embedded Restricted Boltzmann Machines for fusion of mixed data types and applications in social measurements analysis
Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
FUSION | 3 |
| 2012 | Sparse Subspace Representation for Spectral Document ClusteringabstractWe present a novel method for document clustering using sparse representation of documents in conjunction with spectral clustering. An ℓ1-norm optimization formulation is posed to learn the sparse representation of each document, allowing us to characterize the affinity between documents by considering the overall information instead of traditional pair wise similarities. This document affinity is encoded through a graph on which spectral clustering is performed. The decomposition into multiple subspaces allows documents to be part of a sub-group that shares a smaller set of similar vocabulary, thus allowing for cleaner clusters. Extensive experimental evaluations on two real-world datasets from Reuters-21578 and 20Newsgroup corpora show that our proposed method consistently outperforms state-of-the-art algorithms. Significantly, the performance improvement over other methods is prominent for this datasets. Budhaditya Saha, Dinh Q. Phung, Duc-Son Pham 0001, Svetha Venkatesh |
ICDM | 4 |
| 2012 | A Sentiment-Aware Approach to Community Formation in Social Media
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
ICWSM | 4 |
| 2012 | A Bayesian Nonparametric Joint Factor Model for Learning Shared and Individual Subspaces from Multiple Data SourcesabstractJoint analysis of multiple data sources is becoming increasingly popular in transfer learning, multi-task learning and cross-domain data mining.One promising approach to model the data jointly is through learning the shared and individual factor subspaces.However, performance of this approach depends on the subspace dimensionalities and the level of sharing needs to be specified a priori.To this end, we propose a nonparametric joint factor analysis framework for modeling multiple related data sources.Our model utilizes the hierarchical beta process as a nonparametric prior to automatically infer the number of shared and individual factors.For posterior inference, we provide a Gibbs sampling scheme using auxiliary variables.The effectiveness of the proposed framework is validated through its application on two real world problemstransfer learning in text and image retrieval. Sunil Gupta 0001, Dinh Q. Phung, Svetha Venkatesh |
SDM | 3 |
| 2011 | Detection of Cross-Channel Anomalies from Multiple Data ChannelsabstractWe identify and formulate a novel problem: cross channel anomaly detection from multiple data channels. Cross channel anomalies are common amongst the individual channel anomalies, and are often portent of significant events. Using spectral approaches, we propose a two-stage detection method: anomaly detection at a single-channel level, followed by the detection of cross-channel anomalies from the amalgamation of single channel anomalies. Our mathematical analysis shows that our method is likely to reduce the false alarm rate. We demonstrate our method in two applications: document understanding with multiple text corpora, and detection of repeated anomalies in video surveillance. The experimental results consistently demonstrate the superior performance of our method compared with related state-of-art methods, including the one-class SVM and principal component pursuit. In addition, our framework can be deployed in a decentralized manner, lending itself for large scale data stream analysis. Duc-Son Pham 0001, Budhaditya Saha, Dinh Q. Phung, Svetha Venkatesh |
ICDM | 4 |
| 2011 | Towards Discovery of Influence and Personality Traits through Social Link Prediction
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
ICWSM | 4 |
| 2011 | A Bayesian Framework for Learning Shared and Individual Subspaces from Multiple Data Sources
Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
PAKDD (1) | 4 |
| 2011 | Probabilistic Models over Ordered Partitions with Applications in Document Ranking and Collaborative FilteringabstractRanking is an important task for handling a large amount of content. Ideally, training data for supervised ranking would include a complete rank of documents (or other objects such as images or videos) for a particular query. However, this is only possible for small sets of documents. In practice, one often resorts to document rating, in that a subset of documents is assigned with a small number indicating the degree of relevance. This poses a general problem of modelling and learning rank data with ties. In this paper, we propose a probabilistic generative model, that models the process as permutations over partitions. This results in super-exponential combinatorial state space with unknown numbers of partitions and unknown ordering among them. We approach the problem from the discrete choice theory, where subsets are chosen in a stagewise manner, reducing the state space per each stage significantly. Further, we show that with suitable parameterisation, we can still learn the models in linear time. We evaluate the proposed models on two application areas: (i) document ranking with the data from the recently held Yahoo! challenge, and (ii) collaborative filtering with movie data. The results demonstrate that the models are competitive against well-known rivals. Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
SDM | 3 |
| 2011 | Prediction of Age, Sentiment, and Connectivity from Social Media Text
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
WISE | 4 |
| 2010 | Nonnegative shared subspace learning and its application to social media retrievalabstractAlthough tagging has become increasingly popular in online image and video sharing systems, tags are known to be noisy, ambiguous, incomplete and subjective. These factors can seriously affect the precision of a social tag-based web retrieval system. Therefore improving the precision performance of these social tag-based web retrieval systems has become an increasingly important research topic. To this end, we propose a shared subspace learning framework to leverage a secondary source to improve retrieval performance from a primary dataset. This is achieved by learning a shared subspace between the two sources under a joint Nonnegative Matrix Factorization in which the level of subspace sharing can be explicitly controlled. We derive an efficient algorithm for learning the factorization, analyze its complexity, and provide proof of convergence. We validate the framework on image and video retrieval tasks in which tags from the LabelMe dataset are used to improve image retrieval performance from a Flickr dataset and video retrieval performance from a YouTube dataset. This has implications for how to exploit and transfer knowledge from readily available auxiliary tagging resources to improve another social web retrieval system. Our shared subspace learning framework is applicable to a range of problems where one needs to exploit the strengths existing among multiple and heterogeneous datasets. Sunil Gupta 0001, Dinh Q. Phung, Brett Adams, Truyen Tran 0001, Svetha Venkatesh |
KDD | 5 |
| 2010 | Classification and Pattern Discovery of Mood in Weblogs
Thin Nguyen, Dinh Q. Phung, Brett Adams, Truyen Tran 0001, Svetha Venkatesh |
PAKDD (2) | 5 |
| 2010 | Discovery of latent subcommunities in a blog's readershipabstractThe blogosphere has grown to be a mainstream forum of social interaction as well as a commercially attractive source of information and influence. Tools are needed to better understand how communities that adhere to individual blogs are constituted in order to facilitate new personal, socially-focused browsing paradigms, and understand how blog content is consumed, which is of interest to blog authors, big media, and search. We present a novel approach to blog subcommunity characterization by modeling individual blog readers using mixtures of an extension to the LDA family that jointly models phrases and time, Ngram Topic over Time (NTOT), and cluster with a number of similarity measures using Affinity Propagation. We experiment with two datasets: a small set of blogs whose authors provide feedback, and a set of popular, highly commented blogs, which provide indicators of algorithm scalability and interpretability without prior knowledge of a given blog. The results offer useful insight to the blog authors about their commenting community, and are observed to offer an integrated perspective on the topics of discussion and members engaged in those discussions for unfamiliar blogs. Our approach also holds promise as a component of solutions to related problems, such as online entity resolution and role discovery. Brett Adams, Dinh Q. Phung, Svetha Venkatesh |
ACM Trans. Web | 3 |
| 2009 | Effective Anomaly Detection in Sensor Networks Data StreamsabstractThis paper addresses a major challenge in data mining applications where the full information about the underlying processes, such as sensor networks or large online database, cannot be practically obtained due to physical limitations such as low bandwidth or memory, storage, or computing power. Motivated by the recent theory on direct information sampling called compressed sensing (CS), we propose a framework for detecting anomalies from these large-scale data mining applications where the full information is not practically possible to obtain. Exploiting the fact that the intrinsic dimension of the data in these applications are typically small relative to the raw dimension and the fact that compressed sensing is capable of capturing most information with few measurements, our work show that spectral methods that used for volume anomaly detection can be directly applied to the CS data with guarantee on performance. Our theoretical contributions are supported by extensive experimental results on large datasets which show satisfactory performance. Budhaditya Saha, Duc-Son Pham 0001, Mihai M. Lazarescu, Svetha Venkatesh |
ICDM | 4 |
| 2000 | Semantic data modelling and visualisation using Noetica
Stewart Greenhill, Svetha Venkatesh |
Data Knowl. Eng. | 2 |
| 1999 | Symbolic Representation and Distributed Matching Strategies for SchematicsabstractThis paper describes object-centered symbolic representation and distributed matching strategies of 3D objects in a schematic form which occur in engineering drawings and maps. The object-centered representation has a hierarchical structure and is constructed from symbolic representations of schematics. With this representation, two independent schematics representing the same object can be matched. We also consider matching strategies using distributed algorithms. The object recognition is carried out with two matching methods: (1) matching between an object model and observed data at the lowest level of the hierarchy, and (2) constraints propagation. The first is carried out with symbolic Hopfield-type neural networks and the second is achieved via hierarchical winner-takes-all algorithms. Masahiro Takatsuka, Terry Caelli, Geoff A. W. West, Svetha Venkatesh |
ICDAR | 4 |
| 1999 | Learning Other Agents' Preferences in Multi-Agent Negotiation Using the Bayesian ClassifierabstractIn multi-agent systems, most of the time, an agent does not have complete information about the preferences and decision making processes of other agents. This prevents even the cooperative agents from making coordinated choices, purely due to their ignorance of what other want. To overcome this problem, traditional coordination methods rely heavily on inter-agent communication, and thus become very inefficient when communication is costly or simply not desirable (e.g. to preserve privacy). In this paper, we propose the use of learning to complement communication in acquiring knowledge about other agents. We augment the communication-intensive negotiating agent architecture with a learning module, implemented as a Bayesian classifier. This allows our agents to incrementally update models of other agents' preferences from past negotiations with them. Based on these models, the agents can make sound predictions about others' preferences, thus reducing the need for communication in their future interactions. Hung Hai Bui, Svetha Venkatesh, Dorota H. Kieronska |
Int. J. Cooperative Inf. Syst. | 2 |
| 1999 | Constructing and navigating personalised views of the Web
Stewart Greenhill, Svetha Venkatesh |
Inf. Process. Manag. | 2 |
| 1998 | Noetica: A Tool for Semantic Data Modelling
Stewart Greenhill, Svetha Venkatesh |
Inf. Process. Manag. | 2 |