EDBT 2026 Demo / reviewers in the wild / expert
Xuan Vinh Nguyen
dblp:83/1296
· DBLP profile ↗
45ranked-venue papers
20as first author
1since 2021 · last 2021
0000-0002-7275-750XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 12 first-authorDatabases, data management, data science and information retrieval · 14 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Data mining · 79% Information retrieval · 9% Data stream processing · 7% | |
| Theoretical computer science
6 papers |
Information theory · 71% Mathematical optimization · 29% | |
| Artificial intelligence
4 papers |
Trustworthy machine learning · 63% Probabilistic and Bayesian machine learning · 37% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
clustering |
0.6 | 4 | 2016 | Adjusting for Chance Clustering Comparison Measures · J. Mach. Learn. Res. 2016 Standardized Mutual Information for Clustering Comparisons: One Step Further in Adjustment for Chance · ICML 2014 minCEntropy: A Novel Information Theoretic Approach for the Generation of Alternative Clusterings · ICDM 2010 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.5 | 1 | 2021 | High Intrinsic Dimensionality Facilitates Adversarial Attack: Theoretical Evidence · IEEE Trans. Inf. Forensics Secur. 2021 |
Data mining › clustering › clustering evaluation
clustering comparison measure |
0.4 | 2 | 2016 | Adjusting for Chance Clustering Comparison Measures · J. Mach. Learn. Res. 2016 Standardized Mutual Information for Clustering Comparisons: One Step Further in Adjustment for Chance · ICML 2014 |
Data mining › dimensionality reduction
feature selection |
0.4 | 2 | 2014 | Effective global approaches for mutual information based feature selection · KDD 2014 Reconsidering Mutual Information Based Feature Selection: A Statistical Significance View · AAAI 2014 |
Data mining › dimensionality reduction › feature selection › information-theoretic feature selection
mutual information based feature selection |
0.4 | 2 | 2014 | Effective global approaches for mutual information based feature selection · KDD 2014 Reconsidering Mutual Information Based Feature Selection: A Statistical Significance View · AAAI 2014 |
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis › tensor factorization
CP decomposition |
0.2 | 1 | 2016 | Accelerating Online CP Decompositions for Higher Order Tensors · KDD 2016 |
Data stream processing › stream mining
streaming tensor decomposition |
0.2 | 1 | 2016 | Accelerating Online CP Decompositions for Higher Order Tensors · KDD 2016 |
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization |
0.2 | 1 | 2016 | Accelerating Online CP Decompositions for Higher Order Tensors · KDD 2016 |
Information theory › information measures
generalized information measures |
0.2 | 1 | 2016 | Adjusting for Chance Clustering Comparison Measures · J. Mach. Learn. Res. 2016 |
Information theory › information measures › entropy › generalized entropy
tsallis entropy |
0.2 | 1 | 2016 | Adjusting for Chance Clustering Comparison Measures · J. Mach. Learn. Res. 2016 |
Machine learning and data management
overfitting control |
0.2 | 1 | 2014 | Reconsidering Mutual Information Based Feature Selection: A Statistical Significance View · AAAI 2014 |
Mathematical optimization
combinatorial optimization |
0.2 | 1 | 2014 | Effective global approaches for mutual information based feature selection · KDD 2014 |
Information theory › information measures
mutual information |
0.2 | 1 | 2014 | Standardized Mutual Information for Clustering Comparisons: One Step Further in Adjustment for Chance · ICML 2014 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.2 | 2 | 2012 | GlobalMIT: learning globally optimal dynamic bayesian network with the mutual information test criterion · Bioinform. 2011 Local and Global Algorithms for Learning Dynamic Bayesian Networks · ICDM 2012 |
Information retrieval
retrieval models |
0.1 | 1 | 2021 | High Intrinsic Dimensionality Facilitates Adversarial Attack: Theoretical Evidence · IEEE Trans. Inf. Forensics Secur. 2021 |
Information retrieval
similarity search |
0.1 | 1 | 2021 | High Intrinsic Dimensionality Facilitates Adversarial Attack: Theoretical Evidence · IEEE Trans. Inf. Forensics Secur. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning › graphical model learning
bayesian network learning |
0.1 | 1 | 2012 | Local and Global Algorithms for Learning Dynamic Bayesian Networks · ICDM 2012 |
Machine learning › Probabilistic and Bayesian machine learning
clustering |
0.1 | 1 | 2010 | Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance · J. Mach. Learn. Res. 2010 |
Data mining › clustering
alternative clustering |
0.1 | 1 | 2010 | minCEntropy: A Novel Information Theoretic Approach for the Generation of Alternative Clusterings · ICDM 2010 |
Data mining › clustering
information-theoretic clustering |
0.1 | 1 | 2010 | minCEntropy: A Novel Information Theoretic Approach for the Generation of Alternative Clusterings · ICDM 2010 |
Information theory
information measures |
0.1 | 1 | 2010 | Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance · J. Mach. Learn. Res. 2010 |
Data mining › clustering › clustering evaluation
clustering comparison |
0.1 | 1 | 2009 | Information theoretic measures for clusterings comparison: is a correction for chance necessary? · ICML 2009 |
Data mining › clustering
clustering evaluation |
0.1 | 1 | 2009 | Information theoretic measures for clusterings comparison: is a correction for chance necessary? · ICML 2009 |
Mathematical optimization › least squares
alternating least squares |
0.1 | 1 | 2016 | Accelerating Online CP Decompositions for Higher Order Tensors · KDD 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning |
0.0 | 1 | 2011 | GlobalMIT: learning globally optimal dynamic bayesian network with the mutual information test criterion · Bioinform. 2011 |
Methods — techniques the papers use, named apart from their topics
theoretical analysis · 1.0k-nearest neighbor ranking · 1.0mutual information test · 0.5shannon information theory · 0.5incremental tracking · 0.5alternating least squares · 0.5statistical hypothesis testing · 0.4mutual information · 0.4local and global optimization · 0.4hypergeometric model · 0.4markov blanket · 0.3MDL scoring · 0.3pair-counting · 0.2pair counting · 0.2dynamic bayesian network · 0.2variance derivation · 0.2quadratic programming · 0.2normalization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | High Intrinsic Dimensionality Facilitates Adversarial Attack: Theoretical EvidenceabstractMachine learning systems are vulnerable to adversarial attack. By applying to the input object a small, carefully-designed perturbation, a classifier can be tricked into making an incorrect prediction. This phenomenon has drawn wide interest, with many attempts made to explain it. However, a complete understanding is yet to emerge. In this paper we adopt a slightly different perspective, still relevant to classification. We consider retrieval, where the output is a set of objects most similar to a user-supplied query object, corresponding to the set of k-nearest neighbors. We investigate the effect of adversarial perturbation on the ranking of objects with respect to a query. Through theoretical analysis, supported by experiments, we demonstrate that as the intrinsic dimensionality of the data domain rises, the amount of perturbation required to subvert neighborhood rankings diminishes, and the vulnerability to adversarial attack rises. We examine two modes of perturbation of the query: either `closer' to the target point, or `farther' from it. We also consider two perspectives: `query-centric', examining the effect of perturbation on the query's own neighborhood ranking, and `target-centric', considering the ranking of the query point in the target's neighborhood set. All four cases correspond to practical scenarios involving classification and retrieval. Laurent Amsaleg, James Bailey 0001, Amélie Barbe, Sarah M. Erfani, Teddy Furon, Michael E. Houle, Milos Radovanovic 0001, Xuan Vinh Nguyen |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2018 | The randomized information coefficient: assessing dependencies in noisy data
Simone Romano 0003, Xuan Vinh Nguyen, Karin Verspoor, James Bailey 0001 |
Mach. Learn. | 2 |
| 2017 | Topology-regularized universal vector autoregression for traffic forecasting in large urban areas
Florin Schimbinschi, Luís Moreira-Matias, Xuan Vinh Nguyen, James Bailey 0001 |
Expert Syst. Appl. | 3 |
| 2017 | rFILTA: relevant and nonredundant view discovery from collections of clusterings via filtering and ranking
Yang Lei 0003, Xuan Vinh Nguyen, Jeffrey Chan, James Bailey 0001 |
Knowl. Inf. Syst. | 2 |
| 2017 | Ground truth bias in external cluster validity indices
Yang Lei 0003, James C. Bezdek, Simone Romano 0003, Xuan Vinh Nguyen, Jeffrey Chan, James Bailey 0001 |
Pattern Recognit. | 4 |
| 2017 | Extending Information-Theoretic Validity Indices for Fuzzy ClusteringabstractPreviously, eight popular information-theoretic-based cluster validity indices have been generalized and tested for probabilistic partitions built by the expectation-maximization (EM) algorithm for the Gaussian mixture model. However, the analysis was limited to probabilistic clusters, and there were limited explanations for differences in the performance of the indices. In this paper, we extend the tests to partitions found by fuzzy c-means (FCM) and provide further explanations and insights about the performance of these indices. Of the eight generalized indices, we advocate a normalized version of the soft mutual information cluster validity index (NMI sM) as the best overall choice, as it outperforms the other seven indices for both FCM and EM according to our tests on synthetic and real data. The superiority of NMIsM is most pronounced for datasets with overlapped and/or varying-sized clusters. Finally, we provide a theoretical analysis, which helps explain the superior performance of NMIsM compared with the other three normalizations of soft mutual information. Yang Lei 0003, James C. Bezdek, Jeffrey Chan, Xuan Vinh Nguyen, Simone Romano 0003, James Bailey 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2016 | Training robust models using Random ProjectionabstractRegularization plays an important role in machine learning systems. We propose a novel methodology for model regularization using random projection. We demonstrate the technique on neural networks, since such models usually comprise a very large number of parameters, calling for strong regularizers. It has been shown recently that neural networks are sensitive to two kinds of samples: (i) adversarial samples, which are generated by imperceptible perturbations of previously correctly-classified samples—yet the network will misclassify them; and (ii) fooling samples, which are completely unrecognizable, yet the network will classify them with extremely high confidence. In this paper, we show how robust neural networks can be trained using random projection. We show that while random projection acts as a strong regularizer, boosting model accuracy similar to other regularizers, such as weight decay and dropout, it is far more robust to adversarial noise and fooling samples. We further show that random projection also helps to improve the robustness of traditional classifiers, such as Random Forrest and Gradient Boosting Machines. Xuan Vinh Nguyen, Sarah M. Erfani, Sakrapee Paisitkriangkrai, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao |
ICPR | 1 |
| 2016 | Accelerating Online CP Decompositions for Higher Order TensorsabstractTensors are a natural representation for multidimensional data. In recent years, CANDECOMP/PARAFAC (CP) decomposition, one of the most popular tools for analyzing multi-way data, has been extensively studied and widely applied. However, today's datasets are often dynamically changing over time. Tracking the CP decomposition for such dynamic tensors is a crucial but challenging task, due to the large scale of the tensor and the velocity of new data arriving. Traditional techniques, such as Alternating Least Squares (ALS), cannot be directly applied to this problem because of their poor scalability in terms of time and memory. Additionally, existing online approaches have only partially addressed this problem and can only be deployed on third-order tensors. To fill this gap, we propose an efficient online algorithm that can incrementally track the CP decompositions of dynamic tensors with an arbitrary number of dimensions. In terms of effectiveness, our algorithm demonstrates comparable results with the most accurate algorithm, ALS, whilst being computationally much more efficient. Specifically, on small and moderate datasets, our approach is tens to hundreds of times faster than ALS, while for large-scale datasets, the speedup can be more than 3,000 times. Compared to other state-of-the-art online approaches, our method shows not only significantly better decomposition quality, but also better performance in terms of stability, efficiency and scalability. Shuo Zhou 0001, Xuan Vinh Nguyen, James Bailey 0001, Yunzhe Jia, Ian Davidson |
KDD | 2 |
| 2016 | A Framework to Adjust Dependency Measure Estimates for ChanceabstractEstimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are dependent to the target class are selected using Gini gain in random forests. Nonetheless, because dependency measures are estimated on finite samples, the interpretability of their quantification and the accuracy when ranking dependencies become challenging. Dependency estimates are not equal to 0 when variables are independent, cannot be compared if computed on different sample size, and they are inflated by chance on variables with more categories. In this paper, we propose a framework to adjust dependency measure estimates on finite samples. Our adjustments, which are simple and applicable to any dependency measure, are helpful in improving interpretability when quantifying dependency and in improving accuracy on the task of ranking dependencies. In particular, we demonstrate that our approach enhances the interpretability of MIC when used as a proxy for the amount of noise between variables, and to gain accuracy when ranking variables during the splitting procedure in random forests. Simone Romano 0003, Xuan Vinh Nguyen, James Bailey 0001, Karin Verspoor |
SDM | 2 |
| 2016 | Discovering outlying aspects in large datasets
Xuan Vinh Nguyen, Jeffrey Chan, Simone Romano 0003, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jian Pei 0001 |
Data Min. Knowl. Discov. | 1 |
| 2016 | Adjusting for Chance Clustering Comparison MeasuresabstractAdjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information theory are very popular in the clustering community. Nonetheless it is an open problem as to what are the best application scenarios for each measure and guidelines in the literature for their usage are sparse, with the result that users often resort to using both. Generalized Information Theoretic (IT) measures based on the Tsallis entropy have been shown to link pair- counting and Shannon IT measures. In this paper, we aim to bridge the gap between adjustment of measures based on pair- counting and measures based on information theory. We solve the key technical challenge of analytically computing the expected value and variance of generalized IT measures. This allows us to propose adjustments of generalized IT measures, which reduce to well known adjusted clustering comparison measures as special cases. Using the theory of generalized IT measures, we are able to propose the following guidelines for using ARI and AMI as external validation indices: ARI should be used when the reference clustering has large equal sized clusters; AMI should be used when the reference clustering is unbalanced and there exist small clusters. Simone Romano 0003, Xuan Vinh Nguyen, James Bailey 0001, Karin Verspoor |
J. Mach. Learn. Res. | 2 |
| 2016 | Efficient discovery of contrast subspaces for object explanation and characterization
Lei Duan, Guanting Tang, Jian Pei 0001, James Bailey 0001, Guozhu Dong, Xuan Vinh Nguyen, Akiko Campbell, Changjie Tang |
Knowl. Inf. Syst. | 6 |
| 2016 | Can high-order dependencies improve mutual information based feature selection?
Xuan Vinh Nguyen, Shuo Zhou 0001, Jeffrey Chan, James Bailey 0001 |
Pattern Recognit. | 1 |
| 2015 | Traffic forecasting in complex urban networks: Leveraging big data and machine learningabstractAccurate network-wide real time traffic forecasting is essential for next generation smart cities. In this context, we study a novel and complex traffic data set and explore the potential to apply big data and machine learning analysis. We evaluate several hypotheses and find that the availability of big data is able to facilitate more accurate predictions. Furthermore, we find that spatial aspects have more influence than temporal ones and that careful choice of thresholding parameters is crucial for high performance classification. Florin Schimbinschi, Xuan Vinh Nguyen, James Bailey 0001, Christopher Leckie, Hai Le Vu 0001, Kotagiri Ramamohanarao |
IEEE BigData | 2 |
| 2015 | Scalable Outlying-Inlying Aspects Discovery via Feature Ranking
Xuan Vinh Nguyen, Jeffrey Chan, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jian Pei 0001 |
PAKDD (2) | 1 |
| 2014 | Reconsidering Mutual Information Based Feature Selection: A Statistical Significance ViewabstractMutual information (MI) based approaches are a popular feature selection paradigm. Although the stated goal of MI-based feature selection is to identify a subset of features that share the highest mutual information with the class variable, most current MI-based techniques are greedy methods that make use of low dimensional MI quantities. The reason for using low dimensional approximation has been mostly attributed to the difficulty associated with estimating the high dimensional MI from limited samples. In this paper, we argue a different viewpoint that, given a very large amount of data, the high dimensional MI objective is still problematic to be employed as a meaningful optimization criterion, due to its overfitting nature: the MI almost always increases as more features are added, thus leading to a trivial solution which includes all features. We propose a novel approach to the MI-based feature selection problem, in which the overfitting phenomenon is controlled rigourously by means of a statistical test. We develop local and global optimization algorithms for this new feature selection model, and demonstrate its effectiveness in the applications of explaining variables and objects. Xuan Vinh Nguyen, Jeffrey Chan, James Bailey 0001 |
AAAI | 1 |
| 2014 | Generalized information theoretic cluster validity indices for soft clusteringsabstractThere have been a large number of external validity indices proposed for cluster validity. One such class of cluster comparison indices is the information theoretic measures, due to their strong mathematical foundation and their ability to detect non-linear relationships. However, they are devised for evaluating crisp (hard) partitions. In this paper, we generalize eight information theoretic crisp indices to soft clusterings, so that they can be used with partitions of any type (i.e., crisp or soft, with soft including fuzzy, probabilistic and possibilistic cases). We present experimental results to demonstrate the effectiveness of the generalized information theoretic indices. Yang Lei 0003, James C. Bezdek, Jeffrey Chan, Xuan Vinh Nguyen, Simone Romano 0003, James Bailey 0001 |
CIDM | 4 |
| 2014 | Standardized Mutual Information for Clustering Comparisons: One Step Further in Adjustment for ChanceabstractMutual information is a very popular measure for comparing clusterings. Previous work has shown that it is beneficial to make an adjustment for chance to this measure, by subtracting an expected value and normalizing via an upper bound. This yields the constant baseline property that enhances intuitiveness. In this paper, we argue that a further type of statistical adjustment for the mutual information is also beneficial - an adjustment to correct selection bias. This type of adjustment is useful when carrying out many clustering comparisons, to select one or more preferred clusterings. It reduces the tendency for the mutual information to choose clustering solutions i) with more clusters, or ii) induced on fewer data points, when compared to a reference one. We term our new adjusted measure the *standardized mutual information*. It requires computation of the variance of mutual information under a hypergeometric model of randomness, which is technically challenging. We derive an analytical formula for this variance and analyze its complexity. We then experimentally assess how our new measure can address selection bias and also increase interpretability. We recommend using the standardized mutual information when making multiple clustering comparisons in situations where the number of records is small compared to the number of clusters considered. Simone Romano 0003, James Bailey 0001, Xuan Vinh Nguyen, Karin Verspoor |
ICML | 3 |
| 2014 | Effective global approaches for mutual information based feature selectionabstractMost current mutual information (MI) based feature selection techniques are greedy in nature thus are prone to sub-optimal decisions. Potential performance improvements could be gained by systematically posing MI-based feature selection as a global optimization problem. A rare attempt at providing a global solution for the MI-based feature selection is the recently proposed Quadratic Programming Feature Selection (QPFS) approach. We point out that the QPFS formulation faces several non-trivial issues, in particular, how to properly treat feature `self-redundancy' while ensuring the convexity of the objective function. In this paper, we take a systematic approach to the problem of global MI-based feature selection. We show how the resulting NP-hard global optimization problem could be efficiently approximately solved via spectral relaxation and semi-definite programming techniques. We experimentally demonstrate the efficiency and effectiveness of these novel feature selection frameworks. Xuan Vinh Nguyen, Jeffrey Chan, Simone Romano 0003, James Bailey 0001 |
KDD | 1 |
| 2014 | Structure-Aware Distance Measures for Comparing Clusterings in Graphs
Jeffrey Chan, Xuan Vinh Nguyen, Wei Liu 0007, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jian Pei 0001 |
PAKDD (1) | 2 |
| 2014 | FILTA: Better View Discovery from Collections of Clusterings via Filtering
Yang Lei 0003, Xuan Vinh Nguyen, Jeffrey Chan, James Bailey 0001 |
ECML/PKDD (2) | 2 |
| 2013 | Inferring large scale genetic networks with S-system modelabstractGene regulatory network (GRN) reconstruction from high-throughput microarray data is an important problem in systems biology. The S-System model, a differential equation based approach, is among the mainstream approaches for modeling GRNs. It has the ability to represent GRNs accurately with precise regulatory weights. However, the current applications of S-System are limited to small and medium scale network, as inferring large network requires inhibitive computational cost. In this paper, we propose a novel S-System based framework to reconstruct biologically relevant GRNs by exploiting their special topological structure. In GRNs, the complex interactions occurring amongst transcription factors (TFs) and target genes (TGs) are unidirectional, i.e., TFs to TGs, and the vice-versa is biologically irrelevant. In addition, TFs can regulate themselves while only self-regulations may exist for TGs. As such, we decompose GRN into two sub-networks representing TF-TF and TF-TG interactions. We learn the sub-networks separately by adapting the traditional S-System model, and combining the solutions to get the entire network. Our experimental studies indicate that the proposed approach can scale up to larger networks, not achievable with other current S-System based approaches, yet with higher accuracy. Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
GECCO | 3 |
| 2013 | mDBN: motif based learning of gene regulatory networks using dynamic bayesian networksabstractSolutions for deriving the most consistent Bayesian gene regulatory network model from given data sets using evolutionary algorithms typically only result in locally optimal solutions. Further, due to genetic drift, merely increasing the size of the population does not overcome this limitation. In this paper, we propose a two-stage genetic algorithm that systematically searches the whole search space using frequent subgraph mining techniques. The approach finds representative patterns present in different local optimal solutions in the first stage and then combines these frequent subgraphs (motifs) in the second stage to converge to the global optima. We apply the algorithm to both synthetic and real life networks of yeast and E.coli and show the effectiveness of our approach. Nizamul Morshed, Madhu Chetty, Xuan Vinh Nguyen, Terry Caelli |
GECCO | 3 |
| 2013 | On the Analysis of Time-Delayed Interactions in Genetic Network Using S-System Model
Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (2) | 3 |
| 2013 | Reverse Engineering Genetic Networks with Time-Delayed S-System Model and Pearson Correlation Coefficient
Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (2) | 3 |
| 2013 | Incorporating time-delays in S-System model for reverse engineering genetic networksabstractBACKGROUND: In any gene regulatory network (GRN), the complex interactions occurring amongst transcription factors and target genes can be either instantaneous or time-delayed. However, many existing modeling approaches currently applied for inferring GRNs are unable to represent both these interactions simultaneously. As a result, all these approaches cannot detect important interactions of the other type. S-System model, a differential equation based approach which has been increasingly applied for modeling GRNs, also suffers from this limitation. In fact, all S-System based existing modeling approaches have been designed to capture only instantaneous interactions, and are unable to infer time-delayed interactions. RESULTS: In this paper, we propose a novel Time-Delayed S-System (TDSS) model which uses a set of delay differential equations to represent the system dynamics. The ability to incorporate time-delay parameters in the proposed S-System model enables simultaneous modeling of both instantaneous and time-delayed interactions. Furthermore, the delay parameters are not limited to just positive integer values (corresponding to time stamps in the data), but can also take fractional values. Moreover, we also propose a new criterion for model evaluation exploiting the sparse and scale-free nature of GRNs to effectively narrow down the search space, which not only reduces the computation time significantly but also improves model accuracy. The evaluation criterion systematically adapts the max-min in-degrees and also systematically balances the effect of network accuracy and complexity during optimization. CONCLUSION: The four well-known performance measures applied to the experimental studies on synthetic networks with various time-delayed regulations clearly demonstrate that the proposed method can capture both instantaneous and delayed interactions correctly with high precision. The experiments carried out on two well-known real-life networks, namely IRMA and SOS DNA repair network in Escherichia coli show a significant improvement compared with other state-of-the-art approaches for GRN modeling. Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
BMC Bioinform. | 3 |
| 2013 | A model of the circadian clock in the cyanobacterium Cyanothece sp. ATCC 51142abstractBACKGROUND: The over consumption of fossil fuels has led to growing concerns over climate change and global warming. Increasing research activities have been carried out towards alternative viable biofuel sources. Of several different biofuel platforms, cyanobacteria possess great potential, for their ability to accumulate biomass tens of times faster than traditional oilseed crops. The cyanobacterium Cyanothece sp. ATCC 51142 has recently attracted lots of research interest as a model organism for such research. Cyanothece can perform efficiently both photosynthesis and nitrogen fixation within the same cell, and has been recently shown to produce biohydrogen--a byproduct of nitrogen fixation--at very high rates of several folds higher than previously described hydrogen-producing photosynthetic microbes. Since the key enzyme for nitrogen fixation is very sensitive to oxygen produced by photosynthesis, Cyanothece employs a sophisticated temporal separation scheme, where nitrogen fixation occurs at night and photosynthesis at day. At the core of this temporal separation scheme is a robust clocking mechanism, which so far has not been thoroughly studied. Understanding how this circadian clock interacts with and harmonizes global transcription of key cellular processes is one of the keys to realize the inherent potential of this organism. RESULTS: In this paper, we employ several state of the art bioinformatics techniques for studying the core circadian clock in Cyanothece sp. ATCC 51142, and its interactions with other key cellular processes. We employ comparative genomics techniques to map the circadian clock genes and genetic interactions from another cyanobacterial species, namely Synechococcus elongatus PCC 7942, of which the circadian clock has been much more thoroughly investigated. Using time series gene expression data for Cyanothece, we employ gene regulatory network reconstruction techniques to learn this network de novo, and compare the reconstructed network against the interactions currently reported in the literature. Next, we build a computational model of the interactions between the core clock and other cellular processes, and show how this model can predict the behaviour of the system under changing environmental conditions. The constructed models significantly advance our understanding of the Cyanothece circadian clock functional mechanisms. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Sandeep Gaudana, Pramod P. Wangikar |
BMC Bioinform. | 1 |
| 2013 | Comments on supervised feature selection by clustering using conditional mutual information-based distances
Xuan Vinh Nguyen, James Bailey 0001 |
Pattern Recognit. | 1 |
| 2012 | Adaptive regulatory genes cardinality for reconstructing genetic networksabstractWith the advent of microarray technology, researchers are able to determine cellular dynamics for thousands of genes simultaneously, thereby enabling reverse engineering of the gene regulatory network (GRN) from high-throughput time-series gene expression data. Amongst the various currently available models for inferring GRN, the S-System formalism is often considered as an excellent compromise between accuracy and mathematical tractability. In this paper, a novel approach for inferring GRN based on the decoupled S-System model, incorporating the new concept of adaptive regulatory genes cardinality, is proposed. Parameter learning for the S-System is carried out in an evolving manner using a versatile and robust Trigonometric Evolutionary Algorithm. The applicability and efficiency of the proposed method is studied using a well-known and widely studied synthetic network with various levels of noise, and excellent performance observed. Further, investigations of a 5 gene in-vivo synthetic biological network of Saccharomyces cerevisiae called IRMA, has succeeded in detecting higher number of correct regulations compared to other approaches reported earlier. Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Local and Global Algorithms for Learning Dynamic Bayesian NetworksabstractLearning optimal Bayesian networks (BN) from data is NP-hard in general. Nevertheless, certain BN classes with additional topological constraints, such as the dynamic BN (DBN) models, widely applied in specific fields such as systems biology, can be efficiently learned in polynomial time. Such algorithms have been developed for the Bayesian-Dirichlet (BD), Minimum Description Length (MDL), and Mutual Information Test (MIT) scoring metrics. The BD-based algorithm admits a large polynomial bound, hence it is impractical for even modestly sized networks. The MDL-and MIT-based algorithms admit much smaller bounds, but require a very restrictive assumption that all variables have the same cardinality, thus significantly limiting their applicability. In this paper, we first propose an improvement to the MDL-and MIT-based algorithms, dropping the equicardinality constraint, thus significantly enhancing their generality. We also explore local Markov blanket based algorithms for constructing BN in the context of DBN, and show an interesting result: under the faithfulness assumption, the mutual information test based local Markov blanket algorithms yield the same network as learned by the global optimization MIT-based algorithm. Experimental validation on small and large scale genetic networks demonstrates the effectiveness of our proposed approaches. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICDM | 1 |
| 2012 | On the Reconstruction of Genetic Network from Partial Microarray Data
Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (1) | 3 |
| 2012 | FusGP: Bayesian Co-learning of Gene Regulatory Networks and Protein Interaction Networks
Nizamul Morshed, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (5) | 3 |
| 2012 | Data Discretization for Dynamic Bayesian Network Based Modeling of Genetic Networks
Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICONIP (2) | 1 |
| 2012 | Gene regulatory network modeling via global optimization of high-order dynamic Bayesian networkabstractBACKGROUND: Dynamic Bayesian network (DBN) is among the mainstream approaches for modeling various biological networks, including the gene regulatory network (GRN). Most current methods for learning DBN employ either local search such as hill-climbing, or a meta stochastic global optimization framework such as genetic algorithm or simulated annealing, which are only able to locate sub-optimal solutions. Further, current DBN applications have essentially been limited to small sized networks. RESULTS: To overcome the above difficulties, we introduce here a deterministic global optimization based DBN approach for reverse engineering genetic networks from time course gene expression data. For such DBN models that consist only of inter time slice arcs, we show that there exists a polynomial time algorithm for learning the globally optimal network structure. The proposed approach, named GlobalMIT+, employs the recently proposed information theoretic scoring metric named mutual information test (MIT). GlobalMIT+ is able to learn high-order time delayed genetic interactions, which are common to most biological systems. Evaluation of the approach using both synthetic and real data sets, including a 733 cyanobacterial gene expression data set, shows significantly improved performance over other techniques. CONCLUSIONS: Our studies demonstrate that deterministic global optimization approaches can infer large scale genetic networks. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
BMC Bioinform. | 1 |
| 2011 | Simultaneous Learning of Instantaneous and Time-Delayed Genetic Interactions Using Novel Information Theoretic Scoring Technique
Nizamul Morshed, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (2) | 3 |
| 2011 | Dynamic Bayesian Network Modeling of Cyanobacterial Biological Processes via Gene Clustering
Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICONIP (1) | 1 |
| 2011 | Polynomial Time Algorithm for Learning Globally Optimal Dynamic Bayesian Network
Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICONIP (3) | 1 |
| 2011 | GlobalMIT: learning globally optimal dynamic bayesian network with the mutual information test criterionabstractMOTIVATION: Dynamic Bayesian networks (DBN) are widely applied in modeling various biological networks including the gene regulatory network (GRN). Due to the NP-hard nature of learning static Bayesian network structure, most methods for learning DBN also employ either local search such as hill climbing, or a meta stochastic global optimization framework such as genetic algorithm or simulated annealing. RESULTS: This article presents GlobalMIT, a toolbox for learning the globally optimal DBN structure from gene expression data. We propose using a recently introduced information theoretic-based scoring metric named mutual information test (MIT). With MIT, the task of learning the globally optimal DBN is efficiently achieved in polynomial time. AVAILABILITY: The toolbox, implemented in Matlab and C++, is available at http://code.google.com/p/globalmit. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
Bioinform. | 1 |
| 2010 | minCEntropy: A Novel Information Theoretic Approach for the Generation of Alternative ClusteringsabstractTraditional clustering has focused on creating a single good clustering solution, while modern, high dimensional data can often be interpreted, and hence clustered, in different ways. Alternative clustering aims at creating multiple clustering solutions that are both of high quality and distinctive from each other. Methods for alternative clustering can be divided into objective-function-oriented and data-transformation-oriented approaches. This paper presents a novel information theoretic-based, objective-function-oriented approach to generate alternative clusterings, in either an unsupervised or semi-supervised manner. We employ the conditional entropy measure for quantifying both clustering quality and distinctiveness, resulting in an analytically consistent combined criterion. Our approach employs a computationally efficient nonparametric entropy estimator, which does not impose any assumption on the probability distributions. We propose a partitional clustering algorithm, named minCEntropy, to concurrently optimize both clustering quality and distinctiveness. minCEntropy requires setting only some rather intuitive parameters, and performs competitively with existing methods for alternative clustering. Xuan Vinh Nguyen, Julien Epps |
ICDM | 1 |
| 2010 | A Set Correlation Model for Partitional Clustering
Xuan Vinh Nguyen, Michael E. Houle |
PAKDD (1) | 1 |
| 2010 | Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
Xuan Vinh Nguyen, Julien Epps, James Bailey 0001 |
J. Mach. Learn. Res. | 1 |
| 2009 | A Novel Approach for Automatic Number of Clusters Detection in Microarray Data Based on Consensus ClusteringabstractEstimating the true number of clusters in a data set is one of the major challenges in cluster analysis. Yet in certain domains,knowing the true number of clusters is of high importance. For example, in medical research, detecting the true number of groups and sub-groups of cancer would be of utmost importance for their effective treatment. In this paper we propose a novel method to estimate the number of clusters in a micro array data set based on the consensus clustering approach. Although the main objective of consensus clustering is to discover a robust and high quality cluster structure in a data set, closer inspection of the set of clusterings obtained can often give valuable information about the appropriate number of clusters present. More specifically, the set off clusterings obtained when the specified number of clusters coincides with the true number of clusters tends to be less diverse.To quantify this diversity we develop a novel index, namely the Consensus Index (CI), which is built upon a suitable clustering similarity measure such as the well known Adjusted Rand Index (ARI)or our recently developed, information theoretic based index, namely the Adjusted Mutual Information (AMI). Our experiments on both synthetic and real microarray data sets indicate that the CI is a useful indicator for determining the appropriate number of clusters. Xuan Vinh Nguyen, Julien Epps |
BIBE | 1 |
| 2009 | Information theoretic measures for clusterings comparison: is a correction for chance necessary?abstractInformation theoretic based measures form a fundamental class of similarity measures for comparing clusterings, beside the class of pair-counting based and set-matching based measures. In this paper, we discuss the necessity of correction for chance for information theoretic based measures for clusterings comparison. We observe that the baseline for such measures, i.e. average value between random partitions of a data set, does not take on a constant value, and tends to have larger variation when the ratio between the number of data points and the number of clusters is small. This effect is similar in some other non-information theoretic based measures such as the well-known Rand Index. Assuming a hypergeometric model of randomness, we derive the analytical formula for the expected mutual information value between a pair of clusterings, and then propose the adjusted version for several popular information theoretic based measures. Some examples are given to demonstrate the need and usefulness of the adjusted measures. Xuan Vinh Nguyen, Julien Epps, James Bailey 0001 |
ICML | 1 |
| 2008 | An information theoretic divergence for microarray data clusteringabstractThe Kullback-Leibler (KL) divergence used in conjunction with the unsupervised Self-Organizing Map (SOM) algorithm has been previously shown to be effective for the gene clustering problem: the patterns of the gene clusters obtained were found to be superior to those obtained by the hierarchical clustering algorithm using the uncentered Pearson correlation measure. Motivated by this initial finding, in this research we study the effectiveness of the KL-divergence in a more general setting where the data points are not necessarily projected to the unit simplex but to a parallel simplex in the positive orthant. Two novel hard and soft clustering algorithms based on the so-called generalized KL-divergence are proposed. We tested the algorithms on both gene and sample clustering problems. Experimental results on real microarray datasets with known class labels (for genes or samples) show that the generalized KL-divergence based algorithms produce comparable or better results to those obtained by similar algorithms based on popular distance measures for microarray data clustering, such as the squared Euclidean distance and the Pearson correlation. Two validation indices, namely the Adjusted Rand Index and the newly developed Variation of Information metric, have been used to validate the results. Xuan Vinh Nguyen |
BIBE | 1 |
| 2008 | Normalized EM algorithm for tumor clustering using gene expression dataabstractMost of the proposed clustering approaches are heuristic in nature. As a result, it is difficult to interpret the obtained clustering outcomes from a statistical standpoint. Mixture model-based clustering has received much attention from the gene expression community due to its sound statistical background and its flexibility in data modeling. However, current clustering algorithms following the model-based framework suffer from two serious drawbacks. First, the performance of these algorithms critically depends on the starting values for their iterative clustering procedures. And second, they are not capable of working directly with very high dimensional data sets whose dimension might be up to thousands. We propose a novel normalized Expectation-Maximization (EM) algorithm to tackle the two challenges. The normalized EM is stable even with random initializations for its EM iterative procedure. Its stability is demonstrated through the performance comparison with other related clustering algorithms such as the unnormalized EM (The conventional EM algorithm for Gaussian mixture model-based clustering) and spherical k-means. Furthermore, the normalized EM is the first mixture model-based clustering algorithm that is shown to be stable when working directly with very high dimensional microarray data sets in the sample clustering problem, where the number of genes is much larger than the number of samples. Besides, an interesting property of the convergence speed of the normalized EM with respect to the squared radius of the hypersphere in its corresponding statistical model is uncovered. Xuan Vinh Nguyen |
BIBE | 2 |