Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hugh A. Chipman

dblp:74/645 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 90% Kernel, tree and ensemble methods · 10%
Databases, data mining, and information retrieval
1 paper
Data mining · 87% Web and social media mining · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixed membership models
0.112010
Mixed-Membership Stochastic Block-Models for Transactional Networks · ICDM 2010
Data mining › structured data mining › graph mining
community detection
0.112010
Mixed-Membership Stochastic Block-Models for Transactional Networks · ICDM 2010
Data mining › structured data mining
graph mining
0.112010
Mixed-Membership Stochastic Block-Models for Transactional Networks · ICDM 2010
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection
0.112009
A Bayesian segmentation approach to ascertain copy number variations at the population level · Bioinform. 2009
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
bayesian additive regression trees
0.112006
Bayesian Ensemble Learning · NIPS 2006
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian prediction
bayesian ensemble methods
0.112006
Bayesian Ensemble Learning · NIPS 2006
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.112006
Bayesian Ensemble Learning · NIPS 2006
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.112006
Bayesian Ensemble Learning · NIPS 2006
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.012006
Bayesian Ensemble Learning · NIPS 2006
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.012006
Bayesian Ensemble Learning · NIPS 2006

Methods — techniques the papers use, named apart from their topics

variational EM · 0.2stochastic block model · 0.2markov chain monte carlo · 0.1bayesian segmentation · 0.1bayes factor · 0.1weak learners · 0.1sum-of-trees model · 0.1backfitting MCMC · 0.1
YearPublicationVenuePosition
2014 BiomeNet: A Bayesian Model for Inference of Metabolic Divergence among Microbial Communities
abstract
Metagenomics yields enormous numbers of microbial sequences that can be assigned a metabolic function. Using such data to infer community-level metabolic divergence is hindered by the lack of a suitable statistical framework. Here, we describe a novel hierarchical Bayesian model, called BiomeNet (Bayesian inference of metabolic networks), for inferring differential prevalence of metabolic subnetworks among microbial communities. To infer the structure of community-level metabolic interactions, BiomeNet applies a mixed-membership modelling framework to enzyme abundance information. The basic idea is that the mixture components of the model (metabolic reactions, subnetworks, and networks) are shared across all groups (microbiome samples), but the mixture proportions vary from group to group. Through this framework, the model can capture nested structures within the data. BiomeNet is unique in modeling each metagenome sample as a mixture of complex metabolic systems (metabosystems). The metabosystems are composed of mixtures of tightly connected metabolic subnetworks. BiomeNet differs from other unsupervised methods by allowing researchers to discriminate groups of samples through the metabolic patterns it discovers in the data, and by providing a framework for interpreting them. We describe a collapsed Gibbs sampler for inference of the mixture weights under BiomeNet, and we use simulation to validate the inference algorithm. Application of BiomeNet to human gut metagenomes revealed a metabosystem with greater prevalence among inflammatory bowel disease (IBD) patients. Based on the discriminatory subnetworks for this metabosystem, we inferred that the community is likely to be closely associated with the human gut epithelium, resistant to dietary interventions, and interfere with human uptake of an antioxidant connected to IBD. Because this metabosystem has a greater capacity to exploit host-associated glycans, we speculate that IBD-associated communities might arise from opportunist growth of bacteria that can circumvent the host's nutrient-based mechanism for bacterial partner selection.
Mahdi Shafiei, Katherine A. Dunn, Hugh A. Chipman, Joseph P. Bielawski
PLoS Comput. Biol.3
2010 Mixed-Membership Stochastic Block-Models for Transactional Networks
abstract
Transactional network data can be thought of as a list of one-to-many communications (e.g., email) between nodes in a social network. Most social network models convert this type of data into binary relations between pairs of nodes. We develop a latent mixed membership model capable of modeling richer forms of transactional network data, including relations between more than two nodes. The model can cluster nodes and predict transactions. The block-model nature of the model implies that groups can be characterized in very general ways. This flexible notion of group structure enables discovery of rich structure in transactional networks. Estimation and inference are accomplished via a variational EM algorithm. Simulations indicate that the learning algorithm can recover the correct generative model. Interesting structure is discovered in the Enron email dataset and another dataset extracted from the Reddit website. Analysis of the Reddit data is facilitated by a novel performance measure for comparing two soft clusterings. The new model is superior at discovering mixed membership in groups and in predicting transactions.
Mahdi Shafiei, Hugh A. Chipman
ICDM2
2009 A Bayesian segmentation approach to ascertain copy number variations at the population level
abstract
MOTIVATION: Efficient and accurate ascertainment of copy number variations (CNVs) at the population level is essential to understand the evolutionary process and population genetics, and to apply CNVs in population-based genome-wide association studies for complex human diseases. We propose a novel Bayesian segmentation approach to identify CNVs in a defined population of any size. It is computationally efficient and provides statistical evidence for the detected CNVs through the Bayes factor. This approach has the unique feature of carrying out segmentation and assigning copy number status simultaneously-a desirable property that current segmentation methods do not share. RESULTS: In comparisons with popular two-step segmentation methods for a single individual using benchmark simulation studies, we find the new approach to perform competitively with respect to false discovery rate and sensitivity in breakpoint detection. In a simulation study of multiple samples with recurrent copy numbers, the new approach outperforms two leading single sample methods. We further demonstrate the effectiveness of our approach in population-level analysis of previously published HapMap data. We also apply our approach in studying population genetics of CNVs. AVAILABILITY: R programs are available at http://www.mshri.on.ca/mitacs/software/SOFTWARE.HTML
Long Yang Wu, Hugh A. Chipman, Shelley B. Bull, Laurent Briollais, Kesheng Wang
Bioinform.2
2007 Visualization and classification of graph-structured data: the case of the Enron dataset
abstract
Graph-structured networks are often used to represent relationships between persons in organizations or communities. In this paper we investigate the problem of learning a latent space representation of the data in which proximity in the latent space increases the likelihood of a social tie between the nodes. In addition, this latent space representation can be used to classify these data into homogeneous groups in order to identify, for instance, marginal communities of persons. We propose a Bayesian way to select both dimension of the latent space and number of groups. We apply our approach to the Enron dataset and we show interesting representation and clustering of individuals.
Charles Bouveyron, Hugh A. Chipman
IJCNN2
2006 Bayesian Ensemble Learning
abstract
We develop a Bayesian "sum-of-trees" model, named BART, where each tree is constrained by a prior to be a weak learner. Fitting and inference are accomplished via an iterative backfitting MCMC algorithm. This model is motivated by ensemble methods in general, and boosting algorithms in particular. Like boosting, each weak learner (i.e., each weak tree) contributes a small amount to the overall model. However, our procedure is defined by a statistical model: a prior and a likelihood, while boosting is defined by an algorithm. This model-based approach enables a full and accurate assessment of uncertainty in model predictions, while remaining highly competitive in terms of predictive accuracy.
Hugh A. Chipman, Edward I. George, Robert E. McCulloch
NIPS1
2002 Bayesian Treed Models
Hugh A. Chipman, Edward I. George, Robert E. McCulloch
Mach. Learn.1