Jangsun Baek

dblp:03/6712 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
0since 2021 · last 2011
0000-0003-0121-8856ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 75% Representation and self-supervised learning · 25%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 95% Computational science and engineering · 5%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
gene expression analysis
0.112011
Mixtures of common t-factor analyzers for clustering high-dimensional microarray data · Bioinform. 2011
Bioinformatics and computational biology › gene expression analysis › gene expression clustering
microarray data clustering
0.112011
Mixtures of common t-factor analyzers for clustering high-dimensional microarray data · Bioinform. 2011
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.112010
Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
factor analysis
0.112010
Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.112010
Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture of factor analyzers
0.112010
Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Bioinformatics and computational biology › bioimage informatics
microarray image analysis
0.112007
Segmentation and intensity estimation of microarray images using a gamma-t mixture model · Bioinform. 2007
Bioinformatics and computational biology › gene expression analysis
model-based clustering
0.012011
Mixtures of common t-factor analyzers for clustering high-dimensional microarray data · Bioinform. 2011
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.012011
Mixtures of common t-factor analyzers for clustering high-dimensional microarray data · Bioinform. 2011
Data mining
clustering
0.012010
Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Data mining › clustering
model-based clustering
0.012010
Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Computational science and engineering › latent variable model
mixture model
0.012007
Segmentation and intensity estimation of microarray images using a gamma-t mixture model · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

parameter reduction · 0.2maximum likelihood estimation · 0.2EM algorithm · 0.2robust statistical modeling · 0.1mixture of t-factor analyzers · 0.1maximum likelihood · 0.1kernel smoothing · 0.1gamma-t mixture model · 0.1
YearPublicationVenuePosition
2011 Mixtures of common t-factor analyzers for clustering high-dimensional microarray data
abstract
MOTIVATION: Mixtures of factor analyzers enable model-based clustering to be undertaken for high-dimensional microarray data, where the number of observations n is small relative to the number of genes p. Moreover, when the number of clusters is not small, for example, where there are several different types of cancer, there may be the need to reduce further the number of parameters in the specification of the component-covariance matrices. A further reduction can be achieved by using mixtures of factor analyzers with common component-factor loadings (MCFA), which is a more parsimonious model. However, this approach is sensitive to both non-normality and outliers, which are commonly observed in microarray experiments. This sensitivity of the MCFA approach is due to its being based on a mixture model in which the multivariate normal family of distributions is assumed for the component-error and factor distributions. RESULTS: An extension to mixtures of t-factor analyzers with common component-factor loadings is considered, whereby the multivariate t-family is adopted for the component-error and factor distributions. An EM algorithm is developed for the fitting of mixtures of common t-factor analyzers. The model can handle data with tails longer than that of the normal distribution, is robust against outliers and allows the data to be displayed in low-dimensional plots. It is applied here to both synthetic data and some microarray gene expression data for clustering and shows its better performance over several existing methods. AVAILABILITY: The algorithms were implemented in Matlab. The Matlab code is available at http://blog.naver.com/aggie100.
Jangsun Baek, Geoffrey J. McLachlan
Bioinform.1
2010 Mixtures of Factor Analyzers with Common Factor Loadings: Applications to the Clustering and Visualization of High-Dimensional Data
abstract
Mixtures of factor analyzers enable model-based density estimation to be undertaken for high-dimensional data, where the number of observations n is not very large relative to their dimension p. In practice, there is often the need to further reduce the number of parameters in the specification of the component-covariance matrices. To this end, we propose the use of common component-factor loadings, which considerably reduces further the number of parameters. Moreover, it allows the data to be displayed in low--dimensional plots.
Jangsun Baek, Geoffrey J. McLachlan, Lloyd K. Flack
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 A modified correlation coefficient based similarity measure for clustering time-course gene expression data
Young Sook Son, Jangsun Baek
Pattern Recognit. Lett.2
2007 Segmentation and intensity estimation of microarray images using a gamma-t mixture model
abstract
MOTIVATION: We present a new approach to the analysis of images for complementary DNA microarray experiments. The image segmentation and intensity estimation are performed simultaneously by adopting a two-component mixture model. One component of this mixture corresponds to the distribution of the background intensity, while the other corresponds to the distribution of the foreground intensity. The intensity measurement is a bivariate vector consisting of red and green intensities. The background intensity component is modeled by the bivariate gamma distribution, whose marginal densities for the red and green intensities are independent three-parameter gamma distributions with different parameters. The foreground intensity component is taken to be the bivariate t distribution, with the constraint that the mean of the foreground is greater than that of the background for each of the two colors. The degrees of freedom of this t distribution are inferred from the data but they could be specified in advance to reduce the computation time. Also, the covariance matrix is not restricted to being diagonal and so it allows for nonzero correlation between R and G foreground intensities. This gamma-t mixture model is fitted by maximum likelihood via the EM algorithm. A final step is executed whereby nonparametric (kernel) smoothing is undertaken of the posterior probabilities of component membership. The main advantages of this approach are: (1) it enjoys the well-known strengths of a mixture model, namely flexibility and adaptability to the data; (2) it considers the segmentation and intensity simultaneously and not separately as in commonly used existing software, and it also works with the red and green intensities in a bivariate framework as opposed to their separate estimation via univariate methods; (3) the use of the three-parameter gamma distribution for the background red and green intensities provides a much better fit than the normal (log normal) or t distributions; (4) the use of the bivariate t distribution for the foreground intensity provides a model that is less sensitive to extreme observations; (5) as a consequence of the aforementioned properties, it allows segmentation to be undertaken for a wide range of spot shapes, including doughnut, sickle shape and artifacts. RESULTS: We apply our method for gridding, segmentation and estimation to cDNA microarray real images and artificial data. Our method provides better segmentation results in spot shapes as well as intensity estimation than Spot and spotSegmentation R language softwares. It detected blank spots as well as bright artifact for the real data, and estimated spot intensities with high-accuracy for the synthetic data. AVAILABILITY: The algorithms were implemented in Matlab. The Matlab codes implementing both the gridding and segmentation/estimation are available upon request. SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.
Jangsun Baek, Young Sook Son, Geoffrey J. McLachlan
Bioinform.1
2006 Local Linear Logistic Discriminant Analysis with Partial Least Square Components
Jangsun Baek, Young Sook Son
ADMA1
2004 Face recognition using partial least squares components
Jangsun Baek, Min-Soo Kim 0007
Pattern Recognit.1