Rafal Kustra

dblp:59/2055 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2013
0000-0002-5949-8718ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 50% Kernel, tree and ensemble methods · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › multi-view learning
canonical correlation analysis
0.212013
Canonical Correlation Analysis based on Hilbert-Schmidt Independence Criterion and Centered Kernel Target Alignment · ICML (2) 2013
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.212013
Canonical Correlation Analysis based on Hilbert-Schmidt Independence Criterion and Centered Kernel Target Alignment · ICML (2) 2013
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel machines
kernel canonical correlation analysis
0.212013
Canonical Correlation Analysis based on Hilbert-Schmidt Independence Criterion and Centered Kernel Target Alignment · ICML (2) 2013
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.212013
Canonical Correlation Analysis based on Hilbert-Schmidt Independence Criterion and Centered Kernel Target Alignment · ICML (2) 2013
Bioinformatics and computational biology › statistical genetics
genetic linkage analysis
0.112008
EM-random forest and new measures of variable importance for multi-locus quantitative trait linkage analysis · Bioinform. 2008

Methods — techniques the papers use, named apart from their topics

hilbert-schmidt independence criterion · 0.2centered kernel target alignment · 0.2random forest · 0.1haseman-elston regression · 0.1EM algorithm · 0.1
YearPublicationVenuePosition
2013 Canonical Correlation Analysis based on Hilbert-Schmidt Independence Criterion and Centered Kernel Target Alignment
abstract
Canonical correlation analysis (CCA) is a well established technique for identifying linear relationships among two variable sets. Kernel CCA (KCCA) is the most notable nonlinear extension but it lacks interpretability and robustness against irrelevant features. The aim of this article is to introduce two nonlinear CCA extensions that rely on the recently proposed Hilbert-Schmidt independence criterion and the centered kernel target alignment. These extensions determine linear projections that provide maximally dependent projected data pairs. The paper demonstrates that the use of linear projections allows removing irrelevant features, whilst extracting combinations of strongly associated features. This is exemplified through a simulation and the analysis of recorded data that are available in the literature.
Billy Chang, Uwe Krüger 0001, Rafal Kustra, Junping Zhang
ICML (2)3
2010 Data-Fusion in Clustering Microarray Data: Balancing Discovery and Interpretability
abstract
While clustering genes remains one of the most popular exploratory tools for expression data, it often results in a highly variable and biologically uninformative clusters. This paper explores a data fusion approach to clustering microarray data. Our method, which combined expression data and Gene Ontology (GO)-derived information, is applied on a real data set to perform genome-wide clustering. A set of novel tools is proposed to validate the clustering results and pick a fair value of infusion coefficient. These tools measure stability, biological relevance, and distance from the expression-only clustering solution. Our results indicate that a data-fusion clustering leads to more stable, biologically relevant clusters that are still representative of the experimental data.
Rafal Kustra, Adam Zagdanski
IEEE ACM Trans. Comput. Biol. Bioinform.1
2008 EM-random forest and new measures of variable importance for multi-locus quantitative trait linkage analysis
abstract
MOTIVATION: We developed an EM-random forest (EMRF) for Haseman-Elston quantitative trait linkage analysis that accounts for marker ambiguity and weighs each sib-pair according to the posterior identical by descent (IBD) distribution. The usual random forest (RF) variable importance (VI) index used to rank markers for variable selection is not optimal when applied to linkage data because of correlation between markers. We define new VI indices that borrow information from linked markers using the correlation structure inherent in IBD linkage data. RESULTS: Using simulations, we find that the new VI indices in EMRF performed better than the original RF VI index and performed similarly or better than EM-Haseman-Elston regression LOD score for various genetic models. Moreover, tree size and markers subset size evaluated at each node are important considerations in RFs. AVAILABILITY: The source code for EMRF written in C is available at www.infornomics.utoronto.ca/downloads/EMRF.
Sophia S. F. Lee, Rafal Kustra, Shelley B. Bull
Bioinform.3
2006 Incorporating Gene Ontology in Clustering Gene Expression Data
abstract
In this paper we consider a general framework for clustering expression data that permits integration of various biological data sources through combination of corresponding dissimilarity measures. In the paper we briefly review currently published attempts to genomic data fusion and discuss a problem of validating results from clustering expression data. We apply our approach to a real microarray expression dataset which induces a correlation-based dissimilarity matrix, and use gene ontology - biological process annotations to derive GO-based dissimilarity matrix. The proposed procedure is verified using a simple knowledge-based validation measure based on protein-protein interaction database. Obtained results reveal that combining experimental data with comprehensive and reliable biological repository may improve performance of cluster analysis and yield biologically meaningful gene clusters
Rafal Kustra, Adam Zagdanski
CBMS1
2006 A factor analysis model for functional genomics
abstract
BACKGROUND: Expression array data are used to predict biological functions of uncharacterized genes by comparing their expression profiles to those of characterized genes. While biologically plausible, this is both statistically and computationally challenging. Typical approaches are computationally expensive and ignore correlations among expression profiles and functional categories. RESULTS: We propose a factor analysis model (FAM) for functional genomics and give a two-step algorithm, using genome-wide expression data for yeast and a subset of Gene-Ontology Biological Process functional annotations. We show that the predictive performance of our method is comparable to the current best approach while our total computation time was faster by a factor of 4000. We discuss the unique challenges in performance evaluation of algorithms used for genome-wide functions genomics. Finally, we discuss extensions to our method that can incorporate the inherent correlation structure of the functional categories to further improve predictive performance. CONCLUSION: Our factor analysis model is a computationally efficient technique for functional genomics and provides a clear and unified statistical framework with potential for incorporating important gene ontology information to improve predictions.
Rafal Kustra, Romy Shioda
BMC Bioinform.1
2004 An Efficient Method to Estimate Labelled Sample Size for Transductive LDA(QDA/MDA) Based on Bayes Risk
Han Liu 0001, Xiaobin Yuan, Qianying Tang, Rafal Kustra
ECML4
2001 Penalized Discriminant Analysis of [15O]-water PET Brain Images with Prediction Error Selection of Smoothness and Regularization
abstract
We propose a flexible, comprehensive approach for analysis of [15O]-water positron emission tomography (PET) brain images using a penalized version of linear discriminant analysis (PDA). We applied it to scans from 20 subjects (eight scans/subject) performing a finger movement task and analyzed: 1) two classes to obtain a covariance-normalized baseline-activation image, and 2) eight classes for the mean within subject temporal structure which contained baseline-activation and time-dependent changes in a two-dimensional canonical subspace. We imposed spatial smoothness on the resulting image(s) by expanding it in five tensor-product B-spline (TPS) bases of varying smoothness, and further regularized with a ridge-type penalty on the noise covariance matrix. The discrimination approach of PDA provides a probabilistic framework within which prediction error (PE) estimates are derived. We used these to optimize over TPS bases and a ridge hyperparameter (expressed as equivalent degrees of freedom, EDF). We obtained unbiased, low variance PE estimates using modern resampling tools (.632+ Bootstrap and cross validation), and compared PDA of 1) TPS-projected, mean-normalized and unnormalized scans and 2) mean-normalized scans with and without additional presmoothing. By examining the tradeoffs between PE and EDF, as a function of basis selection and image smoothing we demonstrate the utility of PDA, the PE framework, and the relationship between singular value decomposition and smooth TPS bases in the analysis of functional neuroimages.
Rafal Kustra, Stephen C. Strother
IEEE Trans. Medical Imaging1