VLDB 2026 Research / reviewers in the wild / expert
Frank De Smet
dblp:58/1074
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2015
0000-0002-0656-5921ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorArtificial intelligence and machine learning · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Kernel, tree and ensemble methods · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.2 | 1 | 2014 | EnsembleSVM: a library for ensemble learning using support vector machines · J. Mach. Learn. Res. 2014 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.2 | 1 | 2014 | EnsembleSVM: a library for ensemble learning using support vector machines · J. Mach. Learn. Res. 2014 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.1 | 3 | 2005 | M@CBETH: a microarray classification benchmarking tool · Bioinform. 2005 Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction · Bioinform. 2004 Functional bioinformatics of microarray data: from expression to regulation · Proc. IEEE 2002 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 2 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 Adaptive quality-based clustering of gene expression profiles · Bioinform. 2002 |
Bioinformatics and computational biology › cancer genomics
cancer classification |
0.0 | 1 | 2004 | Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction · Bioinform. 2004 |
Bioinformatics and computational biology › gene expression analysis › gene expression classification
microarray classification |
0.0 | 1 | 2004 | Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction · Bioinform. 2004 |
Bioinformatics and computational biology › sequence analysis › motif discovery
binding motif discovery |
0.0 | 1 | 2002 | Functional bioinformatics of microarray data: from expression to regulation · Proc. IEEE 2002 |
Bioinformatics and computational biology › gene expression analysis
gene expression clustering |
0.0 | 1 | 2002 | Functional bioinformatics of microarray data: from expression to regulation · Proc. IEEE 2002 |
Bioinformatics and computational biology
gibbs sampling |
0.0 | 1 | 2002 | Functional bioinformatics of microarray data: from expression to regulation · Proc. IEEE 2002 |
Bioinformatics and computational biology › gene expression analysis › gene expression clustering
microarray data clustering |
0.0 | 1 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 |
Bioinformatics and computational biology › sequence analysis
motif discovery |
0.0 | 1 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.0 | 1 | 2002 | Functional bioinformatics of microarray data: from expression to regulation · Proc. IEEE 2002 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.2ensemble learning · 0.2gibbs sampling · 0.1adaptive quality-based clustering · 0.1randomization-based benchmarking · 0.1cross-validation · 0.1regularization · 0.0radial basis function kernel · 0.0kernel PCA · 0.0motifsampler · 0.0expectation-maximization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | A robust ensemble approach to learn from positive and unlabeled data using SVM base models
Marc Claesen, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
Neurocomputing | 2 |
| 2014 | EnsembleSVM: a library for ensemble learning using support vector machines
Marc Claesen, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
J. Mach. Learn. Res. | 2 |
| 2005 | M@CBETH: a microarray classification benchmarking toolabstractMicroarray classification can be useful to support clinical management decisions for individual patients in, for example, oncology. However, comparing classifiers and selecting the best for each microarray dataset can be a tedious and non-straightforward task. The M@CBETH (a MicroArray Classification BEnchmarking Tool on a Host server) web service offers the microarray community a simple tool for making optimal two-class predictions. M@CBETH aims at finding the best prediction among different classification methods by using randomizations of the benchmarking dataset. The M@CBETH web service intends to introduce an optimal use of clinical microarray data classification. Nathalie Pochet, Frizo A. L. Janssens, Frank De Smet, Kathleen Marchal, Johan A. K. Suykens, Bart De Moor |
Bioinform. | 3 |
| 2004 | Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reductionabstractMOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. The aim of this paper is to systematically benchmark the role of non-linear versus linear techniques and dimensionality reduction methods. RESULTS: A systematic benchmarking study is performed by comparing linear versions of standard classification and dimensionality reduction techniques with their non-linear versions based on non-linear kernel functions with a radial basis function (RBF) kernel. A total of 9 binary cancer classification problems, derived from 7 publicly available microarray datasets, and 20 randomizations of each problem are examined. CONCLUSIONS: Three main conclusions can be formulated based on the performances on independent test sets. (1) When performing classification with least squares support vector machines (LS-SVMs) (without dimensionality reduction), RBF kernels can be used without risking too much overfitting. The results obtained with well-tuned RBF kernels are never worse and sometimes even statistically significantly better compared to results obtained with a linear kernel in terms of test set receiver operating characteristic and test set accuracy performances. (2) Even for classification with linear classifiers like LS-SVM with linear kernel, using regularization is very important. (3) When performing kernel principal component analysis (kernel PCA) before classification, using an RBF kernel for kernel PCA tends to result in overfitting, especially when using supervised feature selection. It has been observed that an optimal selection of a large number of features is often an indication for overfitting. Kernel PCA with linear kernel gives better results. Nathalie Pochet, Frank De Smet, Johan A. K. Suykens, Bart De Moor |
Bioinform. | 2 |
| 2002 | Adaptive quality-based clustering of gene expression profilesabstractMOTIVATION: Microarray experiments generate a considerable amount of data, which analyzed properly help us gain a huge amount of biologically relevant information about the global cellular behaviour. Clustering (grouping genes with similar expression profiles) is one of the first steps in data analysis of high-throughput expression measurements. A number of clustering algorithms have proved useful to make sense of such data. These classical algorithms, though useful, suffer from several drawbacks (e.g. they require the predefinition of arbitrary parameters like the number of clusters; they force every gene into a cluster despite a low correlation with other cluster members). In the following we describe a novel adaptive quality-based clustering algorithm that tackles some of these drawbacks. RESULTS: We propose a heuristic iterative two-step algorithm: First, we find in the high-dimensional representation of the data a sphere where the "density" of expression profiles is locally maximal (based on a preliminary estimate of the radius of the cluster-quality-based approach). In a second step, we derive an optimal radius of the cluster (adaptive approach) so that only the significantly coexpressed genes are included in the cluster. This estimation is achieved by fitting a model to the data using an EM-algorithm. By inferring the radius from the data itself, the biologist is freed from finding an optimal value for this radius by trial-and-error. The computational complexity of this method is approximately linear in the number of gene expression profiles in the data set. Finally, our method is successfully validated using existing data sets. AVAILABILITY: http://www.esat.kuleuven.ac.be/~thijs/Work/Clustering.html Frank De Smet, Janick Mathys, Kathleen Marchal, Gert Thijs, Bart De Moor, Yves Moreau |
Bioinform. | 1 |
| 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif SamplingabstractAbstract Summary: INCLUSive allows automatic multistep analysis of microarray data (clustering and motif finding). The clustering algorithm (adaptive quality-based clustering) groups together genes with highly similar expression profiles. The upstream sequences of the genes belonging to a cluster are automatically retrieved from GenBank and can be fed directly into Motif Sampler, a Gibbs sampling algorithm that retrieves statistically over-represented motifs in sets of sequences, in this case upstream regions of co-expressed genes. Availability: For academic purposes at http://www.esat.kuleuven.ac.be/~dna/BioI/Software.html Contact: [email protected] * To whom correspondence should be addressed. Email: [email protected]. Gert Thijs, Yves Moreau, Frank De Smet, Janick Mathys, Magali Lescot, Stephane Rombauts, Pierre Rouzé, Bart De Moor, Kathleen Marchal |
Bioinform. | 3 |
| 2002 | Functional bioinformatics of microarray data: from expression to regulationabstractUsing microarrays is a powerful technique to monitor the expression of thousands of genes in a single experiment. From series of such experiments, it is possible to identify the mechanisms that govern the activation of genes in an organism. Short deoxyribonucleic acid patterns (called binding sites) near the genes serve as switches that control gene expression. As a result similar patterns of expression can correspond to similar binding site patterns. Here we integrate clustering of coexpressed genes with the discovery of binding motifs. We overview several important clustering techniques and present a clustering algorithm (called adaptive quality-based clustering), which we have developed to address several shortcomings of existing methods. We overview the different techniques for motif finding, in particular the technique of Gibbs sampling, and we present several extensions of this technique in our Motif Sampler. Finally, we present an integrated web tool called INCLUSive (available online at http://www.esat.kuleuven.ac.be//spl sim/dna/BioI/Software.html) that allows the easy analysis of microarray data for motif finding. Yves Moreau, Frank De Smet, Gert Thijs, Kathleen Marchal, Bart De Moor |
Proc. IEEE | 2 |