VLDB 2026 Research / reviewers in the wild / expert
Pierre Rouzé
dblp:71/350
· DBLP profile ↗
12ranked-venue papers
0as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
genome annotation |
0.1 | 2 | 2008 | ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profiles · ISMB 2008 In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists · Bioinform. 2007 |
Bioinformatics and computational biology › genome annotation
gene prediction |
0.1 | 2 | 2007 | In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists · Bioinform. 2007 Evaluation of gene prediction software using a genomic data set: application to <$O_SSF>Arabidopsis thaliana<$C_SSF>sequences · Bioinform. 1999 |
Bioinformatics and computational biology
genomics |
0.1 | 2 | 2003 | Automatic design of gene-specific sequence tags for genome-wide functional studies · Bioinform. 2003 AFLPinSilico, simulating AFLP fingerprints · Bioinform. 2003 |
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis |
0.1 | 1 | 2008 | ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profiles · ISMB 2008 |
Bioinformatics and computational biology › sequence analysis
motif discovery |
0.1 | 2 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling · Bioinform. 2001 |
Bioinformatics and computational biology › genome annotation
gene structure prediction |
0.1 | 1 | 2005 | SpliceMachine: predicting splice sites from high-dimensional local context representations · Bioinform. 2005 |
Bioinformatics and computational biology › sequence analysis
sequence annotation |
0.1 | 1 | 2005 | SpliceMachine: predicting splice sites from high-dimensional local context representations · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
splice site prediction |
0.1 | 1 | 2005 | SpliceMachine: predicting splice sites from high-dimensional local context representations · Bioinform. 2005 |
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
RNA secondary structure prediction |
0.0 | 1 | 2004 | Evidence that microRNA precursors, unlike other non-coding RNAs, have lower folding free energies than random sequences · Bioinform. 2004 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecular fingerprint |
0.0 | 1 | 2003 | AFLPinSilico, simulating AFLP fingerprints · Bioinform. 2003 |
Bioinformatics and computational biology › genomics
primer design |
0.0 | 1 | 2003 | Automatic design of gene-specific sequence tags for genome-wide functional studies · Bioinform. 2003 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 |
Bioinformatics and computational biology › gene expression analysis › gene expression clustering
microarray data clustering |
0.0 | 1 | 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002 |
Bioinformatics and computational biology
gene regulation |
0.0 | 1 | 2001 | A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genes · RECOMB 2001 |
Bioinformatics and computational biology › sequence analysis
motif detection |
0.0 | 1 | 2001 | A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genes · RECOMB 2001 |
Bioinformatics and computational biology › gene regulation
regulatory element discovery |
0.0 | 1 | 2001 | A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling · Bioinform. 2001 |
Bioinformatics and computational biology › genome annotation › gene prediction
gene prediction evaluation |
0.0 | 1 | 1999 | Evaluation of gene prediction software using a genomic data set: application to <$O_SSF>Arabidopsis thaliana<$C_SSF>sequences · Bioinform. 1999 |
Bioinformatics and computational biology › genome annotation › gene prediction
exon prediction |
0.0 | 1 | 2007 | In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists · Bioinform. 2007 |
Methods — techniques the papers use, named apart from their topics
gibbs sampling · 0.1unsupervised clustering · 0.1self-organizing map · 0.1sequence feature combination · 0.1markov model · 0.1higher-order background model · 0.1machine learning · 0.1randomization test · 0.0minimum free energy folding · 0.0in silico simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profilesabstractMOTIVATION: More and more genomes are being sequenced, and to keep up with the pace of sequencing projects, automated annotation techniques are required. One of the most challenging problems in genome annotation is the identification of the core promoter. Because the identification of the transcription initiation region is such a challenging problem, it is not yet a common practice to integrate transcription start site prediction in genome annotation projects. Nevertheless, better core promoter prediction can improve genome annotation and can be used to guide experimental work. RESULTS: Comparing the average structural profile based on base stacking energy of transcribed, promoter and intergenic sequences demonstrates that the core promoter has unique features that cannot be found in other sequences. We show that unsupervised clustering by using self-organizing maps can clearly distinguish between the structural profiles of promoter sequences and other genomic sequences. An implementation of this promoter prediction program, called ProSOM, is available and has been compared with the state-of-the-art. We propose an objective, accurate and biologically sound validation scheme for core promoter predictors. ProSOM performs at least as well as the software currently available, but our technique is more balanced in terms of the number of predicted sites and the number of false predictions, resulting in a better all-round performance. Additional tests on the ENCODE regions of the human genome show that 98% of all predictions made by ProSOM can be associated with transcriptionally active regions, which demonstrates the high precision. AVAILABILITY: Predictions for the human genome, the validation datasets and the program (ProSOM) are available upon request. Thomas Abeel, Yvan Saeys, Pierre Rouzé, Yves Van de Peer |
ISMB | 3 |
| 2007 | In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protistsabstractMOTIVATION: Prediction of the coding potential for stretches of DNA is crucial in gene calling and genome annotation, where it is used to identify potential exons and to position their boundaries in conjunction with functional sites, such as splice sites and translation initiation sites. The ability to discriminate between coding and non-coding sequences relates to the structure of coding sequences, which are organized in codons, and by their biased usage. For statistical reasons, the longer the sequences, the easier it is to detect this codon bias. However, in many eukaryotic genomes, where genes harbour many introns, both introns and exons might be small and hard to distinguish based on coding potential. RESULTS: Here, we present novel approaches that specifically aim at a better detection of coding potential in short sequences. The methods use complementary sequence features, combined with identification of which features are relevant in discriminating between coding and non-coding sequences. These newly developed methods are evaluated on different species, representative of four major eukaryotic kingdoms, and extensively compared to state-of-the-art Markov models, which are often used for predicting coding potential. The main conclusions drawn from our analyses are that (1) combining complementary sequence features clearly outperforms current Markov models for coding potential prediction in short sequence fragments, (2) coding potential prediction benefits from length-specific models, and these models are not necessarily the same for different sequence lengths and (3) comparing the results across several species indicates that, although our combined method consistently performs extremely well, there are important differences across genomes. SUPPLEMENTARY DATA: http://bioinformatics.psb.ugent.be/. Yvan Saeys, Pierre Rouzé, Yves Van de Peer |
Bioinform. | 2 |
| 2005 | SpliceMachine: predicting splice sites from high-dimensional local context representationsabstractMOTIVATION: In this age of complete genome sequencing, finding the location and structure of genes is crucial for further molecular research. The accurate prediction of intron boundaries largely facilitates the correct prediction of gene structure in nuclear genomes. Many tools for localizing these boundaries on DNA sequences have been developed and are available to researchers through the internet. Nevertheless, these tools still make many false positive predictions. RESULTS: This manuscript presents a novel publicly available splice site prediction tool named SpliceMachine that (i) shows state-of-the-art prediction performance on Arabidopsis thaliana and human sequences, (ii) performs a computationally fast annotation and (iii) can be trained by the user on its own data. AVAILABILITY: Results, figures and software are available at http://www.bioinformatics.psb.ugent.be/supplementary_data/ CONTACT: [email protected]; [email protected]. Sven Degroeve, Yvan Saeys, Bernard De Baets, Pierre Rouzé, Yves Van de Peer |
Bioinform. | 4 |
| 2004 | Evidence that microRNA precursors, unlike other non-coding RNAs, have lower folding free energies than random sequencesabstractAbstract Motivation: Most non-coding RNAs are characterized by a specific secondary and tertiary structure that determines their function. Here, we investigate the folding energy of the secondary structure of non-coding RNA sequences, such as microRNA precursors, transfer RNAs and ribosomal RNAs in several eukaryotic taxa. Statistical biases are assessed by a randomization test, in which the predicted minimum free energy of folding is compared with values obtained for structures inferred from randomly shuffling the original sequences. Results: In contrast with transfer RNAs and ribosomal RNAs, the majority of the microRNA sequences clearly exhibit a folding free energy that is considerably lower than that for shuffled sequences, indicating a high tendency in the sequence towards a stable secondary structure. A possible usage of this statistical test in the framework of the detection of genuine miRNA sequences is discussed. Availability: The dataset, software and additional data files are freely available as supplementary information on our Website. Supplementary information: http://www.psb.ugent.be/bioinformatics/ Eric Bonnet, Jan Wuyts, Pierre Rouzé, Yves Van de Peer |
Bioinform. | 3 |
| 2004 | Feature selection for splice site prediction: A new method using EDA-based feature rankingabstractBACKGROUND: The identification of relevant biological features in large and complex datasets is an important step towards gaining insight in the processes underlying the data. Other advantages of feature selection include the ability of the classification system to attain good or even better solutions using a restricted subset of features, and a faster classification. Thus, robust methods for fast feature selection are of key importance in extracting knowledge from complex biological data. RESULTS: In this paper we present a novel method for feature subset selection applied to splice site prediction, based on estimation of distribution algorithms, a more general framework of genetic algorithms. From the estimated distribution of the algorithm, a feature ranking is derived. Afterwards this ranking is used to iteratively discard features. We apply this technique to the problem of splice site prediction, and show how it can be used to gain insight into the underlying biological process of splicing. CONCLUSION: We show that this technique proves to be more robust than the traditional use of estimation of distribution algorithms for feature selection: instead of returning a single best subset of features (as they normally do) this method provides a dynamical view of the feature selection process, like the traditional sequential wrapper methods. However, the method is faster than the traditional techniques, and scales better to datasets described by a large number of features. Yvan Saeys, Sven Degroeve, Dirk Aeyels, Pierre Rouzé, Yves Van de Peer |
BMC Bioinform. | 4 |
| 2003 | AFLPinSilico, simulating AFLP fingerprintsabstractAbstract Summary: A drawback of the Amplified Fragment Length Polymorphism (AFLP) fingerprinting method is the difficulty to correlate the different fragments with their DNA sequence. The AFLPinSilico application presented here simulates AFLP experiments run on either cDNA or genomic sequences, producing virtual fingerprints that allow high throughput identification of AFLP fragments. The program also enables biologists to manage experiments through simulations done beforehand, thereby reducing the number of experiments that have to be run. AFLPinSilico is available through the www or as a stand-alone version, through a command line executable (available upon request, for any platform running PERL). Availability: For academic use http://www.psb.rug.ac.be/bioinformatics/AFLPinSilico.html Contact: [email protected] * To whom correspondence should be addressed. Stephane Rombauts, Yves Van de Peer, Pierre Rouzé |
Bioinform. | 3 |
| 2003 | Automatic design of gene-specific sequence tags for genome-wide functional studiesabstractMOTIVATION: The availability of complete genome sequences allows the identification of short DNA segments that are specific to each annotated gene. Such unique gene sequence tags (GSTs) replace advantageously cDNAs in microarray transcript profiling experiments. In particular, probes corresponding to individual members of multigene families can be chosen carefully to avoid cross-hybridization events. RESULTS: The Specific Primer and Amplicon Design Software (SPADS) was constructed to delineate the more divergent regions in each gene by comparing them with a completely annotated genome sequence and to select optimal primer pairs for the polymerase chain reaction amplification of one divergent region per gene. SPADS is a unique integrated tool to design specific GSTs from any public or private genome sequences and allows the user to fine-tune GST size and specificity. SPADS has been used to obtain probes for whole genome and family-wide transcript profiling, as well as inserts for gene-specific knock-out experiments. AVAILABILITY: The GENOPLANTE SPADS source code and web interface are available upon request. The online version is accessible via http://genoplante-info.infobiogen.fr/spads and via http://oberon.fvms.ugent.be:8080/SPADS/ Vincent Thareau, Patrice Déhais, Carine Serizet, Pierre Hilson, Pierre Rouzé, Sébastien Aubourg |
Bioinform. | 5 |
| 2003 | Orphan gene finding - an exon assembly approach
Philippe Blayo, Pierre Rouzé, Marie-France Sagot |
Theor. Comput. Sci. | 2 |
| 2002 | INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif SamplingabstractAbstract Summary: INCLUSive allows automatic multistep analysis of microarray data (clustering and motif finding). The clustering algorithm (adaptive quality-based clustering) groups together genes with highly similar expression profiles. The upstream sequences of the genes belonging to a cluster are automatically retrieved from GenBank and can be fed directly into Motif Sampler, a Gibbs sampling algorithm that retrieves statistically over-represented motifs in sets of sequences, in this case upstream regions of co-expressed genes. Availability: For academic purposes at http://www.esat.kuleuven.ac.be/~dna/BioI/Software.html Contact: [email protected] * To whom correspondence should be addressed. Email: [email protected]. Gert Thijs, Yves Moreau, Frank De Smet, Janick Mathys, Magali Lescot, Stephane Rombauts, Pierre Rouzé, Bart De Moor, Kathleen Marchal |
Bioinform. | 7 |
| 2001 | A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genesabstractMicroarray experiments can reveal useful information on the transcriptional regulation. We try to find regulatory elements in the region upstream of translation start of coexpressed genes. Here we present a modification to the original Gibbs Sampling algorithm [12]. We introduce a probability distribution to estimate the number of copies of the motif in a sequence. The second modification is the incorporation of a higher-order background model. We have successfully tested our algorithm on several data sets. First we show results on two selected data set: sequences from plants containing the G-box motif and the upstream sequences from bacterial genes regulated by O2-responsive protein FNR. In both cases the motif sampler is able to find the expected motifs. Finally, the sampler is tested on 4 clusters of coexpressed genes from a wounding experiment in Arabidopsis thaliana. We find several putative motifs that are related to the pathways involved in the plant defense mechanism. Gert Thijs, Kathleen Marchal, Magali Lescot, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau |
RECOMB | 6 |
| 2001 | A higher-order background model improves the detection of promoter regulatory elements by Gibbs samplingabstractMOTIVATION: Transcriptome analysis allows detection and clustering of genes that are coexpressed under various biological circumstances. Under the assumption that coregulated genes share cis-acting regulatory elements, it is important to investigate the upstream sequences controlling the transcription of these genes. To improve the robustness of the Gibbs sampling algorithm to noisy data sets we propose an extension of this algorithm for motif finding with a higher-order background model. RESULTS: Simulated data and real biological data sets with well-described regulatory elements are used to test the influence of the different background models on the performance of the motif detection algorithm. We show that the use of a higher-order model considerably enhances the performance of our motif finding algorithm in the presence of noisy data. For Arabidopsis thaliana, a reliable background model based on a set of carefully selected intergenic sequences was constructed. AVAILABILITY: Our implementation of the Gibbs sampler called the Motif Sampler can be used through a web interface: http://www.esat.kuleuven.ac.be/~thijs/Work/MotifSampler.html. CONTACT: [email protected]; [email protected] Gert Thijs, Magali Lescot, Kathleen Marchal, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau |
Bioinform. | 6 |
| 1999 | Evaluation of gene prediction software using a genomic data set: application to <$O_SSF>Arabidopsis thaliana<$C_SSF>sequencesabstractMOTIVATION: The annotation of the Arabidopsis thaliana genome remains a problem in terms of time and quality. To improve the annotation process, we want to choose the most appropriate tools to use inside a computer-assisted annotation platform. We therefore need evaluation of prediction programs with Arabidopsis sequences containing multiple genes. RESULTS: We have developed AraSet, a data set of contigs of validated genes, enabling the evaluation of multi-gene models for the Arabidopsis genome. Besides conventional metrics to evaluate gene prediction at the site and the exon levels, new measures were introduced for the prediction at the protein sequence level as well as for the evaluation of gene models. This evaluation method is of general interest and could apply to any new gene prediction software and to any eukaryotic genome. The GeneMark.hmm program appears to be the most accurate software at all three levels for the Arabidopsis genomic sequences. Gene modeling could be further improved by combination of prediction software. AVAILABILITY: The AraSet sequence set, the Perl programs and complementary results and notes are available at http://sphinx.rug.ac.be:8080/biocomp/napav/. CONTACT: [email protected]. Nathalie Pavy, Stephane Rombauts, Patrice Déhais, Catherine Mathé, Ramana V. Davuluri, Philippe Leroy, Pierre Rouzé |
Bioinform. | 7 |