Pierre Rouzé

dblp:71/350 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
10 papers
Bioinformatics and computational biology · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genome annotation
0.122008
ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profiles · ISMB 2008
In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists · Bioinform. 2007
Bioinformatics and computational biology › genome annotation
gene prediction
0.122007
In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists · Bioinform. 2007
Evaluation of gene prediction software using a genomic data set: application to <$O_SSF>Arabidopsis thaliana<$C_SSF>sequences · Bioinform. 1999
Bioinformatics and computational biology
genomics
0.122003
Automatic design of gene-specific sequence tags for genome-wide functional studies · Bioinform. 2003
AFLPinSilico, simulating AFLP fingerprints · Bioinform. 2003
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.112008
ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profiles · ISMB 2008
Bioinformatics and computational biology › sequence analysis
motif discovery
0.122002
INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002
A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling · Bioinform. 2001
Bioinformatics and computational biology › genome annotation
gene structure prediction
0.112005
SpliceMachine: predicting splice sites from high-dimensional local context representations · Bioinform. 2005
Bioinformatics and computational biology › sequence analysis
sequence annotation
0.112005
SpliceMachine: predicting splice sites from high-dimensional local context representations · Bioinform. 2005
Bioinformatics and computational biology › genome annotation
splice site prediction
0.112005
SpliceMachine: predicting splice sites from high-dimensional local context representations · Bioinform. 2005
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
RNA secondary structure prediction
0.012004
Evidence that microRNA precursors, unlike other non-coding RNAs, have lower folding free energies than random sequences · Bioinform. 2004
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecular fingerprint
0.012003
AFLPinSilico, simulating AFLP fingerprints · Bioinform. 2003
Bioinformatics and computational biology › genomics
primer design
0.012003
Automatic design of gene-specific sequence tags for genome-wide functional studies · Bioinform. 2003
Bioinformatics and computational biology
gene expression analysis
0.012002
INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002
Bioinformatics and computational biology › gene expression analysis › gene expression clustering
microarray data clustering
0.012002
INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling · Bioinform. 2002
Bioinformatics and computational biology
gene regulation
0.012001
A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genes · RECOMB 2001
Bioinformatics and computational biology › sequence analysis
motif detection
0.012001
A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genes · RECOMB 2001
Bioinformatics and computational biology › gene regulation
regulatory element discovery
0.012001
A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling · Bioinform. 2001
Bioinformatics and computational biology › genome annotation › gene prediction
gene prediction evaluation
0.011999
Evaluation of gene prediction software using a genomic data set: application to <$O_SSF>Arabidopsis thaliana<$C_SSF>sequences · Bioinform. 1999
Bioinformatics and computational biology › genome annotation › gene prediction
exon prediction
0.012007
In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

gibbs sampling · 0.1unsupervised clustering · 0.1self-organizing map · 0.1sequence feature combination · 0.1markov model · 0.1higher-order background model · 0.1machine learning · 0.1randomization test · 0.0minimum free energy folding · 0.0in silico simulation · 0.0
YearPublicationVenuePosition
2008 ProSOM: core promoter prediction based on unsupervised clustering of DNA physical profiles
abstract
MOTIVATION: More and more genomes are being sequenced, and to keep up with the pace of sequencing projects, automated annotation techniques are required. One of the most challenging problems in genome annotation is the identification of the core promoter. Because the identification of the transcription initiation region is such a challenging problem, it is not yet a common practice to integrate transcription start site prediction in genome annotation projects. Nevertheless, better core promoter prediction can improve genome annotation and can be used to guide experimental work. RESULTS: Comparing the average structural profile based on base stacking energy of transcribed, promoter and intergenic sequences demonstrates that the core promoter has unique features that cannot be found in other sequences. We show that unsupervised clustering by using self-organizing maps can clearly distinguish between the structural profiles of promoter sequences and other genomic sequences. An implementation of this promoter prediction program, called ProSOM, is available and has been compared with the state-of-the-art. We propose an objective, accurate and biologically sound validation scheme for core promoter predictors. ProSOM performs at least as well as the software currently available, but our technique is more balanced in terms of the number of predicted sites and the number of false predictions, resulting in a better all-round performance. Additional tests on the ENCODE regions of the human genome show that 98% of all predictions made by ProSOM can be associated with transcriptionally active regions, which demonstrates the high precision. AVAILABILITY: Predictions for the human genome, the validation datasets and the program (ProSOM) are available upon request.
Thomas Abeel, Yvan Saeys, Pierre Rouzé, Yves Van de Peer
ISMB3
2007 In search of the small ones: improved prediction of short exons in vertebrates, plants, fungi and protists
abstract
MOTIVATION: Prediction of the coding potential for stretches of DNA is crucial in gene calling and genome annotation, where it is used to identify potential exons and to position their boundaries in conjunction with functional sites, such as splice sites and translation initiation sites. The ability to discriminate between coding and non-coding sequences relates to the structure of coding sequences, which are organized in codons, and by their biased usage. For statistical reasons, the longer the sequences, the easier it is to detect this codon bias. However, in many eukaryotic genomes, where genes harbour many introns, both introns and exons might be small and hard to distinguish based on coding potential. RESULTS: Here, we present novel approaches that specifically aim at a better detection of coding potential in short sequences. The methods use complementary sequence features, combined with identification of which features are relevant in discriminating between coding and non-coding sequences. These newly developed methods are evaluated on different species, representative of four major eukaryotic kingdoms, and extensively compared to state-of-the-art Markov models, which are often used for predicting coding potential. The main conclusions drawn from our analyses are that (1) combining complementary sequence features clearly outperforms current Markov models for coding potential prediction in short sequence fragments, (2) coding potential prediction benefits from length-specific models, and these models are not necessarily the same for different sequence lengths and (3) comparing the results across several species indicates that, although our combined method consistently performs extremely well, there are important differences across genomes. SUPPLEMENTARY DATA: http://bioinformatics.psb.ugent.be/.
Yvan Saeys, Pierre Rouzé, Yves Van de Peer
Bioinform.2
2005 SpliceMachine: predicting splice sites from high-dimensional local context representations
abstract
MOTIVATION: In this age of complete genome sequencing, finding the location and structure of genes is crucial for further molecular research. The accurate prediction of intron boundaries largely facilitates the correct prediction of gene structure in nuclear genomes. Many tools for localizing these boundaries on DNA sequences have been developed and are available to researchers through the internet. Nevertheless, these tools still make many false positive predictions. RESULTS: This manuscript presents a novel publicly available splice site prediction tool named SpliceMachine that (i) shows state-of-the-art prediction performance on Arabidopsis thaliana and human sequences, (ii) performs a computationally fast annotation and (iii) can be trained by the user on its own data. AVAILABILITY: Results, figures and software are available at http://www.bioinformatics.psb.ugent.be/supplementary_data/ CONTACT: [email protected]; [email protected].
Sven Degroeve, Yvan Saeys, Bernard De Baets, Pierre Rouzé, Yves Van de Peer
Bioinform.4
2004 Evidence that microRNA precursors, unlike other non-coding RNAs, have lower folding free energies than random sequences
abstract
Abstract Motivation: Most non-coding RNAs are characterized by a specific secondary and tertiary structure that determines their function. Here, we investigate the folding energy of the secondary structure of non-coding RNA sequences, such as microRNA precursors, transfer RNAs and ribosomal RNAs in several eukaryotic taxa. Statistical biases are assessed by a randomization test, in which the predicted minimum free energy of folding is compared with values obtained for structures inferred from randomly shuffling the original sequences. Results: In contrast with transfer RNAs and ribosomal RNAs, the majority of the microRNA sequences clearly exhibit a folding free energy that is considerably lower than that for shuffled sequences, indicating a high tendency in the sequence towards a stable secondary structure. A possible usage of this statistical test in the framework of the detection of genuine miRNA sequences is discussed. Availability: The dataset, software and additional data files are freely available as supplementary information on our Website. Supplementary information: http://www.psb.ugent.be/bioinformatics/
Eric Bonnet, Jan Wuyts, Pierre Rouzé, Yves Van de Peer
Bioinform.3
2004 Feature selection for splice site prediction: A new method using EDA-based feature ranking
abstract
BACKGROUND: The identification of relevant biological features in large and complex datasets is an important step towards gaining insight in the processes underlying the data. Other advantages of feature selection include the ability of the classification system to attain good or even better solutions using a restricted subset of features, and a faster classification. Thus, robust methods for fast feature selection are of key importance in extracting knowledge from complex biological data. RESULTS: In this paper we present a novel method for feature subset selection applied to splice site prediction, based on estimation of distribution algorithms, a more general framework of genetic algorithms. From the estimated distribution of the algorithm, a feature ranking is derived. Afterwards this ranking is used to iteratively discard features. We apply this technique to the problem of splice site prediction, and show how it can be used to gain insight into the underlying biological process of splicing. CONCLUSION: We show that this technique proves to be more robust than the traditional use of estimation of distribution algorithms for feature selection: instead of returning a single best subset of features (as they normally do) this method provides a dynamical view of the feature selection process, like the traditional sequential wrapper methods. However, the method is faster than the traditional techniques, and scales better to datasets described by a large number of features.
Yvan Saeys, Sven Degroeve, Dirk Aeyels, Pierre Rouzé, Yves Van de Peer
BMC Bioinform.4
2003 AFLPinSilico, simulating AFLP fingerprints
abstract
Abstract Summary: A drawback of the Amplified Fragment Length Polymorphism (AFLP) fingerprinting method is the difficulty to correlate the different fragments with their DNA sequence. The AFLPinSilico application presented here simulates AFLP experiments run on either cDNA or genomic sequences, producing virtual fingerprints that allow high throughput identification of AFLP fragments. The program also enables biologists to manage experiments through simulations done beforehand, thereby reducing the number of experiments that have to be run. AFLPinSilico is available through the www or as a stand-alone version, through a command line executable (available upon request, for any platform running PERL). Availability: For academic use http://www.psb.rug.ac.be/bioinformatics/AFLPinSilico.html Contact: [email protected] * To whom correspondence should be addressed.
Stephane Rombauts, Yves Van de Peer, Pierre Rouzé
Bioinform.3
2003 Automatic design of gene-specific sequence tags for genome-wide functional studies
abstract
MOTIVATION: The availability of complete genome sequences allows the identification of short DNA segments that are specific to each annotated gene. Such unique gene sequence tags (GSTs) replace advantageously cDNAs in microarray transcript profiling experiments. In particular, probes corresponding to individual members of multigene families can be chosen carefully to avoid cross-hybridization events. RESULTS: The Specific Primer and Amplicon Design Software (SPADS) was constructed to delineate the more divergent regions in each gene by comparing them with a completely annotated genome sequence and to select optimal primer pairs for the polymerase chain reaction amplification of one divergent region per gene. SPADS is a unique integrated tool to design specific GSTs from any public or private genome sequences and allows the user to fine-tune GST size and specificity. SPADS has been used to obtain probes for whole genome and family-wide transcript profiling, as well as inserts for gene-specific knock-out experiments. AVAILABILITY: The GENOPLANTE SPADS source code and web interface are available upon request. The online version is accessible via http://genoplante-info.infobiogen.fr/spads and via http://oberon.fvms.ugent.be:8080/SPADS/
Vincent Thareau, Patrice Déhais, Carine Serizet, Pierre Hilson, Pierre Rouzé, Sébastien Aubourg
Bioinform.5
2003 Orphan gene finding - an exon assembly approach
Philippe Blayo, Pierre Rouzé, Marie-France Sagot
Theor. Comput. Sci.2
2002 INCLUSive: INtegrated Clustering, Upstream sequence retrieval and motif Sampling
abstract
Abstract Summary: INCLUSive allows automatic multistep analysis of microarray data (clustering and motif finding). The clustering algorithm (adaptive quality-based clustering) groups together genes with highly similar expression profiles. The upstream sequences of the genes belonging to a cluster are automatically retrieved from GenBank and can be fed directly into Motif Sampler, a Gibbs sampling algorithm that retrieves statistically over-represented motifs in sets of sequences, in this case upstream regions of co-expressed genes. Availability: For academic purposes at http://www.esat.kuleuven.ac.be/~dna/BioI/Software.html Contact: [email protected] * To whom correspondence should be addressed. Email: [email protected].
Gert Thijs, Yves Moreau, Frank De Smet, Janick Mathys, Magali Lescot, Stephane Rombauts, Pierre Rouzé, Bart De Moor, Kathleen Marchal
Bioinform.7
2001 A Gibbs sampling method to detect over-represented motifs in the upstream regions of co-expressed genes
abstract
Microarray experiments can reveal useful information on the transcriptional regulation. We try to find regulatory elements in the region upstream of translation start of coexpressed genes. Here we present a modification to the original Gibbs Sampling algorithm [12]. We introduce a probability distribution to estimate the number of copies of the motif in a sequence. The second modification is the incorporation of a higher-order background model. We have successfully tested our algorithm on several data sets. First we show results on two selected data set: sequences from plants containing the G-box motif and the upstream sequences from bacterial genes regulated by O2-responsive protein FNR. In both cases the motif sampler is able to find the expected motifs. Finally, the sampler is tested on 4 clusters of coexpressed genes from a wounding experiment in Arabidopsis thaliana. We find several putative motifs that are related to the pathways involved in the plant defense mechanism.
Gert Thijs, Kathleen Marchal, Magali Lescot, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau
RECOMB6
2001 A higher-order background model improves the detection of promoter regulatory elements by Gibbs sampling
abstract
MOTIVATION: Transcriptome analysis allows detection and clustering of genes that are coexpressed under various biological circumstances. Under the assumption that coregulated genes share cis-acting regulatory elements, it is important to investigate the upstream sequences controlling the transcription of these genes. To improve the robustness of the Gibbs sampling algorithm to noisy data sets we propose an extension of this algorithm for motif finding with a higher-order background model. RESULTS: Simulated data and real biological data sets with well-described regulatory elements are used to test the influence of the different background models on the performance of the motif detection algorithm. We show that the use of a higher-order model considerably enhances the performance of our motif finding algorithm in the presence of noisy data. For Arabidopsis thaliana, a reliable background model based on a set of carefully selected intergenic sequences was constructed. AVAILABILITY: Our implementation of the Gibbs sampler called the Motif Sampler can be used through a web interface: http://www.esat.kuleuven.ac.be/~thijs/Work/MotifSampler.html. CONTACT: [email protected]; [email protected]
Gert Thijs, Magali Lescot, Kathleen Marchal, Stephane Rombauts, Bart De Moor, Pierre Rouzé, Yves Moreau
Bioinform.6
1999 Evaluation of gene prediction software using a genomic data set: application to <$O_SSF>Arabidopsis thaliana<$C_SSF>sequences
abstract
MOTIVATION: The annotation of the Arabidopsis thaliana genome remains a problem in terms of time and quality. To improve the annotation process, we want to choose the most appropriate tools to use inside a computer-assisted annotation platform. We therefore need evaluation of prediction programs with Arabidopsis sequences containing multiple genes. RESULTS: We have developed AraSet, a data set of contigs of validated genes, enabling the evaluation of multi-gene models for the Arabidopsis genome. Besides conventional metrics to evaluate gene prediction at the site and the exon levels, new measures were introduced for the prediction at the protein sequence level as well as for the evaluation of gene models. This evaluation method is of general interest and could apply to any new gene prediction software and to any eukaryotic genome. The GeneMark.hmm program appears to be the most accurate software at all three levels for the Arabidopsis genomic sequences. Gene modeling could be further improved by combination of prediction software. AVAILABILITY: The AraSet sequence set, the Perl programs and complementary results and notes are available at http://sphinx.rug.ac.be:8080/biocomp/napav/. CONTACT: [email protected].
Nathalie Pavy, Stephane Rombauts, Patrice Déhais, Catherine Mathé, Ramana V. Davuluri, Philippe Leroy, Pierre Rouzé
Bioinform.7