EDBT 2026 Demo / reviewers in the wild / expert
François Radvanyi
dblp:01/2058
· DBLP profile ↗
17ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-5696-6424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 3 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 98% Computational science and engineering · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
transcriptomics |
0.4 | 1 | 2019 | Assessing reproducibility of matrix factorization methods in independent transcriptomes · Bioinform. 2019 |
Bioinformatics and computational biology
cancer genomics |
0.3 | 5 | 2019 | Assessing reproducibility of matrix factorization methods in independent transcriptomes · Bioinform. 2019 CoRegNet: reconstruction and integrated analysis of co-regulatory networks · Bioinform. 2015 Computation of recurrent minimal genomic alterations from array-CGH data · Bioinform. 2006 |
Bioinformatics and computational biology
interpretable classification |
0.3 | 1 | 2025 | MMnc: multi-modal interpretable representation for non-coding RNA classification and class annotation · Bioinform. 2025 |
Bioinformatics and computational biology › gene regulation › gene regulatory network
gene regulatory network analysis |
0.2 | 1 | 2015 | CoRegNet: reconstruction and integrated analysis of co-regulatory networks · Bioinform. 2015 |
Bioinformatics and computational biology › protein analysis
protein complex prediction |
0.2 | 1 | 2014 | Pepper: cytoscape app for protein complex expansion using protein-protein interaction networks · Bioinform. 2014 |
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis |
0.2 | 1 | 2013 | HMCan: a method for detecting chromatin modifications in cancer samples using ChIP-seq data · Bioinform. 2013 |
Bioinformatics and computational biology
epigenomics |
0.2 | 1 | 2013 | HMCan: a method for detecting chromatin modifications in cancer samples using ChIP-seq data · Bioinform. 2013 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis |
0.1 | 2 | 2006 | Computation of recurrent minimal genomic alterations from array-CGH data · Bioinform. 2006 Analysis of array CGH data: from signal ratio to gain and loss of DNA regions · Bioinform. 2004 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.1 | 1 | 2007 | LICORN: learning cooperative regulation networks from gene expression data · Bioinform. 2007 |
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network analysis |
0.1 | 1 | 2014 | Pepper: cytoscape app for protein complex expansion using protein-protein interaction networks · Bioinform. 2014 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2005 | Identifying genes from up-down properties of microarray expression series · Bioinform. 2005 |
Computational science and engineering
pattern recognition |
0.1 | 1 | 2005 | Identifying genes from up-down properties of microarray expression series · Bioinform. 2005 |
Bioinformatics and computational biology › genomics
breakpoint detection |
0.0 | 1 | 2004 | Analysis of array CGH data: from signal ratio to gain and loss of DNA regions · Bioinform. 2004 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number alteration detection |
0.0 | 1 | 2004 | Analysis of array CGH data: from signal ratio to gain and loss of DNA regions · Bioinform. 2004 |
Information retrieval › document retrieval
bibliographic retrieval |
0.0 | 1 | 2011 | Gene List significance at-a-glance with GeneValorization · Bioinform. 2011 |
Bioinformatics and computational biology › cancer genomics
copy number analysis |
0.0 | 1 | 2006 | Computation of recurrent minimal genomic alterations from array-CGH data · Bioinform. 2006 |
Methods — techniques the papers use, named apart from their topics
deep learning · 0.9attention-based multi-modal data integration · 0.9reciprocally best hit graph · 0.4matrix factorization · 0.4independent component analysis · 0.4topological analysis · 0.2multi-objective optimization · 0.2hidden markov model · 0.2copy number bias correction · 0.2GC-content bias correction · 0.2web visualization · 0.1co-occurrence matrix · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MMnc: multi-modal interpretable representation for non-coding RNA classification and class annotationabstractMOTIVATION: As the biological roles and disease implications of non-coding RNAs continue to emerge, the need to thoroughly characterize previously unexplored non-coding RNAs becomes increasingly urgent. These molecules hold potential as biomarkers and therapeutic targets. However, the vast and complex nature of non-coding RNAs data presents a challenge. We introduce MMnc, an interpretable deep-learning approach designed to classify non-coding RNAs into functional groups. MMnc leverages multiple data sources-such as the sequence, secondary structure, and expression-using attention-based multi-modal data integration. This ensures the learning of meaningful representations while accounting for missing sources in some samples. RESULTS: Our findings demonstrate that MMnc achieves high classification accuracy across diverse non-coding RNA classes. The method's modular architecture allows for the consideration of multiple types of modalities, whereas other tools only consider one or two at most. MMnc is resilient to missing data, ensuring that all available information is effectively utilized. Importantly, the generated attention scores offer interpretable insights into the underlying patterns of the different non-coding RNA classes, potentially driving future non-coding RNA research and applications. AVAILABILITY AND IMPLEMENTATION: Data and source code can be found at EvryRNA.ibisc.univ-evry.fr/EvryRNA/MMnc. Constance Creux, Farida Zehraoui, François Radvanyi, Fariza Tahi |
Bioinform. | 3 |
| 2024 | Comparison and benchmark of deep learning methods for non-coding RNA classificationabstractThe involvement of non-coding RNAs in biological processes and diseases has made the exploration of their functions crucial. Most non-coding RNAs have yet to be studied, creating the need for methods that can rapidly classify large sets of non-coding RNAs into functional groups, or classes. In recent years, the success of deep learning in various domains led to its application to non-coding RNA classification. Multiple novel architectures have been developed, but these advancements are not covered by current literature reviews. We present an exhaustive comparison of the different methods proposed in the state-of-the-art and describe their associated datasets. Moreover, the literature lacks objective benchmarks. We perform experiments to fairly evaluate the performance of various tools for non-coding RNA classification on popular datasets. The robustness of methods to non-functional sequences and sequence boundary noise is explored. We also measure computation time and CO2 emissions. With regard to these results, we assess the relevance of the different architectural choices and provide recommendations to consider in future methods. Constance Creux, Farida Zehraoui, François Radvanyi, Fariza Tahi |
PLoS Comput. Biol. | 3 |
| 2023 | Prediction of Secondary Structure for Long Non-Coding RNAs using a Recursive Cutting Method based on Deep LearningabstractAccurately predicting the secondary structure of RNA, particularly for long non-coding RNA, has direct implications in healthcare, where it can be used for diagnostic, therapeutic, and drug discovery purposes. However, the majority of previous approaches are too costly in terms of computation budget to cope with the increasing complexity of long RNAs, and the ones that can scale to long RNAs lack accuracy to reliably predict their structures. We propose a new approach combining recursive cutting and machine learning techniques for predicting the secondary structures of long non-coding RNAs. In comparison, our method proves to be computationally efficient by recursively partitioning a sequence into smaller fragments until they can be easily managed by an existing model. We perform a benchmark of different state-of-the-art models and show that our approach indeed demonstrates better performance for long RNAs and a potential to bring significant improvements in the future, as well as interesting enhancing properties, which we discuss. Loïc Omnes, Eric Angel, Pierre Bartet, François Radvanyi, Fariza Tahi |
BIBE | 4 |
| 2019 | Assessing reproducibility of matrix factorization methods in independent transcriptomesabstractMOTIVATION: Matrix factorization (MF) methods are widely used in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). MF algorithms have never been compared based on the between-datasets reproducibility of their outputs in similar independent datasets. Lack of this knowledge might have a crucial impact when generalizing the predictions made in a study to others. RESULTS: We systematically test widely used MF methods on several transcriptomic datasets collected from the same cancer type (14 colorectal, 8 breast and 4 ovarian cancer transcriptomic datasets). Inspired by concepts of evolutionary bioinformatics, we design a novel framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the MF methods for their ability to produce generalizable components. We show that a particular protocol of application of independent component analysis (ICA), accompanied by a stabilization procedure, leads to a significant increase in the between-datasets reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other standard methods. We developed a user-friendly tool for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors associated to biological processes or to technological artifacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping. AVAILABILITY AND IMPLEMENTATION: The RBH construction tool is available from http://goo.gl/DzpwYp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Laura Cantini, Ulykbek Kairov, Aurélien de Reyniès, Emmanuel Barillot, François Radvanyi, Andrei Yu. Zinovyev |
Bioinform. | 5 |
| 2017 | SegCorr a statistical procedure for the detection of genomic regions of correlated expressionabstractBACKGROUND: Detecting local correlations in expression between neighboring genes along the genome has proved to be an effective strategy to identify possible causes of transcriptional deregulation in cancer. It has been successfully used to illustrate the role of mechanisms such as copy number variation (CNV) or epigenetic alterations as factors that may significantly alter expression in large chromosomal regions (gene silencing or gene activation). RESULTS: The identification of correlated regions requires segmenting the gene expression correlation matrix into regions of homogeneously correlated genes and assessing whether the observed local correlation is significantly higher than the background chromosomal correlation. A unified statistical framework is proposed to achieve these two tasks, where optimal segmentation is efficiently performed using dynamic programming algorithm, and detection of highly correlated regions is then achieved using an exact test procedure. We also propose a simple and efficient procedure to correct the expression signal for mechanisms already known to impact expression correlation. The performance and robustness of the proposed procedure, called SegCorr, are evaluated on simulated data. The procedure is illustrated on cancer data, where the signal is corrected for correlations caused by copy number variation. It permitted the detection of regions with high correlations linked to epigenetic marks like DNA methylation. CONCLUSIONS: SegCorr is a novel method that performs correlation matrix segmentation and applies a test procedure in order to detect highly correlated regions in gene expression. Eleni Ioanna Delatola, Emilie Lebarbier, Tristan Mary-Huard, François Radvanyi, Stéphane Robin, Jennifer Wong |
BMC Bioinform. | 4 |
| 2015 | CoRegNet: reconstruction and integrated analysis of co-regulatory networksabstractUNLABELLED: CoRegNet is an R/Bioconductor package to analyze large-scale transcriptomic data by highlighting sets of co-regulators. Based on a transcriptomic dataset, CoRegNet can be used to: reconstruct a large-scale co-regulatory network, integrate regulation evidences such as transcription factor binding sites and ChIP data, estimate sample-specific regulator activity, identify cooperative transcription factors and analyze the sample-specific combinations of active regulators through an interactive visualization tool. In this study CoRegNet was used to identify driver regulators of bladder cancer. AVAILABILITY: CoRegNet is available at http://bioconductor.org/packages/CoRegNet CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rémy Nicolle, François Radvanyi, Mohamed Elati |
Bioinform. | 2 |
| 2014 | Pepper: cytoscape app for protein complex expansion using protein-protein interaction networksabstractUNLABELLED: We introduce Pepper (Protein complex Expansion using Protein-Protein intERactions), a Cytoscape app designed to identify protein complexes as densely connected subnetworks from seed lists of proteins derived from proteomic studies. Pepper identifies connected subgraph by using multi-objective optimization involving two functions: (i) the coverage, a solution must contain as many proteins from the seed as possible, (ii) the density, the proteins of a solution must be as connected as possible, using only interactions from a proteome-wide interaction network. Comparisons based on gold standard yeast and human datasets showed Pepper's integrative approach as superior to standard protein complex discovery methods. The visualization and interpretation of the results are facilitated by an automated post-processing pipeline based on topological analysis and data integration about the predicted complex proteins. Pepper is a user-friendly tool that can be used to analyse any list of proteins. AVAILABILITY: Pepper is available from the Cytoscape plug-in manager or online (http://apps.cytoscape.org/apps/pepper) and released under GNU General Public License version 3. Charles Winterhalter, Rémy Nicolle, A. Louis, Cuong To, François Radvanyi, Mohamed Elati |
Bioinform. | 5 |
| 2013 | HMCan: a method for detecting chromatin modifications in cancer samples using ChIP-seq dataabstractMOTIVATION: Cancer cells are often characterized by epigenetic changes, which include aberrant histone modifications. In particular, local or regional epigenetic silencing is a common mechanism in cancer for silencing expression of tumor suppressor genes. Though several tools have been created to enable detection of histone marks in ChIP-seq data from normal samples, it is unclear whether these tools can be efficiently applied to ChIP-seq data generated from cancer samples. Indeed, cancer genomes are often characterized by frequent copy number alterations: gains and losses of large regions of chromosomal material. Copy number alterations may create a substantial statistical bias in the evaluation of histone mark signal enrichment and result in underdetection of the signal in the regions of loss and overdetection of the signal in the regions of gain. RESULTS: We present HMCan (Histone modifications in cancer), a tool specially designed to analyze histone modification ChIP-seq data produced from cancer genomes. HMCan corrects for the GC-content and copy number bias and then applies Hidden Markov Models to detect the signal from the corrected data. On simulated data, HMCan outperformed several commonly used tools developed to analyze histone modification data produced from genomes without copy number alterations. HMCan also showed superior results on a ChIP-seq dataset generated for the repressive histone mark H3K27me3 in a bladder cancer cell line. HMCan predictions matched well with experimental data (qPCR validated regions) and included, for example, the previously detected H3K27me3 mark in the promoter of the DLEC1 gene, missed by other tools we tested. Haitham Ashoor, Aurélie Hérault, Aurélie Kamoun, François Radvanyi, Vladimir B. Bajic, Emmanuel Barillot, Valentina Boeva |
Bioinform. | 4 |
| 2013 | Integrative Modelling of the Influence of MAPK Network on Cancer Cell Fate DecisionabstractThe Mitogen-Activated Protein Kinase (MAPK) network consists of tightly interconnected signalling pathways involved in diverse cellular processes, such as cell cycle, survival, apoptosis and differentiation. Although several studies reported the involvement of these signalling cascades in cancer deregulations, the precise mechanisms underlying their influence on the balance between cell proliferation and cell death (cell fate decision) in pathological circumstances remain elusive. Based on an extensive analysis of published data, we have built a comprehensive and generic reaction map for the MAPK signalling network, using CellDesigner software. In order to explore the MAPK responses to different stimuli and better understand their contributions to cell fate decision, we have considered the most crucial components and interactions and encoded them into a logical model, using the software GINsim. Our logical model analysis particularly focuses on urinary bladder cancer, where MAPK network deregulations have often been associated with specific phenotypes. To cope with the combinatorial explosion of the number of states, we have applied novel algorithms for model reduction and for the compression of state transition graphs, both implemented into the software GINsim. The results of systematic simulations for different signal combinations and network perturbations were found globally coherent with published data. In silico experiments further enabled us to delineate the roles of specific components, cross-talks and regulatory feedbacks in cell fate decision. Finally, tentative proliferative or anti-proliferative mechanisms can be connected with established bladder cancer deregulations, namely Epidermal Growth Factor Receptor (EGFR) over-expression and Fibroblast Growth Factor Receptor 3 (FGFR3) activating mutations. Luca Grieco, Laurence Calzone, Isabelle Bernard-Pierrot, François Radvanyi, Brigitte Kahn-Perlès, Denis Thieffry |
PLoS Comput. Biol. | 4 |
| 2012 | Network Transformation of Gene Expression for Feature ExtractionabstractClassical approaches to analyze transcriptomic data usually produce average classification models that have very low reproducibility. In this work, genome wide gene expression is considered through the activity of large regulatory networks. We introduce a new measure of regulatory influence based on the variations of expression of genes in a large inferred regulatory network. This methodology can be used to transform transcriptomic data into a smaller influence data set on which feature selection and classification models show similar predictive performance and increased stability and reproducibility, especially when comparing models trained on different datasets. The methodology was tested on two distinct bladder cancer data sets. Rémy Nicolle, Mohamed Elati, François Radvanyi |
ICMLA (1) | 3 |
| 2011 | Gene List significance at-a-glance with GeneValorizationabstractMOTIVATION: High-throughput technologies provide fundamental informations concerning thousands of genes. Many of the current research laboratories daily use one or more of these technologies and end-up with lists of genes. Assessing the originality of the results obtained includes being aware of the number of publications available concerning individual or multiple genes and accessing information about these publications. Faced with the exponential growth of publications avaliable and number of genes involved in a study, this task is becoming particularly difficult to achieve. RESULTS: We introduce GeneValorization, a web-based tool that gives a clear and handful overview of the bibliography available corresponding to the user input formed by (i) a gene list (expressed by gene names or ids from EntrezGene) and (ii) a context of study (expressed by keywords). From this input, GeneValorization provides a matrix containing the number of publications with co-occurrences of gene names and keywords. Graphics are automatically generated to assess the relative importance of genes within various contexts. Links to publications and other databases offering information on genes and keywords are also available. To illustrate how helpful GeneValorization is, we will consider the gene list of the OncotypeDX prognostic marker test. AVAILABILITY: http://bioguide-project.net/gv CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bryan Brancotte, Anne Biton, Isabelle Bernard-Pierrot, François Radvanyi, Fabien Reyal, Sarah Cohen Boulakia |
Bioinform. | 4 |
| 2007 | LICORN: learning cooperative regulation networks from gene expression dataabstractMOTIVATION: One of the most challenging tasks in the post-genomic era is the reconstruction of transcriptional regulation networks. The goal is to identify, for each gene expressed in a particular cellular context, the regulators affecting its transcription, and the co-ordination of several regulators in specific types of regulation. DNA microarrays can be used to investigate relationships between regulators and their target genes, through simultaneous observations of their RNA levels. RESULTS: We propose a data mining system for inferring transcriptional regulation relationships from RNA expression values. This system is particularly suitable for the detection of cooperative transcriptional regulation. We model regulatory relationships as labelled two-layer gene regulatory networks, and describe a method for the efficient learning of these bipartite networks from discretized expression data sets. We also evaluate the statistical significance of such inferred networks and validate our methods on two public yeast expression data sets. AVAILABILITY: http://www.lri.fr/~elati/licorn.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mohamed Elati, Pierre Neuvial, Monique Bolotin-Fukuhara, Emmanuel Barillot, François Radvanyi, Céline Rouveirol |
Bioinform. | 5 |
| 2006 | VAMP: Visualization and analysis of array-CGH, transcriptome and other molecular profilesabstractMOTIVATION: Microarray-based CGH (Comparative Genomic Hybridization), transcriptome arrays and other large-scale genomic technologies are now routinely used to generate a vast amount of genomic profiles. Exploratory analysis of this data is crucial in helping to understand the data and to help form biological hypotheses. This step requires visualization of the data in a meaningful way to visualize the results and to perform first level analyses. RESULTS: We have developed a graphical user interface for visualization and first level analysis of molecular profiles. It is currently in use at the Institut Curie for cancer research projects involving CGH arrays, transcriptome arrays, SNP (single nucleotide polymorphism) arrays, loss of heterozygosity results (LOH), and Chromatin ImmunoPrecipitation arrays (ChIP chips). The interface offers the possibility of studying these different types of information in a consistent way. Several views are proposed, such as the classical CGH karyotype view or genome-wide multi-tumor comparison. Many functionalities for analyzing CGH data are provided by the interface, including looking for recurrent regions of alterations, confrontation to transcriptome data or clinical information, and clustering. Our tool consists of PHP scripts and of an applet written in Java. It can be run on public datasets at http://bioinfo.curie.fr/vamp AVAILABILITY: The VAMP software (Visualization and Analysis of array-CGH,transcriptome and other Molecular Profiles) is available upon request. It can be tested on public datasets at http://bioinfo.curie.fr/vamp. The documentation is available at http://bioinfo.curie.fr/vamp/doc. Philippe La Rosa, Eric Viara, Philippe Hupé, Gaëlle Pierron, Stéphane Liva, Pierre Neuvial, Isabel Brito 0002, Séverine Lair, Nicolas Servant, Nicolas Robine, Elodie Manié, Caroline Brennetot, Isabelle Janoueix-Lerosey, Virginie Raynal, Nadège Gruel, Céline Rouveirol, Nicolas Stransky, Marc-Henri Stern, Olivier Delattre, Alain Aurias, François Radvanyi, Emmanuel Barillot |
Bioinform. | 21 |
| 2006 | Computation of recurrent minimal genomic alterations from array-CGH dataabstractMOTIVATION: The identification of recurrent genomic alterations can provide insight into the initiation and progression of genetic diseases, such as cancer. Array-CGH can identify chromosomal regions that have been gained or lost, with a resolution of approximately 1 mb, for the cutting-edge techniques. The extraction of discrete profiles from raw array-CGH data has been studied extensively, but subsequent steps in the analysis require flexible, efficient algorithms, particularly if the number of available profiles exceeds a few tens or the number of array probes exceeds a few thousands. RESULTS: We propose two algorithms for computing minimal and minimal constrained regions of gain and loss from discretized CGH profiles. The second of these algorithms can handle additional constraints describing relevant regions of copy number change. We have validated these algorithms on two public array-CGH datasets. AVAILABILITY: From the authors, upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Céline Rouveirol, Nicolas Stransky, Philippe Hupé, Philippe La Rosa, Eric Viara, Emmanuel Barillot, François Radvanyi |
Bioinform. | 7 |
| 2006 | Spatial normalization of array-CGH dataabstractBACKGROUND: Array-based comparative genomic hybridization (array-CGH) is a recently developed technique for analyzing changes in DNA copy number. As in all microarray analyses, normalization is required to correct for experimental artifacts while preserving the true biological signal. We investigated various sources of systematic variation in array-CGH data and identified two distinct types of spatial effect of no biological relevance as the predominant experimental artifacts: continuous spatial gradients and local spatial bias. Local spatial bias affects a large proportion of arrays, and has not previously been considered in array-CGH experiments. RESULTS: We show that existing normalization techniques do not correct these spatial effects properly. We therefore developed an automatic method for the spatial normalization of array-CGH data. This method makes it possible to delineate and to eliminate and/or correct areas affected by spatial bias. It is based on the combination of a spatial segmentation algorithm called NEM (Neighborhood Expectation Maximization) and spatial trend estimation. We defined quality criteria for array-CGH data, demonstrating significant improvements in data quality with our method for three data sets coming from two different platforms (198, 175 and 26 BAC-arrays). CONCLUSION: We have designed an automatic algorithm for the spatial normalization of BAC CGH-array data, preventing the misinterpretation of experimental artifacts as biologically relevant outliers in the genomic profile. This algorithm is implemented in the R package MANOR (Micro-Array NORmalization), which is described at http://bioinfo.curie.fr/projects/manor and available from the Bioconductor site http://www.bioconductor.org. It can also be tested on the CAPweb bioinformatics platform at http://bioinfo.curie.fr/CAPweb. Pierre Neuvial, Philippe Hupé, Isabel Brito 0002, Stéphane Liva, Elodie Manié, Caroline Brennetot, François Radvanyi, Alain Aurias, Emmanuel Barillot |
BMC Bioinform. | 7 |
| 2005 | Identifying genes from up-down properties of microarray expression seriesabstractMOTIVATION: We consider any collection of microarrays that can be ordered to form a progression; for example, as a function of time, severity of disease or dose of a stimulant. By plotting the expression level of each gene as a function of time, or severity, or dose, we form an expression series, or curve, for each gene. While most of these curves will exhibit random fluctuations, some will contain a pattern, and these are the genes that are most likely associated with the quantity used to order them. RESULTS: We introduce a method of identifying the pattern and hence genes in microarray expression curves without knowing what kind of pattern to look for. Key to our approach is the sequence of ups and downs formed by pairs of consecutive data points in each curve. As a benchmark, we blindly identified genes from yeast cell cycles without selecting for periodic or any other anticipated behaviour. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: The complete versions of Table 2 and Figure 4, as well as other material, can be found at http://www.lps.ens.fr/~willbran/up-down/ or http://www.tcm.phy.cam.ac.uk/~tmf20/up-down/ Karen Willbrand, François Radvanyi, Jean-Pierre Nadal, Jean-Paul Thiery, Thomas M. A. Fink |
Bioinform. | 2 |
| 2004 | Analysis of array CGH data: from signal ratio to gain and loss of DNA regionsabstractMOTIVATION: Genomic DNA regions are frequently lost or gained during tumor progression. Array Comparative Genomic Hybridization (array CGH) technology makes it possible to assess these changes in DNA in cancers, by comparison with a normal reference. The identification of systematically deleted or amplified genomic regions in a set of tumors enables biologists to identify genes involved in cancer progression because tumor suppressor genes are thought to be located in lost genomic regions and oncogenes, in gained regions. Array CGH profiles should also improve the classification of tumors. The achievement of these goals requires a methodology for detecting the breakpoints delimiting altered regions in genomic patterns and assigning a status (normal, gained or lost) to each chromosomal region. RESULTS: We have developed a methodology for the automatic detection of breakpoints from array CGH profile, and the assignment of a status to each chromosomal region. The breakpoint detection step is based on the Adaptive Weights Smoothing (AWS) procedure and provides highly convincing results: our algorithm detects 97, 100 and 94% of breakpoints in simulated data, karyotyping results and manually analyzed profiles, respectively. The percentage of correctly assigned statuses ranges from 98.9 to 99.8% for simulated data and is 100% for karyotyping results. Our algorithm also outperforms other solutions on a public reference dataset. AVAILABILITY: The R package GLAD (Gain and Loss Analysis of DNA) is available upon request. Philippe Hupé, Nicolas Stransky, Jean-Paul Thiery, François Radvanyi, Emmanuel Barillot |
Bioinform. | 4 |