Géraldine Jean

dblp:46/3857 · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0002-1534-2682ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 9 since 2021Theory of computation · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GSI: A New Approach to the Protein Inference Problem
abstract
The protein inference problem, i.e., determining which proteins are present in a biological sample, is key to understanding the roles of proteins and, more broadly, many biological processes. Protein identification is typically achieved by first cleaving proteins into smaller sequences called peptides. Peptides are then identified using tandem mass spectrometry, a process that produces mass spectra, and in which peptide identification consists of associating, via dedicated tools, a mass spectrum to a peptide sequence. Protein inference consists of identifying, from a list of identified peptides, the proteins that most likely produced them, and were therefore present in the original sample. Usually, peptide identification and protein inference are two separate steps, which are sequentially achieved. However, by proceeding in such a way, a significant amount of potentially useful information contained in the spectra may be discarded in the second step. Moreover, AI-based tools can now predict the likelihood of a peptide’s identification when its parent protein is present in the sample. In this paper, we present the Global Spectrum Interpretation (GSI) model, a protein inference model that integrates all this information to produce more accurate protein identifications. We show that GSI is NP-hard and provide a Mixed Integer Linear Program (MILP) formulation for it. This MILP is then benchmarked against state-of-the-art protein inference models on several datasets. Our results show that GSI’s promising and original approach achieves performance comparable to current models and outperforms other widely used ones, while being more explainable.
Aurélien Berthier, Emile Benoist, Guillaume Fertin, Géraldine Jean
WABI4
2025 Partition Based Algorithms for Rearrangement Distances With Flexible Intergenic Regions
abstract
Genome Rearrangement distance problems are used in Computational Biology to estimate the evolutionary distance between genomes. These problems consist of minimizing the number of rearrangement events necessary to transform one genome into another. Two commonly used rearrangement events are reversal and transposition. The first studied problems ignored nucleotides outside genes (called intergenic regions), or assumed that genomes have a single copy of each gene. Recent works made advancements in more general problems considering the number of nucleotides in intergenic regions, and replicated genes. Nevertheless, genomes tend to have wildly different quantities of nucleotides on their intergenic regions, which poses a problem when comparing these regions exactly. To overcome this limitation, our work considers some flexibility when matching intergenic regions that do not have the same number of nucleotides. We propose new problems seeking the minimum number of reversals, or reversals and transpositions, necessary to transform one genome into another, while considering flexible intergenic region information. We show approximations for these problems by exploring their relationship with the Signed Minimum Common Flexible Intergenic String Partition problem. We also present different heuristics for the partition problem, and conduct experimental tests on simulated genomes to assess the performance of our algorithms.
Gabriel Siqueira, Alexsandro Oliveira Alexandrino, Andre Rodrigues Oliveira, Géraldine Jean, Guillaume Fertin, Zanoni Dias
IEEE Trans. Comput. Biol. Bioinform.4
2025 The Exact Subset MultiCover problem
Emile Benoist, Guillaume Fertin, Géraldine Jean
Theor. Comput. Sci.3
2025 Sorting genomes by prefix double-cut-and-joins
abstract
In this paper, we study the problem of sorting unichromosomal linear genomes by prefix double-cut-and-joins (or DCJs) in both the signed and the unsigned settings. Prefix DCJs cut the leftmost segment of a genome and any other segment, and recombine the severed endpoints in one of two possible ways: one of these options corresponds to a prefix reversal, which reverses the order of elements between the two cuts (as well as their signs in the signed case). Our main results are: (1) new structural lower bounds based on the breakpoint graph for sorting by unsigned prefix reversals, unsigned prefix DCJs, and signed prefix DCJs; (2) two polynomial-time algorithms for sorting by prefix DCJs, both in the signed case (which answers an open question of Labarre [1] ) and in the unsigned case; (3) a 1-absolute approximation algorithm for sorting by unsigned prefix reversals for a specific class of permutations.
Guillaume Fertin, Géraldine Jean, Anthony Labarre
Theor. Comput. Sci.2
2024 The Maximum Zero-Sum Partition problem
abstract
We study the Maximum Zero-Sum Partition problem (or MZSP ), defined as follows: given a multiset S = { a 1 , a 2 , … , a n } of integers a i ∈ Z ⁎ (where Z ⁎ denotes the set of non-zero integers) such that ∑ i = 1 n a i = 0 , find a maximum cardinality partition { S 1 , S 2 , … , S k } of S such that, for every 1 ≤ i ≤ k , ∑ a j ∈ S i a j = 0 . Solving MZSP is useful in genomics for computing evolutionary distances between pairs of species. Our contributions are a series of algorithmic results concerning MZSP , in terms of complexity, (in)approximability, with a particular focus on the fixed-parameter tractability of MZSP with respect to either (i) the size k of the solution, (ii) the number of negative (resp. positive) values in S and (iii) the largest integer in S .
Guillaume Fertin, Oscar Fontaine, Géraldine Jean, Stéphane Vialette
Theor. Comput. Sci.3
2023 Approximating Rearrangement Distances with Replicas and Flexible Intergenic Regions
Gabriel Siqueira, Alexsandro Oliveira Alexandrino, Andre Rodrigues Oliveira, Géraldine Jean, Guillaume Fertin, Zanoni Dias
ISBRA4
2023 Fast alignment of mass spectra in large proteomics datasets, capturing dissimilarities arising from multiple complex modifications of peptides
abstract
BACKGROUND: In proteomics, the interpretation of mass spectra representing peptides carrying multiple complex modifications remains challenging, as it is difficult to strike a balance between reasonable execution time, a limited number of false positives, and a huge search space allowing any number of modifications without a priori. The scientific community needs new developments in this area to aid in the discovery of novel post-translational modifications that may play important roles in disease. RESULTS: To make progress on this issue, we implemented SpecGlobX (SpecGlob eXTended to eXperimental spectra), a standalone Java application that quickly determines the best spectral alignments of a (possibly very large) list of Peptide-to-Spectrum Matches (PSMs) provided by any open modification search method, or generated by the user. As input, SpecGlobX reads a file containing spectra in MGF or mzML format and a semicolon-delimited spreadsheet describing the PSMs. SpecGlobX returns the best alignment for each PSM as output, splitting the mass difference between the spectrum and the peptide into one or more shifts while considering the possibility of non-aligned masses (a phenomenon resulting from many situations including neutral losses). SpecGlobX is fast, able to align one million PSMs in about 1.5 min on a standard desktop. Firstly, we remind the foundations of the algorithm and detail how we adapted SpecGlob (the method we previously developed following the same aim, but limited to the interpretation of perfect simulated spectra) to the interpretation of imperfect experimental spectra. Then, we highlight the interest of SpecGlobX as a complementary tool downstream to three open modification search methods on a large simulated spectra dataset. Finally, we ran SpecGlobX on a proteome-wide dataset downloaded from PRIDE to demonstrate that SpecGlobX functions just as well on simulated and experimental spectra. We then carefully analyzed a limited set of interpretations. CONCLUSIONS: SpecGlobX is helpful as a decision support tool, providing keys to interpret peptides carrying complex modifications still poorly considered by current open modification search software. Better alignment of PSMs enhances confidence in the identification of spectra provided by open modification search methods and should improve the interpretation rate of spectra.
Grégoire Prunier, Mehdi Cherkaoui, Albane Lysiak, Olivier Langella, Mélisande Blein-Nicolas, Virginie Lollier, Emile Benoist, Géraldine Jean, Guillaume Fertin, Hélène Rogniaux, Dominique Tessier
BMC Bioinform.8
2022 Transposition Distance Considering Intergenic Regions for Unbalanced Genomes
Alexsandro Oliveira Alexandrino, Andre Rodrigues Oliveira, Géraldine Jean, Guillaume Fertin, Ulisses Dias, Zanoni Dias
ISBRA3
2022 Sorting Genomes by Prefix Double-Cut-and-Joins
Guillaume Fertin, Géraldine Jean, Anthony Labarre
SPIRE2
2022 The Exact Subset MultiCover Problem
Emile Benoist, Guillaume Fertin, Géraldine Jean
TAMC3
2021 Sorting by Multi-cut Rearrangements
Laurent Bulteau, Guillaume Fertin, Géraldine Jean, Christian Komusiewicz
SOFSEM3
2021 Evaluation of open search methods based on theoretical mass spectra comparison
abstract
BACKGROUND: Mass spectrometry remains the privileged method to characterize proteins. Nevertheless, most of the spectra generated by an experiment remain unidentified after their analysis, mostly because of the modifications they carry. Open Modification Search (OMS) methods offer a promising answer to this problem. However, assessing the quality of OMS identifications remains a difficult task. METHODS: Aiming at better understanding the relationship between (1) similarity of pairs of spectra provided by OMS methods and (2) relevance of their corresponding peptide sequences, we used a dataset composed of theoretical spectra only, on which we applied two OMS strategies. We also introduced two appropriately defined measures for evaluating the above mentioned spectra/sequence relevance in this context: one is a color classification representing the level of difficulty to retrieve the proper sequence of the peptide that generated the identified spectrum ; the other, called LIPR, is the proportion of common masses, in a given Peptide Spectrum Match (PSM), that represent dissimilar sequences. These two measures were also considered in conjunction with the False Discovery Rate (FDR). RESULTS: According to our measures, the strategy that selects the best candidate by taking the mass difference between two spectra into account yields better quality results. Besides, although the FDR remains an interesting indicator in OMS methods (as shown by LIPR), it is questionable: indeed, our color classification shows that a non negligible proportion of relevant spectra/sequence interpretations corresponds to PSMs coming from the decoy database. CONCLUSIONS: The three above mentioned measures allowed us to clearly determine which of the two studied OMS strategies outperformed the other, both in terms of number of identifications and of accuracy of these identifications. Even though quality evaluation of PSMs in OMS methods remains challenging, the study of theoretical spectra is a favorable framework for going further in this direction.
Albane Lysiak, Guillaume Fertin, Géraldine Jean, Dominique Tessier
BMC Bioinform.3
2021 Sorting Signed Permutations by Intergenic Reversals
abstract
Genome rearrangements are mutations affecting large portions of a genome, and a reversal is one of the most studied genome rearrangements in the literature through the Sorting by Reversals (SbR) problem. SbR is solvable in polynomial time on signed permutations (i.e., the gene orientation is known), and it is NP-hard on unsigned permutations. This problem (and many others considering genome rearrangements) models genome as a list of its genes in the order they appear, ignoring all other information present in the genome. Recent works claimed that the incorporation of the size of intergenic regions, i.e., sequences of nucleotides between genes, may result in better estimators for the real distance between genomes. Here we introduce the Sorting Signed Permutations by Intergenic Reversals problem, that sorts a signed permutation using reversals both on gene order and intergenic sizes. We show that this problem is NP-hard by a reduction from the 3-partition problem. Then, we propose a 2-approximation algorithm for it. Finally, we also incorporate intergenic indels (i.e., insertions or deletions of intergenic regions) to overcome a limitation of sorting by conservative events (such as reversals) and propose two approximation algorithms.
Andre Rodrigues Oliveira, Géraldine Jean, Guillaume Fertin, Klairton Lima Brito, Laurent Bulteau, Ulisses Dias, Zanoni Dias
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Sorting Permutations by Intergenic Operations
abstract
Genome Rearrangements are events that affect large stretches of genomes during evolution. Many mathematical models have been used to estimate the evolutionary distance between two genomes based on genome rearrangements. However, most of them focused on the (order of the) genes of a genome, disregarding other important elements in it. Recently, researchers have shown that considering regions between each pair of genes, called intergenic regions, can enhance distance estimation in realistic data. Two of the most studied genome rearrangements are the reversal, which inverts a sequence of genes, and the transposition, which occurs when two adjacent gene sequences swap their positions inside the genome. In this work, we study the transposition distance between two genomes, but we also consider intergenic regions, a problem we name Sorting by Intergenic Transpositions. We show that this problem is NP-hard and propose two approximation algorithms, with factors 3.5 and 2.5, considering two distinct definitions for the problem. We also investigate the signed reversal and transposition distance between two genomes considering their intergenic regions. This second problem is called Sorting by Signed Intergenic Reversals and Intergenic Transpositions. We show that this problem is NP-hard and develop two approximation algorithms, with factors 3 and 2.5. We check how these algorithms behave when assigning weights for genome rearrangements. Finally, we implemented all these algorithms and tested them on real and simulated data.
Andre Rodrigues Oliveira, Géraldine Jean, Guillaume Fertin, Klairton Lima Brito, Ulisses Dias, Zanoni Dias
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 The Maximum Colorful Arborescence problem: How (computationally) hard can it be?
Guillaume Fertin, Julien Fradin, Géraldine Jean
Theor. Comput. Sci.3
2019 Sorting by Reversals, Transpositions, and Indels on Both Gene Order and Intergenic Sizes
Klairton Lima Brito, Géraldine Jean, Guillaume Fertin, Andre Rodrigues Oliveira, Ulisses Dias, Zanoni Dias
ISBRA2
2018 Prefix and suffix reversals on strings
Guillaume Fertin, Loïc Jankowiak, Géraldine Jean
Discret. Appl. Math.3
2017 Algorithmic Aspects of the Maximum Colorful Arborescence Problem
Guillaume Fertin, Julien Fradin, Géraldine Jean
TAMC3
2016 Genome Rearrangements on Both Gene Order and Intergenic Regions
Guillaume Fertin, Géraldine Jean, Eric Tannier
WABI2
2015 Prefix and Suffix Reversals on Strings
Guillaume Fertin, Loïc Jankowiak, Géraldine Jean
SPIRE3
2014 DExTaR: Detection of exact tandem repeats based on the de Bruijn graph
abstract
Genomes present various types of repeated structures having important roles in the mechanism of evolution. In particular, tandem repeats are analysed for their impact on genetic backgrounds of inherited diseases. However, the main objective of today's de novo assemblers is to output long, high-quality, assembled sequences; to this end, they use heuristic-based assembling procedures, which can leave many repeated regions unassembled - and in particular exact tandem repeats - due to the genomes complexity. In this paper, we propose an effective method, called DExTaR, that improves the detection of exact tandem repeats (ETRs) in any de novo de Bruijn assembly. DExTaR is based on a de Bruijn graph constructed by an assembler and retrieves ETRs left unassembled. When used with the well-known assembler ABySS, we show that DExTaR is able to obtain high quality results in terms of number and length of the detected ETRs.
Guillaume Fertin, Géraldine Jean, Andreea Radulescu, Irena Rusu
BIBM2
2014 Oqtans: the RNA-seq workbench in the cloud for complete and reproducible quantitative transcriptome analysis
abstract
We present Oqtans, an open-source workbench for quantitative transcriptome analysis, that is integrated in Galaxy. Its distinguishing features include customizable computational workflows and a modular pipeline architecture that facilitates comparative assessment of tool and data quality. Oqtans integrates an assortment of machine learning-powered tools into Galaxy, which show superior or equal performance to state-of-the-art tools. Implemented tools comprise a complete transcriptome analysis workflow: short-read alignment, transcript identification/quantification and differential expression analysis. Oqtans and Galaxy facilitate persistent storage, data exchange and documentation of intermediate results and analysis workflows. We illustrate how Oqtans aids the interpretation of data from different experiments in easy to understand use cases. Users can easily create their own workflows and extend Oqtans by integrating specific tools. Oqtans is available as (i) a cloud machine image with a demo instance at cloud.oqtans.org, (ii) a public Galaxy instance at galaxy.cbio.mskcc.org, (iii) a git repository containing all installed software (oqtans.org/git); most of which is also available from (iv) the Galaxy Toolshed and (v) a share string to use along with Galaxy CloudMan.
Vipin T. Sreedharan, Sebastian J. Schultheiß, Géraldine Jean, André Kahles, Regina Bohnert, Philipp Drewe, Pramod Mudrakarta, Nico Görnitz, Georg Zeller, Gunnar Rätsch
Bioinform.3
2014 Oqtans: a multifunctional workbench for RNA-seq data analysis
abstract
The current revolution in sequencing technologies allows us to obtain a much more detailed picture of transcriptomes via deep RNA Sequencing (RNA-Seq). In considering the full complement of RNA transcripts that comprise the transcriptome, two important analytical questions emerge: what is the abundance of RNA transcripts and which genes or transcripts are differentially expressed. In parallel with developing sequencing technologies, data analysis software is also constantly updated to improve accuracy and sensitivity while minimizing run times. The abundance of software programs, however, can be prohibitive and confusing for researchers evaluating RNA-Seq analysis pipelines. We present an open-source workbench, Oqtans , that can be integrated into the Galaxy framework that enables researchers to set up a computational pipeline for quantitative transcriptome analysis. Its distinguishing features include a modular pipeline architecture, which facilitates comparative assessment of tool and data quality. Within Oqtans , the Galaxy ’s workflow architecture enables direct comparison of several tools. Furthermore, it is straightforward to compare the performance of different programs and parameter settings on the same data and choose the best suited for the task. Oqtans analysis pipelines are easy to set up, modify, and (re-)use without significant computational skill. Oqtans integrates more than twenty sophisticated tools that perform very well compared to the state-of-the-art for transcript identification, quantification and differential expression analysis. The toolsuite contains several tools developed in the Rätsch Laboratory, but the majority of the tools were developed by other groups. In particular, we provide tools for read alignment (bwa, STAR, TopHat, PALMapper, …), transcript prediction (cufflinks, Trinity, Scripture, …) and quantitative analyses (DESeq2, edgeR, rDiff, rQuant, …). In addition, we provide tools for alignment filtering (RNA-geeq toolbox), GFF file processing (GFF toolbox) and tools for predictive sequence analysis (EasySVM, ASP, ARTS, …). See http://oqtans.org/tools for more details on included tools. Oqtans is integrated into the publicly available Galaxy server http://galaxy.cbio.mskcc.org which is maintained by the Rätsch Laboratory. It is also available as source code in a public GitHub repository http://bioweb.me/oqtans/git and as a machine image (managed by Galaxy CloudMan) for the Amazon Web Service cloud environment (instructions available at http://oqtans.org ). Oqtans sets a new standard in terms of reproducibility and builds upon Galaxy’s features to facilitate persistent storage, exchange and documentation of intermediate results and analysis workflows. Support: [email protected] Contact: [email protected] Details: http://oqtans.org Public Computing Server: http://galaxy.cbio.mskcc.org oqtans Demo Server: http://cloud.oqtans.org oqtans Amazon Machine Image: ami-65376a0c License: GPL http://www.gnu.org/licenses/gpl.html
Vipin T. Sreedharan, Sebastian J. Schultheiß, Géraldine Jean, André Kahles, Regina Bohnert, Philipp Drewe, Pramod Mudrakarta, Nico Görnitz, Georg Zeller, Gunnar Rätsch
BMC Bioinform.3
2011 Oqtans: a Galaxy-integrated workflow for quantitative transcriptome analysis from NGS Data
abstract
The current revolution in sequencing technologies allows us to obtain a much more detailed picture of transcriptomes via RNA-Sequencing. We have developed the first integrative online platform, oqtans, for quantitatively analyzing RNA-Seq experiments. Our approach of providing a self-contained machine image with the accessible, transparent Galaxy framework [ 1 ] minimizes the risk of using a third-party web service for data analysis. These services often disappear a few years after publication and render results irreproducible [ 2 ]. With oqtans, bioinformatics becomes reproducible by providing analysis building blocks for a customized workflow of read mapping, transcript reconstruction and quantitation as well as differential expression analysis. Oqtans includes a comprehensive machine-learning-powered toolsuite developed by the authors for NGS data analysis. PALMapper is a short-read mapper which efficiently computes both unspliced and spliced alignments at high accuracy by taking advantage of base quality information and computational splice site predictions [ 3 ]. mTIM is a transcript reconstruction method, which exploits features derived from RNA-seq read alignments and from computational splice site predictions to infer the exon-intron structure of the corresponding transcripts. rQuant is based on quadratic programming. It simultaneously estimates biases inherent in library preparation, sequencing, and read mapping, and accurately determines the abundances of given transcripts [ 4 ]. rDiff is a set of statistical test techniques that determine significant differences between two RNA-seq experiments to find differentially expressed regions with or without knowledge of transcripts. We compare predictions to the published annotation at the intron and transcript levels. The performance of read aligners is shown in Fig. 1A from D. melanogaster data , and transcript segmentation tools in Fig. 1B, on C. elegans . Our tools, PALMapper and mTIM, outperform TopHat [ 5 ] and Cufflinks [ 6 ]. Oqtans is available free and open-source, from http://oqtans.org as a virtual machine for cloud computing environments, and ready to use on our public compute cluster at http://bioweb.me/mlb-galaxy . A) Accuracy (F-score) of intron predictions in 3-day-old adults of D. melanogaster with aligners PALMapper (green) and TopHat (blue). B) Accuracy of intron predictions with the same aligners and transcript predictions with mTIM (green) and Cufflinks (blue) on C. elegans RNA-seq transcriptome data.
Sebastian J. Schultheiß, Géraldine Jean, Jonas Behr, Regina Bohnert, Philipp Drewe, Nico Görnitz, André Kahles, Pramod Mudrakarta, Vipin T. Sreedharan, Georg Zeller, Gunnar Rätsch
BMC Bioinform.2
2007 Genome rearrangements: a correct algorithm for optimal capping
Géraldine Jean, Macha Nikolski
Inf. Process. Lett.1