EDBT 2026 Demo / reviewers in the wild / expert
Mathieu Blanchette
dblp:b/MathieuBlanchette
· DBLP profile ↗
50ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-9555-860XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence modelabstractMOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on ∼149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818. Dongjoon Lim, Mathieu Blanchette |
Bioinform. | 2 |
| 2025 | dyAb: Flow Matching for Flexible Antibody Design with AlphaFold-driven Pre-binding AntigenabstractThe development of therapeutic antibodies heavily relies on accurate predictions of how antigens will interact with antibodies. Existing computational methods in antibody design often overlook crucial conformational changes that antigens undergo during the binding process, significantly impacting the reliability of the resulting antibodies. To bridge this gap, we introduce dyAb, a flexible framework that incorporates AlphaFold2-driven predictions to model pre-binding antigen structures and specifically addresses the dynamic nature of antigen conformation changes. Our dyAb model leverages a unique combination of coarse-grained interface alignment and fine-grained flow matching techniques to simulate the interaction dynamics and structural evolution of the antigen-antibody complex, providing a realistic representation of the binding process. Extensive experiments show that dyAb significantly outperforms existing models in antibody design involving changing antigen conformations. These results highlight dyAb's potential to streamline the design process for therapeutic antibodies, promising more efficient development cycles and improved outcomes in clinical applications. Cheng Tan 0012, Zhangyang Gao, Yufei Huang 0002, Lirong Wu, Fandi Wu, Mathieu Blanchette, Stan Z. Li |
AAAI | 8 |
| 2025 | Meta Flow Matching: Integrating Vector Fields on the Wasserstein ManifoldabstractNumerous biological and physical processes can be modeled as systems of interacting entities evolving continuously over time, e.g. the dynamics of communicating cells or physical particles. Learning the dynamics of such systems is essential for predicting the temporal evolution of populations across novel samples and unseen environments. Flow-based models allow for learning these dynamics at the population level - they model the evolution of the entire distribution of samples. However, current flow-based models are limited to a single initial population and a set of predefined conditions which describe different dynamics. We argue that multiple processes in natural sciences have to be represented as vector fields on the Wasserstein manifold of probability densities. That is, the change of the population at any moment in time depends on the population itself due to the interactions between samples. In particular, this is crucial for personalized medicine where the development of diseases and their respective treatment response depend on the microenvironment of cells specific to each patient. We propose *Meta Flow Matching* (MFM), a practical approach to integrate along these vector fields on the Wasserstein manifold by amortizing the flow model over the initial populations. Namely, we embed the population of samples using a Graph Neural Network (GNN) and use these embeddings to train a Flow Matching model. This gives MFM the ability to generalize over the initial distributions, unlike previously proposed methods. We demonstrate the ability of MFM to improve the prediction of individual treatment responses on a large-scale multi-patient single-cell drug screen dataset. Lazar Atanackovic, Brandon Amos, Mathieu Blanchette, Leo J. Lee, Yoshua Bengio, Alexander Tong 0001, Kirill Neklyudov |
ICLR | 4 |
| 2025 | R3Design: deep tertiary structure-based RNA sequence design and beyondabstractThe rational design of Ribonucleic acid (RNA) molecules is crucial for advancing therapeutic applications, synthetic biology, and understanding the fundamental principles of life. Traditional RNA design methods have predominantly focused on secondary structure-based sequence design, often neglecting the intricate and essential tertiary interactions. We introduce R3Design, a tertiary structure-based RNA sequence design method that shifts the paradigm to prioritize tertiary structure in the RNA sequence design. R3Design significantly enhances sequence design on native RNA backbones, achieving high sequence recovery and Macro-F1 score, and outperforming traditional secondary structure-based approaches by substantial margins. We demonstrate that R3Design can design RNA sequences that fold into the desired tertiary structures by validating these predictions using advanced structure prediction models. This method, which is available through standalone software, provides a comprehensive toolkit for designing, folding, and evaluating RNA at the tertiary level. Our findings demonstrate R3Design's superior capability in designing RNA sequences, which achieves around $44\%$ in terms of both recovery score and Macro-F1 score in multiple datasets. This not only denotes the accuracy and fairness of the model but also underscores its potential to drive forward the development of innovative RNA-based therapeutics and to deepen our understanding of RNA biology. Cheng Tan 0012, Zhangyang Gao, Hanqun Cao, Siyuan Li 0002, Mathieu Blanchette, Stan Z. Li |
Briefings Bioinform. | 7 |
| 2024 | Learning the Game: Decoding the Differences between Novice and Expert Players in a Citizen Science Game with Millions of PlayersabstractIn recent years, video games have surged in popularity, attracting millions of players across platforms. Citizen science games (CSGs) leverage the processing power of gamers to solve computational and scientific problems. Borderlands Science (BLS) is a mini-game within the mass market game Borderlands 3 that turns multiple sequence alignment (MSA) problems into puzzles. Parallel research demonstrated that BLS players outperformed classical approaches solving small sequence alignment tasks. This study aims to analyze the strategical differences in player solutions in BLS as they gain experience. Through the many collected player solutions from players of different experience level, we gained insights into players’ strategies, differences between expert and non-expert players, and how strategies evolve. We developed a Markov chain trained on solutions from players of different experience levels to understand their actions and outcomes. Results indicate that expert players utilize more gaps and achieve more matches, gradually improving and converging toward unique strategies. Our findings reveal distinct and evolving player strategies. For future citizen science projects, it will be important to consider the identification of player strategies and their evolution over time to improve the game design and data processing. Eddie Cai, Roman Sarrazin-Gendron, Renata Mutalova, Parham Ghasemloo Gheidari, Alexander Butyaev, Gabriel Richard, Sébastien Caisse, Rob Knight 0001, Mathieu Blanchette, Attila Szantner, Jérôme Waldispühl |
FDG | 9 |
| 2024 | PhyloGFN: Phylogenetic inference with generative flow networksabstractPhylogenetics is a branch of computational biology that studies the evolutionary relationships among biological entities. Its long history and numerous applications notwithstanding, inference of phylogenetic trees from sequence data remains challenging: the high complexity of tree space poses a significant obstacle for the current combinatorial and probabilistic techniques. In this paper, we adopt the framework of generative flow networks (GFlowNets) to tackle two core problems in phylogenetics: parsimony-based and Bayesian phylogenetic inference. Because GFlowNets are well-suited for sampling complex combinatorial structures, they are a natural choice for exploring and sampling from the multimodal posterior distribution over tree topologies and evolutionary distances. We demonstrate that our amortized posterior sampler, PhyloGFN, produces diverse and high-quality evolutionary hypotheses on real benchmark datasets. PhyloGFN is competitive with prior works in marginal likelihood estimation and achieves a closer fit to the target distribution than state-of-the-art variational inference methods. Zichao Yan, Elliot Layne, Nikolay Malkin, Dinghuai Zhang, Moksh Jain, Mathieu Blanchette, Yoshua Bengio |
ICLR | 7 |
| 2024 | PERFUMES: pipeline to extract RNA functional motifs and exposed structuresabstractMOTIVATION: Up to 75% of the human genome encodes RNAs. The function of many non-coding RNAs relies on their ability to fold into 3D structures. Specifically, nucleotides inside secondary structure loops form non-canonical base pairs that help stabilize complex local 3D structures. These RNA 3D motifs can promote specific interactions with other molecules or serve as catalytic sites. RESULTS: We introduce PERFUMES, a computational pipeline to identify 3D motifs that can be associated with observable features. Given a set of RNA sequences with associated binary experimental measurements, PERFUMES searches for RNA 3D motifs using BayesPairing2 and extracts those that are over-represented in the set of positive sequences. It also conducts a thermodynamics analysis of the structural context that can support the interpretation of the predictions. We illustrate PERFUMES' usage on the SNRPA protein binding site, for which the tool retrieved both previously known binder motifs and new ones. AVAILABILITY AND IMPLEMENTATION: PERFUMES is an open-source Python package (https://jwgitlab.cs.mcgill.ca/arnaud_chol/perfumes). Arnaud Chol, Roman Sarrazin-Gendron, Eric Lécuyer, Mathieu Blanchette, Jérôme Waldispühl |
Bioinform. | 4 |
| 2024 | ARGV: 3D genome structure exploration using augmented realityabstractOver the past two decades, scientists have increasingly realized the importance of the three-dimensional (3D) genome organization in regulating cellular activity. Hi-C and related experiments yield 2D contact matrices that can be used to infer 3D models of chromosome structure. Visualizing and analyzing genomes in 3D space remains challenging. Here, we present ARGV, an augmented reality 3D Genome Viewer. ARGV contains more than 350 pre-computed and annotated genome structures inferred from Hi-C and imaging data. It offers interactive and collaborative visualization of genomes in 3D space, using standard mobile phones or tablets. A user study comparing ARGV to existing tools demonstrates its benefits. Chrisostomos Drogaris, Yanlin Zhang, Elena Nazarova, Roman Sarrazin-Gendron, Sélik Wilhelm-Landry, Yan Cyr, Jacek Majewski, Mathieu Blanchette, Jérôme Waldispühl |
BMC Bioinform. | 9 |
| 2023 | Playing the System: Can Puzzle Players Teach us How to Solve Hard Problems?abstractWith nearly three billion players, video games are more popular than ever. Casual puzzle games are among the most played categories. These games capitalize on the players’ analytical and problem-solving skills. Can we leverage these abilities to teach ourselves how to solve complex combinatorial problems? In this study, we harness the collective wisdom of millions of players to tackle the classical NP-hard problem of multiple sequence alignment, relevant to many areas of biology and medicine. We show that Borderlands Science players propose solutions to multiple sequence alignment tasks that perform as well or better than standard approaches, while exploring a much larger area of the Pareto-optimal solution space. We also show the strategies of the players, although highly heterogeneous, follow a collective logic that can be mimicked with Behavioral Cloning with minimal performance loss, allowing the players’ collective wisdom to be leveraged for alignment of any sequences. Renata Mutalova, Roman Sarrazin-Gendron, Eddie Cai, Gabriel Richard, Parham Ghasemloo Gheidari, Sébastien Caisse, Rob Knight 0001, Mathieu Blanchette, Attila Szantner, Jérôme Waldispühl |
CHI | 8 |
| 2023 | Reference panel-guided super-resolution inference of Hi-C dataabstractMOTIVATION: Accurately assessing contacts between DNA fragments inside the nucleus with Hi-C experiment is crucial for understanding the role of 3D genome organization in gene regulation. This challenging task is due in part to the high sequencing depth of Hi-C libraries required to support high-resolution analyses. Most existing Hi-C data are collected with limited sequencing coverage, leading to poor chromatin interaction frequency estimation. Current computational approaches to enhance Hi-C signals focus on the analysis of individual Hi-C datasets of interest, without taking advantage of the facts that (i) several hundred Hi-C contact maps are publicly available and (ii) the vast majority of local spatial organizations are conserved across multiple cell types. RESULTS: Here, we present RefHiC-SR, an attention-based deep learning framework that uses a reference panel of Hi-C datasets to facilitate the enhancement of Hi-C data resolution of a given study sample. We compare RefHiC-SR against tools that do not use reference samples and find that RefHiC-SR outperforms other programs across different cell types, and sequencing depths. It also enables high-accuracy mapping of structures such as loops and topologically associating domains. AVAILABILITY AND IMPLEMENTATION: https://github.com/BlanchetteLab/RefHiC. Yanlin Zhang, Mathieu Blanchette |
Bioinform. | 2 |
| 2022 | PhyloPGM: boosting regulatory function prediction accuracy using evolutionary informationabstractMOTIVATION: The computational prediction of regulatory function associated with a genomic sequence is of utter importance in -omics study, which facilitates our understanding of the underlying mechanisms underpinning the vast gene regulatory network. Prominent examples in this area include the binding prediction of transcription factors in DNA regulatory regions, and predicting RNA-protein interaction in the context of post-transcriptional gene expression. However, existing computational methods have suffered from high false-positive rates and have seldom used any evolutionary information, despite the vast amount of available orthologous data across multitudes of extant and ancestral genomes, which readily present an opportunity to improve the accuracy of existing computational methods. RESULTS: In this study, we present a novel probabilistic approach called PhyloPGM that leverages previously trained TFBS or RNA-RBP binding predictors by aggregating their predictions from various orthologous regions, in order to boost the overall prediction accuracy on human sequences. Throughout our experiments, PhyloPGM has shown significant improvement over baselines such as the sequence-based RNA-RBP binding predictor RNATracker and the sequence-based TFBS predictor that is known as FactorNet. PhyloPGM is simple in principle, easy to implement and yet, yields impressive results. AVAILABILITY AND IMPLEMENTATION: The PhyloPGM package is available at https://github.com/BlanchetteLab/PhyloPGM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Faizy Ahsan, Zichao Yan, Doina Precup, Mathieu Blanchette |
Bioinform. | 4 |
| 2021 | Neural representation and generation for RNA secondary structures
Zichao Yan, William L. Hamilton, Mathieu Blanchette |
ICLR | 3 |
| 2020 | Phylogenetic Manifold Regularization: A semi-supervised approach to predict transcription factor binding sitesabstractThe computational prediction of transcription factor binding sites remains a challenging problems in bioinformatics, despite significant methodological developments from the field of machine learning. Such computational models are essential to help interpret the non-coding portion of human genomes, and to learn more about the regulatory mechanisms controlling gene expression. In parallel, massive genome sequencing efforts have produced assembled genomes for hundred of vertebrate species, but this data is underused. We present PhyloReg, a new semi-supervised learning approach that can be used for a wide variety of sequence-to-function prediction problems, and that takes advantage of hundreds of millions of years of evolution to regularize predictors and improve accuracy. We demonstrate that PhyloReg can be used to better train a previously proposed deep learning model of transcription factor binding. Simulation studies further help delineate the benefits of the a pproach. G ains in prediction accuracy are obtained over a broad set of transcription factors and cell types. Faizy Ahsan, Alexandre Drouin, François Laviolette, Doina Precup, Mathieu Blanchette |
BIBM | 5 |
| 2020 | Mycorrhiza: genotype assignment using phylogenetic networksabstractMOTIVATION: The genotype assignment problem consists of predicting, from the genotype of an individual, which of a known set of populations it originated from. The problem arises in a variety of contexts, including wildlife forensics, invasive species detection and biodiversity monitoring. Existing approaches perform well under ideal conditions but are sensitive to a variety of common violations of the assumptions they rely on. RESULTS: In this article, we introduce Mycorrhiza, a machine learning approach for the genotype assignment problem. Our algorithm makes use of phylogenetic networks to engineer features that encode the evolutionary relationships among samples. Those features are then used as input to a Random Forests classifier. The classification accuracy was assessed on multiple published empirical SNP, microsatellite or consensus sequence datasets with wide ranges of size, geographical distribution and population structure and on simulated datasets. It compared favorably against widely used assessment tests or mixture analysis methods such as STRUCTURE and Admixture, and against another machine-learning based approach using principal component analysis for dimensionality reduction. Mycorrhiza yields particularly significant gains on datasets with a large average fixation index (FST) or deviation from the Hardy-Weinberg equilibrium. Moreover, the phylogenetic network approach estimates mixture proportions with good accuracy. AVAILABILITY AND IMPLEMENTATION: Mycorrhiza is released as an easy to use open-source python package at github.com/jgeofil/mycorrhiza. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jeremy Georges-Filteau, Richard C. Hamelin, Mathieu Blanchette |
Bioinform. | 3 |
| 2020 | Supervised learning on phylogenetically distributed dataabstractMOTIVATION: The ability to develop robust machine-learning (ML) models is considered imperative to the adoption of ML techniques in biology and medicine fields. This challenge is particularly acute when data available for training is not independent and identically distributed (iid), in which case trained models are vulnerable to out-of-distribution generalization problems. Of particular interest are problems where data correspond to observations made on phylogenetically related samples (e.g. antibiotic resistance data). RESULTS: We introduce DendroNet, a new approach to train neural networks in the context of evolutionary data. DendroNet explicitly accounts for the relatedness of the training/testing data, while allowing the model to evolve along the branches of the phylogenetic tree, hence accommodating potential changes in the rules that relate genotypes to phenotypes. Using simulated data, we demonstrate that DendroNet produces models that can be significantly better than non-phylogenetically aware approaches. DendroNet also outperforms other approaches at two biological tasks of significant practical importance: antiobiotic resistance prediction in bacteria and trophic level prediction in fungi. AVAILABILITY AND IMPLEMENTATION: https://github.com/BlanchetteLab/DendroNet. Elliot Layne, Erika N. Dort, Richard C. Hamelin, Yue Li 0017, Mathieu Blanchette |
Bioinform. | 5 |
| 2020 | EvoLSTM: context-dependent models of sequence evolution using a sequence-to-sequence LSTMabstractMOTIVATION: Accurate probabilistic models of sequence evolution are essential for a wide variety of bioinformatics tasks, including sequence alignment and phylogenetic inference. The ability to realistically simulate sequence evolution is also at the core of many benchmarking strategies. Yet, mutational processes have complex context dependencies that remain poorly modeled and understood. RESULTS: We introduce EvoLSTM, a recurrent neural network-based evolution simulator that captures mutational context dependencies. EvoLSTM uses a sequence-to-sequence long short-term memory model trained to predict mutation probabilities at each position of a given sequence, taking into consideration the 14 flanking nucleotides. EvoLSTM can realistically simulate mammalian and plant DNA sequence evolution and reveals unexpectedly strong long-range context dependencies in mutation probabilities. EvoLSTM brings modern machine-learning approaches to bear on sequence evolution. It will serve as a useful tool to study and simulate complex mutational processes. AVAILABILITY AND IMPLEMENTATION: Code and dataset are available at https://github.com/DongjoonLim/EvoLSTM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dongjoon Lim, Mathieu Blanchette |
Bioinform. | 2 |
| 2020 | Graph neural representational learning of RNA secondary structures for predicting RNA-protein interactionsabstractMOTIVATION: RNA-protein interactions are key effectors of post-transcriptional regulation. Significant experimental and bioinformatics efforts have been expended on characterizing protein binding mechanisms on the molecular level, and on highlighting the sequence and structural traits of RNA that impact the binding specificity for different proteins. Yet our ability to predict these interactions in silico remains relatively poor. RESULTS: In this study, we introduce RPI-Net, a graph neural network approach for RNA-protein interaction prediction. RPI-Net learns and exploits a graph representation of RNA molecules, yielding significant performance gains over existing state-of-the-art approaches. We also introduce an approach to rectify an important type of sequence bias caused by the RNase T1 enzyme used in many CLIP-Seq experiments, and we show that correcting this bias is essential in order to learn meaningful predictors and properly evaluate their accuracy. Finally, we provide new approaches to interpret the trained models and extract simple, biologically interpretable representations of the learned sequence and structural motifs. AVAILABILITY AND IMPLEMENTATION: Source code can be accessed at https://www.github.com/HarveyYan/RNAonGraph. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zichao Yan, William L. Hamilton, Mathieu Blanchette |
Bioinform. | 3 |
| 2019 | Large-scale mammalian genome rearrangements coincide with chromatin interactionsabstractMOTIVATION: Genome rearrangements drastically change gene order along great stretches of a chromosome. There has been initial evidence that these apparently non-local events in the 1D sense may have breakpoints that are close in the 3D sense. We harness the power of the Double Cut and Join model of genome rearrangement, along with Hi-C chromosome conformation capture data to test this hypothesis between human and mouse. RESULTS: We devise novel statistical tests that show that indeed, rearrangement scenarios that transform the human into the mouse gene order are enriched for pairs of breakpoints that have frequent chromosome interactions. This is observed for both intra-chromosomal breakpoint pairs, as well as for inter-chromosomal pairs. For intra-chromosomal rearrangements, the enrichment exists from close (<20 Mb) to very distant (100 Mb) pairs. Further, the pattern exists across multiple cell lines in Hi-C data produced by different laboratories and at different stages of the cell cycle. We show that similarities in the contact frequencies between these many experiments contribute to the enrichment. We conclude that either (i) rearrangements usually involve breakpoints that are spatially close or (ii) there is selection against rearrangements that act on spatially distant breakpoints. AVAILABILITY AND IMPLEMENTATION: Our pipeline is freely available at https://bitbucket.org/thekswenson/locality. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Krister M. Swenson, Mathieu Blanchette |
Bioinform. | 2 |
| 2019 | Prediction of mRNA subcellular localization using deep recurrent neural networksabstractMOTIVATION: Messenger RNA subcellular localization mechanisms play a crucial role in post-transcriptional gene regulation. This trafficking is mediated by trans-acting RNA-binding proteins interacting with cis-regulatory elements called zipcodes. While new sequencing-based technologies allow the high-throughput identification of RNAs localized to specific subcellular compartments, the precise mechanisms at play, and their dependency on specific sequence elements, remain poorly understood. RESULTS: We introduce RNATracker, a novel deep neural network built to predict, from their sequence alone, the distributions of mRNA transcripts over a predefined set of subcellular compartments. RNATracker integrates several state-of-the-art deep learning techniques (e.g. CNN, LSTM and attention layers) and can make use of both sequence and secondary structure information. We report on a variety of evaluations showing RNATracker's strong predictive power, which is significantly superior to a variety of baseline predictors. Despite its complexity, several aspects of the model can be isolated to yield valuable, testable mechanistic hypotheses, and to locate candidate zipcode sequences within transcripts. AVAILABILITY AND IMPLEMENTATION: Code and data can be accessed at https://www.github.com/HarveyYan/RNATracker. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zichao Yan, Eric Lécuyer, Mathieu Blanchette |
Bioinform. | 3 |
| 2018 | RLALIGN: A Reinforcement Learning Approach for Multiple Sequence AlignmentabstractMultiple sequence alignment (MSA) is one of the best studied problems in bioinformatics because of the broad set of genomics, proteomics, and evolutionary analyses that rely on it. Yet the problem is NP-hard and existing heuristics are imperfect. Reinforcement learning (RL) techniques have emerged recently as a potential solution to a wide diversity of computational problems, but have yet to be applied to MSA. In this paper, we describe RLALIGN, a method to solve the MSA problem using RL. RLALIGN is based on Asynchronous Advantage Actor Critic (A3C), a cutting-edge RL framework. Due to the absence of a goal state, however, it required several important modifications. RLALIGN can be trained to accurately align moderate-length sequences, and various heuristics allow it to scale to longer sequences. The accuracy of the alignments produced is on par with, and often better than those of well established alignment algorithms. Overall, our work demonstrates the potential of RL approaches for complex combinatorial problems such as MSA. RLALIGN will prove useful for realignment tasks, where portions of a larger alignment need to be optimized. Unlike classical algorithms, RLALIGN is incognizant to the nature of the scoring scheme, leading to easy generalization to a variety of problem variants. Ramchalam Kinattinkara Ramakrishnan, Jaspal Singh, Mathieu Blanchette |
BIBE | 3 |
| 2018 | Detection of Errors in Multi-genome Alignments Using Machine Learning ApproachesabstractWhole-genome multiple alignments are widely used in genomics and evolution, and yet their accuracy is imperfect, due in part to the computational complexity of the task at hand. Identifying portions of these alignments that are likely to be incorrect would allow researchers to either work on improving them or flagging them for exclusion from downstream analyses. We introduce MSA-ED, a machine learning tool for the detection of errors in whole-genome multiple alignments. MSA-ED uses random forests or artificial neural networks to identify and classify several types of alignment errors. It is trained on labeled data obtained by using an evolution simulator to generate fake orthologous sequences and their correct alignment, and comparing it to the alignment produced by Multiz, a popular whole-genome aligner. Key to the success of MSA-ED is the engineering of several types of evolutionarily-inspired features that boost prediction accuracy. MSA-ED is shown to be able to detect certain types of errors with good accuracy. It is then applied to actual genomic alignments to identify putative alignment errors. Availability: https://github.com/jaspal1329/MSA-ED Jaspal Singh, Ramchalam Kinattinkara Ramakrishnan, Mathieu Blanchette |
BIBE | 3 |
| 2017 | Lessons from an Online Massive Genomics Computer GameabstractCrowdsourcing through human-computing games is an increasingly popular practice for classifying and analyzing scientific data. Early contributions such as Phylo have now been running for several years. The analysis of the performance of these systems enables us to identify patterns that contributed to their successes, but also possible pitfalls. In this paper, we review the results and user statistics collected since 2010 by our platform Phylo, which aims to engage citizens in comparative genome analysis through a casual tile matching computer game. We also identify features that allow predicting a task difficulty, which is essential for channeling them to human players with the appropriate skill level. Finally, we show how our platform has been used to quickly improve a reference alignment of Ebola virus sequences. Faizy Ahsan, Mathieu Blanchette, Jérôme Waldispühl |
HCOMP | 3 |
| 2017 | CoreTracker: accurate codon reassignment prediction, applied to mitochondrial genomesabstractMOTIVATION: Codon reassignments have been reported across all domains of life. With the increasing number of sequenced genomes, the development of systematic approaches for genetic code detection is essential for accurate downstream analyses. Three automated prediction tools exist so far: FACIL, GenDecoder and Bagheera; the last two respectively restricted to metazoan mitochondrial genomes and CUG reassignments in yeast nuclear genomes. These tools can only analyze a single genome at a time and are often not followed by a validation procedure, resulting in a high rate of false positives. RESULTS: We present CoreTracker, a new algorithm for the inference of sense-to-sense codon reassignments. CoreTracker identifies potential codon reassignments in a set of related genomes, then uses statistical evaluations and a random forest classifier to predict those that are the most likely to be correct. Predicted reassignments are then validated through a phylogeny-aware step that evaluates the impact of the new genetic code on the protein alignment. Handling simultaneously a set of genomes in a phylogenetic framework, allows tracing back the evolution of each reassignment, which provides information on its underlying mechanism. Applied to metazoan and yeast genomes, CoreTracker significantly outperforms existing methods on both precision and sensitivity. AVAILABILITY AND IMPLEMENTATION: CoreTracker is written in Python and available at https://github.com/UdeM-LBIT/CoreTracker. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Emmanuel Noutahi, Virginie Calderon, Mathieu Blanchette, B. Franz Lang, Nadia El-Mabrouk |
Bioinform. | 3 |
| 2015 | Models and Algorithms for Genome Rearrangement with Positional Constraints
Krister M. Swenson, Mathieu Blanchette |
WABI | 2 |
| 2015 | BigDataScript: a scripting language for data pipelinesabstractMOTIVATION: The analysis of large biological datasets often requires complex processing pipelines that run for a long time on large computational infrastructures. We designed and implemented a simple script-like programming language with a clean and minimalist syntax to develop and manage pipeline execution and provide robustness to various types of software and hardware failures as well as portability. RESULTS: We introduce the BigDataScript (BDS) programming language for data processing pipelines, which improves abstraction from hardware resources and assists with robustness. Hardware abstraction allows BDS pipelines to run without modification on a wide range of computer architectures, from a small laptop to multi-core servers, server farms, clusters and clouds. BDS achieves robustness by incorporating the concepts of absolute serialization and lazy processing, thus allowing pipelines to recover from errors. By abstracting pipeline concepts at programming language level, BDS simplifies implementation, execution and management of complex bioinformatics pipelines, resulting in reduced development and debugging cycles as well as cleaner code. AVAILABILITY AND IMPLEMENTATION: BigDataScript is available under open-source license at http://pcingola.github.io/BigDataScript. Pablo Cingolani, Rob Sladek, Mathieu Blanchette |
Bioinform. | 3 |
| 2014 | Phylo and Open-Phylo: A Human-Computing Platform for Comparative GenomicsabstractComparative genomics is a field of research that aims to provide us accurate mappings between the genetic material of multiple species. These techniques are useful for biomedical research and evolutionary studies. We present Phylo and Open-Phylo, an open citizen science platform and human computing-game for comparative genomic studies that reached more than 300,000 users. Jérôme Waldispühl, Mathieu Blanchette |
HCOMP | 2 |
| 2012 | Clique Cover on Sparse NetworksabstractWe consider the problem of edge clique cover on sparse networks and study an application to the identification of overlapping protein complexes for a network of binary protein-protein interactions. We first give an algorithm whose running time is linear in the size of the graph, provided the treewidth is bounded. We then provide an algorithm for planar graphs with bounded branchwidth upon which we build a PTAS for planar graphs. Empirical studies show that our algorithms are both efficient and practical on actual simulated and biological networks, and that the clique covers obtained on real networks yield biological insights. Mathieu Blanchette, Ethan Kim, Adrian Vetta |
ALENEX | 1 |
| 2012 | Exploiting ancestral mammalian genomes for the prediction of human transcription factor binding sitesabstractBACKGROUND: The computational prediction of Transcription Factor Binding Sites (TFBS) remains a challenge due to their short length and low information content. Comparative genomics approaches that simultaneously consider several related species and favor sites that have been conserved throughout evolution improve the accuracy (specificity) of the predictions but are limited due to a phenomenon called binding site turnover, where sequence evolution causes one TFBS to replace another in the same region. In parallel to this development, an increasing number of mammalian genomes are now sequenced and it is becoming possible to infer, to a surprisingly high degree of accuracy, ancestral mammalian sequences. RESULTS: We propose a TFBS prediction approach that makes use of the availability of inferred ancestral mammalian genomes to improve its accuracy. This method aims to identify binding loci, which are regions of a few hundred base pairs that have preserved their potential to bind a given transcription factor over evolutionary time. After proposing a neutral evolutionary model of predicted TFBS counts in a DNA region of a given length, we use it to identify regions that have preserved the number of predicted TFBS they contain to an unexpected degree given their divergence. The approach is applied to human chromosome 1 and shows significant gains in accuracy as compared to both existing single-species and multi-species TFBS prediction approaches, in particular for transcription factors that are subject to high turnover rates. AVAILABILITY: The source code and predictions made by the program are available at http://www.cs.mcgill.ca/~blanchem/bindingLoci. Mathieu Blanchette |
BMC Bioinform. | 1 |
| 2012 | A flexible ancestral genome reconstruction method based on gapped adjacenciesabstractBACKGROUND: The "small phylogeny" problem consists in inferring ancestral genomes associated with each internal node of a phylogenetic tree of a set of extant species. Existing methods can be grouped into two main categories: the distance-based methods aiming at minimizing a total branch length, and the synteny-based (or mapping) methods that first predict a collection of relations between ancestral markers in term of "synteny", and then assemble this collection into a set of Contiguous Ancestral Regions (CARs). The predicted CARs are likely to be more reliable as they are more directly deduced from observed conservations in extant species. However the challenge is to end up with a completely assembled genome. RESULTS: We develop a new synteny-based method that is flexible enough to handle a model of evolution involving whole genome duplication events, in addition to rearrangements, gene insertions, and losses. Ancestral relationships between markers are defined in term of Gapped Adjacencies, i.e. pairs of markers separated by up to a given number of markers. It improves on a previous restricted to direct adjacencies, which revealed a high accuracy for adjacency prediction, but with the drawback of being overly conservative, i.e. of generating a large number of CARs. Applying our algorithm on various simulated data sets reveals good performance as we usually end up with a completely assembled genome, while keeping a low error rate. AVAILABILITY: All source code is available at http://www.iro.umontreal.ca/~mabrouk. Yves Gagnon, Mathieu Blanchette, Nadia El-Mabrouk |
BMC Bioinform. | 2 |
| 2011 | A Probabilistic Model for Sequence Alignment with Context-Sensitive Indels
Glenn Hickey, Mathieu Blanchette |
RECOMB | 2 |
| 2011 | Predicting site-specific human selective pressure using evolutionary signaturesabstractMOTIVATION: The identification of non-coding functional regions of the human genome remains one of the main challenges of genomics. By observing how a given region evolved over time, one can detect signs of negative or positive selection hinting that the region may be functional. With the quickly increasing number of vertebrate genomes to compare with our own, this type of approach is set to become extremely powerful, provided the right analytical tools are available. RESULTS: A large number of approaches have been proposed to measure signs of past selective pressure, usually in the form of reduced mutation rate. Here, we propose a radically different approach to the detection of non-coding functional region: instead of measuring past evolutionary rates, we build a machine learning classifier to predict current substitution rates in human based on the inferred evolutionary events that affected the region during vertebrate evolution. We show that different types of evolutionary events, occurring along different branches of the phylogenetic tree, bring very different amounts of information. We propose a number of simple machine learning classifiers and show that a Support-Vector Machine (SVM) predictor clearly outperforms existing tools at predicting human non-coding functional sites. Comparison to external evidences of selection and regulatory function confirms that these SVM predictions are more accurate than those of other approaches. AVAILABILITY: The predictor and predictions made are available at http://www.mcb.mcgill.ca/~blanchem/sadri. CONTACT: [email protected]. Javad Sadri, Abdoulaye Baniré Diallo, Mathieu Blanchette |
Bioinform. | 3 |
| 2011 | Three-dimensional modeling of chromatin structure from interaction frequency data using Markov chain Monte Carlo samplingabstractBACKGROUND: Long-range interactions between regulatory DNA elements such as enhancers, insulators and promoters play an important role in regulating transcription. As chromatin contacts have been found throughout the human genome and in different cell types, spatial transcriptional control is now viewed as a general mechanism of gene expression regulation. Chromosome Conformation Capture Carbon Copy (5C) and its variant Hi-C are techniques used to measure the interaction frequency (IF) between specific regions of the genome. Our goal is to use the IF data generated by these experiments to computationally model and analyze three-dimensional chromatin organization. RESULTS: We formulate a probabilistic model linking 5C/Hi-C data to physical distances and describe a Markov chain Monte Carlo (MCMC) approach called MCMC5C to generate a representative sample from the posterior distribution over structures from IF data. Structures produced from parallel MCMC runs on the same dataset demonstrate that our MCMC method mixes quickly and is able to sample from the posterior distribution of structures and find subclasses of structures. Structural properties (base looping, condensation, and local density) were defined and their distribution measured across the ensembles of structures generated. We applied these methods to a biological model of human myelomonocyte cellular differentiation and identified distinct chromatin conformation signatures (CCSs) corresponding to each of the cellular states. We also demonstrate the ability of our method to run on Hi-C data and produce a model of human chromosome 14 at 1Mb resolution that is consistent with previously observed structural properties as measured by 3D-FISH. CONCLUSIONS: We believe that tools like MCMC5C are essential for the reliable analysis of data from the 3C-derived techniques such as 5C and Hi-C. By integrating complex, high-dimensional and noisy datasets into an easy to interpret ensemble of three-dimensional conformations, MCMC5C allows researchers to reliably interpret the result of their assay and contrast conformations under different conditions. AVAILABILITY: http://Dostielab.biochem.mcgill.ca. Mathieu Rousseau, James Fraser, Maria A. Ferraiuolo, Josee Dostie, Mathieu Blanchette |
BMC Bioinform. | 5 |
| 2011 | An Approximation Algorithm for the Noah's Ark Problem with Random Feature LossabstractThe phylogenetic diversity (PD) of a set of species is a measure of their evolutionary distinctness based on a phylogenetic tree. PD is increasingly being adopted as an index of biodiversity in ecological conservation projects. The Noah's Ark Problem (NAP) is an NP-Hard optimization problem that abstracts a fundamental conservation challenge in asking to maximize the expected PD of a set of taxa given a fixed budget, where each taxon is associated with a cost of conservation and a probability of extinction. Only simplified instances of the problem, where one or more parameters are fixed as constants, have as of yet been addressed in the literature. Furthermore, it has been argued that PD is not an appropriate metric for models that allow information to be lost along paths in the tree. We therefore generalize the NAP to incorporate a proposed model of feature loss according to an exponential distribution and term this problem NAP with Loss (NAPL). In this paper, we present a pseudopolynomial time approximation scheme for NAPL. Glenn Hickey, Mathieu Blanchette, Paz Carmi, Anil Maheshwari, Norbert Zeh |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Reconstruction of Ancestral Genome Subject to Whole Genome Duplication, Speciation, Rearrangement and Loss
Denis Bertrand, Yves Gagnon, Mathieu Blanchette, Nadia El-Mabrouk |
WABI | 3 |
| 2010 | Ancestors 1.0: a web server for ancestral sequence reconstructionabstractSUMMARY: The computational inference of ancestral genomes consists of five difficult steps: identifying syntenic regions, inferring ancestral arrangement of syntenic regions, aligning multiple sequences, reconstructing the insertion and deletion history and finally inferring substitutions. Each of these steps have received lot of attention in the past years. However, there currently exists no framework that integrates all of the different steps in an easy workflow. Here, we introduce Ancestors 1.0, a web server allowing one to easily and quickly perform the last three steps of the ancestral genome reconstruction procedure. It implements several alignment algorithms, an indel maximum likelihood solver and a context-dependent maximum likelihood substitution inference algorithm. The results presented by the server include the posterior probabilities for the last two steps of the ancestral genome reconstruction and the expected error rate of each ancestral base prediction. AVAILABILITY: The Ancestors 1.0 is available at http://ancestors.bioinfo.uqam.ca/ancestorWeb/. Abdoulaye Baniré Diallo, Vladimir Makarenkov, Mathieu Blanchette |
Bioinform. | 3 |
| 2010 | Computational Analysis of Whole-Genome Differential Allelic Expression Data in HumanabstractAllelic imbalance (AI) is a phenomenon where the two alleles of a given gene are expressed at different levels in a given cell, either because of epigenetic inactivation of one of the two alleles, or because of genetic variation in regulatory regions. Recently, Bing et al. have described the use of genotyping arrays to assay AI at a high resolution (approximately 750,000 SNPs across the autosomes). In this paper, we investigate computational approaches to analyze this data and identify genomic regions with AI in an unbiased and robust statistical manner. We propose two families of approaches: (i) a statistical approach based on z-score computations, and (ii) a family of machine learning approaches based on Hidden Markov Models. Each method is evaluated using previously published experimental data sets as well as with permutation testing. When applied to whole genome data from 53 HapMap samples, our approaches reveal that allelic imbalance is widespread (most expressed genes show evidence of AI in at least one of our 53 samples) and that most AI regions in a given individual are also found in at least a few other individuals. While many AI regions identified in the genome correspond to known protein-coding transcripts, others overlap with recently discovered long non-coding RNAs. We also observe that genomic regions with AI not only include complete transcripts with consistent differential expression levels, but also more complex patterns of allelic expression such as alternative promoters and alternative 3' end. The approaches developed not only shed light on the incidence and mechanisms of allelic expression, but will also help towards mapping the genetic causes of allelic expression and identify cases where this variation may be linked to diseases. James R. Wagner, Bing Ge, Dmitry Pokholok, Kevin L. Gunderson, Tomi Pastinen, Mathieu Blanchette |
PLoS Comput. Biol. | 6 |
| 2009 | Detection of Locally Over-Represented GO Terms in Protein-Protein Interaction Networks
Mathieu Lavallée-Adam, Benoit Coulombe, Mathieu Blanchette |
RECOMB | 3 |
| 2008 | Seeder: discriminative seeding DNA motif discoveryabstractAbstract Motivation: The computational identification of transcription factor binding sites is a major challenge in bioinformatics and an important complement to experimental approaches. Results: We describe a novel, exact discriminative seeding DNA motif discovery algorithm designed for fast and reliable prediction of cis-regulatory elements in eukaryotic promoters. The algorithm is tested on biological benchmark data and shown to perform equally or better than other motif discovery tools. The algorithm is applied to the analysis of plant tissue-specific promoter sequences and successfully identifies key regulatory elements. Availability: The Seeder Perl distribution includes four modules. It is available for download on the Comprehensive Perl Archive Network (CPAN) at http://www.cpan.org. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. François Fauteux, Mathieu Blanchette, Martina V. Stromvik |
Bioinform. | 2 |
| 2008 | Improving the prediction of mRNA extremities in the parasitic protozoan LeishmaniaabstractBACKGROUND: Leishmania and other members of the Trypanosomatidae family diverged early on in eukaryotic evolution and consequently display unique cellular properties. Their apparent lack of transcriptional regulation is compensated by complex post-transcriptional control mechanisms, including the processing of polycistronic transcripts by means of coupled trans-splicing and polyadenylation. Trans-splicing signals are often U-rich polypyrimidine (poly(Y)) tracts, which precede AG splice acceptor sites. However, as opposed to higher eukaryotes there is no consensus polyadenylation signal in trypanosomatid mRNAs. RESULTS: We refined a previously reported method to target 5' splice junctions by incorporating the pyrimidine content of query sequences into a scoring function. We also investigated a novel approach for predicting polyadenylation (poly(A)) sites in-silico, by comparing query sequences to polyadenylated expressed sequence tags (ESTs) using position-specific scanning matrices (PSSMs). An additional analysis of the distribution of putative splice junction to poly(A) distances helped to increase prediction rates by limiting the scanning range. These methods were able to simplify splice junction prediction without loss of precision and to increase polyadenylation site prediction from 22% to 47% within 100 nucleotides. CONCLUSION: We propose a simplified trans-splicing prediction tool and a novel poly(A) prediction tool based on comparative sequence analysis. We discuss the impact of certain regions surrounding the poly(A) sites on prediction rates and contemplate correlating biological mechanisms. This work aims to sharpen the identification of potentially functional untranslated regions (UTRs) in a large-scale, comparative genomics framework. Martin A. Smith, Mathieu Blanchette, Barbara Papadopoulou |
BMC Bioinform. | 2 |
| 2007 | Prediction of tissue-specific cis-regulatory modules using Bayesian networks and regression treesabstractBACKGROUND: In vertebrates, a large part of gene transcriptional regulation is operated by cis-regulatory modules. These modules are believed to be regulating much of the tissue-specificity of gene expression. RESULT: We develop a Bayesian network approach for identifying cis-regulatory modules likely to regulate tissue-specific expression. The network integrates predicted transcription factor binding site information, transcription factor expression data, and target gene expression data. At its core is a regression tree modeling the effect of combinations of transcription factors bound to a module. A new unsupervised EM-like algorithm is developed to learn the parameters of the network, including the regression tree structure. CONCLUSION: Our approach is shown to accurately identify known human liver and erythroid-specific modules. When applied to the prediction of tissue-specific modules in 10 different tissues, the network predicts a number of important transcription factor combinations whose concerted binding is associated to specific expression. Mathieu Blanchette |
BMC Bioinform. | 2 |
| 2006 | Common Substrings in Random Strings
Eric Blais, Mathieu Blanchette |
CPM | 2 |
| 2004 | Reconstructing Ancestral Gene Orders Using Conserved Intervals
Anne Bergeron, Mathieu Blanchette, Annie Chateau, Cédric Chauve |
WABI | 2 |
| 2004 | PhyME: A probabilistic algorithm for finding motifs in sets of orthologous sequencesabstractBACKGROUND: This paper addresses the problem of discovering transcription factor binding sites in heterogeneous sequence data, which includes regulatory sequences of one or more genes, as well as their orthologs in other species. RESULTS: We propose an algorithm that integrates two important aspects of a motif's significance - overrepresentation and cross-species conservation - into one probabilistic score. The algorithm allows the input orthologous sequences to be related by any user-specified phylogenetic tree. It is based on the Expectation-Maximization technique, and scales well with the number of species and the length of input sequences. We evaluate the algorithm on synthetic data, and also present results for data sets from yeast, fly, and human. CONCLUSIONS: The results demonstrate that the new approach improves motif discovery by exploiting multiple species information. Mathieu Blanchette, Martin Tompa |
BMC Bioinform. | 2 |
| 2003 | An Empirical Comparison of Tools for Phylogenetic FootprintingabstractPhylogenetic footprinting is an increasingly popular comparative genomics method for detecting regulatory elements in DNA sequences. With the profusion of possible methods to use for phylogenetic footprinting, the biologist needs some guidance to choose the most appropriate tool. We present methods for comparing tools on phylogenetic footprinting data. More specifically, we discuss two different classes of comparative experiments: those on simulated data and those on real orthologous promoter regions. We then report the results of a series of such empirical comparisons. The tools compared are the alignment-based methods using ClustalW and Dialign, and the motif-finding programs MEME and FootPrinter. Our results show that methods taking the species' phylogenetic relationships into consideration obtain better accuracy. Mathieu Blanchette, Samson Kwong, Martin Tompa |
BIBE | 1 |
| 2003 | A comparative analysis method for detecting binding sites in coding regionsabstractWhile the problem of predicting transcription factor binding sites in a gene's promoter region has been extensively studied, binding sites located in coding regions are also crucial for regulating gene expression but are more difficult to detect. Coding region binding sites are mostly involved in splicing regulation, but also in transcriptional and post-transcriptional regulation. We consider the problem of predicting such binding sites by comparative analysis. Comparative analysis is based on the idea that functional sequences tend to evolve at slower rate than nonfunctional sequence, making unusually well conserved regions likely to be of interest. The difficulty in applying comparative analysis to the detection of binding sites located in coding sequence is that the whole sequence is under selective pressure, because it needs to code for a functional protein. We present a technique to distinguish between conservation due to constraints on the amino acid product and conservation due to constraints imposed by regulatory factors. More precisely, we show how to calculate the probability of observing a certain degree of conservation among the nucleotides of given set of orthologous codons, given a set of constraints on the amino acids they need to encode. The algorithms described are implemented in a program called Cosmo, available at http://bio.cs.washington.edu. We ran Cosmo on several genes known to contain exonic splicing enhancers and report the results. Mathieu Blanchette |
RECOMB | 1 |
| 2001 | Algorithms for phylogenetic footprintingabstractPhylogenetic footprinting is a technique that identifies regulatory elements by finding unusually well conserved regions in a set of orthologous non-coding DNA sequences from multiple species. In an earlier paper, we presented an exact algorithm that identifies the most conserved region of a set of sequences. Here, we present a number of algorithmic improvements that produce a 1000 fold speedup over the original algorithm. We also show how prior knowledge can be used to identify weaker motifs, and how to handle data sets in which only an unknown subset of the sequences contain the regulatory element. Each technique is implemented and successfully identifies a large number of known binding sites, as well as several highly conserved but uncharacterized regions. Mathieu Blanchette |
RECOMB | 1 |
| 2000 | An Exact Algorithm to Identify Motifs in Orthologous Sequences from Multiple Species
Mathieu Blanchette, Benno Schwikowski, Martin Tompa |
ISMB | 1 |
| 1999 | Probability models for genome rearrangement and linear invariants for phylogenetic inferenceabstractWe review the combinatorial optimization problems in calculating edit distances between genomes and phylogenetic inference based on minimizing gene order changes.With a view to avoiding the computational cost and the "long branches attract" artifact of some tree-building methods, we explore the probabiization of genome rearrangment models prior to developing a methodology based on branch-length invariants.We characterize probabilistically the evolution of the structure of the gene adjacency set for inversions on unsigned circular genomes and, using a non-trivial recurrence relation, inversions on signed genomes.Concepts from the theory of invariants developed for the phylogenetics of ho mologous gene sequences can be used to derive a complete set of linear invariants for unsigned inversions, as well as for a mixed rearrangement model for signed genomes, though not for pure transposition nor pure signed inversion models.The invariants are based on an extended Jukes-Cantor semigroup.We ilhrstrate the use of these invariants to relate mitochondrial genomes from a number of invertebrate animals. David Sankoff, Mathieu Blanchette |
RECOMB | 2 |
| 1998 | Multiple genome rearrangements
David Sankoff, Mathieu Blanchette |
RECOMB | 2 |
| 1997 | The Median Problem for Breakpoints in Comparative Genomics
David Sankoff, Mathieu Blanchette |
COCOON | 2 |