VLDB 2026 Research / reviewers in the wild / expert
Anthony Mathelier
dblp:75/8712
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-5127-5459ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | inMOTIFin: a lightweight end-to-end simulation software for regulatory sequencesabstractSUMMARY: The accurate development, assessment, interpretation, and benchmarking of bioinformatics frameworks for analyzing transcriptional regulatory grammars rely on controlled simulations to validate the underlying methods. However, existing simulators often lack end-to-end flexibility or ease of integration, which limits their practical use. We present inMOTIFin, a lightweight, modular, and user-friendly Python-based software that addresses these gaps by providing versatile and efficient simulation and modification of DNA regulatory sequences. inMOTIFin enables users to simulate or modify regulatory sequences efficiently for the customizable generation of motifs and insertion of motif instances with precise control over their positions, co-occurrences, and spacing, as well as direct modification of real sequences, facilitating a comprehensive evaluation of motif-based methods and interpretation tools. We demonstrate inMOTIFin applications for the assessment of de novo motif discovery, the analysis of transcription factor cooperativity, and the support of explainability analyses for deep learning models. inMOTIFin ensures robust and reproducible analyses for studying transcriptional regulatory grammars. AVAILABILITY AND IMPLEMENTATION: inMOTIFin is available at PyPI https://pypi.org/project/inMOTIFin/ and Docker Hub https://hub.docker.com/r/cbgr/inmotifin. Detailed documentation is available at https://inmotifin.readthedocs.io/en/latest/. The code for use case analyses is available at https://bitbucket.org/CBGR/inmotifin_evaluation/src/main/. The version of the code used for this article has been uploaded to Zenodo with DOI: 10.5281/zenodo.17638579. Katalin Ferenc, Lorenzo Martini, Ieva Rauluseviciute, Geir Kjetil Sandve, Anthony Mathelier |
Bioinform. | 5 |
| 2025 | Gene regulatory network integration with multi-omics data enhances survival predictions in cancerabstractThe emergence of high-throughput omics technologies has resulted in their wide application to cancer studies, greatly increasing our understanding of the disruptions occurring at different molecular levels. To fully harness these data, integrative approaches have emerged as essential tools, enabling the combination of multiple omics modalities to uncover disease mechanisms. However, many such approaches overlook gene regulatory mechanisms, which play a central role in the development and progression of cancer. Patient-specific gene regulatory networks (GRNs), representing interactions between regulators (such as transcription factors) and their target genes in each individual tumour, offer a powerful framework to bridge this gap and investigate the regulatory landscape of cancer. In this study, we introduce a novel approach for integrating patient-specific GRNs with multi-omic data and assess whether their inclusion in joint dimensionality reduction models improves survival prediction across multiple cancer types. By applying our method on ten cancer datasets from The Cancer Genome Atlas, we demonstrate that incorporating GRNs enhances associations with patient survival in several cancer types. Focusing on liver cancer, with validation in independent data, our methodology identifies potential mechanisms of gene regulatory dysregulation associated with cancer progression. These were linked to dysregulated fatty acid metabolism, and identified JUND as a potential novel transcriptional regulator driving these processes. Our findings highlight the value of network-based multi-omics integration for uncovering clinically relevant regulatory mechanisms and improving our understanding of cancer biology at the patient-specific level. Romana T. Pop, Ping-Han Hsieh, Tatiana Belova, Anthony Mathelier, Marieke L. Kuijjer |
Briefings Bioinform. | 4 |
| 2024 | Improving bioinformatics software quality through teamworkabstractSUMMARY: Since high-throughput techniques became a staple in biological science laboratories, computational algorithms, and scientific software have boomed. However, the development of bioinformatics software usually lacks software development quality standards. The resulting software code is hard to test, reuse, and maintain. We believe that the root of inefficiency in implementing the best software development practices in academic settings is the individualistic approach, which has traditionally been the norm for recognizing scientific achievements and, by extension, for developing specialized software. Software development is a collective effort in most software-heavy endeavors. Indeed, the literature suggests teamwork directly impacts code quality through knowledge sharing, collective software development, and established coding standards. In our computational biology research groups, we sustainably involve all group members in learning, sharing, and discussing software development while maintaining the personal ownership of research projects and related software products. We found that group members involved in this endeavor improved their coding skills, became more efficient bioinformaticians, and obtained detailed knowledge about their peers' work, triggering new collaborative projects. We strongly advocate for improving software development culture within bioinformatics through collective effort in computational biology groups or institutes with three or more bioinformaticians. AVAILABILITY AND IMPLEMENTATION: Additional information and guidance on how to get started is available at https://ferenckata.github.io/ImprovingSoftwareTogether.github.io/. Katalin Ferenc, Ieva Rauluseviciute, Ladislav Hovan, Marieke L. Kuijjer, Anthony Mathelier |
Bioinform. | 6 |
| 2024 | Integrative pan-cancer analysis reveals a common architecture of dysregulated transcriptional networks characterized by loss of enhancer methylationabstractAberrant DNA methylation contributes to gene expression deregulation in cancer. However, these alterations' precise regulatory role and clinical implications are still not fully understood. In this study, we performed expression-methylation Quantitative Trait Loci (emQTL) analysis to identify deregulated cancer-driving transcriptional networks linked to CpG demethylation pan-cancer. By analyzing 33 cancer types from The Cancer Genome Atlas, we identified and confirmed significant correlations between CpG methylation and gene expression (emQTL) in cis and trans, both across and within cancer types. Bipartite network analysis of the emQTL revealed groups of CpGs and genes related to important biological processes involved in carcinogenesis including proliferation, metabolism and hormone-signaling. These bipartite communities were characterized by loss of enhancer methylation in specific transcription factor binding regions (TFBRs) and the CpGs were topologically linked to upregulated genes through chromatin loops. Penalized Cox regression analysis showed a significant prognostic impact of the pan-cancer emQTL in many cancer types. Taken together, our integrative pan-cancer analysis reveals a common architecture where hallmark cancer-driving functions are affected by the loss of enhancer methylation and may be epigenetically regulated. Jørgen Ankill, Xavier Tekpli, Elin H. Kure, Vessela N. Kristensen, Anthony Mathelier, Thomas Fleischer |
PLoS Comput. Biol. | 6 |
| 2021 | BiasAway: command-line and web server to generate nucleotide composition-matched DNA background sequencesabstractMOTIVATION: Accurate motif enrichment analyses depend on the choice of background DNA sequences used, which should ideally match the sequence composition of the foreground sequences. It is important to avoid false positive enrichment due to sequence biases in the genome, such as GC-bias. Therefore, relying on an appropriate set of background sequences is crucial for enrichment analysis. RESULTS: We developed BiasAway, a command line tool and its dedicated easy-to-use web server to generate synthetic sequences matching any k-mer nucleotide composition or select genomic DNA sequences matching the mononucleotide composition of the foreground sequences through four different models. For genomic sequences, we provide precomputed partitions of genomes from nine species with five different bin sizes to generate appropriate genomic background sequences. AVAILABILITY AND IMPLEMENTATION: BiasAway source code is freely available from Bitbucket (https://bitbucket.org/CBGR/biasaway) and can be easily installed using bioconda or pip. The web server is available at https://biasaway.uio.no and a detailed documentation is available at https://biasaway.readthedocs.io. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aziz Khan 0001, Rafael Riudavets Puig, Paul Boddie, Anthony Mathelier |
Bioinform. | 4 |
| 2020 | Beware the Jaccard: the choice of similarity measure is important and non-trivial in genomic colocalisation analysisabstractThe generation and systematic collection of genome-wide data is ever-increasing. This vast amount of data has enabled researchers to study relations between a variety of genomic and epigenomic features, including genetic variation, gene regulation and phenotypic traits. Such relations are typically investigated by comparatively assessing genomic co-occurrence. Technically, this corresponds to assessing the similarity of pairs of genome-wide binary vectors. A variety of similarity measures have been proposed for this problem in other fields like ecology. However, while several of these measures have been employed for assessing genomic co-occurrence, their appropriateness for the genomic setting has never been investigated. We show that the choice of similarity measure may strongly influence results and propose two alternative modelling assumptions that can be used to guide this choice. On both simulated and real genomic data, the Jaccard index is strongly altered by dataset size and should be used with caution. The Forbes coefficient (fold change) and tetrachoric correlation are less influenced by dataset size, but one should be aware of increased variance for small datasets. All results on simulated and real data can be inspected and reproduced at https://hyperbrowser.uio.no/sim-measure. Stefania Salvatore, Knut D. Rand, Ivar Grytten, Egil Ferkingstad, Diana Domanska, Lars Holden, Marius Gheorghe, Anthony Mathelier, Ingrid Kristine Glad, Geir Kjetil Sandve |
Briefings Bioinform. | 8 |
| 2018 | JASPAR RESTful API: accessing JASPAR data from any programming languageabstractSummary: JASPAR is a widely used open-access database of curated, non-redundant transcription factor binding profiles. Currently, data from JASPAR can be retrieved as flat files or by using programming language-specific interfaces. Here, we present a programming language-independent application programming interface (API) to access JASPAR data using the Representational State Transfer (REST) architecture. The REST API enables programmatic access to JASPAR by most programming languages and returns data in eight widely used formats. Several endpoints are available to access the data and an endpoint is available to infer the TF binding profile(s) likely bound by a given DNA binding domain protein sequence. Additionally, it provides an interactive browsable interface for bioinformatics tool developers. Availability and implementation: This REST API is implemented in Python using the Django REST Framework. It is accessible at http://jaspar.genereg.net/api/ and the source code is freely available at https://bitbucket.org/CBGR/jaspar under GPL v3 license. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Aziz Khan 0001, Anthony Mathelier |
Bioinform. | 2 |
| 2017 | Intervene: a tool for intersection and visualization of multiple gene or genomic region setsabstractBACKGROUND: A common task for scientists relies on comparing lists of genes or genomic regions derived from high-throughput sequencing experiments. While several tools exist to intersect and visualize sets of genes, similar tools dedicated to the visualization of genomic region sets are currently limited. RESULTS: To address this gap, we have developed the Intervene tool, which provides an easy and automated interface for the effective intersection and visualization of genomic region or list sets, thus facilitating their analysis and interpretation. Intervene contains three modules: venn to generate Venn diagrams of up to six sets, upset to generate UpSet plots of multiple sets, and pairwise to compute and visualize intersections of multiple sets as clustered heat maps. Intervene, and its interactive web ShinyApp companion, generate publication-quality figures for the interpretation of genomic region and list sets. CONCLUSIONS: Intervene and its web application companion provide an easy command line and an interactive web interface to compute intersections of multiple genomic and list sets. They have the capacity to plot intersections using easy-to-interpret visual approaches. Intervene is developed and designed to meet the needs of both computer scientists and biologists. The source code is freely available at https://bitbucket.org/CBGR/intervene , with the web application available at https://asntech.shinyapps.io/intervene . Aziz Khan 0001, Anthony Mathelier |
BMC Bioinform. | 2 |
| 2016 | CAGEd-oPOSSUM: motif enrichment analysis from CAGE-derived TSSsabstractUNLABELLED: With the emergence of large-scale Cap Analysis of Gene Expression (CAGE) datasets from individual labs and the FANTOM consortium, one can now analyze the cis-regulatory regions associated with gene transcription at an unprecedented level of refinement. By coupling transcription factor binding site (TFBS) enrichment analysis with CAGE-derived genomic regions, CAGEd-oPOSSUM can identify TFs that act as key regulators of genes involved in specific mammalian cell and tissue types. The webtool allows for the analysis of CAGE-derived transcription start sites (TSSs) either provided by the user or selected from ∼1300 mammalian samples from the FANTOM5 project with pre-computed TFBS predicted with JASPAR TF binding profiles. The tool helps power insights into the regulation of genes through the study of the specific usage of TSSs within specific cell types and/or under specific conditions. AVAILABILITY AND IMPLEMENTATION: The CAGEd-oPOSUM web tool is implemented in Perl, MySQL and Apache and is available at http://cagedop.cmmt.ubc.ca/CAGEd_oPOSSUM CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David J. Arenillas, Alistair R. R. Forrest, Hideya Kawaji, Timo Lassmann, Wyeth W. Wasserman, Anthony Mathelier |
Bioinform. | 6 |
| 2013 | The Next Generation of Transcription Factor Binding Site PredictionabstractFinding where transcription factors (TFs) bind to the DNA is of key importance to decipher gene regulation at a transcriptional level. Classically, computational prediction of TF binding sites (TFBSs) is based on basic position weight matrices (PWMs) which quantitatively score binding motifs based on the observed nucleotide patterns in a set of TFBSs for the corresponding TF. Such models make the strong assumption that each nucleotide participates independently in the corresponding DNA-protein interaction and do not account for flexible length motifs. We introduce transcription factor flexible models (TFFMs) to represent TF binding properties. Based on hidden Markov models, TFFMs are flexible, and can model both position interdependence within TFBSs and variable length motifs within a single dedicated framework. The availability of thousands of experimentally validated DNA-TF interaction sequences from ChIP-seq allows for the generation of models that perform as well as PWMs for stereotypical TFs and can improve performance for TFs with flexible binding characteristics. We present a new graphical representation of the motifs that convey properties of position interdependence. TFFMs have been assessed on ChIP-seq data sets coming from the ENCODE project, revealing that they can perform better than both PWMs and the dinucleotide weight matrix extension in discriminating ChIP-seq from background sequences. Under the assumption that ChIP-seq signal values are correlated with the affinity of the TF-DNA binding, we find that TFFM scores correlate with ChIP-seq peak signals. Moreover, using available TF-DNA affinity measurements for the Max TF, we demonstrate that TFFMs constructed from ChIP-seq data correlate with published experimentally measured DNA-binding affinities. Finally, TFFMs allow for the straightforward computation of an integrated TF occupancy score across a sequence. These results demonstrate the capacity of TFFMs to accurately model DNA-protein interactions, while providing a single unified framework suitable for the next generation of TFBS prediction. Anthony Mathelier, Wyeth W. Wasserman |
PLoS Comput. Biol. | 1 |
| 2010 | MIReNA: finding microRNAs with high accuracy and no learning at genome scale and from deep sequencing dataabstractMOTIVATION: MicroRNAs (miRNAs) are a class of endogenes derived from a precursor (pre-miRNA) and involved in post-transcriptional regulation. Experimental identification of novel miRNAs is difficult because they are often transcribed under specific conditions and cell types. Several computational methods were developed to detect new miRNAs starting from known ones or from deep sequencing data, and to validate their pre-miRNAs. RESULTS: We present a genome-wide search algorithm, called MIReNA, that looks for miRNA sequences by exploring a multidimensional space defined by only five (physical and combinatorial) parameters characterizing acceptable pre-miRNAs. MIReNA validates pre-miRNAs with high sensitivity and specificity, and detects new miRNAs by homology from known miRNAs or from deep sequencing data. A performance comparison between MIReNA and four available predictive systems has been done. MIReNA approach is strikingly simple but it turns out to be powerful at least as much as more sophisticated algorithmic methods. MIReNA obtains better results than three known algorithms that validate pre-miRNAs. It demonstrates that machine-learning is not a necessary algorithmic approach for pre-miRNAs computational validation. In particular, machine learning algorithms can only confirm pre-miRNAs that look alike known ones, this being a limitation while exploring species with no known pre-miRNAs. The possibility to adapt the search to specific species, possibly characterized by specific properties of their miRNAs and pre-miRNAs, is a major feature of MIReNA. A parameter adjustment calibrates specificity and sensitivity in MIReNA, a key feature for predictive systems, which is not present in machine learning approaches. Comparison of MIReNA with miRDeep using deep sequencing data to predict miRNAs highlights a highly specific predictive power of MIReNA. AVAILABILITY: At the address http://www.ihes.fr/carbone/data8/. Anthony Mathelier, Alessandra Carbone |
Bioinform. | 1 |