María Rodríguez Martínez

dblp:178/8685 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
11since 2021 · last 2025
0000-0003-3766-4233ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 SurvBoard: standardized benchmarking for multi-omics cancer survival models
abstract
Multi-omics data, which include genomic, transcriptomic, epigenetic, and proteomic data, are gaining increasing importance for determining the clinical outcomes of cancer patients. Several recent studies have evaluated various multimodal integration strategies for cancer survival prediction, highlighting the need for standardizing model performance results. Addressing this issue, we introduce SurvBoard, a benchmark framework that standardizes key experimental design choices. SurvBoard enables comparisons between single-cancer and pan-cancer data models and assesses the benefits of using patient data with missing modalities. We also address common pitfalls in preprocessing and validating multi-omics cancer survival models. We apply SurvBoard to several exemplary use cases, further confirming that statistical models tend to outperform deep learning methods, especially for metrics measuring survival function calibration. Moreover, most models exhibit better performance when trained in a pan-cancer context and can benefit from leveraging samples for which data of some omics modalities are missing. We provide a web service for model evaluation and to make our benchmark results easily accessible and viewable: https://www.survboard.science/. All code is available on GitHub: https://github.com/BoevaLab/survboard/. All benchmark outputs are available on Zenodo: 10.5281/zenodo.11066226. A video tutorial on how to use the Survboard leaderboard is available on YouTube at https://youtu.be/HJrdpJP8Vvk.
David Wissel, Nikita Janakarajan, Aayush Grover, Enrico Toniato, María Rodríguez Martínez, Valentina Boeva
Briefings Bioinform.5
2024 Conformal Autoregressive Generation: Beam Search with Coverage Guarantees
abstract
We introduce two new extensions to the beam search algorithm based on conformal predictions (CP) to produce sets of sequences with theoretical coverage guarantees. The first method is very simple and proposes dynamically-sized subsets of beam search results but, unlike typical CP proceedures, has an upper bound on the achievable guarantee depending on a post-hoc calibration measure. Our second algorithm introduces the conformal set prediction procedure as part of the decoding process, producing a variable beam width which adapts to the current uncertainty. While more complex, this procedure can achieve coverage guarantees selected a priori. We provide marginal coverage bounds as well as calibration-conditional guarantees for each method, and evaluate them empirically on a selection of tasks drawing from natural language processing and chemistry.
Nicolas Deutschmann, Marvin Alberts, María Rodríguez Martínez
AAAI3
2023 Aligned Diffusion Schrödinger Bridges
abstract
Diffusion Schrödinger bridges (DSBs) have recently emerged as a powerful framework for recovering stochastic dynamics via their marginal observations at different time points. Despite numerous successful applications, existing algorithms for solving DSBs have so far failed to utilize the structure of aligned data, which naturally arises in many biological phenomena. In this paper, we propose a novel algorithmic framework that, for the first time, solves DSBs while respecting the data alignment. Our approach hinges on a combination of two decades-old ideas: The classical Schrödinger bridge theory and Doob’s $h$-transform. Compared to prior methods, our approach leads to a simpler training procedure with lower variance, which we further augment with principled regularization schemes. This ultimately leads to sizeable improvements across experiments on synthetic and real data, including the tasks of predicting conformational changes in proteins and temporal evolution of cellular differentiation processes.
Vignesh Ram Somnath, Matteo Pariset, Ya-Ping Hsieh, María Rodríguez Martínez, Andreas Krause 0001, Charlotte Bunne
UAI4
2023 FLAN: feature-wise latent additive neural models for biological applications
abstract
MOTIVATION: Interpretability has become a necessary feature for machine learning models deployed in critical scenarios, e.g. legal system, healthcare. In these situations, algorithmic decisions may have (potentially negative) long-lasting effects on the end-user affected by the decision. While deep learning models achieve impressive results, they often function as a black-box. Inspired by linear models, we propose a novel class of structurally constrained deep neural networks, which we call FLAN (Feature-wise Latent Additive Networks). Crucially, FLANs process each input feature separately, computing for each of them a representation in a common latent space. These feature-wise latent representations are then simply summed, and the aggregated representation is used for the prediction. These feature-wise representations allow a user to estimate the effect of each individual feature independently from the others, similarly to the way linear models are interpreted. RESULTS: We demonstrate FLAN on a series of benchmark datasets in different biological domains. Our experiments show that FLAN achieves good performances even in complex datasets (e.g. TCR-epitope binding prediction), despite the structural constraint we imposed. On the other hand, this constraint enables us to interpret FLAN by deciphering its decision process, as well as obtaining biological insights (e.g. by identifying the marker genes of different cell populations). In supplementary experiments, we show similar performances also on non-biological datasets. CODE AND DATA AVAILABILITY: Code and example data are available at https://github.com/phineasng/flan_bio.
An-Phi Nguyen, Stefania Vasilaki, María Rodríguez Martínez
Briefings Bioinform.3
2022 PCfun: a hybrid computational framework for systematic characterization of protein complex function
abstract
In molecular biology, it is a general assumption that the ensemble of expressed molecules, their activities and interactions determine biological function, cellular states and phenotypes. Stable protein complexes-or macromolecular machines-are, in turn, the key functional entities mediating and modulating most biological processes. Although identifying protein complexes and their subunit composition can now be done inexpensively and at scale, determining their function remains challenging and labor intensive. This study describes Protein Complex Function predictor (PCfun), the first computational framework for the systematic annotation of protein complex functions using Gene Ontology (GO) terms. PCfun is built upon a word embedding using natural language processing techniques based on 1 million open access PubMed Central articles. Specifically, PCfun leverages two approaches for accurately identifying protein complex function, including: (i) an unsupervised approach that obtains the nearest neighbor (NN) GO term word vectors for a protein complex query vector and (ii) a supervised approach using Random Forest (RF) models trained specifically for recovering the GO terms of protein complex queries described in the CORUM protein complex database. PCfun consolidates both approaches by performing a hypergeometric statistical test to enrich the top NN GO terms within the child terms of the GO terms predicted by the RF models. The documentation and implementation of the PCfun package are available at https://github.com/sharmavaruns/PCfun. We anticipate that PCfun will serve as a useful tool and novel paradigm for the large-scale characterization of protein complex function.
Varun S. Sharma, Andrea Fossati, Rodolfo Ciuffa, Marija Buljan, Evan G. Williams, Zhen Chen 0009, Wenguang Shao, Patrick G. A. Pedrioli, Anthony W. Purcell, María Rodríguez Martínez, Jiangning Song, Matteo Manica, Ruedi Aebersold, Chen Li 0021
Briefings Bioinform.10
2022 Computational modelling in health and disease: highlights of the 6th annual SysMod meeting
abstract
Abstract Summary The Community of Special Interest (COSI) in Computational Modelling of Biological Systems (SysMod) brings together interdisciplinary scientists interested in combining data-driven computational modelling, multi-scale mechanistic frameworks, large-scale -omics data and bioinformatics. SysMod’s main activity is an annual meeting at the Intelligent Systems for Molecular Biology (ISMB) conference, a meeting for computer scientists, biologists, mathematicians, engineers and computational and systems biologists. The 2021 SysMod meeting was conducted virtually due to the ongoing COVID-19 pandemic (coronavirus disease 2019). During the 2-day meeting, the development of computational tools, approaches and predictive models was discussed, along with their application to biological systems, emphasizing disease mechanisms. This report summarizes the meeting. Availability and implementation All resources and further information are freely accessible at https://sysmod.info.
Anna Niarakis, Juilee Thakar, Matteo Barberis, María Rodríguez Martínez, Tomás Helikar, Marc R. Birtwistle, Claudine Chaouiya, Laurence Calzone, Andreas Dräger
Bioinform.4
2022 DECODE: a computational pipeline to discover T cell receptor binding rules
abstract
MOTIVATION: Understanding the mechanisms underlying T cell receptor (TCR) binding is of fundamental importance to understanding adaptive immune responses. A better understanding of the biochemical rules governing TCR binding can be used, e.g. to guide the design of more powerful and safer T cell-based therapies. Advances in repertoire sequencing technologies have made available millions of TCR sequences. Data abundance has, in turn, fueled the development of many computational models to predict the binding properties of TCRs from their sequences. Unfortunately, while many of these works have made great strides toward predicting TCR specificity using machine learning, the black-box nature of these models has resulted in a limited understanding of the rules that govern the binding of a TCR and an epitope. RESULTS: We present an easy-to-use and customizable computational pipeline, DECODE, to extract the binding rules from any black-box model designed to predict the TCR-epitope binding. DECODE offers a range of analytical and visualization tools to guide the user in the extraction of such rules. We demonstrate our pipeline on a recently published TCR-binding prediction model, TITAN, and show how to use the provided metrics to assess the quality of the computed rules. In conclusion, DECODE can lead to a better understanding of the sequence motifs that underlie TCR binding. Our pipeline can facilitate the investigation of current immunotherapeutic challenges, such as cross-reactive events due to off-target TCR binding. AVAILABILITY AND IMPLEMENTATION: Code is available publicly at https://github.com/phineasng/DECODE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Iliana Papadopoulou, An-Phi Nguyen, Anna Weber, María Rodríguez Martínez
Bioinform.4
2021 On the feasibility of deep learning applications using raw mass spectrometry data
abstract
SUMMARY: In recent years, SWATH-MS has become the proteomic method of choice for data-independent-acquisition, as it enables high proteome coverage, accuracy and reproducibility. However, data analysis is convoluted and requires prior information and expert curation. Furthermore, as quantification is limited to a small set of peptides, potentially important biological information may be discarded. Here we demonstrate that deep learning can be used to learn discriminative features directly from raw MS data, eliminating hence the need of elaborate data processing pipelines. Using transfer learning to overcome sample sparsity, we exploit a collection of publicly available deep learning models already trained for the task of natural image classification. These models are used to produce feature vectors from each mass spectrometry (MS) raw image, which are later used as input for a classifier trained to distinguish tumor from normal prostate biopsies. Although the deep learning models were originally trained for a completely different classification task and no additional fine-tuning is performed on them, we achieve a highly remarkable classification performance of 0.876 AUC. We investigate different types of image preprocessing and encoding. We also investigate whether the inclusion of the secondary MS2 spectra improves the classification performance. Throughout all tested models, we use standard protein expression vectors as gold standards. Even with our naïve implementation, our results suggest that the application of deep learning and transfer learning techniques might pave the way to the broader usage of raw mass spectrometry data in real-time diagnosis. AVAILABILITY AND IMPLEMENTATION: The open source code used to generate the results from MS images is available on GitHub: https://ibm.biz/mstransc. The raw MS data underlying this article cannot be shared publicly for the privacy of individuals that participated in the study. Processed data including the MS images, their encodings, classification labels and results can be accessed at the following link: https://ibm.box.com/v/mstc-supplementary. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joris Cadow, Matteo Manica, Roland Mathis, Tiannan Guo, Ruedi Aebersold, María Rodríguez Martínez
Bioinform.6
2021 SysMod: the ISCB community for data-driven computational modelling and multi-scale analysis of biological systems
abstract
Computational models of biological systems can exploit a broad range of rapidly developing approaches, including novel experimental approaches, bioinformatics data analysis, emerging modelling paradigms, data standards and algorithms. A discussion about the most recent advances among experts from various domains is crucial to foster data-driven computational modelling and its growing use in assessing and predicting the behaviour of biological systems. Intending to encourage the development of tools, approaches and predictive models, and to deepen our understanding of biological systems, the Community of Special Interest (COSI) was launched in Computational Modelling of Biological Systems (SysMod) in 2016. SysMod's main activity is an annual meeting at the Intelligent Systems for Molecular Biology (ISMB) conference, which brings together computer scientists, biologists, mathematicians, engineers, computational and systems biologists. In the five years since its inception, SysMod has evolved into a dynamic and expanding community, as the increasing number of contributions and participants illustrate. SysMod maintains several online resources to facilitate interaction among the community members, including an online forum, a calendar of relevant meetings and a YouTube channel with talks and lectures of interest for the modelling community. For more than half a decade, the growing interest in computational systems modelling and multi-scale data integration has inspired and supported the SysMod community. Its members get progressively more involved and actively contribute to the annual COSI meeting and several related community workshops and meetings, focusing on specific topics, including particular techniques for computational modelling or standardisation efforts.
Andreas Dräger, Tomás Helikar, Matteo Barberis, Marc R. Birtwistle, Laurence Calzone, Claudine Chaouiya, Jan Hasenauer, Jonathan R. Karr, Anna Niarakis, María Rodríguez Martínez, Julio Saez-Rodriguez, Juilee Thakar
Bioinform.10
2021 COSIFER: a Python package for the consensus inference of molecular interaction networks
abstract
SUMMARY: The advent of high-throughput technologies has provided researchers with measurements of thousands of molecular entities and enable the investigation of the internal regulatory apparatus of the cell. However, network inference from high-throughput data is far from being a solved problem. While a plethora of different inference methods have been proposed, they often lead to non-overlapping predictions, and many of them lack user-friendly implementations to enable their broad utilization. Here, we present Consensus Interaction Network Inference Service (COSIFER), a package and a companion web-based platform to infer molecular networks from expression data using state-of-the-art consensus approaches. COSIFER includes a selection of state-of-the-art methodologies for network inference and different consensus strategies to integrate the predictions of individual methods and generate robust networks. AVAILABILITY AND IMPLEMENTATION: COSIFER Python source code is available at https://github.com/PhosphorylatedRabbits/cosifer. The web service is accessible at https://ibm.biz/cosifer-aas. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matteo Manica, Charlotte Bunne, Roland Mathis, Joris Cadow, Mehmet Eren Ahsen, Gustavo Stolovitzky, María Rodríguez Martínez
Bioinform.7
2021 TITAN: T-cell receptor specificity prediction with bimodal attention networks
abstract
MOTIVATION: The activity of the adaptive immune system is governed by T-cells and their specific T-cell receptors (TCR), which selectively recognize foreign antigens. Recent advances in experimental techniques have enabled sequencing of TCRs and their antigenic targets (epitopes), allowing to research the missing link between TCR sequence and epitope binding specificity. Scarcity of data and a large sequence space make this task challenging, and to date only models limited to a small set of epitopes have achieved good performance. Here, we establish a k-nearest-neighbor (K-NN) classifier as a strong baseline and then propose Tcr epITope bimodal Attention Networks (TITAN), a bimodal neural network that explicitly encodes both TCR sequences and epitopes to enable the independent study of generalization capabilities to unseen TCRs and/or epitopes. RESULTS: By encoding epitopes at the atomic level with SMILES sequences, we leverage transfer learning and data augmentation to enrich the input data space and boost performance. TITAN achieves high performance in the prediction of specificity of unseen TCRs (ROC-AUC 0.87 in 10-fold CV) and surpasses the results of the current state-of-the-art (ImRex) by a large margin. Notably, our Levenshtein-based K-NN classifier also exhibits competitive performance on unseen TCRs. While the generalization to unseen epitopes remains challenging, we report two major breakthroughs. First, by dissecting the attention heatmaps, we demonstrate that the sparsity of available epitope data favors an implicit treatment of epitopes as classes. This may be a general problem that limits unseen epitope performance for sufficiently complex models. Second, we show that TITAN nevertheless exhibits significantly improved performance on unseen epitopes and is capable of focusing attention on chemically meaningful molecular structures. AVAILABILITY AND IMPLEMENTATION: The code as well as the dataset used in this study is publicly available at https://github.com/PaccMann/TITAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anna Weber, Jannis Born, María Rodríguez Martínez
Bioinform.3
2020 PaccMannRL: Designing Anticancer Drugs From Transcriptomic Data via Reinforcement Learning
Jannis Born, Matteo Manica, Ali Oskooei, Joris Cadow, María Rodríguez Martínez
RECOMB5
2020 FPGA Accelerated Analysis of Boolean Gene Regulatory Networks
abstract
Boolean models are a powerful abstraction for qualitative modeling of gene regulatory networks. With the recent availability of advanced high-throughput technologies, Boolean models have increasingly grown in size and complexity, posing a challenge for existing software simulation tools that have not scaled at the same speed. Field Programmable Gate Arrays (FPGAs) are powerful reconfigurable integrated circuits that can offer massive performance improvements. Due to their highly parallel nature, FPGAs are well suited to simulate complex molecular networks. We present here a new simulation framework for Boolean models, which first converts the model to Verilog, a standardized hardware description language, and then connects it to an execution core that runs on an FPGA coherently attached to a POWER8 processor. We report an order of magnitude speedup over a multi-threaded software simulation tool running on the same processor on a selection of Boolean models. Analysis on a T-cell large granular lymphocyte leukemia (T-LGL) demonstrates that our framework achieves consistent performance improvements resulting in new biological insights. In addition, we show that our solution allows to perform attractor detection at an unprecedented speed, exhibiting a speedup ranging from one to three orders of magnitude compared to alternative software solutions.
Matteo Manica, Raphael Polig, Mitra Purandare, Roland Mathis, Christoph Hagleitner, María Rodríguez Martínez
IEEE ACM Trans. Comput. Biol. Bioinform.6
2016 Marginalized Continuous Time Bayesian Networks for Network Reconstruction from Incomplete Observations
abstract
Continuous Time Bayesian Networks (CTBNs) provide a powerful means to model complex network dynamics. How- ever, their inference is computationally demanding — especially if one considers incomplete and noisy time-series data. The latter gives rise to a joint state- and parameter estimation problem, which can only be solved numerically. Yet, finding the exact parameterization of the CTBN has often only secondary importance in practical scenarios. We therefore focus on the structure learning problem and present a way to analytically marginalize the Markov chain underlying the CTBN model with respect its parameters. Since the resulting stochastic process is parameter-free, its inference reduces to an optimal filtering problem. We solve the latter using an efficient parallel implementation of a sequential Monte Carlo scheme. Our framework enables CTBN inference to be applied to incomplete noisy time-series data frequently found in molecular biology and other disciplines.
Lukas Studer, Loïc Paulevé, Christoph Zechner, Matthias Reumann, María Rodríguez Martínez, Heinz Koeppl
AAAI5