Vítor A. P. Martins dos Santos

dblp:71/10996 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-2352-9017ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Response to Letter to Editor by A. Derbalah et al.: the role of automation in enhancing reproducibility and interoperability of PBPK models
abstract
Dear colleagues, We thank you for your comments on our manuscript [1] and for raising the important question of automation [2]. The publications by Sepp et al. (2019) [3] and Liu et al. (2024) [4] focus on one type of PBPK models, denominated ‘for biologics’. In these models, each compartment (tissue/organ) is divided into vascular, endothelial endosomal and interstitial (sub-)compartments, and lymphatic flows are also represented [3, 4]. In these examples, the pharmacokinetics of proteins are represented. Notably, the protein distribution within tissues includes processes of passive transport across two types of pores via diffusion or fluid convection, processes of pinocytosis, binding, recycling and degradation [4]. This type of PBPK models is complex and their authors successfully used specific tools for automated code generation. In particular, Liu et al. (2024) used ‘mathematical sets’ available in the Ubiquity package, which facilitates assembling model components in R-language [4]. Sepp et al. (2019) used the MATLAB script PBPKassembler.m for automated code generation, where files in Simbiology (MATLAB) and Excel formats were combined [3]. Briefly, automation can help build models, which is highly valuable notably for complex models. Having a robust, thoroughly validated platform for automated code generation will indeed contribute to accessibility and reproducibility. Also, automation of model building seems to be a natural way to address the case of large models, in which manual scripting will be prone to introducing errors. Smaller models or those with fewer equations do not necessarily require automation. From our experience, writing the script manually provides a deep understanding of the underlying model mechanics. Manually written scripts also allow to give a clear view of how equations and parameters are applied, which we consider especially valuable for teaching and knowledge transfer. To conclude, as observed in other niches of systems biology, automation brings significant benefits to the modelling field, e.g. there is a plethora of tools that reconstruct and benchmark genome-scale metabolic models. As the correspondence authors argue, there is more scope for automation in the current practices of PBPK modelling. To facilitate this, we have opted to begin the journey of standardization via open collaboration through ELIXIR. We are looking forward to continuing to engage the research community towards automatic construction, validation and deposition of PBPK models. Thank you. Best regards, Elena Domínguez-Romero, Stanislav Mazurenko, Martin Scheringer, Vítor Martins dos Santos, Chris Evelo, Mihail Anton, John M. Hancock, Anže Županič, and Maria Suarez-Diez. None declared.
Elena Domínguez Romero, Stanislav Mazurenko, Martin Scheringer, Vítor A. P. Martins dos Santos, Chris T. A. Evelo, Mihail Anton, John M. Hancock, Anze Zupanic, María Suárez-Diez
Briefings Bioinform.4
2024 Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping
abstract
Unsupervised learning, particularly clustering, plays a pivotal role in disease subtyping and patient stratification, especially with the abundance of large-scale multi-omics data. Deep learning models, such as variational autoencoders (VAEs), can enhance clustering algorithms by leveraging inter-individual heterogeneity. However, the impact of confounders-external factors unrelated to the condition, e.g. batch effect or age-on clustering is often overlooked, introducing bias and spurious biological conclusions. In this work, we introduce four novel VAE-based deconfounding frameworks tailored for clustering multi-omics data. These frameworks effectively mitigate confounding effects while preserving genuine biological patterns. The deconfounding strategies employed include (i) removal of latent features correlated with confounders, (ii) a conditional VAE, (iii) adversarial training, and (iv) adding a regularization term to the loss function. Using real-life multi-omics data from The Cancer Genome Atlas, we simulated various confounding effects (linear, nonlinear, categorical, mixed) and assessed model performance across 50 repetitions based on reconstruction error, clustering stability, and deconfounding efficacy. Our results demonstrate that our novel models, particularly the conditional multi-omics VAE (cXVAE), successfully handle simulated confounding effects and recover biologically driven clustering structures. cXVAE accurately identifies patient labels and unveils meaningful pathological associations among cancer types, validating deconfounded representations. Furthermore, our study suggests that some of the proposed strategies, such as adversarial training, prove insufficient in confounder removal. In summary, our study contributes by proposing innovative frameworks for simultaneous multi-omics data integration, dimensionality reduction, and deconfounding in clustering. Benchmarking on open-access data offers guidance to end-users, facilitating meaningful patient stratification for optimized precision medicine.
Zuqi Li, Sonja Katz, Edoardo Saccenti, David W. Fardo, Peter Claes, Vítor A. P. Martins dos Santos, Kristel Van Steen, Gennady Roshchupkin
Briefings Bioinform.6
2024 Making PBPK models more reproducible in practice
abstract
Systems biology aims to understand living organisms through mathematically modeling their behaviors at different organizational levels, ranging from molecules to populations. Modeling involves several steps, from determining the model purpose to developing the mathematical model, implementing it computationally, simulating the model's behavior, evaluating, and refining the model. Importantly, model simulation results must be reproducible, ensuring that other researchers can obtain the same results after writing the code de novo and/or using different software tools. Guidelines to increase model reproducibility have been published. However, reproducibility remains a major challenge in this field. In this paper, we tackle this challenge for physiologically-based pharmacokinetic (PBPK) models, which represent the pharmacokinetics of chemicals following exposure in humans or animals. We summarize recommendations for PBPK model reporting that should apply during model development and implementation, in order to ensure model reproducibility and comprehensibility. We make a proposal aiming to harmonize abbreviations used in PBPK models. To illustrate these recommendations, we present an original and reproducible PBPK model code in MATLAB, alongside an example of MATLAB code converted to Systems Biology Markup Language format using MOCCASIN. As directions for future improvement, more tools to convert computational PBPK models from different software platforms into standard formats would increase the interoperability of these models. The application of other systems biology standards to PBPK models is encouraged. This work is the result of an interdisciplinary collaboration involving the ELIXIR systems biology community. More interdisciplinary collaborations like this would facilitate further harmonization and application of good modeling practices in different systems biology fields.
Elena Domínguez Romero, Stanislav Mazurenko, Martin Scheringer, Vítor A. P. Martins dos Santos, Chris T. A. Evelo, Mihail Anton, John M. Hancock, Anze Zupanic, María Suárez-Diez
Briefings Bioinform.4
2022 The Digitalization of Bioassays in the Open Research Knowledge Graph
Jennifer D'Souza 0001, Anita Monteverdi, Muhammad Haris 0001, Marco Anteghini, Kheir Eddine Farfar, Markus Stocker, Vítor A. P. Martins dos Santos, Sören Auer
DEXA (1)7
2022 SALARECON connects the Atlantic salmon genome to growth and feed efficiency
abstract
Atlantic salmon (Salmo salar) is the most valuable farmed fish globally and there is much interest in optimizing its genetics and rearing conditions for growth and feed efficiency. Marine feed ingredients must be replaced to meet global demand, with challenges for fish health and sustainability. Metabolic models can address this by connecting genomes to metabolism, which converts nutrients in the feed to energy and biomass, but such models are currently not available for major aquaculture species such as salmon. We present SALARECON, a model focusing on energy, amino acid, and nucleotide metabolism that links the Atlantic salmon genome to metabolic fluxes and growth. It performs well in standardized tests and captures expected metabolic (in)capabilities. We show that it can explain observed hypoxic growth in terms of metabolic fluxes and apply it to aquaculture by simulating growth with commercial feed ingredients. Predicted limiting amino acids and feed efficiencies agree with data, and the model suggests that marine feed efficiency can be achieved by supplementing a few amino acids to plant- and insect-based feeds. SALARECON is a high-quality model that makes it possible to simulate Atlantic salmon metabolism and growth. It can be used to explain Atlantic salmon physiology and address key challenges in aquaculture such as development of sustainable feeds.
Maksim V. Zakhartsev, Filip Rotnes, Marie Gulla, Ove Øyås, Jesse C. J. van Dam, María Suárez-Diez, Fabian Grammes, Róbert Anton Hafþórsson, Wout van Helvoirt, Jasper J. Koehorst, Peter J. Schaap, Liv Torunn Mydland, Arne B. Gjuvsland, Simen R. Sandve, Vítor A. P. Martins dos Santos, Jon Olav Vik
PLoS Comput. Biol.16
2021 Exploring the associations between transcript levels and fluxes in constraint-based models of metabolism
abstract
BACKGROUND: Several computational methods have been developed that integrate transcriptomics data with genome-scale metabolic reconstructions to increase accuracy of inferences of intracellular metabolic flux distributions. Even though existing methods use transcript abundances as a proxy for enzyme activity, each method uses a different hypothesis and assumptions. Most methods implicitly assume a proportionality between transcript levels and flux through the corresponding function, although these proportionality constant(s) are often not explicitly mentioned nor discussed in any of the published methods. E-Flux is one such method and, in this algorithm, flux bounds are related to expression data, so that reactions associated with highly expressed genes are allowed to carry higher flux values. RESULTS: Here, we extended E-Flux and systematically evaluated the impact of an assumed proportionality constant on model predictions. We used data from published experiments with Escherichia coli and Saccharomyces cerevisiae and we compared the predictions of the algorithm to measured extracellular and intracellular fluxes. CONCLUSION: We showed that detailed modelling using a proportionality constant can greatly impact the outcome of the analysis. This increases accuracy and allows for extraction of better physiological information.
Neeraj Sinha, Evert M. van Schothorst, Guido J. E. K. Hooiveld, Jaap Keijer, Vítor A. P. Martins dos Santos, María Suárez-Diez
BMC Bioinform.5
2018 SAPP: functional genome annotation and analysis through a semantic framework using FAIR principles
abstract
Summary: To unlock the full potential of genome data and to enhance data interoperability and reusability of genome annotations we have developed SAPP, a Semantic Annotation Platform with Provenance. SAPP is designed as an infrastructure supporting FAIR de novo computational genomics but can also be used to process and analyze existing genome annotations. SAPP automatically predicts, tracks and stores structural and functional annotations and associated dataset- and element-wise provenance in a Linked Data format, thereby enabling information mining and retrieval with Semantic Web technologies. This greatly reduces the administrative burden of handling multiple analysis tools and versions thereof and facilitates multi-level large scale comparative analysis. Availability and implementation: SAPP is written in JAVA and freely available at https://gitlab.com/sapp and runs on Unix-like operating systems. The documentation, examples and a tutorial are available at https://sapp.gitlab.io. Contact: [email protected] or [email protected].
Jasper J. Koehorst, Jesse C. J. van Dam, Edoardo Saccenti, Vítor A. P. Martins dos Santos, María Suárez-Diez, Peter J. Schaap
Bioinform.4
2018 SyNDI: synchronous network data integration framework
abstract
BACKGROUND: Systems biology takes a holistic approach by handling biomolecules and their interactions as big systems. Network based approach has emerged as a natural way to model these systems with the idea of representing biomolecules as nodes and their interactions as edges. Very often the input data come from various sorts of omics analyses. Those resulting networks sometimes describe a wide range of aspects, for example different experiment conditions, species, tissue types, stimulating factors, mutants, or simply distinct interaction features of the same network produced by different algorithms. For these scenarios, synchronous visualization of more than one distinct network is an excellent mean to explore all the relevant networks efficiently. In addition, complementary analysis methods are needed and they should work in a workflow manner in order to gain maximal biological insights. RESULTS: In order to address the aforementioned needs, we have developed a Synchronous Network Data Integration (SyNDI) framework. This framework contains SyncVis, a Cytoscape application for user-friendly synchronous and simultaneous visualization of multiple biological networks, and it is seamlessly integrated with other bioinformatics tools via the Galaxy platform. We demonstrated the functionality and usability of the framework with three biological examples - we analyzed the distinct connectivity of plasma metabolites in networks associated with high or low latent cardiovascular disease risk; deeper insights were obtained from a few similar inflammatory response pathways in Staphylococcus aureus infection common to human and mouse; and regulatory motifs which have not been reported associated with transcriptional adaptations of Mycobacterium tuberculosis were identified. CONCLUSIONS: Our SyNDI framework couples synchronous network visualization seamlessly with additional bioinformatics tools. The user can easily tailor the framework for his/her needs by adding new tools and datasets to the Galaxy platform.
Erno Lindfors, Jesse C. J. van Dam, Carolyn M. C. Lam, Niels A. Zondervan, Vítor A. P. Martins dos Santos, María Suárez-Diez
BMC Bioinform.5
2016 Efficient Reconstruction of Predictive Consensus Metabolic Network Models
abstract
Understanding cellular function requires accurate, comprehensive representations of metabolism. Genome-scale, constraint-based metabolic models (GSMs) provide such representations, but their usability is often hampered by inconsistencies at various levels, in particular for concurrent models. COMMGEN, our tool for COnsensus Metabolic Model GENeration, automatically identifies inconsistencies between concurrent models and semi-automatically resolves them, thereby contributing to consolidate knowledge of metabolic function. Tests of COMMGEN for four organisms showed that automatically generated consensus models were predictive and that they substantially increased coherence of knowledge representation. COMMGEN ought to be particularly useful for complex scenarios in which manual curation does not scale, such as for eukaryotic organisms, microbial communities, and host-pathogen interactions.
Ruben G. A. van Heck, Mathias Ganter, Vítor A. P. Martins dos Santos, Jörg Stelling
PLoS Comput. Biol.3
2011 Reconciliation of Genome-Scale Metabolic Reconstructions for Comparative Systems Analysis
abstract
In the past decade, over 50 genome-scale metabolic reconstructions have been built for a variety of single- and multi- cellular organisms. These reconstructions have enabled a host of computational methods to be leveraged for systems-analysis of metabolism, leading to greater understanding of observed phenotypes. These methods have been sparsely applied to comparisons between multiple organisms, however, due mainly to the existence of differences between reconstructions that are inherited from the respective reconstruction processes of the organisms to be compared. To circumvent this obstacle, we developed a novel process, termed metabolic network reconciliation, whereby non-biological differences are removed from genome-scale reconstructions while keeping the reconstructions as true as possible to the underlying biological data on which they are based. This process was applied to two organisms of great importance to disease and biotechnological applications, Pseudomonas aeruginosa and Pseudomonas putida, respectively. The result is a pair of revised genome-scale reconstructions for these organisms that can be analyzed at a systems level with confidence that differences are indicative of true biological differences (to the degree that is currently known), rather than artifacts of the reconstruction process. The reconstructions were re-validated with various experimental data after reconciliation. With the reconciled and validated reconstructions, we performed a genome-wide comparison of metabolic flexibility between P. aeruginosa and P. putida that generated significant new insight into the underlying biology of these important organisms. Through this work, we provide a novel methodology for reconciling models, present new genome-scale reconstructions of P. aeruginosa and P. putida that can be directly compared at a network level, and perform a network-wide comparison of the two species. These reconstructions provide fresh insights into the metabolic similarities and differences between these important Pseudomonads, and pave the way towards full comparative analysis of genome-scale metabolic reconstructions of multiple species.
Matthew A. Oberhardt, Jacek Puchalka, Vítor A. P. Martins dos Santos, Jason A. Papin
PLoS Comput. Biol.3
2008 Genome-Scale Reconstruction and Analysis of the Pseudomonas putida KT2440 Metabolic Network Facilitates Applications in Biotechnology
abstract
A cornerstone of biotechnology is the use of microorganisms for the efficient production of chemicals and the elimination of harmful waste. Pseudomonas putida is an archetype of such microbes due to its metabolic versatility, stress resistance, amenability to genetic modifications, and vast potential for environmental and industrial applications. To address both the elucidation of the metabolic wiring in P. putida and its uses in biocatalysis, in particular for the production of non-growth-related biochemicals, we developed and present here a genome-scale constraint-based model of the metabolism of P. putida KT2440. Network reconstruction and flux balance analysis (FBA) enabled definition of the structure of the metabolic network, identification of knowledge gaps, and pin-pointing of essential metabolic functions, facilitating thereby the refinement of gene annotations. FBA and flux variability analysis were used to analyze the properties, potential, and limits of the model. These analyses allowed identification, under various conditions, of key features of metabolism such as growth yield, resource distribution, network robustness, and gene essentiality. The model was validated with data from continuous cell cultures, high-throughput phenotyping data, (13)C-measurement of internal flux distributions, and specifically generated knock-out mutants. Auxotrophy was correctly predicted in 75% of the cases. These systematic analyses revealed that the metabolic network structure is the main factor determining the accuracy of predictions, whereas biomass composition has negligible influence. Finally, we drew on the model to devise metabolic engineering strategies to improve production of polyhydroxyalkanoates, a class of biotechnologically useful compounds whose synthesis is not coupled to cell survival. The solidly validated model yields valuable insights into genotype-phenotype relationships and provides a sound framework to explore this versatile bacterium and to capitalize on its vast biotechnological potential.
Jacek Puchalka, Matthew A. Oberhardt, Miguel Godinho, Agata Bielecka, Daniela Regenhardt, Kenneth N. Timmis, Jason A. Papin, Vítor A. P. Martins dos Santos
PLoS Comput. Biol.8