VLDB 2026 Research / reviewers in the wild / expert
Carlos G. Oliver
dblp:205/3100
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-8742-8795ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Endowing protein language models with structural knowledgeabstractMOTIVATION: Protein language models (PLMs) have transformed protein research by learning rich representations from sequence data alone, yet they largely ignore the wealth of structural information now available through advances in structure prediction. Current methods that incorporate structural data often require substantial computational resources and complex architectures, limiting their practical adoption. We present a novel joint sequence and structure embedding method that achieves computational and parameter efficiency while maintaining high performance. Our approach introduces a lightweight integration framework that combines pretrained sequence transformers' self-attention with specialized structural adapters, enabling seamless incorporation of structural knowledge into existing PLMs through these enhanced self-attention mechanisms. RESULTS: The method demonstrates remarkable efficiency, requiring only modest pretraining on 542K protein structures, three orders of magnitude less than the data used to train PLMs, using standard masked language modeling objectives. Despite this lightweight approach, our joint embeddings consistently outperform sequence-only models like ESM-2 while achieving comparable results to more complex structure-based methods that use significantly more parameters and computational resources. This work establishes a new paradigm for protein representation learning that balances performance with practical constraints. By providing computationally efficient joint sequence-structure embeddings, we offer the scientific community an accessible tool that captures both sequential and structural protein information without the computational overhead typically associated with structure-aware models. AVAILABILITY AND IMPLEMENTATION: code and links to checkpoints are available at https://github.com/BorgwardtLab/PST. Philip Hartout, Dexiong Chen, Paolo Pellizzoni, Carlos G. Oliver, Karsten M. Borgwardt |
Bioinform. | 4 |
| 2024 | Structure- and Function-Aware Substitution Matrices via Learnable Graph Matching
Paolo Pellizzoni, Carlos G. Oliver, Karsten M. Borgwardt |
RECOMB | 2 |
| 2023 | Unsupervised Manifold Alignment with Joint Multidimensional Scaling
Dexiong Chen, Bowen Fan, Carlos G. Oliver, Karsten M. Borgwardt |
ICLR | 3 |
| 2023 | ProteinShake: Building datasets and benchmarks for deep learning on protein structuresabstractWe present ProteinShake, a Python software package that simplifies datasetcreation and model evaluation for deep learning on protein structures. Users cancreate custom datasets or load an extensive set of pre-processed datasets fromthe Protein Data Bank (PDB) and AlphaFoldDB. Each dataset is associated withprediction tasks and evaluation functions covering a broad array of biologicalchallenges. A benchmark on these tasks shows that pre-training almost alwaysimproves performance, the optimal data modality (graphs, voxel grids, or pointclouds) is task-dependent, and models struggle to generalize to new structures.ProteinShake makes protein structure data easily accessible and comparisonamong models straightforward, providing challenging benchmark settings withreal-world implications.ProteinShake is available at: https://proteinshake.ai Tim Kucera, Carlos G. Oliver, Dexiong Chen, Karsten M. Borgwardt |
NeurIPS | 2 |
| 2023 | Multimodal learning in clinical proteomics: enhancing antimicrobial resistance prediction models with chemical informationabstractMOTIVATION: Large-scale clinical proteomics datasets of infectious pathogens, combined with antimicrobial resistance outcomes, have recently opened the door for machine learning models which aim to improve clinical treatment by predicting resistance early. However, existing prediction frameworks typically train a separate model for each antimicrobial and species in order to predict a pathogen's resistance outcome, resulting in missed opportunities for chemical knowledge transfer and generalizability. RESULTS: We demonstrate the effectiveness of multimodal learning over proteomic and chemical features by exploring two clinically relevant tasks for our proposed deep learning models: drug recommendation and generalized resistance prediction. By adopting this multi-view representation of the pathogenic samples and leveraging the scale of the available datasets, our models outperformed the previous single-drug and single-species predictive models by statistically significant margins. We extensively validated the multi-drug setting, highlighting the challenges in generalizing beyond the training data distribution, and quantitatively demonstrate how suitable representations of antimicrobial drugs constitute a crucial tool in the development of clinically relevant predictive models. AVAILABILITY AND IMPLEMENTATION: The code used to produce the results presented in this article is available at https://github.com/BorgwardtLab/MultimodalAMR. Giovanni Visonà, Diane Duroux, Lucas Miranda 0002, Emese Sükei, Karsten M. Borgwardt, Carlos G. Oliver |
Bioinform. | 7 |
| 2022 | RNAglib: a python package for RNA 2.5 D graphsabstractSUMMARY: RNA 3D architectures are stabilized by sophisticated networks of (non-canonical) base pair interactions, which can be conveniently encoded as multi-relational graphs and efficiently exploited by graph theoretical approaches and recent progresses in machine learning techniques. RNAglib is a library that eases the use of this representation, by providing clean data, methods to load it in machine learning pipelines and graph-based deep learning models suited for this representation. RNAglib also offers other utilities to model RNA with 2.5 D graphs, such as drawing tools, comparison functions or baseline performances on RNA applications. AVAILABILITY AND IMPLEMENTATION: The method is distributed as a pip package, RNAglib. Data are available in a repository and can be accessed on rnaglib's web page. The source code, data and documentation are available at https://rnaglib.cs.mcgill.ca. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vincent Mallet, Carlos G. Oliver, Jonathan Broadbent, William L. Hamilton, Jérôme Waldispühl |
Bioinform. | 2 |
| 2022 | Vernal: a tool for mining fuzzy network motifs in RNAabstractMOTIVATION: RNA 3D motifs are recurrent substructures, modeled as networks of base pair interactions, which are crucial for understanding structure-function relationships. The task of automatically identifying such motifs is computationally hard, and remains a key challenge in the field of RNA structural biology and network analysis. State-of-the-art methods solve special cases of the motif problem by constraining the structural variability in occurrences of a motif, and narrowing the substructure search space. RESULTS: Here, we relax these constraints by posing the motif finding problem as a graph representation learning and clustering task. This framing takes advantage of the continuous nature of graph representations to model the flexibility and variability of RNA motifs in an efficient manner. We propose a set of node similarity functions, clustering methods and motif construction algorithms to recover flexible RNA motifs. Our tool, Vernal can be easily customized by users to desired levels of motif flexibility, abundance and size. We show that Vernal is able to retrieve and expand known classes of motifs, as well as to propose novel motifs. AVAILABILITY AND IMPLEMENTATION: The source code, data and a webserver are available at vernal.cs.mcgill.ca. We also provide a flexible interface and a user-friendly webserver to browse and download our results. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Carlos G. Oliver, Vincent Mallet, Pericles Philippopoulos, William L. Hamilton, Jérôme Waldispühl |
Bioinform. | 1 |
| 2020 | Stochastic Sampling of Structural Contexts Improves the Scalability and Accuracy of RNA 3D Module Identification
Roman Sarrazin-Gendron, Hua-Ting Yao, Vladimir Reinharz, Carlos G. Oliver, Yann Ponty, Jérôme Waldispühl |
RECOMB | 4 |