Vincent Mallet

dblp:241/9856 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Phenotypic Profile of Subclinical Alcohol Use Disorder and Liver Cancer in Type 2 Diabetes
Melissa Larbi, Edouard Audit, Joël Chavas, Simplice Donfack, Vincent Mallet
AIME (1)5
2025 AI-Driven Identification of Subclinical Alcohol Use Disorder in Type 2 Diabetes Patients
Melissa Larbi, Edouard Audit, Joël Chavas, Simplice Donfack, Vincent Mallet
AIME (2)5
2025 AtomSurf: Surface Representation for Learning on Protein Structures
abstract
While there has been significant progress in evaluating and comparing different representations for learning on protein data, the role of surface-based learning approaches remains not well-understood. In particular, there is a lack of direct and fair benchmark comparison between the best available surface-based learning methods against alternative representations such as graphs. Moreover, the few existing surface-based approaches either use surface information in isolation or, at best, perform global pooling between surface and graph-based architectures. In this work, we fill this gap by first adapting a state-of-the-art surface encoder for protein learning tasks. We then perform a direct and fair comparison of the resulting method against alternative approaches within the Atom3D benchmark, highlighting the limitations of pure surface-based learning. Finally, we propose an integrated approach, which allows learned feature sharing between graphs and surface representations on the level of nodes and vertices \textit{across all layers}. We demonstrate that the resulting architecture achieves state-of-the-art results on all tasks in the Atom3D benchmark, while adhering to the strict benchmark protocol, as well as more broadly on binding site identification and binding pocket classification. Furthermore, we use coarsened surfaces and optimize our approach for efficiency, making our tool competitive in training and inference time with existing techniques. Code can be found online: https://github.com/Vincentx15/atomsurf
Vincent Mallet, Yangyang Miao, Souhaib Attaiki, Bruno Correia, Maks Ovsjanikov
ICLR1
2025 Finding antibodies in cryo-EM maps with <tt>CrAI</tt>
abstract
MOTIVATION: Therapeutic antibodies have emerged as a prominent class of new drugs due to their high specificity and their ability to bind to several protein targets. Once an initial antibody has been identified, its design and characteristics are refined using structural information, when it is available. Cryo-EM is currently the most effective method to obtain 3D structures. It relies on well-established methods to process raw data into a 3D map, which may, however, be noisy and contain artifacts. To fully interpret these maps the number, position, and structure of antibodies and other proteins present must be determined. Unfortunately, existing automated methods addressing this step have limited accuracy, require additional inputs and high-resolution maps, and exhibit long running times. RESULTS: We propose the first fully automatic and efficient method dedicated to finding antibodies in cryo-EM maps: CrAI. This machine learning approach leverages the conserved structure of antibodies and a dedicated novel database that we built to solve this problem. Running a prediction takes only a few seconds, instead of hours, and requires nothing but the cryo-EM map, seamlessly integrating within automated analysis pipelines. Our method can find the location and pose of both Fabs and VHHs at resolutions up to 10 Å and is significantly more reliable than existing approaches. AVAILABILITY AND IMPLEMENTATION: We make our method available both in open source github.com/Sanofi-Public/crai and as a ChimeraX bundle (crai).
Vincent Mallet, Chiara Rapisarda, Hervé Minoux, Maks Ovsjanikov
Bioinform.1
2022 RNAglib: a python package for RNA 2.5 D graphs
abstract
SUMMARY: RNA 3D architectures are stabilized by sophisticated networks of (non-canonical) base pair interactions, which can be conveniently encoded as multi-relational graphs and efficiently exploited by graph theoretical approaches and recent progresses in machine learning techniques. RNAglib is a library that eases the use of this representation, by providing clean data, methods to load it in machine learning pipelines and graph-based deep learning models suited for this representation. RNAglib also offers other utilities to model RNA with 2.5 D graphs, such as drawing tools, comparison functions or baseline performances on RNA applications. AVAILABILITY AND IMPLEMENTATION: The method is distributed as a pip package, RNAglib. Data are available in a repository and can be accessed on rnaglib's web page. The source code, data and documentation are available at https://rnaglib.cs.mcgill.ca. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vincent Mallet, Carlos G. Oliver, Jonathan Broadbent, William L. Hamilton, Jérôme Waldispühl
Bioinform.1
2022 InDeep: 3D fully convolutional neural networks to assist in silico drug design on protein-protein interactions
abstract
MOTIVATION: Protein-protein interactions (PPIs) are key elements in numerous biological pathways and the subject of a growing number of drug discovery projects including against infectious diseases. Designing drugs on PPI targets remains a difficult task and requires extensive efforts to qualify a given interaction as an eligible target. To this end, besides the evident need to determine the role of PPIs in disease-associated pathways and their experimental characterization as therapeutics targets, prediction of their capacity to be bound by other protein partners or modulated by future drugs is of primary importance. RESULTS: We present InDeep, a tool for predicting functional binding sites within proteins that could either host protein epitopes or future drugs. Leveraging deep learning on a curated dataset of PPIs, this tool can proceed to enhanced functional binding site predictions either on experimental structures or along molecular dynamics trajectories. The benchmark of InDeep demonstrates that our tool outperforms state-of-the-art ligandable binding sites predictors when assessing PPI targets but also conventional targets. This offers new opportunities to assist drug design projects on PPIs by identifying pertinent binding pockets at or in the vicinity of PPI interfaces. AVAILABILITY AND IMPLEMENTATION: The tool is available on GitLab at https://gitlab.pasteur.fr/InDeep/InDeep. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vincent Mallet, Luis Checa Ruano, Alexandra Moine-Franel, Michael Nilges, Karen Druart, Guillaume Bouvier, Olivier Sperandio
Bioinform.1
2022 Vernal: a tool for mining fuzzy network motifs in RNA
abstract
MOTIVATION: RNA 3D motifs are recurrent substructures, modeled as networks of base pair interactions, which are crucial for understanding structure-function relationships. The task of automatically identifying such motifs is computationally hard, and remains a key challenge in the field of RNA structural biology and network analysis. State-of-the-art methods solve special cases of the motif problem by constraining the structural variability in occurrences of a motif, and narrowing the substructure search space. RESULTS: Here, we relax these constraints by posing the motif finding problem as a graph representation learning and clustering task. This framing takes advantage of the continuous nature of graph representations to model the flexibility and variability of RNA motifs in an efficient manner. We propose a set of node similarity functions, clustering methods and motif construction algorithms to recover flexible RNA motifs. Our tool, Vernal can be easily customized by users to desired levels of motif flexibility, abundance and size. We show that Vernal is able to retrieve and expand known classes of motifs, as well as to propose novel motifs. AVAILABILITY AND IMPLEMENTATION: The source code, data and a webserver are available at vernal.cs.mcgill.ca. We also provide a flexible interface and a user-friendly webserver to browse and download our results. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Carlos G. Oliver, Vincent Mallet, Pericles Philippopoulos, William L. Hamilton, Jérôme Waldispühl
Bioinform.2
2021 Reverse-Complement Equivariant Networks for DNA Sequences
abstract
As DNA sequencing technologies keep improving in scale and cost, there is a growing need to develop machine learning models to analyze DNA sequences, e.g., to decipher regulatory signals from DNA fragments bound by a particular protein of interest. As a double helix made of two complementary strands, a DNA fragment can be sequenced as two equivalent, so-called reverse complement (RC) sequences of nucleotides. To take into account this inherent symmetry of the data in machine learning models can facilitate learning. In this sense, several authors have recently proposed particular RC-equivariant convolutional neural networks (CNNs). However, it remains unknown whether other RC-equivariant architecture exist, which could potentially increase the set of basic models adapted to DNA sequences for practitioners. Here, we close this gap by characterizing the set of all linear RC-equivariant layers, and show in particular that new architectures exist beyond the ones already explored. We further discuss RC-equivariant pointwise nonlinearities adapted to different architectures, as well as RC-equivariant embeddings of $k$-mers as an alternative to one-hot encoding of nucleotides. We show experimentally that the new architectures can outperform existing ones.
Vincent Mallet, Jean-Philippe Vert
NeurIPS1
2021 quicksom: Self-Organizing Maps on GPUs for clustering of molecular dynamics trajectories
abstract
SUMMARY: We implemented the Self-Organizing Maps algorithm running efficiently on GPUs, and also provide several clustering methods of the resulting maps. We provide scripts and a use case to cluster macro-molecular conformations generated by molecular dynamics simulations. AVAILABILITY AND IMPLEMENTATION: The method is available on GitHub and distributed as a pip package.
Vincent Mallet, Michael Nilges, Guillaume Bouvier
Bioinform.1