VLDB 2026 Research / reviewers in the wild / expert
Francisco Vargas 0001
dblp:79/7431-1
· DBLP profile ↗
13ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-2714-3357ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NMA-tune: Generating Highly Designable and Dynamics Aware Protein BackbonesabstractProtein’s backbone flexibility is a crucial property that heavily influences its functionality. Recent work in the field of protein diffusion probabilistic modelling has leveraged Normal Mode Analysis (NMA) and, for the first time, introduced information about large scale protein motion into the generative process. However, obtaining molecules with both the desired dynamics and designable quality has proven challenging. In this work, we present NMA-tune, a new method that introduces the dynamics information to the protein design stage. NMA-tune uses a trainable component to condition the backbone generation on the lowest normal mode of oscillation. We implement NMA-tune as a plug-and-play extension to RFdiffusion, show that the proportion of samples with high quality structure and the desired dynamics is improved as compared to other methods without the trainable component, and we show the presence of the targeted modes in the Molecular Dynamics simulations. Urszula Julia Komorowska, Francisco Vargas 0001, Alessandro Rondina, Pietro Liò, Mateja Jamnik |
ICML | 2 |
| 2025 | FEAT: Free energy Estimators with Adaptive TransportabstractWe present Free energy Estimators with Adaptive Transport (FEAT), a novel framework for free energy estimation---a critical challenge across scientific domains.
FEAT leverages learned transports implemented via stochastic interpolants and provides consistent, minimum-variance estimators based on escorted Jarzynski equality and controlled Crooks theorem, alongside variational upper and lower bounds on free energy differences.
Unifying equilibrium and non-equilibrium methods under a single theoretical framework, FEAT establishes a principled foundation for neural free energy calculations.
Experimental validation on toy examples, molecular simulations, and quantum field theory demonstrates promising improvements over existing learning-based methods.
Our PyTorch implementation is available at https://github.com/jiajunhe98/FEAT. Yuanqi Du, Jiajun He 0003, Francisco Vargas 0001, Carla P. Gomes, José Miguel Hernández-Lobato, Eric Vanden-Eijnden |
NeurIPS | 3 |
| 2025 | Gradient Variance Reveals Failure Modes in Flow-Based Generative ModelsabstractRectified Flows learn ODE vector fields whose trajectories are straight between source and target distributions, enabling near one-step inference. We show that this straight-path objective reveals fundamental failure modes: under deterministic training, low gradient variance drives memorization of arbitrary training pairings, even when interpolant lines between training pairs intersect. To analyze this mechanism, we study Gaussian-to-Gaussian transport and use the loss gradient variance across stochastic and deterministic regimes to characterize which vector fields optimization favors in each setting. We then show that, in a setting where all interpolating lines intersect, applying Rectified Flow yields the same specific pairings at inference as during training. More generally, we prove that a memorizing vector field exists even when training interpolants intersect, and that optimizing the straight-path objective converges to this ill-defined field. At inference, deterministic integration reproduces the exact training pairings. We validate our findings empirically on the CelebA dataset, confirming that deterministic interpolants induce memorization, while the injection of small noise restores generalization. Teodora Reu, Sixtine Dromigny, Michael M. Bronstein, Francisco Vargas 0001 |
NeurIPS | 4 |
| 2024 | Transport meets Variational Inference: Controlled Monte Carlo DiffusionsabstractConnecting optimal transport and variational inference, we present a principled and systematic framework for sampling and generative modelling centred around divergences on path space. Our work culminates in the development of the Controlled Monte Carlo Diffusion sampler (CMCD) for Bayesian computation, a score-based annealing technique that crucially adapts both forward and backward dynamics in a diffusion model. On the way, we clarify the relationship between the EM-algorithm and iterative proportional fitting (IPF) for Schroedinger bridges, deriving as well a regularised objective that bypasses the iterative bottleneck of standard IPF-updates. Finally, we show that CMCD has a strong foundation in the Jarzinsky and Crooks identities from statistical physics, and that it convincingly outperforms competing approaches across a wide array of experiments. Francisco Vargas 0001, Shreyas Padhy, Denis Blessing, Nikolas Nüsken |
ICLR | 1 |
| 2024 | Dynamics-Informed Protein Design with Structure ConditioningabstractCurrent protein generative models are able to design novel backbones with desired shapes or functional motifs. However, despite the importance of a protein’s dynamical properties for its function, conditioning on dynamical properties remains elusive. We present a new approach to protein generative modeling by leveraging Normal Mode Analysis that enables us to capture dynamical properties too. We introduce a method for conditioning the diffusion probabilistic models on protein dynamics, specifically on the lowest non-trivial normal mode of oscillation. Our method, similar to the classifier guidance conditioning, formulates the sampling process as being driven by conditional and unconditional terms. However, unlike previous works, we approximate the conditional term with a simple analytical function rather than an external neural network, thus making the eigenvector calculations approachable. We present the corresponding SDE theory as a formal justification of our approach. We extend our framework to conditioning on structure and dynamics at the same time, enabling scaffolding of the dynamical motifs. We demonstrate the empirical effectiveness of our method by turning the open-source unconditional protein diffusion model Genie into the conditional model with no retraining. Generated proteins exhibit the desired dynamical and structural properties while still being biologically plausible. Our work represents a first step towards incorporating dynamical behaviour in protein design and may open the door to designing more flexible and functional proteins in the future. Urszula Julia Komorowska, Simon V. Mathis, Kieran Didi, Francisco Vargas 0001, Pietro Liò, Mateja Jamnik |
ICLR | 4 |
| 2024 | Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for SamplingabstractMonte Carlo methods, Variational Inference, and their combinations play a pivotal role in sampling from intractable probability distributions. However, current studies lack a unified evaluation framework, relying on disparate performance measures and limited method comparisons across diverse tasks, complicating the assessment of progress and hindering the decision-making of practitioners. In response to these challenges, our work introduces a benchmark that evaluates sampling methods using a standardized task suite and a broad range of performance criteria. Moreover, we study existing metrics for quantifying mode collapse and introduce novel metrics for this purpose. Our findings provide insights into strengths and weaknesses of existing sampling methods, serving as a valuable reference for future developments. Denis Blessing, Xiaogang Jia, Johannes Esslinger, Francisco Vargas 0001, Gerhard Neumann |
ICML | 4 |
| 2024 | DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transformabstractGenerative modelling paradigms based on denoising diffusion processes have emerged as a leading candidate for conditional sampling in inverse problems.
In many real-world applications, we often have access to large, expensively trained unconditional diffusion models, which we aim to exploit for improving conditional sampling.
Most recent approaches are motivated heuristically and lack a unifying framework, obscuring connections between them. Further, they often suffer from issues such as being very sensitive to hyperparameters, being expensive to train or needing access to weights hidden behind a closed API. In this work, we unify conditional training and sampling using the mathematically well-understood Doob's h-transform. This new perspective allows us to unify many existing methods under a common umbrella. Under this framework, we propose DEFT (Doob's h-transform Efficient FineTuning), a new approach for conditional generation that simply fine-tunes a very small network to quickly learn the conditional $h$-transform, while keeping the larger unconditional network unchanged. DEFT is much faster than existing baselines while achieving state-of-the-art performance across a variety of linear and non-linear benchmarks. On image reconstruction tasks, we achieve speedups of up to 1.6$\times$, while having the best perceptual quality on natural images and reconstruction performance on medical images. Further, we also provide initial experiments on protein motif scaffolding and outperform reconstruction guidance methods. Alexander Denker, Francisco Vargas 0001, Shreyas Padhy, Kieran Didi, Simon V. Mathis, Riccardo Barbano, Vincent Dutordoir, Emile Mathieu, Urszula Julia Komorowska, Pietro Liò |
NeurIPS | 2 |
| 2024 | To smooth a cloud or to pin it down: Expressiveness guarantees and insights on score matching in denoising diffusion modelsabstractDenoising diffusion models are a class of generative models that have recently achieved state-of-the-art results across many domains. Gradual noise is added to the data using a diffusion process, which transforms the data distribution into a Gaussian. Samples from the generative model are then obtained by simulating an approximation of the time reversal of this diffusion initialized by Gaussian samples. Recent research has explored the sampling error achieved by diffusion models under the assumption of an absolute error $\epsilon$ achieved via a neural approximation of the score. To the best of our knowledge, no work formally quantifies the error of such neural approximation to the score. In this paper, we close the gap and present quantitative error bounds for approximating the score of denoising diffusion models using neural networks leveraging ideas from stochastic control. Finally, through simulation, we explore some of the insights that arise from our results confirming that diffusion models based on the Ornstein-Uhlenbeck (OU) process require fewer parameters to better approximate the score than those based on the Fölmer drift / Pinned Brownian Motion. Teodora Reu, Francisco Vargas 0001, Anna Böhm, Michael M. Bronstein |
UAI | 2 |
| 2023 | Denoising Diffusion Samplers
Francisco Vargas 0001, Will Grathwohl, Arnaud Doucet |
ICLR | 1 |
| 2022 | Adversarial Concept Erasure in Kernel SpaceabstractThe representation space of neural models for textual data emerges in an unsupervised manner during training.Understanding how those representations encode human-interpretable concepts is a fundamental problem.One prominent approach for the identification of concepts in neural representations is searching for a linear subspace whose erasure prevents the prediction of the concept from the representations.However, while many linear erasure algorithms are tractable and interpretable, neural networks do not necessarily represent concepts in a linear manner.To identify non-linearly encoded concepts, we propose a kernelization of a linear minimax game for concept erasure.We demonstrate that it is possible to prevent specific nonlinear adversaries from predicting the concept.However, the protection does not transfer to different nonlinear adversaries.Therefore, exhaustively erasing a non-linearly encoded concept remains an open problem. Shauli Ravfogel, Francisco Vargas 0001, Yoav Goldberg, Ryan Cotterell |
EMNLP | 2 |
| 2020 | Exploring the Linear Subspace Hypothesis in Gender Bias MitigationabstractBolukbasi et al. (2016) presents one of the first gender bias mitigation techniques for word embeddings.Their method takes pre-trained word embeddings as input and attempts to isolate a linear subspace that captures most of the gender bias in the embeddings.As judged by an analogical evaluation task, their method virtually eliminates gender bias in the embeddings.However, an implicit and untested assumption of their method is that the bias subspace is actually linear.In this work, we generalize their method to a kernelized, non-linear version.We take inspiration from kernel principal component analysis and derive a nonlinear bias isolation technique.We discuss and overcome some of the practical drawbacks of our method for non-linear gender bias mitigation in word embeddings and analyze empirically whether the bias subspace is actually linear.Our analysis shows that gender bias is in fact well captured by a linear subspace, justifying the assumption of Bolukbasi et al. (2016). Francisco Vargas 0001, Ryan Cotterell |
EMNLP (1) | 1 |
| 2019 | Multilingual Factor AnalysisabstractIn this work we approach the task of learning multilingual word representations in an offline manner by fitting a generative latent variable model to a multilingual dictionary.We model equivalent words in different languages as different views of the same word generated by a common latent variable representing their latent lexical meaning.We explore the task of alignment by querying the fitted model for multilingual embeddings achieving competitive results across a variety of tasks.The proposed model is robust to noise in the embedding space making it a suitable method for distributed representations learned from noisy corpora. Francisco Vargas 0001, Kamen Brestnichki, Alex Papadopoulos-Korfiatis, Nils Y. Hammerla |
ACL (1) | 1 |
| 2019 | Model Comparison for Semantic GroupingabstractWe introduce a probabilistic framework for quantifying the semantic similarity between two groups of embeddings. We formulate the task of semantic similarity as a model comparison task in which we contrast a generative model which jointly models two sentences versus one that does not. We illustrate how this framework can be used for the Semantic Textual Similarity tasks using clear assumptions about how the embeddings of words are generated. We apply model comparison that utilises information criteria to address some of the shortcomings of Bayesian model comparison, whilst still penalising model complexity. We achieve competitive results by applying the proposed framework with an appropriate choice of likelihood on the STS datasets. Francisco Vargas 0001, Kamen Brestnichki, Nils Y. Hammerla |
ICML | 1 |