Adrián Javaloy

dblp:259/2011 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 34% Generative modeling · 27% Learning paradigms · 19%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
causal inference
1.522025
DeCaFlow: A deconfounding causal generative model · NeurIPS 2025
Causal normalizing flows: from theory to practice · NeurIPS 2023
Machine learning › Learning paradigms
multi-task learning
1.122022
Mitigating Modality Collapse in Multimodal VAEs via Impartial Optimization · ICML 2022
RotoGrad: Gradient Homogenization in Multitask Learning · ICLR 2022
Machine learning › Generative modeling › generative model › probabilistic generative model
causal generative model
0.912025
DeCaFlow: A deconfounding causal generative model · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders
0.912025
DeCaFlow: A deconfounding causal generative model · NeurIPS 2025
Machine learning › Learning paradigms › multi-task learning
gradient conflict
0.612022
Mitigating Modality Collapse in Multimodal VAEs via Impartial Optimization · ICML 2022
Machine learning › Generative modeling › variational autoencoder
multimodal variational autoencoder
0.612022
Mitigating Modality Collapse in Multimodal VAEs via Impartial Optimization · ICML 2022
Machine learning › Generative modeling
variational autoencoder
0.612022
Mitigating Modality Collapse in Multimodal VAEs via Impartial Optimization · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.412020
Relative gradient optimization of the Jacobian term in unsupervised deep learning · NeurIPS 2020
Machine learning › Representation and self-supervised learning › blind source separation
independent component analysis
0.412020
Relative gradient optimization of the Jacobian term in unsupervised deep learning · NeurIPS 2020
Machine learning › Representation and self-supervised learning › blind source separation › independent component analysis
nonlinear ICA
0.412020
Relative gradient optimization of the Jacobian term in unsupervised deep learning · NeurIPS 2020
Machine learning › Generative modeling
normalizing flow
0.412020
Relative gradient optimization of the Jacobian term in unsupervised deep learning · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › causal inference
counterfactual prediction
0.312025
DeCaFlow: A deconfounding causal generative model · NeurIPS 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.212023
Learnable Graph Convolutional Attention Networks · ICLR 2023

Methods — techniques the papers use, named apart from their topics

proxy variables · 0.9do-calculus · 0.9normalizing flow · 0.7nonlinear ICA · 0.7graph convolution · 0.7autoregressive model · 0.7attention · 0.7impartial optimization · 0.6gradient-conflict solutions · 0.6gradient homogenization · 0.6
YearPublicationVenuePosition
2025 DeCaFlow: A deconfounding causal generative model
abstract
We introduce DeCaFlow, a deconfounding causal generative model. Training once per dataset using just observational data and the underlying causal graph, DeCaFlow enables accurate causal inference on continuous variables under the presence of hidden confounders. Specifically, we extend previous results on causal estimation under hidden confounding to show that a single instance of DeCaFlow provides correct estimates for all causal queries identifiable with do-calculus, leveraging proxy variables to adjust for the causal effects when do-calculus alone is insufficient. Moreover, we show that counterfactual queries are identifiable as long as their interventional counterparts are identifiable, and thus are also correctly estimated by DeCaFlow. Our empirical results on diverse settings—including the Ecoli70 dataset, with 3 independent hidden confounders, tens of observed variables and hundreds of causal queries—show that DeCaFlow outperforms existing approaches, while demonstrating its out-of-the-box applicability to any given causal graph.
Alejandro Almodóvar, Adrián Javaloy, Juan Parras, Santiago Zazo, Isabel Valera
NeurIPS2
2023 Learnable Graph Convolutional Attention Networks
Adrián Javaloy, Pablo Sánchez-Martín, Amit Levi 0001, Isabel Valera
ICLR1
2023 Causal normalizing flows: from theory to practice
abstract
In this work, we deepen on the use of normalizing flows for causal reasoning. Specifically, we first leverage recent results on non-linear ICA to show that causal models are identifiable from observational data given a causal ordering, and thus can be recovered using autoregressive normalizing flows (NFs). Second, we analyze different design and learning choices for *causal normalizing flows* to capture the underlying causal data-generating process. Third, we describe how to implement the *do-operator* in causal NFs, and thus, how to answer interventional and counterfactual questions. Finally, in our experiments, we validate our design and training choices through a comprehensive ablation study; compare causal NFs to other approaches for approximating causal models; and empirically demonstrate that causal NFs can be used to address real-world problems—where the presence of mixed discrete-continuous data and partial knowledge on the causal graph is the norm. The code for this work can be found at https://github.com/psanch21/causal-flows.
Adrián Javaloy, Pablo Sánchez-Martín, Isabel Valera
NeurIPS1
2022 RotoGrad: Gradient Homogenization in Multitask Learning
Adrián Javaloy, Isabel Valera
ICLR1
2022 Mitigating Modality Collapse in Multimodal VAEs via Impartial Optimization
abstract
A number of variational autoencoders (VAEs) have recently emerged with the aim of modeling multimodal data, e.g., to jointly model images and their corresponding captions. Still, multimodal VAEs tend to focus solely on a subset of the modalities, e.g., by fitting the image while neglecting the caption. We refer to this limitation as modality collapse. In this work, we argue that this effect is a consequence of conflicting gradients during multimodal VAE training. We show how to detect the sub-graphs in the computational graphs where gradients conflict (impartiality blocks), as well as how to leverage existing gradient-conflict solutions from multitask learning to mitigate modality collapse. That is, to ensure impartial optimization across modalities. We apply our training framework to several multimodal VAE models, losses and datasets from the literature, and empirically show that our framework significantly improves the reconstruction performance, conditional generation, and coherence of the latent space across modalities.
Adrián Javaloy, Maryam Meghdadi, Isabel Valera
ICML1
2020 Relative gradient optimization of the Jacobian term in unsupervised deep learning
abstract
Learning expressive probabilistic models correctly describing the data is a ubiquitous problem in machine learning. A popular approach for solving it is mapping the observations into a representation space with a simple joint distribution, which can typically be written as a product of its marginals — thus drawing a connection with the field of nonlinear independent component analysis. Deep density models have been widely used for this task, but their maximum likelihood based training requires estimating the log-determinant of the Jacobian and is computationally expensive, thus imposing a trade-off between computation and expressive power. In this work, we propose a new approach for exact training of such neural networks. Based on relative gradients, we exploit the matrix structure of neural network parameters to compute updates efficiently even in high-dimensional spaces; the computational cost of the training is quadratic in the input size, in contrast with the cubic scaling of naive approaches. This allows fast training with objective functions involving the log-determinant of the Jacobian, without imposing constraints on its structure, in stark contrast to autoregressive normalizing flows.
Luigi Gresele, Giancarlo Fissore, Adrián Javaloy, Bernhard Schölkopf, Aapo Hyvärinen
NeurIPS3