EDBT 2026 Demo / reviewers in the wild / expert
Gaël Varoquaux
dblp:36/7585
· DBLP profile ↗
61ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0003-1076-5122ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Survival Models: Proper Scoring Rule and Stochastic Optimization with Competing RisksabstractWhen dealing with right-censored data, where some outcomes are missing due to a limited observation period, survival analysis —known as \emph{time-to-event analysis}— focuses on predicting the time until an event of interest occurs. Multiple classes of outcomes lead to a classification variant: predicting the most likely event, a less explored area known as \emph{competing risks}. Classic competing risks models couple architecture and loss, limiting scalability. To address these issues, we design a strictly proper censoring-adjusted separable scoring rule, allowing optimization on a subset of the data as each observation is evaluated independently. The loss estimates outcome probabilities and enables stochastic optimization for competing risks, which we use for efficient gradient boosting trees. \textbf{SurvivalBoost} not only outperforms 12 state-of-the-art models across several metrics on 4 real-life datasets, both in competing risks and survival settings, but also provides great calibration, the ability to predict across any time horizon, and computation times faster than existing methods. Julie Alberge, Vincent Maladière, Olivier Grisel, Judith Abécassis, Gaël Varoquaux |
AISTATS | 5 |
| 2025 | Decision from Suboptimal Classifiers: Excess Risk Pre- and Post-CalibrationabstractProbabilistic classifiers are central for making informed decisions under uncertainty. Based on the maximum expected utility principle, optimal decision rules can be derived using the posterior class probabilities and misclassification costs. Yet, in practice only learned approximations of the oracle posterior probabilities are available. In this work, we quantify the excess risk (a.k.a. regret) incurred using approximate posterior probabilities in batch binary decision-making. We provide analytical expressions for miscalibration-induced regret ($R^{CL}$), as well as tight and informative upper and lower bounds on the regret of calibrated classifiers ($R^{GL}$). These expressions allow us to identify regimes where recalibration alone addresses most of the regret, and regimes where the regret is dominated by the grouping loss, which calls for post-training beyond recalibration. Crucially, both $R^{CL}$ and $R^{Gl}$ can be estimated in practice using a calibration curve and a recent grouping loss estimator. On NLP experiments, we show that these quantities identify when the expected gain of more advanced post-training is worth the operational cost. Finally, we highlight the potential of multicalibration approaches as efficient alternatives to costlier fine-tuning approaches. Alexandre Perez-Lebel, Gaël Varoquaux, Oluwasanmi Koyejo, Matthieu Doutreligne, Marine Le Morvan |
AISTATS | 2 |
| 2025 | Imputation for prediction: beware of diminishing returnsabstractMissing values are prevalent across various fields, posing challenges for training and deploying predictive models. In this context, imputation is a common practice, driven by the hope that accurate imputations will enhance predictions. However, recent theoretical and empirical studies indicate that simple constant imputation can be consistent and competitive. This empirical study aims at clarifying
*if* and *when* investing in advanced imputation methods yields significantly better predictions. Relating imputation and predictive accuracies across combinations of imputation and predictive models on 19 datasets, we show that imputation accuracy matters less i) when using expressive models, ii) when incorporating missingness indicators as complementary inputs, iii) matters much more for generated linear outcomes than for real-data outcomes. Interestingly, we also show that the use of the missingness indicator is beneficial to the prediction performance, even in MCAR scenarios. Overall, on real-data with powerful models, imputation quality has only a minor effect on prediction performance. Thus, investing in better imputations for improved predictions often offers limited benefits. Marine Le Morvan, Gaël Varoquaux |
ICLR | 2 |
| 2025 | TabICL: A Tabular Foundation Model for In-Context Learning on Large DataabstractThe long-standing dominance of gradient-boosted decision trees on tabular data is currently challenged by tabular foundation models using In-Context Learning (ICL): setting the training data as context for the test data and predicting in a single forward pass without parameter updates. While TabPFNv2 foundation model excels on tables with up to 10K samples, its alternating column- and row-wise attentions make handling large training sets computationally prohibitive. So, can ICL be effectively scaled and deliver a benefit for larger tables? We introduce TabICL, a tabular foundation model for classification, pretrained on synthetic datasets with up to 60K samples and capable of handling 500K samples on affordable resources. This is enabled by a novel two-stage architecture: a column-then-row attention mechanism to build fixed-dimensional embeddings of rows, followed by a transformer for efficient ICL. Across 200 classification datasets from the TALENT benchmark, TabICL is on par with TabPFNv2 while being systematically faster (up to 10 times), and significantly outperforms all other approaches. On 53 datasets with over 10K samples, TabICL surpasses both TabPFNv2 and CatBoost, demonstrating the potential of ICL for large data. Pretraining code, inference code, and pre-trained models are available at https://github.com/soda-inria/tabicl. Jingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le Morvan |
ICML | 3 |
| 2025 | Scalable Feature Learning on Huge Knowledge Graphs for Downstream Machine LearningabstractMany machine learning tasks can benefit from external knowledge. Large knowledge graphs store such knowledge, and embedding methods can be used to distill it into ready-to-use vector representations for downstream applications. For this purpose, current models have however two limitations: they are primarily optimized for link prediction, via local contrastive learning, and their application to the largest graphs requires significant engineering effort due to GPU memory limits. To address these, we introduce SEPAL: a Scalable Embedding Propagation ALgorithm for large knowledge graphs designed to produce high-quality embeddings for downstream tasks at scale. The key idea of SEPAL is to ensure global embedding consistency by optimizing embeddings only on a small core of entities, and then propagating them to the rest of the graph with message passing. We evaluate SEPAL on 7 large-scale knowledge graphs and 46 downstream machine learning tasks. Our results show that SEPAL significantly outperforms previous methods on downstream tasks. In addition, SEPAL scales up its base embedding model, enabling fitting huge knowledge graphs on commodity hardware. Our code is available at: <https://github.com/soda-inria/sepal>. Félix Lefebvre, Gaël Varoquaux |
NeurIPS | 2 |
| 2025 | Confidence intervals for performance estimates in brain MRI segmentation
Rosana El Jurdi, Gaël Varoquaux, Olivier Colliot |
Medical Image Anal. | 2 |
| 2024 | CARTE: Pretraining and Transfer for Tabular LearningabstractPretrained deep-learning models are the go-to solution for images or text. However, for tabular data the standard is still to train tree-based models. Indeed, transfer learning on tables hits the challenge of data integration: finding correspondences, correspondences in the entries (entity matching) where different words may denote the same entity, correspondences across columns (schema matching), which may come in different orders, names... We propose a neural architecture that does not need such correspondences. As a result, we can pretrain it on background data that has not been matched. The architecture –CARTE for Context Aware Representation of Table Entries– uses a graph representation of tabular (or relational) data to process tables with different columns, string embedding of entries and columns names to model an open vocabulary, and a graph-attentional network to contextualize entries with column names and neighboring entries. An extensive benchmark shows that CARTE facilitates learning, outperforming a solid set of baselines including the best tree-based models. CARTE also enables joint learning across tables with unmatched columns, enhancing a small table with bigger ones. CARTE opens the door to large pretrained models for tabular data. Myung Jun Kim, Léo Grinsztajn, Gaël Varoquaux |
ICML | 3 |
| 2024 | Confidence Intervals Uncovered: Are We Ready for Real-World Medical Imaging AI?
Evangelia Christodoulou, Annika Reinke, Rola Houhou, Piotr Kalinowski, Selen Erkan, Carole H. Sudre, Ninon Burgos, Sofiène Boutaj, Sophie Loizillon, Maëlys Solal, Nicola Rieke, Veronika Cheplygina, Michela Antonelli, Leon D. Mayer, Minu Tizabi, Manuel Jorge Cardoso, Amber L. Simpson, Paul F. Jaeger, Annette Kopp-Schneider, Gaël Varoquaux, Olivier Colliot, Lena Maier-Hein |
MICCAI (10) | 20 |
| 2023 | GLADIS: A General and Large Acronym Disambiguation BenchmarkabstractAcronym Disambiguation (AD) is crucial for natural language understanding on various sources, including biomedical reports, scientific papers, and search engine queries.However, existing acronym disambiguation benchmarks and tools are limited to specific domains, and the size of prior benchmarks is rather small.To accelerate the research on acronym disambiguation, we construct a new benchmark named GLADIS with three components: (1) a much larger acronym dictionary with 1.5M acronyms and 6.4M long forms;(2) a pre-training corpus with 160 million sentences; (3) three datasets that cover the general, scientific, and biomedical domains.We then pre-train a language model, AcroBERT, on our constructed corpus for general acronym disambiguation, and show the challenges and values of our new benchmark. Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek |
EACL | 2 |
| 2023 | Beyond calibration: estimating the grouping loss of modern neural networks
Alexandre Perez-Lebel, Marine Le Morvan, Gaël Varoquaux |
ICLR | 3 |
| 2023 | Relational data embeddings for feature enrichment with background information
Alexis Cvetkov-Iliev, Alexandre Allauzen, Gaël Varoquaux |
Mach. Learn. | 3 |
| 2022 | Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little CostabstractState-of-the-art NLP systems represent inputs with word embeddings, but these are brittle when faced with Out-of-Vocabulary (OOV) words.To address this issue, we follow the principle of mimick-like models to generate vectors for unseen words, by learning the behavior of pre-trained embeddings using only the surface form of words.We present a simple contrastive learning framework, LOVE, which extends the word representation of an existing pre-trained language model (such as BERT), and makes it robust to OOV with few additional parameters.Extensive evaluations demonstrate that our lightweight model achieves similar or even better performances than prior competitors, both on original datasets and on corrupted variants.Moreover, it can be used in a plug-and-play fashion with FastText and BERT, where it significantly improves their robustness. Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek |
ACL (1) | 2 |
| 2022 | Why do tree-based models still outperform deep learning on typical tabular data?abstractWhile deep learning has enabled tremendous progress on text and image datasets, its superiority on tabular data is not clear. We contribute extensive benchmarks of standard and novel deep learning methods as well as tree-based models such as XGBoost and Random Forests, across a large number of datasets and hyperparameter combinations. We define a standard set of 45 datasets from varied domains with clear characteristics of tabular data and a benchmarking methodology accounting for both fitting models and finding good hyperparameters. Results show that tree-based models remain state-of-the-art on medium-sized data ($\sim$10K samples) even without accounting for their superior speed. To understand this gap, we conduct an empirical investigation into the differing inductive biases of tree-based models and neural networks. This leads to a series of challenges which should guide researchers aiming to build tabular-specific neural network: 1) be robust to uninformative features, 2) preserve the orientation of the data, and 3) be able to easily learn irregular functions. To stimulate research on tabular architectures, we contribute a standard benchmark and raw data for baselines: every point of a 20\,000 compute hours hyperparameter search for each learner. Léo Grinsztajn, Edouard Oyallon, Gaël Varoquaux |
NeurIPS | 3 |
| 2022 | Encoding High-Cardinality String Categorical VariablesabstractStatistical models usually require vector representations of categorical variables, using for instanceone-hot encoding. This strategy breaks down when the number of categories grows, as it creates high-dimensional feature vectors. Additionally, for string entries, one-hot encoding does not capture morphological information in their representation. Here, we seek low-dimensional encoding of high-cardinality string categorical variables. Ideally, these should be: scalable to many categories; interpretable to end users; and facilitate statistical analysis. We introduce two encoding approaches for string categories: aGamma-Poisson matrix factorizationon substring counts, and amin-hash encoder, for fast approximation of string similarities. We show that min-hash turns set inclusions into inequality relations that are easier to learn. Both approaches are scalable and streamable. Experiments on real and simulated data show that these methods improve supervised learning with high-cardinality categorical variables. We recommend the following: if scalability is central, the min-hash encoder is the best option as it does not require any data fit; if interpretability is important, the Gamma-Poisson factorization is the best alternative, as it can be interpreted as one-hot encoding on inferred categories with informative feature names. Both models enable autoML on string entries as they remove the need for feature engineering or data cleaning. Patricio Cerda-Mardini, Gaël Varoquaux |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | A Lightweight Neural Model for Biomedical Entity LinkingabstractBiomedical entity linking aims to map biomedical mentions, such as diseases and drugs, to standard entities in a given knowledge base. The specific challenge in this context is that the same biomedical entity can have a wide range of names, including synonyms, morphological variations, and names with different word orderings. Recently, BERT-based methods have advanced the state-of-the-art by allowing for rich representations of word sequences. However, they often have hundreds of millions of parameters and require heavy computing resources, which limits their applications in resource-limited scenarios. Here, we propose a lightweight neural method for biomedical entity linking, which needs just a fraction of the parameters of a BERT model and much less computing resources. Our method uses a simple alignment layer with attention mechanisms to capture the variations between mention and entity names. Yet, we show that our model is competitive with previous work on standard evaluation benchmarks. Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek |
AAAI | 2 |
| 2021 | What's a good imputation to predict with missing values?abstractHow to learn a good predictor on data with missing values? Most efforts focus on first imputing as well as possible and second learning on the completed data to predict the outcome. Yet, this widespread practice has no theoretical grounding. Here we show that for almost all imputation functions, an impute-then-regress procedure with a powerful learner is Bayes optimal. This result holds for all missing-values mechanisms, in contrast with the classic statistical results that require missing-at-random settings to use imputation in probabilistic modeling. Moreover, it implies that perfect conditional imputation is not needed for good prediction asymptotically. In fact, we show that on perfectly imputed data the best regression function will generally be discontinuous, which makes it hard to learn. Crafting instead the imputation so as to leave the regression function unchanged simply shifts the problem to learning discontinuous imputations. Rather, we suggest that it is easier to learn imputation and regression jointly. We propose such a procedure, adapting NeuMiss, a neural network capturing the conditional links across observed and unobserved variables whatever the missing-value pattern. Our experiments confirm that joint imputation and regression through NeuMiss is better than various two step procedures in a finite-sample regime. Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël Varoquaux |
NeurIPS | 4 |
| 2021 | Extracting representations of cognition across neuroimaging studies improves brain decodingabstractCognitive brain imaging is accumulating datasets about the neural substrate of many different mental processes. Yet, most studies are based on few subjects and have low statistical power. Analyzing data across studies could bring more statistical power; yet the current brain-imaging analytic framework cannot be used at scale as it requires casting all cognitive tasks in a unified theoretical framework. We introduce a new methodology to analyze brain responses across tasks without a joint model of the psychological processes. The method boosts statistical power in small studies with specific cognitive focus by analyzing them jointly with large studies that probe less focal mental processes. Our approach improves decoding performance for 80% of 35 widely-different functional-imaging studies. It finds commonalities across tasks in a data-driven way, via common brain representations that predict mental processes. These are brain networks tuned to psychological manipulations. They outline interpretable and plausible brain structures. The extracted networks have been made available; they can be readily reused in new neuro-imaging studies. We provide a multi-study decoding tool to adapt to new data. Arthur Mensch, Julien Mairal, Bertrand Thirion, Gaël Varoquaux |
PLoS Comput. Biol. | 4 |
| 2020 | Linear predictor on linearly-generated data with missing values: non consistency and solutionsabstractWe consider building predictors when the data have missing values. We study the seemingly-simple case where the target to predict is a linear function of the fully observed data and we show that, in the presence of missing values, the optimal predictor is not linear in general. In the particular Gaussian case, it can be written as a linear function of multiway interactions between the observed data and the various missing value indicators. Due to its intrinsic complexity, we study a simple approximation and prove generalization bounds with finite samples, highlighting regimes for which each method performs best. We then show that multilayer perceptrons with ReLU activation functions can be consistent, and can explore good trade-offs between the true model and approximations. Our study highlights the interesting family of models that are beneficial to fit with missing values depending on the amount of data available. Marine Le Morvan, Nicolas Prost, Julie Josse, Erwan Scornet, Gaël Varoquaux |
AISTATS | 5 |
| 2020 | NeuMiss networks: differentiable programming for supervised learning with missing valuesabstractThe presence of missing values makes supervised learning much more challenging. Indeed, previous work has shown that even when the response is a linear function of the complete data, the optimal predictor is a complex function of the observed entries and the missingness indicator. As a result, the computational or sample complexities of consistent approaches depend on the number of missing patterns, which can be exponential in the number of dimensions. In this work, we derive the analytical form of the optimal predictor under a linearity assumption and various missing data mechanisms including Missing at Random (MAR) and self-masking (Missing Not At Random). Based on a Neumann-series approximation of the optimal predictor, we propose a new principled architecture, named NeuMiss networks. Their originality and strength come from the use of a new type of non-linearity: the multiplication by the missingness indicator. We provide an upper bound on the Bayes risk of NeuMiss networks, and show that they have good predictive accuracy with both a number of parameters and a computational complexity independent of the number of missing data patterns. As a result they scale well to problems with many features, and remain statistically efficient for medium-sized samples. Moreover, we show that, contrary to procedures using EM or imputation, they are robust to the missing data mechanism, including difficult MNAR settings such as self-masking. Marine Le Morvan, Julie Josse, Thomas Moreau 0001, Erwan Scornet, Gaël Varoquaux |
NeurIPS | 5 |
| 2019 | Feature Grouping as a Stochastic Regularizer for High-Dimensional Structured DataabstractIn many applications where collecting data is expensive, for example neuroscience or medical imaging, the sample size is typically small compared to the feature dimension. These datasets call for intelligent regularization that exploits known structure, such as correlations between the features arising from the measurement device. However, existing structured regularizers need specially crafted solvers, which are difficult to apply to complex models. We propose a new regularizer specifically designed to leverage structure in the data in a way that can be applied efficiently to complex models. Our approach relies on feature grouping, using a fast clustering algorithm inside a stochastic gradient descent loop: given a family of feature groupings that capture feature covariations, we randomly select these groups at each iteration. Experiments on two real-world datasets demonstrate that the proposed approach produces models that generalize better than those trained with conventional regularizers, and also improves convergence speed, and has a linear computational cost. Sergül Aydöre, Bertrand Thirion, Gaël Varoquaux |
ICML | 3 |
| 2019 | Manifold-regression to predict from MEG/EEG brain signals without source modelingabstractMagnetoencephalography and electroencephalography (M/EEG) can reveal neuronal dynamics non-invasively in real-time and are therefore appreciated methods in medicine and neuroscience. Recent advances in modeling brain-behavior relationships have highlighted the effectiveness of Riemannian geometry for summarizing the spatially correlated time-series from M/EEG in terms of their covariance. However, after artefact-suppression, M/EEG data is often rank deficient which limits the application of Riemannian concepts. In this article, we focus on the task of regression with rank-reduced covariance matrices. We study two Riemannian approaches that vectorize the M/EEG covariance between sensors through projection into a tangent space. The Wasserstein distance readily applies to rank-reduced data but lacks affine-invariance. This can be overcome by finding a common subspace in which the covariance matrices are full rank, enabling the affine-invariant geometric distance. We investigated the implications of these two approaches in synthetic generative models, which allowed us to control estimation bias of a linear model for prediction. We show that Wasserstein and geometric distances allow perfect out-of-sample prediction on the generative models. We then evaluated the methods on real data with regard to their effectiveness in predicting age from M/EEG covariance matrices. The findings suggest that the data-driven Riemannian methods outperform different sensor-space estimators and that they get close to the performance of biophysics-driven source-localization model that requires MRI acquisitions and tedious data processing. Our study suggests that the proposed Riemannian methods can serve as fundamental building-blocks for automated large-scale analysis of M/EEG. David Sabbagh, Pierre Ablin, Gaël Varoquaux, Alexandre Gramfort, Denis A. Engemann |
NeurIPS | 3 |
| 2019 | Comparing distributions: 퓁1 geometry improves kernel two-sample testing
Meyer Scetbon, Gaël Varoquaux |
NeurIPS | 2 |
| 2019 | Population shrinkage of covariance (PoSCE) for better individual brain functional-connectivity estimation
Mehdi Rahim, Bertrand Thirion, Gaël Varoquaux |
Medical Image Anal. | 3 |
| 2019 | Recursive Nearest Agglomeration (ReNA): Fast Clustering for Approximation of Structured SignalsabstractIn this work, we revisit fast dimension reduction approaches, as with random projections and random sampling. Our goal is to summarize the data to decrease computational costs and memory footprint of subsequent analysis. Such dimension reduction can be very efficient when the signals of interest have a strong structure, such as with images. We focus on this setting and investigate feature clustering schemes for data reductions that capture this structure. An impediment to fast dimension reduction is then that good clustering comes with large algorithmic costs. We address it by contributing a linear-time agglomerative clustering scheme, Recursive Nearest Agglomeration (ReNA). Unlike existing fast agglomerative schemes, it avoids the creation of giant clusters. We empirically validate that it approximates the data as well as traditional variance-minimizing clustering schemes that have a quadratic complexity. In addition, we analyze signal approximation with feature clustering and show that it can remove noise, improving subsequent analysis steps. As a consequence, data reduction by clustering features with ReNA yields very fast and accurate models, enabling to process large datasets on budget. Our theoretical analysis is backed by extensive experiments on publicly-available data that illustrate the computation efficiency and the denoising properties of the resulting dimension reduction scheme. Andrés Hoyos Idrobo, Gaël Varoquaux, Jonas Kahn, Bertrand Thirion |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Text to Brain: Predicting the Spatial Distribution of Neuroimaging Observations from Text Reports
Jérôme Dockès, Demian Wassermann, Russell A. Poldrack, Fabian M. Suchanek, Bertrand Thirion, Gaël Varoquaux |
MICCAI (3) | 6 |
| 2018 | Similarity encoding for learning with dirty categorical variables
Patricio Cerda-Mardini, Gaël Varoquaux, Balázs Kégl |
Mach. Learn. | 2 |
| 2018 | Atlases of cognition with large-scale human brain mappingabstractTo map the neural substrate of mental function, cognitive neuroimaging relies on controlled psychological manipulations that engage brain systems associated with specific cognitive processes. In order to build comprehensive atlases of cognitive function in the brain, it must assemble maps for many different cognitive processes, which often evoke overlapping patterns of activation. Such data aggregation faces contrasting goals: on the one hand finding correspondences across vastly different cognitive experiments, while on the other hand precisely describing the function of any given brain region. Here we introduce a new analysis framework that tackles these difficulties and thereby enables the generation of brain atlases for cognitive function. The approach leverages ontologies of cognitive concepts and multi-label brain decoding to map the neural substrate of these concepts. We demonstrate the approach by building an atlas of functional brain organization based on 30 diverse functional neuroimaging studies, totaling 196 different experimental conditions. Unlike conventional brain mapping, this functional atlas supports robust reverse inference: predicting the mental processes from brain activity in the regions delineated by the atlas. To establish that this reverse inference is indeed governed by the corresponding concepts, and not idiosyncrasies of experimental designs, we show that it can accurately decode the cognitive concepts recruited in new tasks. These results demonstrate that aggregating independent task-fMRI studies can provide a more precise global atlas of selective associations between brain and cognition. Gaël Varoquaux, Yannick Schwartz, Russell A. Poldrack, Baptiste Gauthier, Danilo Bzdok, Jean-Baptiste Poline, Bertrand Thirion |
PLoS Comput. Biol. | 1 |
| 2017 | Learning to Discover Sparse Graphical ModelsabstractWe consider structure discovery of undirected graphical models from observational data. Inferring likely structures from few examples is a complex task often requiring the formulation of priors and sophisticated inference procedures. Popular methods rely on estimating a penalized maximum likelihood of the precision matrix. However, in these approaches structure recovery is an indirect consequence of the data-fit term, the penalty can be difficult to adapt for domain-specific knowledge, and the inference is computationally demanding. By contrast, it may be easier to generate training samples of data that arise from graphs with the desired structure properties. We propose here to leverage this latter source of information as training data to learn a function, parametrized by a neural network, that maps empirical covariance matrices to estimated graph structures. Learning this function brings two benefits: it implicitly models the desired structure or sparsity properties to form suitable priors, and it can be tailored to the specific problem of edge structure discovery, rather than maximizing data likelihood. Applying this framework, we find our learnable graph-discovery method trained on synthetic data generalizes well: identifying relevant edges in both synthetic and real data, completely unknown at training time. We find that on genetics, brain imaging, and simulation data we obtain performance generally superior to analytical methods. Eugene Belilovsky, Kyle Kastner, Gaël Varoquaux, Matthew B. Blaschko |
ICML | 3 |
| 2017 | Population-Shrinkage of Covariance to Estimate Better Brain Functional Connectivity
Mehdi Rahim, Bertrand Thirion, Gaël Varoquaux |
MICCAI (1) | 3 |
| 2017 | Learning Neural Representations of Human Cognition across Many fMRI StudiesabstractCognitive neuroscience is enjoying rapid increase in extensive public brain-imaging datasets. It opens the door to large-scale statistical models. Finding a unified perspective for all available data calls for scalable and automated solutions to an old challenge: how to aggregate heterogeneous information on brain function into a universal cognitive system that relates mental operations/cognitive processes/psychological tasks to brain networks? We cast this challenge in a machine-learning approach to predict conditions from statistical brain maps across different studies. For this, we leverage multi-task learning and multi-scale dimension reduction to learn low-dimensional representations of brain images that carry cognitive information and can be robustly associated with psychological stimuli. Our multi-dataset classification model achieves the best prediction performance on several large reference datasets, compared to models without cognitive-aware low-dimension representations; it brings a substantial performance boost to the analysis of small datasets, and can be introspected to identify universal template cognitive concepts. Arthur Mensch, Julien Mairal, Danilo Bzdok, Bertrand Thirion, Gaël Varoquaux |
NIPS | 5 |
| 2017 | BIDS apps: Improving ease of use, accessibility, and reproducibility of neuroimaging data analysis methodsabstractThe rate of progress in human neurosciences is limited by the inability to easily apply a wide range of analysis methods to the plethora of different datasets acquired in labs around the world. In this work, we introduce a framework for creating, testing, versioning and archiving portable applications for analyzing neuroimaging data organized and described in compliance with the Brain Imaging Data Structure (BIDS). The portability of these applications (BIDS Apps) is achieved by using container technologies that encapsulate all binary and other dependencies in one convenient package. BIDS Apps run on all three major operating systems with no need for complex setup and configuration and thanks to the comprehensiveness of the BIDS standard they require little manual user input. Previous containerized data processing solutions were limited to single user environments and not compatible with most multi-tenant High Performance Computing systems. BIDS Apps overcome this limitation by taking advantage of the Singularity container technology. As a proof of concept, this work is accompanied by 22 ready to use BIDS Apps, packaging a diverse set of commonly used neuroimaging algorithms. Krzysztof J. Gorgolewski, Fidel Alfaro-Almagro, Tibor Auer, Lune Bellec, Mihai Capota, M. Mallar Chakravarty, Nathan William Churchill, Alexander Li Cohen, R. Cameron Craddock, Gabriel A. Devenyi, Anders Eklund 0002, Oscar Esteban, Guillaume Flandin, Satrajit S. Ghosh, J. Swaroop Guntupalli, Mark Jenkinson, Anisha Keshavan, Gregory Kiar, Franziskus Liem, Pradeep Reddy Raamana, David Raffelt, Christopher John Steele, Pierre-Olivier Quirion, Robert E. Smith 0002, Stephen C. Strother, Gaël Varoquaux, Yida Wang 0003, Tal Yarkoni, Russell A. Poldrack |
PLoS Comput. Biol. | 26 |
| 2016 | Local Q-linear convergence and finite-time active set identification of ADMM on a class of penalized regression problemsabstractWe study the convergence of the ADMM (Alternating Direction Method of Multipliers) algorithm on a broad range of penalized regression problems including the Lasso, Group-Lasso and Graph-Lasso,(isotropic) TV-L1, Sparse Variation, and others. First, we establish a fixed-point iterationvia a nonlinear operator-which is equivalent to the ADMM iterates. We then show that this nonlinear operator is Fréchet-differentiable almost everywhere and that around each fixed point, Q-linear convergence is guaranteed, provided the spectral radius of the Jacobian of the operator at the fixed point is less than 1 (a classical result on stability). Moreover, this spectral radius is then a rate of convergence for the ADMM algorithm. Also, we show that the support of the split variable can be identified after finitely many iterations. In the anisotropic cases, we show that for sufficiently large values of the tuning parameter, we recover the optimal rates in terms of Friedrichs angles, that have appeared recently in the literature. Empirical results on various problems are also presented and discussed. Elvis Dohmatob, Michael Eickenberg, Bertrand Thirion, Gaël Varoquaux |
ICASSP | 4 |
| 2016 | Dictionary Learning for Massive Matrix FactorizationabstractSparse matrix factorization is a popular tool to obtain interpretable data decompositions, which are also effective to perform data completion or denoising. Its applicability to large datasets has been addressed with online and randomized methods, that reduce the complexity in one of the matrix dimension, but not in both of them. In this paper, we tackle very large matrices in both dimensions. We propose a new factorization method that scales gracefully to terabyte-scale datasets. Those could not be processed by previous algorithms in a reasonable amount of time. We demonstrate the efficiency of our approach on massive functional Magnetic Resonance Imaging (fMRI) data, and on matrix completion problems for recommender systems, where we obtain significant speed-ups compared to state-of-the art coordinate descent methods. Arthur Mensch, Julien Mairal, Bertrand Thirion, Gaël Varoquaux |
ICML | 4 |
| 2016 | Testing for Differences in Gaussian Graphical Models: Applications to Brain ConnectivityabstractFunctional brain networks are well described and estimated from data with Gaussian Graphical Models (GGMs), e.g.\ using sparse inverse covariance estimators. Comparing functional connectivity of subjects in two populations calls for comparing these estimated GGMs. Our goal is to identify differences in GGMs known to have similar structure. We characterize the uncertainty of differences with confidence intervals obtained using a parametric distribution on parameters of a sparse estimator. Sparse penalties enable statistical guarantees and interpretable models even in high-dimensional and low-sample settings. Characterizing the distributions of sparse models is inherently challenging as the penalties produce a biased estimator. Recent work invokes the sparsity assumptions to effectively remove the bias from a sparse estimator such as the lasso. These distributions can be used to give confidence intervals on edges in GGMs, and by extension their differences. However, in the case of comparing GGMs, these estimators do not make use of any assumed joint structure among the GGMs. Inspired by priors from brain functional connectivity we derive the distribution of parameter differences under a joint penalty when parameters are known to be sparse in the difference. This leads us to introduce the debiased multi-task fused lasso, whose distribution can be characterized in an efficient manner. We then show how the debiased lasso and multi-task fused lasso can be used to obtain confidence intervals on edge differences in GGMs. We validate the techniques proposed on a set of synthetic examples as well as neuro-imaging dataset created for the study of autism. Eugene Belilovsky, Gaël Varoquaux, Matthew B. Blaschko |
NIPS | 2 |
| 2016 | Learning brain regions via large-scale online structured sparse dictionary learningabstractWe propose a multivariate online dictionary-learning method for obtaining decompositions of brain images with structured and sparse components (aka atoms). Sparsity is to be understood in the usual sense: the dictionary atoms are constrained to contain mostly zeros. This is imposed via an $\ell_1$-norm constraint. By "structured", we mean that the atoms are piece-wise smooth and compact, thus making up blobs, as opposed to scattered patterns of activation. We propose to use a Sobolev (Laplacian) penalty to impose this type of structure. Combining the two penalties, we obtain decompositions that properly delineate brain structures from functional images. This non-trivially extends the online dictionary-learning work of Mairal et al. (2010), at the price of only a factor of 2 or 3 on the overall running time. Just like the Mairal et al. (2010) reference method, the online nature of our proposed algorithm allows it to scale to arbitrarily sized datasets. Experiments on brain data show that our proposed method extracts structured and denoised dictionaries that are more intepretable and better capture inter-subject variability in small medium, and large-scale regimes alike, compared to state-of-the-art models. Elvis Dohmatob, Arthur Mensch, Gaël Varoquaux, Bertrand Thirion |
NIPS | 3 |
| 2016 | Formal Models of the Network Co-occurrence Underlying Mental OperationsabstractSystems neuroscience has identified a set of canonical large-scale networks in humans. These have predominantly been characterized by resting-state analyses of the task-unconstrained, mind-wandering brain. Their explicit relationship to defined task performance is largely unknown and remains challenging. The present work contributes a multivariate statistical learning approach that can extract the major brain networks and quantify their configuration during various psychological tasks. The method is validated in two extensive datasets (n = 500 and n = 81) by model-based generation of synthetic activity maps from recombination of shared network topographies. To study a use case, we formally revisited the poorly understood difference between neural activity underlying idling versus goal-directed behavior. We demonstrate that task-specific neural activity patterns can be explained by plausible combinations of resting-state networks. The possibility of decomposing a mental task into the relative contributions of major brain networks, the "network co-occurrence architecture" of a given task, opens an alternative access to the neural substrates of human cognition. Danilo Bzdok, Gaël Varoquaux, Olivier Grisel, Michael Eickenberg, Cyril Poupon, Bertrand Thirion |
PLoS Comput. Biol. | 2 |
| 2016 | Transport on Riemannian Manifold for Connectivity-Based Brain DecodingabstractThere is a recent interest in using functional magnetic resonance imaging (fMRI) for decoding more naturalistic, cognitive states, in which subjects perform various tasks in a continuous, self-directed manner. In this setting, the set of brain volumes over the entire task duration is usually taken as a single sample with connectivity estimates, such as Pearson's correlation, employed as features. Since covariance matrices live on the positive semidefinite cone, their elements are inherently inter-related. The assumption of uncorrelated features implicit in most classifier learning algorithms is thus violated. Coupled with the usual small sample sizes, the generalizability of the learned classifiers is limited, and the identification of significant brain connections from the classifier weights is nontrivial. In this paper, we present a Riemannian approach for connectivity-based brain decoding. The core idea is to project the covariance estimates onto a common tangent space to reduce the statistical dependencies between their elements. For this, we propose a matrix whitening transport, and compare it against parallel transport implemented via the Schild's ladder algorithm. To validate our classification approach, we apply it to fMRI data acquired from twenty four subjects during four continuous, self-driven tasks. We show that our approach provides significantly higher classification accuracy than directly using Pearson's correlation and its regularized variants as features. To facilitate result interpretation, we further propose a non-parametric scheme that combines bootstrapping and permutation testing for identifying significantly discriminative brain connections from the classifier weights. Using this scheme, a number of neuro-anatomically meaningful connections are detected, whereas no significant connections are found with pure permutation testing. Bernard Ng, Gaël Varoquaux, Jean-Baptiste Poline, Michael D. Greicius, Bertrand Thirion |
IEEE Trans. Medical Imaging | 2 |
| 2015 | Grouping Total Variation and Sparsity: Statistical Learning with Segmenting Penalties
Michael Eickenberg, Elvis Dohmatob, Bertrand Thirion, Gaël Varoquaux |
MICCAI (1) | 4 |
| 2015 | Integrating Multimodal Priors in Predictive Models for the Functional Characterization of Alzheimer's Disease
Mehdi Rahim, Bertrand Thirion, Alexandre Abraham, Michael Eickenberg, Elvis Dohmatob, Claude Comtat, Gaël Varoquaux |
MICCAI (1) | 7 |
| 2015 | Semi-Supervised Factored Logistic Regression for High-Dimensional Neuroimaging DataabstractImaging neuroscience links human behavior to aspects of brain biology in ever-increasing datasets. Existing neuroimaging methods typically perform either discovery of unknown neural structure or testing of neural structure associated with mental tasks. However, testing hypotheses on the neural correlates underlying larger sets of mental tasks necessitates adequate representations for the observations. We therefore propose to blend representation modelling and task classification into a unified statistical learning problem. A multinomial logistic regression is introduced that is constrained by factored coefficients and coupled with an autoencoder. We show that this approach yields more accurate and interpretable neural models of psychological tasks in a reference dataset, as well as better generalization to other datasets. Danilo Bzdok, Michael Eickenberg, Olivier Grisel, Bertrand Thirion, Gaël Varoquaux |
NIPS | 5 |
| 2015 | Convex relaxations of penalties for sparse correlated variables with bounded total variation
Eugene Belilovsky, Andreas Argyriou, Gaël Varoquaux, Matthew B. Blaschko |
Mach. Learn. | 3 |
| 2014 | Transport on Riemannian Manifold for Functional Connectivity-Based Classification
Bernard Ng, Martin Dressler, Gaël Varoquaux, Jean-Baptiste Poline, Michael D. Greicius, Bertrand Thirion |
MICCAI (2) | 3 |
| 2014 | Deriving a Multi-subject Functional-Connectivity Atlas to Inform Connectome Estimation
Ronald Phlypo, Bertrand Thirion, Gaël Varoquaux |
MICCAI (3) | 3 |
| 2014 | Principal Component Regression Predicts Functional Responses across Individuals
Bertrand Thirion, Gaël Varoquaux, Olivier Grisel, Cyril Poupon, Philippe Pinel |
MICCAI (2) | 2 |
| 2013 | Extracting Brain Regions from Rest fMRI with Total-Variation Constrained Dictionary Learning
Alexandre Abraham, Elvis Dohmatob, Bertrand Thirion, Dimitris Samaras, Gaël Varoquaux |
MICCAI (2) | 5 |
| 2013 | Enhancing the Reproducibility of Group Analysis with Randomized Brain Parcellations
Benoit Da Mota, Virgile Fritsch, Gaël Varoquaux, Vincent Frouin, Jean-Baptiste Poline, Bertrand Thirion |
MICCAI (2) | 3 |
| 2013 | Implications of Inconsistencies between fMRI and dMRI on Multimodal Connectivity Estimation
Bernard Ng, Gaël Varoquaux, Jean-Baptiste Poline, Bertrand Thirion |
MICCAI (3) | 2 |
| 2013 | Mapping paradigm ontologies to and from the brainabstractImaging neuroscience links brain activation maps to behavior and cognition via correlational studies. Due to the nature of the individual experiments, based on eliciting neural response from a small number of stimuli, this link is incomplete, and unidirectional from the causal point of view. To come to conclusions on the function implied by the activation of brain regions, it is necessary to combine a wide exploration of the various brain functions and some inversion of the statistical inference. Here we introduce a methodology for accumulating knowledge towards a bidirectional link between observed brain activity and the corresponding function. We rely on a large corpus of imaging studies and a predictive engine. Technically, the challenges are to find commonality between the studies without denaturing the richness of the corpus. The key elements that we contribute are labeling the tasks performed with a cognitive ontology, and modeling the long tail of rare paradigms in the corpus. To our knowledge, our approach is the first demonstration of predicting the cognitive content of completely new brain images. To that end, we propose a method that predicts the experimental paradigms across different studies. Yannick Schwartz, Bertrand Thirion, Gaël Varoquaux |
NIPS | 3 |
| 2013 | A Framework for Inter-Subject Prediction of Functional Connectivity From Structural NetworksabstractFunctional connections between brain regions are supported by structural connectivity. Both functional and structural connectivity are estimated from in vivo magnetic resonance imaging and offer complementary information on brain organization and function. However, imaging only provides noisy measures, and we lack a good neuroscientific understanding of the links between structure and function. Therefore, inter-subject joint modeling of structural and functional connectivity, the key to multimodal biomarkers, is an open challenge. We present a probabilistic framework to learn across subjects a mapping from structural to functional brain connectivity. Expanding on our previous work [1], our approach is based on a predictive framework with multiple sparse linear regression. We rely on the randomized LASSO to identify relevant anatomo-functional links with some confidence interval. In addition, we describe resting-state functional magnetic resonance imaging in the setting of Gaussian graphical models, on the one hand imposing conditional independences from structural connectivity and on the other hand parameterizing the problem in terms of multivariate autoregressive models. We introduce an intrinsic measure of prediction error for functional connectivity that is independent of the parameterization chosen and provides the means for robust model selection. We demonstrate our methodology with regions within the default mode and the salience network as well as, atlas-based cortical parcellation. Fani Deligianni, Gaël Varoquaux, Bertrand Thirion, David J. Sharp, Christian Ledig, Robert Leech, Daniel Rueckert |
IEEE Trans. Medical Imaging | 2 |
| 2012 | Small-sample brain mapping: sparse recovery on spatially correlated designs with randomization and clustering
Gaël Varoquaux, Alexandre Gramfort, Bertrand Thirion |
ICML | 1 |
| 2012 | A Novel Sparse Graphical Approach for Multimodal Brain Connectivity Inference
Bernard Ng, Gaël Varoquaux, Jean-Baptiste Poline, Bertrand Thirion |
MICCAI (1) | 2 |
| 2012 | Improving Accuracy and Power with Transfer Learning Using a Meta-analytic Database
Yannick Schwartz, Gaël Varoquaux, Christophe Pallier, Philippe Pinel, Jean-Baptiste Poline, Bertrand Thirion |
MICCAI (3) | 2 |
| 2012 | Detecting outliers in high-dimensional neuroimaging datasets with robust covariance estimators
Virgile Fritsch, Gaël Varoquaux, Benjamin Thyreau, Jean-Baptiste Poline, Bertrand Thirion |
Medical Image Anal. | 2 |
| 2012 | A supervised clustering approach for fMRI-based inference of brain states
Vincent Michel, Alexandre Gramfort, Gaël Varoquaux, Evelyn Eger, Christine Keribin, Bertrand Thirion |
Pattern Recognit. | 3 |
| 2011 | Detecting Outlying Subjects in High-Dimensional Neuroimaging Datasets with Regularized Minimum Covariance Determinant
Virgile Fritsch, Gaël Varoquaux, Benjamin Thyreau, Jean-Baptiste Poline, Bertrand Thirion |
MICCAI (3) | 2 |
| 2011 | Connectivity-Informed fMRI Activation Detection
Bernard Ng, Rafeef Abugharbieh, Gaël Varoquaux, Jean-Baptiste Poline, Bertrand Thirion |
MICCAI (2) | 3 |
| 2011 | Scikit-learn: Machine Learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jacob VanderPlas, Alexandre Tachard Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, Edouard Duchesnay |
J. Mach. Learn. Res. | 2 |
| 2011 | Total Variation Regularization for fMRI-Based Prediction of BehaviorabstractWhile medical imaging typically provides massive amounts of data, the extraction of relevant information for predictive diagnosis remains a difficult challenge. Functional magnetic resonance imaging (fMRI) data, that provide an indirect measure of task-related or spontaneous neuronal activity, are classically analyzed in a mass-univariate procedure yielding statistical parametric maps. This analysis framework disregards some important principles of brain organization: population coding, distributed and overlapping representations. Multivariate pattern analysis, i.e., the prediction of behavioral variables from brain activation patterns better captures this structure. To cope with the high dimensionality of the data, the learning method has to be regularized. However, the spatial structure of the image is not taken into account in standard regularization methods, so that the extracted features are often hard to interpret. More informative and interpretable results can be obtained with the l(1) norm of the image gradient, also known as its total variation (TV), as regularization. We apply for the first time this method to fMRI data, and show that TV regularization is well suited to the purpose of brain mapping while being a powerful tool for brain decoding. Moreover, this article presents the first use of TV regularization for classification. Vincent Michel, Alexandre Gramfort, Gaël Varoquaux, Evelyn Eger, Bertrand Thirion |
IEEE Trans. Medical Imaging | 3 |
| 2010 | Accurate Definition of Brain Regions Position through the Functional Landmark Approach
Bertrand Thirion, Gaël Varoquaux, Jean-Baptiste Poline |
MICCAI (2) | 2 |
| 2010 | Detection of Brain Functional-Connectivity Difference in Post-stroke Patients Using Group-Level Covariance Modeling
Gaël Varoquaux, Flore Baronnet, Andreas Kleinschmidt, Pierre Fillard, Bertrand Thirion |
MICCAI (1) | 1 |
| 2010 | Brain covariance selection: better individual functional connectivity models using population priorabstractSpontaneous brain activity, as observed in functional neuroimaging, has been shown to display reproducible structure that expresses brain architecture and carries markers of brain pathologies. An important view of modern neuroscience is that such large-scale structure of coherent activity reflects modularity properties of brain connectivity graphs. However, to date, there has been no demonstration that the limited and noisy data available in spontaneous activity observations could be used to learn full-brain probabilistic models that generalize to new data. Learning such models entails two main challenges: i) modeling full brain connectivity is a difficult estimation problem that faces the curse of dimensionality and ii) variability between subjects, coupled with the variability of functional signals between experimental runs, makes the use of multiple datasets challenging. We describe subject-level brain functional connectivity structure as a multivariate Gaussian process and introduce a new strategy to estimate it from group data, by imposing a common structure on the graphical model in the population. We show that individual models learned from functional Magnetic Resonance Imaging (fMRI) data using this population prior generalize better to unseen data than models based on alternative regularization schemes. To our knowledge, this is the first report of a cross-validated model of spontaneous brain activity. Finally, we use the estimated graphical model to explore the large-scale characteristics of functional architecture and show for the first time that known cognitive networks appear as the integrated communities of functional connectivity graph. Gaël Varoquaux, Alexandre Gramfort, Jean-Baptiste Poline, Bertrand Thirion |
NIPS | 1 |