Anthony L. Caterini

dblp:167/4383 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Deep learning architectures and training · 19% Trustworthy machine learning · 16% Generative modeling · 16%

Topics — the 20 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
normalizing flow
1.032021
Rectangular Flows for Manifold Learning · NeurIPS 2021
Relaxing Bijectivity Constraints with Continuously Indexed Normalising Flows · ICML 2020
Hamiltonian Variational Auto-Encoder · NeurIPS 2018
Machine learning › Deep learning architectures and training
foundation model
0.912025
TabDPT: Scaling Tabular Foundation Models on Real Data · NeurIPS 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
TabDPT: Scaling Tabular Foundation Models on Real Data · NeurIPS 2025
Machine learning › Deep learning architectures and training › foundation model
tabular foundation model
0.912025
TabDPT: Scaling Tabular Foundation Models on Real Data · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
Retrieval & Fine-Tuning for In-Context Tabular Models · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
in-context learning for tabular data
0.812024
Retrieval & Fine-Tuning for In-Context Tabular Models · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › intrinsic dimension
intrinsic dimension estimation
0.812024
A Geometric Explanation of the Likelihood OOD Detection Paradox · ICML 2024
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.812024
A Geometric Explanation of the Likelihood OOD Detection Paradox · ICML 2024
Machine learning › Deep learning architectures and training
tabular deep learning
0.812024
Retrieval & Fine-Tuning for In-Context Tabular Models · NeurIPS 2024
Machine learning › Trustworthy machine learning
uncertainty and robustness
0.812024
A Geometric Explanation of the Likelihood OOD Detection Paradox · ICML 2024
Machine learning › Generative modeling
generative model evaluation
0.712023
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models · NeurIPS 2023
Machine learning › Learning theory › inductive bias
manifold hypothesis
0.712023
Verifying the Union of Manifolds Hypothesis for Image Data · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning
representation geometry
0.712023
Verifying the Union of Manifolds Hypothesis for Image Data · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.512021
Rectangular Flows for Manifold Learning · NeurIPS 2021
Machine learning › Reinforcement learning
value function estimation
0.512021
C-Learning: Horizon-Aware Cumulative Accessibility Estimation · ICLR 2021
Machine learning › Generative modeling
variational autoencoder
0.312018
Hamiltonian Variational Auto-Encoder · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.312018
Hamiltonian Variational Auto-Encoder · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.112021
Rectangular Flows for Manifold Learning · NeurIPS 2021
Machine learning › Generative modeling › normalizing flow
residual flows
0.112020
Relaxing Bijectivity Constraints with Continuously Indexed Normalising Flows · ICML 2020

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 0.9retrieval-based in-context learning · 0.9transformer · 0.8score-based diffusion models · 0.8retrieval · 0.8normalizing flow · 0.8local intrinsic dimension estimation · 0.8in-context learning · 0.8fine-tuning · 0.8manifold analysis · 0.7
YearPublicationVenuePosition
2025 TabDPT: Scaling Tabular Foundation Models on Real Data
abstract
Tabular data is one of the most ubiquitous sources of information worldwide, spanning a wide variety of domains. This inherent heterogeneity has slowed the development of Tabular Foundation Models (TFMs) capable of fast generalization to unseen datasets. In-Context Learning (ICL) has recently emerged as a promising solution for TFMs, enabling dynamic adaptation to new tasks without additional tuning. While many studies have attempted to re-purpose large language models for tabular ICL, they have had limited success, so recent works have focused on developing tabular-specific foundation models. In this work, we propose an approach to combine ICL-based retrieval with self supervised learning to train tabular foundation models. We also investigate the utility of real vs. synthetic data for model pre-training, and show that real data can contain useful signal not easily captured in synthetic training. Specifically, we show that incorporating real data during the pre-training phase can lead to significantly faster training and better downstream generalization to unseen data. Our resulting model, **TabDPT**, achieves strong performance on both regression (CTR23) and classification (CC18) benchmarks. Importantly, we also demonstrate that with our pre-training procedure, scaling both model and data size leads to consistent performance improvements that follow power laws. This echoes scaling laws in LLMs and other foundation models, and suggests that large-scale TFMs can be achievable. We open-source our full pipeline: inference code including trained model weights can be found [here](https://github.com/layer6ai-labs/TabDPT-inference), and the training code to reproduce experiments can be found [here](https://github.com/layer6ai-labs/TabDPT-training).
Junwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach, Jesse C. Cresswell, Keyvan Golestan, Guangwei Yu, Anthony L. Caterini, Maksims Volkovs
NeurIPS8
2024 A Geometric Explanation of the Likelihood OOD Detection Paradox
abstract
Likelihood-based deep generative models (DGMs) commonly exhibit a puzzling behaviour: when trained on a relatively complex dataset, they assign higher likelihood values to out-of-distribution (OOD) data from simpler sources. Adding to the mystery, OOD samples are never generated by these DGMs despite having higher likelihoods. This two-pronged paradox has yet to be conclusively explained, making likelihood-based OOD detection unreliable. Our primary observation is that high-likelihood regions will not be generated if they contain minimal probability mass. We demonstrate how this seeming contradiction of large densities yet low probability mass can occur around data confined to low-dimensional manifolds. We also show that this scenario can be identified through local intrinsic dimension (LID) estimation, and propose a method for OOD detection which pairs the likelihoods and LID estimates obtained from a *pre-trained* DGM. Our method can be applied to normalizing flows and score-based diffusion models, and obtains results which match or surpass state-of-the-art OOD detection benchmarks using the same DGM backbones. Our code is available at our [GitHub repository](https://github.com/layer6ai-labs/dgm_ood_detection).
Hamidreza Kamkari, Brendan Leigh Ross, Jesse C. Cresswell, Anthony L. Caterini, Rahul G. Krishnan, Gabriel Loaiza-Ganem
ICML4
2024 Retrieval & Fine-Tuning for In-Context Tabular Models
abstract
Tabular data is a pervasive modality spanning a wide range of domains, and this inherent diversity poses a considerable challenge for deep learning. Recent advancements using transformer-based in-context learning have shown promise on smaller and less complex tabular datasets, but have struggled to scale to larger and more complex ones. To address this limitation, we propose a combination of retrieval and fine-tuning: we can adapt the transformer to a local subset of the data by collecting nearest neighbours, and then perform task-specific fine-tuning with this retrieved set of neighbours in context. Using TabPFN as the base model -- currently the best tabular in-context learner -- and applying our retrieval and fine-tuning scheme on top results in what we call a locally-calibrated PFN, or LoCalPFN. We conduct extensive evaluation on 95 datasets curated by TabZilla from OpenML, upon which we establish a new state-of-the-art with LoCalPFN -- even with respect to tuned tree-based models. Notably, we show a significant boost in performance compared to the base in-context model, demonstrating the efficacy of our approach and advancing the frontier of deep learning in tabular data.
Valentin Thomas, Junwei Ma, Rasa Hosseinzadeh, Keyvan Golestan, Guangwei Yu, Maksims Volkovs, Anthony L. Caterini
NeurIPS7
2023 Verifying the Union of Manifolds Hypothesis for Image Data
Bradley C. A. Brown, Anthony L. Caterini, Brendan Leigh Ross, Jesse C. Cresswell, Gabriel Loaiza-Ganem
ICLR2
2023 Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
abstract
We systematically study a wide variety of generative models spanning semantically-diverse image datasets to understand and improve the feature extractors and metrics used to evaluate them. Using best practices in psychophysics, we measure human perception of image realism for generated samples by conducting the largest experiment evaluating generative models to date, and find that no existing metric strongly correlates with human evaluations. Comparing to 17 modern metrics for evaluating the overall performance, fidelity, diversity, rarity, and memorization of generative models, we find that the state-of-the-art perceptual realism of diffusion models as judged by humans is not reflected in commonly reported metrics such as FID. This discrepancy is not explained by diversity in generated samples, though one cause is over-reliance on Inception-V3. We address these flaws through a study of alternative self-supervised feature extractors, find that the semantic information encoded by individual networks strongly depends on their training procedure, and show that DINOv2-ViT-L/14 allows for much richer evaluation of generative models. Next, we investigate data memorization, and find that generative models do memorize training examples on simple, smaller datasets like CIFAR10, but not necessarily on more complex datasets like ImageNet. However, our experiments show that current metrics do not properly detect memorization: none in the literature is able to separate memorization from other phenomena such as underfitting or mode shrinkage. To facilitate further development of generative models and their evaluation we release all generated image datasets, human evaluation data, and a modular library to compute 17 common metrics for 9 different encoders at https://github.com/layer6ai-labs/dgm-eval.
George Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui 0001, Brendan Leigh Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L. Caterini, J. Eric T. Taylor, Gabriel Loaiza-Ganem
NeurIPS8
2021 C-Learning: Horizon-Aware Cumulative Accessibility Estimation
Panteha Naderian, Gabriel Loaiza-Ganem, Harry J. Braviner, Anthony L. Caterini, Jesse C. Cresswell, Animesh Garg
ICLR4
2021 Rectangular Flows for Manifold Learning
abstract
Normalizing flows are invertible neural networks with tractable change-of-volume terms, which allow optimization of their parameters to be efficiently performed via maximum likelihood. However, data of interest are typically assumed to live in some (often unknown) low-dimensional manifold embedded in a high-dimensional ambient space. The result is a modelling mismatch since -- by construction -- the invertibility requirement implies high-dimensional support of the learned distribution. Injective flows, mappings from low- to high-dimensional spaces, aim to fix this discrepancy by learning distributions on manifolds, but the resulting volume-change term becomes more challenging to evaluate. Current approaches either avoid computing this term entirely using various heuristics, or assume the manifold is known beforehand and therefore are not widely applicable. Instead, we propose two methods to tractably calculate the gradient of this term with respect to the parameters of the model, relying on careful use of automatic differentiation and techniques from numerical linear algebra. Both approaches perform end-to-end nonlinear manifold learning and density estimation for data projected onto this manifold. We study the trade-offs between our proposed methods, empirically verify that we outperform approaches ignoring the volume-change term by more accurately learning manifolds and the corresponding distributions on them, and show promising results on out-of-distribution detection. Our code is available at https://github.com/layer6ai-labs/rectangular-flows.
Anthony L. Caterini, Gabriel Loaiza-Ganem, Geoff Pleiss, John P. Cunningham
NeurIPS1
2021 Variational inference with continuously-indexed normalizing flows
abstract
Continuously-indexed flows (CIFs) have recently achieved improvements over baseline normalizing flows on a variety of density estimation tasks. CIFs do not possess a closed-form marginal density, and so, unlike standard flows, cannot be plugged in directly to a variational inference (VI) scheme in order to produce a more expressive family of approximate posteriors. However, we show here how CIFs can be used as part of an auxiliary VI scheme to formulate and train expressive posterior approximations in a natural way. We exploit the conditional independence structure of multi-layer CIFs to build the required auxiliary inference models, which we show empirically yield low-variance estimators of the model evidence. We then demonstrate the advantages of CIFs over baseline flows in VI problems when the posterior distribution of interest possesses a complicated topology, obtaining improved results in both the Bayesian inference and surrogate maximum likelihood settings.
Anthony L. Caterini, Robert Cornish, Dino Sejdinovic, Arnaud Doucet
UAI1
2020 Relaxing Bijectivity Constraints with Continuously Indexed Normalising Flows
abstract
We show that normalising flows become pathological when used to model targets whose supports have complicated topologies. In this scenario, we prove that a flow must become arbitrarily numerically noninvertible in order to approximate the target closely. This result has implications for all flow-based models, and especially residual flows (ResFlows), which explicitly control the Lipschitz constant of the bijection used. To address this, we propose continuously indexed flows (CIFs), which replace the single bijection used by normalising flows with a continuously indexed family of bijections, and which can intuitively "clean up" mass that would otherwise be misplaced by a single bijection. We show theoretically that CIFs are not subject to the same topological limitations as normalising flows, and obtain better empirical performance on a variety of models and benchmarks.
Robert Cornish, Anthony L. Caterini, George Deligiannidis, Arnaud Doucet
ICML2
2018 Hamiltonian Variational Auto-Encoder
abstract
Variational Auto-Encoders (VAE) have become very popular techniques to perform inference and learning in latent variable models as they allow us to leverage the rich representational power of neural networks to obtain flexible approximations of the posterior of latent variables as well as tight evidence lower bounds (ELBO). Com- bined with stochastic variational inference, this provides a methodology scaling to large datasets. However, for this methodology to be practically efficient, it is neces- sary to obtain low-variance unbiased estimators of the ELBO and its gradients with respect to the parameters of interest. While the use of Markov chain Monte Carlo (MCMC) techniques such as Hamiltonian Monte Carlo (HMC) has been previously suggested to achieve this [23, 26], the proposed methods require specifying reverse kernels which have a large impact on performance. Additionally, the resulting unbiased estimator of the ELBO for most MCMC kernels is typically not amenable to the reparameterization trick. We show here how to optimally select reverse kernels in this setting and, by building upon Hamiltonian Importance Sampling (HIS) [17], we obtain a scheme that provides low-variance unbiased estimators of the ELBO and its gradients using the reparameterization trick. This allows us to develop a Hamiltonian Variational Auto-Encoder (HVAE). This method can be re-interpreted as a target-informed normalizing flow [20] which, within our context, only requires a few evaluations of the gradient of the sampled likelihood and trivial Jacobian calculations at each iteration.
Anthony L. Caterini, Arnaud Doucet, Dino Sejdinovic
NeurIPS1
2015 Algorithmic Acceleration of Parallel ALS for Collaborative Filtering: Speeding up Distributed Big Data Recommendation in Spark
abstract
Collaborative filtering algorithms are important building blocks in many practical recommendation systems. For example, many large-scale data processing environments include collaborative filtering models for which the Alternating Least Squares (ALS) algorithm is used to compute latent factor matrix decompositions. In this paper, we propose an approach to accelerate the convergence of parallel ALS-based optimization methods for collaborative filtering using a nonlinear conjugate gradient (NCG) wrapper around the ALS iterations. We also provide a parallel implementation of the accelerated ALS-NCG algorithm in the Apache Spark distributed data processing environment, and an efficient line search technique as part of the ALS-NCG implementation that requires only one pass over the data on distributed datasets. In serial numerical experiments on a linux workstation and parallel numerical experiments on a 16 node cluster with 256 computing cores, we demonstrate that the combined ALS-NCG method requires many fewer iterations and less time than standalone ALS to reach movie rankings with high accuracy on the MovieLens 20M dataset. In parallel, ALS-NCG can achieve an acceleration factor of 4 or greater in clock time when an accurate solution is desired; furthermore, the acceleration factor increases as greater numerical precision is required in the solution. Furthermore, the NCG acceleration mechanism is efficient in parallel and scales linearly with problem size on synthetic datasets with up to nearly 1 billion ratings. The acceleration mechanism is general and may also be applicable to other optimization methods for collaborative filtering.
Manda Winlaw, Michael B. Hynes, Anthony L. Caterini, Hans De Sterck
ICPADS3