EDBT 2026 Demo / reviewers in the wild / expert
Antônio H. Ribeiro
dblp:202/1699
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-3632-8529ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning theory · 24% Kernel, tree and ensemble methods · 23% Trustworthy machine learning · 22% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
1.5 | 2 | 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization · NeurIPS 2025 Regularization properties of adversarially-trained linear regression · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
1.5 | 2 | 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization · NeurIPS 2025 Regularization properties of adversarially-trained linear regression · NeurIPS 2023 |
Computer vision › 3D vision › brain decoding
brain visual decoding |
0.9 | 1 | 2025 | Human-Aligned Image Models Improve Visual Decoding from the Brain · ICML 2025 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning |
0.9 | 1 | 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization · NeurIPS 2025 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.9 | 1 | 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization · NeurIPS 2025 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression |
0.9 | 1 | 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization · NeurIPS 2025 |
Machine learning › Optimization for machine learning
minimax optimization |
0.9 | 1 | 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive Regularization · NeurIPS 2025 |
Machine learning › Learning theory › over-parameterization
double descent |
0.8 | 1 | 2024 | No Double Descent in Principal Component Regression: A High-Dimensional Analysis · ICML 2024 |
Machine learning › Learning theory
generalization bounds |
0.8 | 1 | 2024 | No Double Descent in Principal Component Regression: A High-Dimensional Analysis · ICML 2024 |
Machine learning › Learning theory
high-dimensional regression |
0.8 | 1 | 2024 | No Double Descent in Principal Component Regression: A High-Dimensional Analysis · ICML 2024 |
Machine learning › Kernel, tree and ensemble methods › linear model
principal component regression |
0.8 | 1 | 2024 | No Double Descent in Principal Component Regression: A High-Dimensional Analysis · ICML 2024 |
Machine learning › Learning theory › over-parameterization › interpolation
minimum-norm interpolation |
0.7 | 1 | 2023 | Regularization properties of adversarially-trained linear regression · NeurIPS 2023 |
Machine learning › Learning theory
over-parameterization |
0.7 | 1 | 2023 | Regularization properties of adversarially-trained linear regression · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
regularization |
0.7 | 1 | 2023 | Regularization properties of adversarially-trained linear regression · NeurIPS 2023 |
Computer vision › Image recognition and object detection
image retrieval |
0.3 | 1 | 2025 | Human-Aligned Image Models Improve Visual Decoding from the Brain · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.2 | 1 | 2024 | No Double Descent in Principal Component Regression: A High-Dimensional Analysis · ICML 2024 |
Bioinformatics and computational biology › molecular informatics
cheminformatics |
0.2 | 1 | 2024 | Can Transformers Smell Like Humans? · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.5self-supervised pretraining · 1.5reproducing kernel hilbert space · 0.9representation alignment · 0.9multiple kernel learning · 0.9EEG decoding · 0.9spiked covariance model · 0.8random matrix theory · 0.8lasso · 0.7convex optimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contrasting by Augmented Patient Electrocardiograms to Learn Representations for a Foundation Model
Gul Rukhkhattak, Konstantinos Patlatzoglou, Yixiu Liang, Libor Pastika, Boroumand Zeidaabadi, Joseph Barker, Mehak Gurnani, Antônio H. Ribeiro, Jeffrey Annis, Antônio L. P. Ribeiro, Nicholas S. Peters, Junbo Ge, Daniel B. Kramer, Jonathan W. Waks, Evan Brittain, Arunashis Sau, Fu Siong Ng |
AIME (2) | 8 |
| 2025 | Efficient Optimization Algorithms for Linear Adversarial TrainingabstractAdversarial training can be used to learn models that are robust against perturbations. For linear models, it can be formulated as a convex optimization problem. Compared to methods proposed in the context of deep learning, leveraging the optimization structure allows significantly faster convergence rates. Still, the use of generic convex solvers can be inefficient for large-scale problems. Here, we propose tailored optimization algorithms for the adversarial training of linear models, which render large-scale regression and classification problems more tractable. For regression problems, we propose a family of solvers based on iterative ridge regression and, for classification, a family of solvers based on projected gradient descent. The methods are based on extended variable reformulations of the original problem. We illustrate their efficiency in numerical examples. Antônio H. Ribeiro, Thomas B. Schön, Dave Zachariah, Francis R. Bach |
AISTATS | 1 |
| 2025 | Deep Learning Amplified Early Stopping Bias: Overestimating Performance on Small DatasetsabstractCross-validation is commonly used to estimate machine learning model performance on new samples. However, using it for both hyperparameter selection and error estimation can lead to overestimating model performance, especially with extensive hyperparameter searches that overly tailor models to validation data. We demonstrate that deep learning further amplifies this bias, with even minor model adjustments causing significant overestimation. Our extensive experiments on simulated and real data focus on the bias from early stopping during cross-validation. We find that overestimation intensifies with network depth and is especially severe in small datasets, which are common in physiological signal processing applications. Selecting the early stopping point during cross-validation can result in ROC-AUC estimates exceeding 90% on random data, and this effect persists across various sample sizes, architectures, and network sizes1. Nona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Danica Kragic |
ICASSP | 2 |
| 2025 | Human-Aligned Image Models Improve Visual Decoding from the BrainabstractDecoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain activity to enable visual decoding. In this paper, we introduce the use of human-aligned image encoders to map brain signals to images. We hypothesize that these models more effectively capture perceptual attributes associated with the rapid visual stimuli presentations commonly used in visual brain data recording experiments. Our empirical results support this hypothesis, demonstrating that this simple modification improves image retrieval accuracy by up to 21\% compared to state-of-the-art methods. Comprehensive experiments confirm consistent performance improvements across diverse EEG architectures, image encoders, alignment methods, participants, and brain imaging modalities. Nona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Farzaneh Taleb, Mårten Björkman, Danica Kragic |
ICML | 2 |
| 2025 | Kernel Learning with Adversarial Features: Numerical Efficiency and Adaptive RegularizationabstractAdversarial training has emerged as a key technique to enhance model robustness against adversarial input perturbations. Many of the existing methods rely on computationally expensive min-max problems that limit their application in practice. We propose a novel formulation of adversarial training in reproducing kernel Hilbert spaces, shifting from input to feature-space perturbations. This reformulation enables the exact solution of inner maximization and efficient optimization. It also provides a regularized estimator that naturally adapts to the noise level and the smoothness of the underlying function. We establish conditions under which the feature-perturbed formulation is a relaxation of the original problem and propose an efficient optimization algorithm based on iterative kernel ridge regression. We provide generalization bounds that help to understand the properties of the method. We also extend the formulation to multiple kernel learning. Empirical evaluation shows good performance in both clean and adversarial settings. Antônio H. Ribeiro, David Vävinggren, Dave Zachariah, Thomas B. Schön, Francis R. Bach |
NeurIPS | 1 |
| 2024 | No Double Descent in Principal Component Regression: A High-Dimensional AnalysisabstractUnderstanding the generalization properties of large-scale models necessitates incorporating realistic data assumptions into the analysis. Therefore, we consider Principal Component Regression (PCR)—combining principal component analysis and linear regression—on data from a low-dimensional manifold. We present an analysis of PCR when the data is sampled from a spiked covariance model, obtaining fundamental asymptotic guarantees for the generalization risk of this model. Our analysis is based on random matrix theory and allows us to provide guarantees for high-dimensional data. We additionally present an analysis of the distribution shift between training and test data. The results allow us to disentangle the effects of (1) the number of parameters, (2) the data-generating model and, (3) model misspecification on the generalization risk. The use of PCR effectively regularizes the model and prevents the interpolation peak of the double descent. Our theoretical findings are empirically validated in simulation, demonstrating their practical relevance. Daniel Gedon, Antônio H. Ribeiro, Thomas B. Schön |
ICML | 2 |
| 2024 | Can Transformers Smell Like Humans?abstractThe human brain encodes stimuli from the environment into representations that form a sensory perception of the world. Despite recent advances in understanding visual and auditory perception, olfactory perception remains an under-explored topic in the machine learning community due to the lack of large-scale datasets annotated with labels of human olfactory perception. In this work, we ask the question of whether pre-trained transformer models of chemical structures encode representations that are aligned with human olfactory perception, i.e., can transformers smell like humans? We demonstrate that representations encoded from transformers pre-trained on general chemical structures are highly aligned with human olfactory perception. We use multiple datasets and different types of perceptual representations to show that the representations encoded by transformer models are able to predict: (i) labels associated with odorants provided by experts; (ii) continuous ratings provided by human participants with respect to pre-defined descriptors; and (iii) similarity ratings between odorants provided by human participants. Finally, we evaluate the extent to which this alignment is associated with physicochemical features of odorants known to be relevant for olfactory decoding. Farzaneh Taleb, Miguel Vasco, Antônio H. Ribeiro, Mårten Björkman, Danica Kragic |
NeurIPS | 3 |
| 2023 | Regularization properties of adversarially-trained linear regressionabstractState-of-the-art machine learning models can be vulnerable to very small input perturbations that are adversarially constructed. Adversarial training is an effective approach to defend against it. Formulated as a min-max problem, it searches for the best solution when the training data were corrupted by the worst-case attacks. Linear models are among the simple models where vulnerabilities can be observed and are the focus of our study. In this case, adversarial training leads to a convex optimization problem which can be formulated as the minimization of a finite sum. We provide a comparative analysis between the solution of adversarial training in linear regression and other regularization methods. Our main findings are that: (A) Adversarial training yields the minimum-norm interpolating solution in the overparameterized regime (more parameters than data), as long as the maximum disturbance radius is smaller than a threshold. And, conversely, the minimum-norm interpolator is the solution to adversarial training with a given radius. (B) Adversarial training can be equivalent to parameter shrinking methods (ridge regression and Lasso). This happens in the underparametrized region, for an appropriate choice of adversarial radius and zero-mean symmetrically distributed covariates. (C) For $\ell_\infty$-adversarial training---as in square-root Lasso---the choice of adversarial radius for optimal bounds does not depend on the additive noise variance. We confirm our theoretical findings with numerical examples. Antônio H. Ribeiro, Dave Zachariah, Francis R. Bach, Thomas B. Schön |
NeurIPS | 1 |
| 2023 | Invertible Kernel PCA With Random Fourier FeaturesabstractKernel principal component analysis (kPCA) is a widely studied method to construct a low-dimensional data representation after a nonlinear transformation. The prevailing method to reconstruct the original input signal from kPCA—an important task for denoising—requires us to solve a supervised learning problem. In this paper, we present an alternative method where the reconstruction follows naturally from the compression step. We first approximate the kernel with random Fourier features. Then, we exploit the fact that the nonlinear transformation is invertible in a certain subdomain. Hence, the nameinvertible kernel PCA (ikPCA). We experiment with different data modalities and show that ikPCA performs similarly to kPCA with supervised reconstruction on denoising tasks, making it a strong alternative. Daniel Gedon, Antônio H. Ribeiro, Niklas Wahlstrom, Thomas B. Schön |
IEEE Signal Process. Lett. | 2 |
| 2021 | How Convolutional Neural Networks Deal with AliasingabstractThe convolutional neural network (CNN) remains an essential tool in solving computer vision problems. Standard convolutional architectures consist of stacked layers of operations that progressively downscale the image. Aliasing is a well-known side-effect of downsampling that may take place: it causes high-frequency components of the original signal to become indistinguishable from its low-frequency components. While downsampling takes place in the max-pooling layers or in the strided-convolutions in these models, there is no explicit mechanism that prevents aliasing from taking place in these layers. Due to the impressive performance of these models, it is natural to suspect that they, somehow, implicitly deal with this distortion. The question we aim to answer in this paper is simply: "how and to what extent do CNNs counteract aliasing?" We explore the question by means of two examples: In the first, we assess the CNNs capability of distinguishing oscillations at the input, showing that the redundancies in the intermediate channels play an important role in succeeding at the task; In the second, we show that an image classifier CNN while, in principle, capable of implementing anti-aliasing filters, does not prevent aliasing from taking place in the intermediate layers. Antônio H. Ribeiro, Thomas B. Schön |
ICASSP | 1 |
| 2020 | Beyond exploding and vanishing gradients: analysing RNN training using attractors and smoothnessabstractThe exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle, while powerful, might need some refinement to explain recent developments. We refine the concept of exploding gradients by reformulating the problem in terms of the cost function smoothness, which gives insight into higher-order derivatives and the existence of regions with many close local minima. We also clarify the distinction between vanishing gradients and the need for the RNN to learn attractors to fully use its expressive power. Through the lens of these refinements, we shed new light on recent developments in the RNN field, namely stable RNN and unitary (or orthogonal) RNNs. Antônio H. Ribeiro, Koen Tiels, Luis Antonio Aguirre, Thomas B. Schön |
AISTATS | 1 |
| 2018 | "Parallel Training Considered Harmful?": Comparing series-parallel and parallel feedforward network training
Antônio H. Ribeiro, Luis Antonio Aguirre |
Neurocomputing | 1 |