VLDB 2026 Research / reviewers in the wild / expert
Ievgen Redko
dblp:150/3980
· DBLP profile ↗
29ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0002-3860-5502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal TransportabstractIn the last decade, we have witnessed the introduction of several novel deep neural network (DNN) architectures exhibiting ever-increasing performance across diverse tasks. Explaining the upward trend of their performance, however, remains difficult as different DNN architectures of comparable depth and width – common factors associated with their expressive power – may exhibit a drastically different performance even when trained on the same dataset. In this paper, we introduce the concept of the non-linearity signature of DNN, the first theoretically sound solution for approximately measuring the non-linearity of deep neural networks. Built upon a score derived from closed-form optimal transport mappings, this signature provides a better understanding of the inner workings of a wide range of DNN architectures and learning paradigms, with a particular emphasis on the computer vision task. We provide extensive experimental results that highlight the practical usefulness of the proposed non-linearity signature and its potential for long-reaching implications. The code for our work is available at https://github.com/qbouniot/AffScoreDeep. Quentin Bouniot, Ievgen Redko, Anton Mallasto, Charlotte Laclau, Oliver Struckmeier, Karol Arndt, Markus Heinonen, Ville Kyrki, Samuel Kaski |
CVPR | 2 |
| 2025 | Zero-shot Model-based Reinforcement Learning using Large Language ModelsabstractThe emerging zero-shot capabilities of Large Language Models (LLMs) have led to their applications in areas extending well beyond natural language processing tasks.
In reinforcement learning, while LLMs have been extensively used in text-based environments, their integration with continuous state spaces remains understudied.
In this paper, we investigate how pre-trained LLMs can be leveraged to predict in context the dynamics of continuous Markov decision processes.
We identify handling multivariate data and incorporating the control signal as key challenges that limit the potential of LLMs' deployment in this setup and propose Disentangled In-Context Learning (DICL) to address them.
We present proof-of-concept applications in two reinforcement learning settings: model-based policy evaluation and data-augmented off-policy reinforcement learning, supported by theoretical analysis of the proposed methods.
Our experiments further demonstrate that our approach produces well-calibrated uncertainty estimates. We release the code at https://github.com/abenechehab/dicl. Abdelhakim Benechehab, Youssef Attia El Hili, Ambroise Odonnat, Oussama Zekri, Albert Thomas 0001, Giuseppe Paolo, Maurizio Filippone, Ievgen Redko, Balázs Kégl |
ICLR | 8 |
| 2024 | Breaking isometric ties and introducing priors in Gromov-Wasserstein distancesabstractGromov-Wasserstein distance has many applications in machine learning due to its ability to compare measures across metric spaces and its invariance to isometric transformations. However, in certain applications, this invariant property can be too flexible, thus undesirable. Moreover, the Gromov-Wasserstein distance solely considers pairwise sample similarities in input datasets, disregarding the raw feature representations. We propose a new optimal transport formulation, called Augmented Gromov-Wasserstein (AGW), that allows for some control over the level of rigidity to transformations. It also incorporates feature alignments, enabling us to better leverage prior knowledge on the input data for improved performance. We first present theoretical insights into the proposed method. We then demonstrate its usefulness for single-cell multi-omic alignment tasks and heterogeneous domain adaptation in machine learning. Pinar Demetci, Quang Huy Tran, Ievgen Redko, Ritambhara Singh |
AISTATS | 3 |
| 2024 | Leveraging Ensemble Diversity for Robust Self-Training in the Presence of Sample Selection BiasabstractSelf-training is a well-known approach for semi-supervised learning. It consists of iteratively assigning pseudo-labels to unlabeled data for which the model is confident and treating them as labeled examples. For neural networks, \texttt{softmax} prediction probabilities are often used as a confidence measure, although they are known to be overconfident, even for wrong predictions. This phenomenon is particularly intensified in the presence of sample selection bias, i.e., when data labeling is subject to some constraints. To address this issue, we propose a novel confidence measure, called $\mathcal{T}$-similarity, built upon the prediction diversity of an ensemble of linear classifiers. We provide the theoretical analysis of our approach by studying stationary points and describing the relationship between the diversity of the individual members and their performance. We empirically demonstrate the benefit of our confidence measure for three different pseudo-labeling policies on classification datasets of various data modalities. The code is available at https://github.com/ambroiseodt/tsim. Ambroise Odonnat, Vasilii Feofanov, Ievgen Redko |
AISTATS | 3 |
| 2024 | SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise AttentionabstractTransformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting. To better understand this phenomenon, we start by studying a toy linear forecasting problem for which we show that transformers are incapable of converging to their true solution despite their high expressive power. We further identify the attention of transformers as being responsible for this low generalization capacity. Building upon this insight, we propose a shallow lightweight transformer model that successfully escapes bad local minima when optimized with sharpness-aware optimization. We empirically demonstrate that this result extends to all commonly used real-world multivariate time series datasets. In particular, SAMformer surpasses current state-of-the-art methods and is on par with the biggest foundation model MOIRAI while having significantly fewer parameters. The code is available at https://github.com/romilbert/samformer. Romain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux, Giuseppe Paolo, Themis Palpanas, Ievgen Redko |
ICML | 7 |
| 2024 | Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series ForecastingabstractIn this paper, we introduce a novel theoretical framework for multi-task regression, applying random matrix theory to provide precise performance estimations, under high-dimensional, non-Gaussian data distributions. We formulate a multi-task optimization problem as a regularization technique to enable single-task models to leverage multi-task learning information. We derive a closed-form solution for multi-task optimization in the context of linear models. Our analysis provides valuable insights by linking the multi-task learning performance to various model statistics such as raw data covariances, signal-generating hyperplanes, noise levels, as well as the size and number of datasets. We finally propose a consistent estimation of training and testing errors, thereby offering a robust foundation for hyperparameter optimization in multi-task regression scenarios. Experimental validations on both synthetic and real-world datasets in regression and multivariate time series forecasting demonstrate improvements on univariate models, incorporating our method into the training loss and thus leveraging multivariate information. Romain Ilbert, Malik Tiomoko, Cosme Louart, Ambroise Odonnat, Vasilii Feofanov, Themis Palpanas, Ievgen Redko |
NeurIPS | 7 |
| 2023 | Unbalanced CO-optimal TransportabstractOptimal transport (OT) compares probability distributions by computing a meaningful alignment between their samples. CO-optimal transport (COOT) takes this comparison further by inferring an alignment between features as well. While this approach leads to better alignments and generalizes both OT and Gromov-Wasserstein distances, we provide a theoretical result showing that it is sensitive to outliers that are omnipresent in real-world data. This prompts us to propose unbalanced COOT for which we provably show its robustness to noise in the compared datasets. To the best of our knowledge, this is the first such result for OT methods in incomparable spaces. With this result in hand, we provide empirical evidence of this robustness for the challenging tasks of heterogeneous domain adaptation with and without varying proportions of classes and simultaneous alignment of samples and features across two single-cell measurements. Quang Huy Tran, Hicham Janati, Nicolas Courty, Rémi Flamary, Ievgen Redko, Pinar Demetci, Ritambhara Singh |
AAAI | 5 |
| 2023 | Meta Optimal TransportabstractWe study the use of amortized optimization to predict optimal transport (OT) maps from the input measures, which we call Meta OT. This helps repeatedly solve similar OT problems between different measures by leveraging the knowledge and information present from past problems to rapidly predict and solve new problems. Otherwise, standard methods ignore the knowledge of the past solutions and suboptimally re-solve each problem from scratch. We instantiate Meta OT models in discrete and continuous settings between grayscale images, spherical data, classification labels, and color palettes and use them to improve the computational time of standard OT solvers. Our source code is available at http://github.com/facebookresearch/meta-ot Brandon Amos, Giulia Luise, Samuel Cohen, Ievgen Redko |
ICML | 4 |
| 2022 | Improving Few-Shot Learning Through Multi-task Representation Learning Theory
Quentin Bouniot, Ievgen Redko, Romaric Audigier, Angélique Loesch, Amaury Habrard |
ECCV (20) | 2 |
| 2021 | All of the Fairness for Edge Prediction with Optimal TransportabstractMachine learning and data mining algorithms have been increasingly used recently to support decision-making systems in many areas of high societal importance such as healthcare, education, or security. While being very efficient in their predictive abilities, the deployed algorithms sometimes tend to learn an inductive model with a discriminative bias due to the presence of this latter in the learning sample. This problem gave rise to a new field of algorithmic fairness where the goal is to correct the discriminative bias introduced by a certain attribute in order to decorrelate it from the model’s output. In this paper, we study the problem of fairness for the task of edge prediction in graphs, a largely underinvestigated scenario compared to a more popular setting of fair classification. To this end, we formulate the problem of fair edge prediction, analyze it theoretically, and propose an embedding-agnostic repairing procedure for the adjacency matrix of an arbitrary graph with a trade-off between the group and individual fairness. We experimentally show the versatility of our approach and its capacity to provide explicit control over different notions of fairness and prediction accuracy. Charlotte Laclau, Ievgen Redko, Manvi Choudhary, Christine Largeron |
AISTATS | 2 |
| 2021 | Deep Neural Networks Are Congestion Games: From Loss Landscape to Wardrop Equilibrium and BeyondabstractThe theoretical analysis of deep neural networks (DNN) is arguably among the most challenging research directions in machine learning (ML) right now, as it requires from scientists to lay novel statistical learning foundations to explain their behaviour in practice. While some success has been achieved recently in this endeavour, the question on whether DNNs can be analyzed using the tools from other scientific fields outside the ML community has not received the attention it may well have deserved. In this paper, we explore the interplay between DNNs and game theory (GT), and show how one can benefit from the classic readily available results from the latter when analyzing the former. In particular, we consider the widely studied class of congestion games, and illustrate their intrinsic relatedness to both linear and non-linear DNNs and to the properties of their loss surface. Beyond retrieving the state-of-the-art results from the literature, we argue that our work provides a very promising novel tool for analyzing the DNNs and support this claim by proposing concrete open problems that can advance significantly our understanding of DNNs when solved. Nina Vesseron, Ievgen Redko, Charlotte Laclau |
AISTATS | 2 |
| 2021 | POT: Python Optimal TransportabstractOptimal transport has recently been reintroduced to the machine learning community thanks in part to novel efficient optimization procedures allowing for medium to large scale applications. We propose a Python toolbox that implements several key optimal transport ideas for the machine learning community. The toolbox contains implementations of a number of founding works of OT for machine learning such as Sinkhorn algorithm and Wasserstein barycenters, but also provides generic solvers that can be used for conducting novel fundamental research. This toolbox, named POT for Python Optimal Transport, is open source with an MIT license. Rémi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z. Alaya, Aurelie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, Léo Gautheron, Nathalie T. H. Gayraud, Hicham Janati, Alain Rakotomamonjy, Ievgen Redko, Antoine Rolet, Antony Schutz, Vivien Seguy, Danica J. Sutherland, Romain Tavenard, Alexander Tong 0001, Titouan Vayer |
J. Mach. Learn. Res. | 15 |
| 2020 | A Swiss Army Knife for Minimax Optimal TransportabstractThe Optimal transport (OT) problem and its associated Wasserstein distance have recently become a topic of great interest in the machine learning community. However, the underlying optimization problem is known to have two major restrictions: (i) it largely depends on the choice of the cost function and (ii) its sample complexity scales exponentially with the dimension. In this paper, we propose a general formulation of a minimax OT problem that can tackle these restrictions by jointly optimizing the cost matrix and the transport plan, allowing us to define a robust distance between distributions. We propose to use a cutting-set method to solve this general problem and show its links and advantages compared to other existing minimax OT approaches. Additionally, we use this method to define a notion of stability allowing us to select the most robust cost matrix. Finally, we provide an experimental study highlighting the efficiency of our approach. Sofien Dhouib, Ievgen Redko, Tanguy Kerdoncuff, Rémi Emonet, Marc Sebban |
ICML | 2 |
| 2020 | Margin-aware Adversarial Domain Adaptation with Optimal TransportabstractIn this paper, we propose a new theoretical analysis of unsupervised domain adaptation that relates notions of large margin separation, adversarial learning and optimal transport. This analysis generalizes previous work on the subject by providing a bound on the target margin violation rate, thus reflecting a better control of the quality of separation between classes in the target domain than bounding the misclassification rate. The bound also highlights the benefit of a large margin separation on the source domain for adaptation and introduces an optimal transport (OT) based distance between domains that has the virtue of being task-dependent, contrary to other approaches. From the obtained theoretical results, we derive a novel algorithmic solution for domain adaptation that introduces a novel shallow OT-based adversarial approach and outperforms other OT-based DA baselines on several simulated and real-world classification tasks. Sofien Dhouib, Ievgen Redko, Carole Lartizien |
ICML | 2 |
| 2020 | CO-Optimal TransportabstractOptimal transport (OT) is a powerful geometric and probabilistic tool for finding correspondences and measuring similarity between two distributions. Yet, its original formulation relies on the existence of a cost function between the samples of the two distributions, which makes it impractical when they are supported on different spaces. To circumvent this limitation, we propose a novel OT problem, named COOT for CO-Optimal Transport, that simultaneously optimizes two transport maps between both samples and features, contrary to other approaches that either discard the individual features by focusing on pairwise distances between samples or need to model explicitly the relations between them. We provide a thorough theoretical analysis of our problem, establish its rich connections with other OT-based distances and demonstrate its versatility with two machine learning applications in heterogeneous domain adaptation and co-clustering/data summarization, where COOT leads to performance improvements over the state-of-the-art methods. Titouan Vayer, Ievgen Redko, Rémi Flamary, Nicolas Courty |
NeurIPS | 2 |
| 2019 | On Fair Cost Sharing Games in Machine Learning
Ievgen Redko, Charlotte Laclau |
AAAI | 1 |
| 2019 | Optimal Transport for Multi-source Domain Adaptation under Target ShiftabstractIn this paper, we tackle the problem of reducing discrepancies between multiple domains, i.e. multi-source domain adaptation, and consider it under the target shift assumption: in all domains we aim to solve a classification problem with the same output classes, but with different labels proportions. This problem, generally ignored in the vast majority of domain adaptation papers, is nevertheless critical in real-world applications, and we theoretically show its impact on the success of the adaptation. Our proposed method is based on optimal transport, a theory that has been successfully used to tackle adaptation problems in machine learning. The introduced approach, Joint Class Proportion and Optimal Transport (JCPOT), performs multi-source adaptation and target shift correction simultaneously by learning the class probabilities of the unlabeled target sample and the coupling allowing to align two (or more) probability distributions. Experiments on both synthetic and real-world data (satellite image pixel classification) task show the superiority of the proposed method over the state-of-the-art. Ievgen Redko, Nicolas Courty, Rémi Flamary, Devis Tuia |
AISTATS | 1 |
| 2019 | On the analysis of adaptability in multi-source domain adaptation
Ievgen Redko, Amaury Habrard, Marc Sebban |
Mach. Learn. | 1 |
| 2018 | Cross-Lingual Document Retrieval Using Regularized Wasserstein Distance
Georgios Balikas, Charlotte Laclau, Ievgen Redko, Massih-Reza Amini |
ECIR | 3 |
| 2018 | Revisiting (\epsilon, \gamma, \tau)-similarity learning for domain adaptationabstractSimilarity learning is an active research area in machine learning that tackles the problem of finding a similarity function tailored to an observable data sample in order to achieve efficient classification. This learning scenario has been generally formalized by the means of a $(\epsilon, \gamma, \tau)-$good similarity learning framework in the context of supervised classification and has been shown to have strong theoretical guarantees. In this paper, we propose to extend the theoretical analysis of similarity learning to the domain adaptation setting, a particular situation occurring when the similarity is learned and then deployed on samples following different probability distributions. We give a new definition of an $(\epsilon, \gamma)-$good similarity for domain adaptation and prove several results quantifying the performance of a similarity function on a target domain after it has been trained on a source domain. We particularly show that if the source distribution dominates the target one, then principally new domain adaptation learning bounds can be proved. Sofiane Dhouib, Ievgen Redko |
NeurIPS | 2 |
| 2018 | Feature Selection for Unsupervised Domain Adaptation Using Optimal Transport
Léo Gautheron, Ievgen Redko, Carole Lartizien |
ECML/PKDD (2) | 2 |
| 2017 | Co-clustering through Optimal TransportabstractIn this paper, we present a novel method for co-clustering, an unsupervised learning approach that aims at discovering homogeneous groups of data instances and features by grouping them simultaneously. The proposed method uses the entropy regularized optimal transport between empirical measures defined on data instances and features in order to obtain an estimated joint probability density function represented by the optimal coupling matrix. This matrix is further factorized to obtain the induced row and columns partitions using multiscale representations approach. To justify our method theoretically, we show how the solution of the regularized optimal transport can be seen from the variational inference perspective thus motivating its use for co-clustering. The algorithm derived for the proposed method and its kernelized version based on the notion of Gromov-Wasserstein distance are fast, accurate and can determine automatically the number of both row and column clusters. These features are vividly demonstrated through extensive experimental evaluations. Charlotte Laclau, Ievgen Redko, Basarab Matei, Younès Bennani, Vincent Brault |
ICML | 2 |
| 2017 | Theoretical Analysis of Domain Adaptation with Optimal Transport
Ievgen Redko, Amaury Habrard, Marc Sebban |
ECML/PKDD (2) | 1 |
| 2016 | Kernel alignment for unsupervised transfer learningabstractThe ability of a human being to extrapolate previously gained knowledge to other domains inspired a new family of methods in machine learning called transfer learning. Transfer learning is often based on the assumption that objects in both target and source domains share some common feature and/or data space. In this paper, we propose a simple and intuitive approach that minimizes iteratively the distance between source and target task distributions by optimizing the kernel target alignment (KTA). We show that this procedure is suitable for transfer learning by relating it to Hilbert-Schmidt Independence Criterion (HSIC) and Quadratic Mutual Information (QMI) maximization. We run our method on benchmark computer vision data sets and show that it can outperform some state-of-art methods. Ievgen Redko, Younès Bennani |
ICPR | 1 |
| 2016 | Non-negative embedding for fully unsupervised domain adaptation
Ievgen Redko, Younès Bennani |
Pattern Recognit. Lett. | 1 |
| 2015 | Sparsity analysis of learned factors in Multilayer NMFabstractThe concept of nonnegative matrix factorization is a recent machine learning technique that is used to decompose large data matrices imposing the non-negativity constraints on the factors. This technique is now used in many data mining applications and thus remains a topic of ongoing interest. In this paper we are particularly interested in the Multilayer NMF - a model that can be seen as a pretraining step of Deep NMF model for learning hidden representations. We analyze the factors obtained using Multilayer NMF and show that the process of building layers can be seen as a repeated application of the Hoyer's projection operator applied sequentially to the factor of the second layer. We also provide the sparsity analysis for matrices obtained during the optimization procedure at each layer. We conclude that the overall sparsity decreases with the increasing number of layers despite the general assumption that Multilayer NMF is efficient due to the fact that it increases the sparsity of learned factors. Ievgen Redko, Younès Bennani |
IJCNN | 1 |
| 2014 | Non-negative Matrix Factorization with Schatten p-norms Reguralization
Ievgen Redko, Younès Bennani |
ICONIP (2) | 1 |
| 2014 | Controlling orthogonality constraints for better NMF clusteringabstractIn this paper we study a variation of a Non-negative Matrix Factorization (NMF) called the Orthogonal NMF(ONMF). This special type of NMF was proposed in order to increase the quality of clustering results of standard NMF by imposing orthogonality on clustering indicator matrix and/or the matrix of basis vectors. We develop an extension of ONMF which we call Weighted ONMF and propose a novel approach for imposing orthogonality on the matrix of basis vectors obtained via NMF using Gram-Schmidt process. Ievgen Redko, Younès Bennani |
IJCNN | 1 |
| 2014 | Random subspaces NMF for unsupervised transfer learningabstractIn this paper we propose a new unsupervised transfer learning approach which aims at finding a partition of unlabeled data in target domain using the knowledge obtained from clustering a source domain unlabeled data. The key idea behind our method is that finding partitions in different feature's subspaces of a source task can help to obtain a more accurate partition in a target one. From the set of source partitions we select only k nearest neighbors using some measure of similarity. Finally, multi-layer non-negative matrix factorization is performed to obtain a partition of objects in target domain. Experimental results show high potential and effectiveness of the proposed technique. Ievgen Redko, Younès Bennani |
IJCNN | 1 |