Antoine Cornuéjols

dblp:c/AntoineCornuejols · DBLP profile ↗
← Back
42ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-2979-3521ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Estimating the Learning Capacity of Bacterial Metabolic Networks
Bastien Mollet, Paul Ahavi, Antoine Cornuéjols, Jean-Loup Faulon, Evelyne Lutton, Alberto Paolo Tonda
IDA3
2024 Synergies between machine learning and reasoning - An introduction by the Kay R. Amel group
abstract
This paper proposes a tentative and original survey of meeting points between Knowledge Representation and Reasoning (KRR) and Machine Learning (ML), two areas which have been developed quite separately in the last four decades. First, some common concerns are identified and discussed such as the types of representation used, the roles of knowledge and data, the lack or the excess of information, or the need for explanations and causal understanding. Then, the survey is organised in seven sections covering most of the territory where KRR and ML meet. We start with a section dealing with prototypical approaches from the literature on learning and reasoning: Inductive Logic Programming, Statistical Relational Learning, and Neurosymbolic AI, where ideas from rule-based reasoning are combined with ML. Then we focus on the use of various forms of background knowledge in learning, ranging from additional regularisation terms in loss functions, to the problem of aligning symbolic and vector space representations, or the use of knowledge graphs for learning. Then, the next section describes how KRR notions may benefit to learning tasks. For instance, constraints can be used as in declarative data mining for influencing the learned patterns; or semantic features are exploited in low-shot learning to compensate for the lack of data; or yet we can take advantage of analogies for learning purposes. Conversely, another section investigates how ML methods may serve KRR goals. For instance, one may learn special kinds of rules such as default rules, fuzzy rules or threshold rules, or special types of information such as constraints, or preferences. The section also covers formal concept analysis and rough sets-based methods. Yet another section reviews various interactions between Automated Reasoning and ML, such as the use of ML methods in SAT solving to make reasoning faster. Then a section deals with works related to model accountability, including explainability and interpretability, fairness and robustness. Finally, a section covers works on handling imperfect or incomplete data, including the problem of learning from uncertain or coarse data, the use of belief functions for regression, a revision-based view of the EM algorithm, the use of possibility theory in statistics, or the learning of imprecise models. This paper thus aims at a better mutual understanding of research in KRR and ML, and how they can cooperate. The paper is completed by an abundant bibliography.
Ismaïl Baaj, Zied Bouraoui, Antoine Cornuéjols, Thierry Denoeux, Sébastien Destercke, Didier Dubois, Marie-Jeanne Lesot, João Marques-Silva 0001, Jérôme Mengin, Henri Prade, Steven Schockaert, Mathieu Serrurier, Olivier Strauss, Christel Vrain
Int. J. Approx. Reason.3
2024 Some thoughts about transfer learning. What role for the source domain?
Antoine Cornuéjols
Int. J. Approx. Reason.1
2024 Reprint of: Some thoughts about transfer learning. What role for the source domain?
Antoine Cornuéjols
Int. J. Approx. Reason.1
2023 Biquality learning: a framework to design algorithms dealing with closed-set distribution shifts
Pierre Nodet, Vincent Lemaire 0001, Alexis Bondu, Antoine Cornuéjols
Mach. Learn.4
2022 When to Classify Events in Open Times Series?
Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
ACML3
2022 Early and Revocable Time Series Classification
abstract
International audience
Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
IJCNN3
2021 Early Classification of Time Series: Cost-based multiclass Algorithms
abstract
Early classification of time series assigns each time series to one of a set of pre-defined classes using as few measurements as possible while preserving a high accuracy. This implies solving online the trade-off between the earliness and the prediction accuracy. This has been formalized in previous work where a cost-based framework taking into account both the cost of misclassification and the cost of delaying the decision has been proposed. The best resulting method, called Economy-$\gamma$, is unfortunately so far limited to binary classification problems. This paper presents a set of six new methods that extend the Economy-$\gamma$method in order to solve multiclass classification problems. Extensive experiments on 33 datasets allowed us to compare the performance of the six proposed approaches to the state-of-the-art one. The results show that: (i) all proposed methods perform significantly better than the state of the art one; (ii) the best way to extend Economy-$\gamma$to multiclass problems is to use a confidence score, either the Gini index or the maximum probability.
Paul-Emile Zafar, Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
DSAA4
2021 Using Agents and Unsupervised Learning for Counting Objects in Images with Spatial Organization
abstract
International audience
Eliott Jacopin, Naomie Berda, Léa Courteille, William Grison, Lucas Mathieu, Antoine Cornuéjols, Christine Martin
ICAART (2)6
2021 Importance Reweighting for Biquality Learning
abstract
The field of Weakly Supervised Learning (WSL) has recently seen a surge of popularity, with numerous papers addressing different types of “supervision deficiencies”, namely: poor quality, non adaptability, and insufficient quantity of labels. Regarding quality, label noise can be of different types, including completely-at-random, at-random or even not-at-random. All these kinds of label noise are addressed separately in the literature, leading to highly specialized approaches. This paper proposes an original, encompassing, view of Weakly Supervised Learning, which results in the design of generic approaches capable of dealing with any kind of label noise. For this purpose, an alternative setting called “Biquality data” is used. It assumes that a small trusted dataset of correctly labeled examples is available, in addition to an untrusted dataset of noisy examples. In this paper, we propose a new reweigthing scheme capable of identifying noncorrupted examples in the untrusted dataset. This allows one to learn classifiers using both datasets. Extensive experiments that simulate several types of label noise and that vary the quality and quantity of untrusted examples, demonstrate that the proposed approach outperforms baselines and state-of-the-art approaches.
Pierre Nodet, Vincent Lemaire 0001, Alexis Bondu, Antoine Cornuéjols, Adam Ouorou
IJCNN4
2021 From Weakly Supervised Learning to Biquality Learning: an Introduction
abstract
The field of Weakly Supervised Learning (WSL) has recently seen a surge of popularity, with numerous papers addressing different types of “supervision deficiencies”. In WSL use cases, a variety of situations exists where the collected “information” is imperfect. The paradigm of WSL attempts to list and cover these problems with associated solutions. In this paper, we review the research progress on WSL with the aim to make it as a brief introduction to this field. We present the three axis of WSL cube and an overview of most of all the elements of their facets. We propose three measurable quantities that acts as coordinates in the previously defined cube namely: Quality, Adaptability and Quantity of information. Thus we suggest that Biquality Learning framework can be defined as a plan of the WSL cube and propose to re-discover previously unrelated patches in WSL literature as a unified Biquality Learning literature.
Pierre Nodet, Vincent Lemaire 0001, Alexis Bondu, Antoine Cornuéjols, Adam Ouorou
IJCNN4
2021 Early classification of time series
Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Asma Dachraoui
Mach. Learn.3
2020 Transfer Learning by Learning Projections from Target to Source
abstract
Using transfer learning to help in solving a new classification task where labeled data is scarce is becoming popular. Numerous experiments with deep neural networks, where the representation learned on a source task is transferred to learn a target neural network, have shown the benefits of the approach. This paper, similarly, deals with hypothesis transfer learning. However, it presents a new approach where, instead of transferring a representation, the source hypothesis is kept and this is a translation from the target domain to the source domain that is learned. In a way, a change of representation is learned. We show how this method performs very well on a classification of time series task where the space of time series is changed between source and target.
Antoine Cornuéjols, Pierre-Alexandre Murena, Raphaël Olivier
IDA1
2020 Solving Analogies on Words based on Minimal Complexity Transformation
abstract
Analogies are 4-ary relations of the form "A is to B as C is to D". When A, B and C are fixed, we call analogical equation the problem of finding the correct D. A direct applicative domain is Natural Language Processing, in which it has been shown successful on word inflections, such as conjugation or declension. If most approaches rely on the axioms of proportional analogy to solve these equations, these axioms are known to have limitations, in particular in the nature of the considered flections. In this paper, we propose an alternative approach, based on the assumption that optimal word inflections are transformations of minimal complexity. We propose a rough estimation of complexity for word analogies and an algorithm to find the optimal transformations. We illustrate our method on a large-scale benchmark dataset and compare with state-of-the-art approaches to demonstrate the interest of using complexity to solve analogies on words.
Pierre-Alexandre Murena, Marie Al-Ghossein, Jean-Louis Dessalles, Antoine Cornuéjols
IJCAI4
2018 Opening the Parallelogram: Considerations on Non-Euclidean Analogies
Pierre-Alexandre Murena, Antoine Cornuéjols, Jean-Louis Dessalles
ICCBR2
2018 An Information Theory based Approach to Multisource Clustering
abstract
Clustering is a compression task which consists in grouping similar objects into clusters. In real-life applications, the system may have access to several views of the same data and each view may be processed by a specific clustering algorithm: this framework is called multi-view clustering and can benefit from algorithms capable of exchanging information between the different views. In this paper, we consider this type of unsupervised ensemble learning as a compression problem and develop a theoretical framework based on algorithmic theory of information suitable for multi-view clustering and collaborative clustering applications. Using this approach, we propose a new algorithm based on solid theoretical basis, and test it on several real and artificial data sets.
Pierre-Alexandre Murena, Jérémie Sublime, Basarab Matei, Antoine Cornuéjols
IJCAI4
2018 Adaptive Window Strategy for Topic Modeling in Document Streams
abstract
Extracting global themes from a written text has recently become a major issue for computational intelligence, in particular in Natural Language Processing communities. Among all proposed solutions, Latent Dirichlet Allocation (LDA) has gained a vast interest and several variants have been proposed to adapt to changing environments. With the emergence of data streams, for instance from social media, the domain faces a new challenge: topic extraction in real time. In this paper, we propose a simple approach called Adaptive Window based Incremental LDA (AWILDA) originating from the cross-over between LDA and state-of-the-art methods in data stream mining. We train new topic models only when a drift is detected and select training data on the fly using ADWIN algorithm. We provide both theoretical guarantees for our method and experimental validation on artificial and real-world data.
Pierre-Alexandre Murena, Marie Al-Ghossein, Talel Abdessalem, Antoine Cornuéjols
IJCNN4
2018 Adaptive collaborative topic modeling for online recommendation
abstract
Collaborative filtering (CF) mainly suffers from rating sparsity and from the cold-start problem. Auxiliary information like texts and images has been leveraged to alleviate these problems, resulting in hybrid recommender systems (RS). Due to the abundance of data continuously generated in real-world applications, it has become essential to design online RS that are able to handle user feedback and the availability of new items in real-time. These systems are also required to adapt to drifts when a change in the data distribution is detected. In this paper, we propose an adaptive collaborative topic modeling approach, CoAWILDA, as a hybrid system relying on adaptive online Latent Dirichlet Allocation (AWILDA) to model newly available items arriving as a document stream and incremental matrix factorization for CF. The topic model is maintained up-to-date in an online fashion and is retrained in batch when a drift is detected using documents automatically selected by an adaptive windowing technique. Our experiments on real-world datasets prove the effectiveness of our approach for online recommendation.
Marie Al-Ghossein, Pierre-Alexandre Murena, Talel Abdessalem, Anthony Barré, Antoine Cornuéjols
RecSys5
2017 Data Collection and Analysis of Usages from Connected Objects: Some Lessons
Sara Meftah, Antoine Cornuéjols, Juliette Dibie, Mariette Sicard
IEA/AIE (2)2
2017 Incremental learning with the minimum description length principle
abstract
Whereas a large number of machine learning methods focus on offline learning over a single batch of data called training data set, the increasing number of automatically generated data leads to the emergence of new issues that offline learning cannot cope with. Incremental learning designates online learning of a model from streaming data. In non-stationary environments, the process generating these data may change over time, hence the learned concept becomes invalid. Adaptation to this non-stationary nature, called concept drift, is an intensively studied topic and can be reached algorithmically by two opposite approaches: active or passive approaches. We propose a formal framework to deal with concept drift, both in active and passive ways. Our framework is derived from the Minimum Description Length principle and exploits the algorithmic theory of information to quantify the model adaptation. We show that this approach is consistent with state of the art techniques and has a valid probabilistic counterpart. We propose two simple algorithms to use our framework in practice and tested both of them on real and simulated data.
Pierre-Alexandre Murena, Antoine Cornuéjols, Jean-Louis Dessalles
IJCNN2
2017 Entropy based probabilistic collaborative clustering
abstract
Unsupervised machine learning approaches involving several clustering algorithms working together to tackle difficult data sets are a recent area of research with a large number of applications such as clustering of distributed data, multi-expert clustering, multi-scale clustering analysis or multi-view clustering. Most of these frameworks can be regrouped under the umbrella of collaborative clustering, the aim of which is to reveal the common underlying structures found by the different algorithms while analyzing the data. Within this context, the purpose of this article is to propose a collaborative framework lifting the limitations of many of the previously proposed methods: Our proposed collaborative learning method makes possible for a wide range of clustering algorithms from different families to work together based solely on their clustering solutions, thus lifting previous limitation requiring identical prototypes between the different collaborators. Our proposed framework uses a variational EM as its theoretical basis for the collaboration process and can be applied to any of the previously mentioned collaborative contexts. In this article, we give the main ideas and theoretical foundations of our method, and we demonstrate its effectiveness in a series of experiments on real data sets as well as data sets from the literature.
Jérémie Sublime, Basarab Matei, Guénaël Cabanes, Nistor Grozavu, Younès Bennani, Antoine Cornuéjols
Pattern Recognit.6
2016 Collaborative-Based Multi-scale Clustering in Very High Resolution Satellite Images
Jérémie Sublime, Antoine Cornuéjols, Younès Bennani
ICONIP (3)2
2016 Denoising 3D Microscopy Images of Cell Nuclei using Shape Priors on an Anisotropic Grid
abstract
This paper presents a new multiscale method to denoise three-dimensional images of cell nuclei. The specificity of this method is its awareness of the noise distribution and object shapes. It combines a multiscale representation called Isotropic Undecimated Wavelet Transform (IUWT) with a nonlinear transform, a statistical test and a variational method, to retrieve spherical shapes in the image. Beyond extending an existing 2D approach to a 3D problem, our algorithm takes the sampling grid dimensions into account. We compare our method to the two algorithms from which it is derived on a representative image analysis task, and show that it is superior to both of them. It brings a slight improvement in the signal-to-noise ratio and a significant improvement in cell detection.
Mathieu Bouyrie, Cristina E. Manfredotti, Nadine Peyriéras, Antoine Cornuéjols
ICPRAM4
2016 Minimum Description Length Principle applied to structure adaptation for classification under concept drift
abstract
Traditional supervised machine learning tests the learned classifiers on data which are drawn from the same distribution as the data used for the learning. In practice, this hypothesis does not always hold and the learned classifier has to be transferred from the space of learning data (also called source data) to the space of test data (also called target data) where it is not directly applicable. To operate this transfer, several methods aim at extracting common structural features in the source and target. Our approach employs a neural model to encode the structure of data: such a model is shown to compress the information in the sense of Kolmogorov theory of information. To perform transfer from source to target, we adapt a result shown for analogy reasoning: the structure of the source and target models are learned by applying the Minimum Description Length Principle which assumes that the chosen transformation has the shortest symbolic description on a universal Turing machine. We encounter a minimization problem over the source and target models. To describe the transfer, we develop a multi-level description of the model transformation which is used directly in the minimization of the description length. Our approach has been tested on toy examples, the difficulty of which can be controlled easily by a one-dimensional parameter and is shown to work efficiently on a wide range of problems.
Pierre-Alexandre Murena, Antoine Cornuéjols
IJCNN2
2015 An initialization scheme for supervized K-means
abstract
Over the last years, researchers have focused their attention on a new approach, supervised clustering, that combines the main characteristics of both traditional clustering and supervised classification tasks. Motivated by the importance of the initialization in the traditional clustering context, this paper explores to what extent supervised initialization step could help traditional clustering to obtain better performances on supervised clustering tasks. This paper reports experiments which show that the simple proposed approach yields a good solution together with significant reduction of the computational cost.
Vincent Lemaire 0001, Oumaima Alaoui Ismaili, Antoine Cornuéjols
IJCNN3
2015 Collaborative clustering with heterogeneous algorithms
abstract
The aim of collaborative clustering is to reveal the common underlying structures found by different algorithms while analyzing data. The fundamental concept of collaboration is that the clustering algorithms operate locally but collaborate by exchanging information about the local structures found by each algorithm. In this framework, the one purpose of this article is to introduce a new method which allows to reinforce the clustering process by exchanging information between several results acquired by different clustering algorithms. The originality of our proposed approach is that the collaboration step can use clustering results obtained from any type of algorithm during the local phase. This article gives the theoretical foundations of our approach as well as some experimental results. The proposed approach has been validated on several data sets and the results have shown to be very competitive.
Jérémie Sublime, Nistor Grozavu, Younès Bennani, Antoine Cornuéjols
IJCNN4
2015 Early Classification of Time Series as a Non Myopic Sequential Decision Making Problem
Asma Dachraoui, Alexis Bondu, Antoine Cornuéjols
ECML/PKDD (1)3
2014 Evaluation Protocol of Early Classifiers over Multiple Data Sets
Asma Dachraoui, Alexis Bondu, Antoine Cornuéjols
ICONIP (2)3
2014 A Supervised Methodology to Measure the Variables Contribution to a Clustering
Oumaima Alaoui Ismaili, Vincent Lemaire 0001, Antoine Cornuéjols
ICONIP (1)3
2014 A New Energy Model for the Hidden Markov Random Fields
Jérémie Sublime, Antoine Cornuéjols, Younès Bennani
ICONIP (2)2
2013 Online Learning: Searching for the Best Forgetting Strategy under Concept Drift
Ghazal Jaber, Antoine Cornuéjols, Philippe Tarroux
ICONIP (2)2
2013 A New On-Line Learning Method for Coping with Recurring Concepts: The ADACC System
Ghazal Jaber, Antoine Cornuéjols, Philippe Tarroux
ICONIP (2)2
2011 Unsupervised Object Ranking Using Not Even Weak Experts
Antoine Cornuéjols, Christine Martin
ICONIP (1)1
2011 Predicting Concept Changes Using a Committee of Experts
Ghazal Jaber, Antoine Cornuéjols, Philippe Tarroux
ICONIP (1)2
2008 Identifying Interface Elements Implied in Protein-Protein Interactions Using Statistical Tests and Frequent Item Sets
abstract
Understanding what are the characteristics of protein-protein interfaces is at the core of numerous applications.This paper introduces a method in which the proteins are described with surfacic geometrical elements. Starting froma database of known interfaces, the method produces the elementsand combinations thereof that are characteristic ofthe interfaces. This is done thanks to a frequent item settechnique and the use of statistical tests to ensure a markeddifference with a null hypothesis. This approach allows oneto easily interpret the results, as compared to techniquesthat operate as “black-boxes”. Furthermore, it is naturallyadapted to discover disjunctive concepts, i.e. differentunderlying processes. The results obtained on a set of459 protein-protein interfaces from the PDB database confirmthat the findings are consistent with current knowledgeabout protein-protein interfaces.
Christine Martin, Antoine Cornuéjols
BIBM2
2008 A note on phase transitions and computational pitfalls of learning from sequences
Antoine Cornuéjols, Michèle Sebag
J. Intell. Inf. Syst.1
2007 A Phase Transition-Based Perspective on Multiple Instance Kernels
Romaric Gaudel, Michèle Sebag, Antoine Cornuéjols
ILP3
2005 Phase Transitions within Grammatical Inference
Nicolas Pernot, Antoine Cornuéjols, Michèle Sebag
IJCAI2
2004 Ensemble Feature Ranking
Kees Jong, Jérémie Mary, Antoine Cornuéjols, Elena Marchiori, Michèle Sebag
PKDD3
1993 Getting Order Independence in Incremental Learning
Antoine Cornuéjols
ECML1
1989 An Exploration Into Incremental Learning: the INFLUENCE System
Antoine Cornuéjols
ML1
1988 Machine learning research at the Laboratoire de Recherche en Informatique at Orsay, France
abstract
This article provides a brief account with sketchy technical details of the major directions in machine learning research done at the Laboratoire de recherche en informatique (LRI) at Orsay University in France. References contain publications giving details on the projects described in this paper and on closely related works. Our research has several objectives: looking for a sound basis of the process of generalization from examples, using this to study conceptual clustering with automatic synthesis of descriptors; studying the nature and goodness of an explanation in the context of apprentice systems; and developing experimental learning systems based on these principles applied to various practical domains. The approach taken by our research group has evolved with time but is still mainly based on learning of concepts from examples using logic representations and techniques. It corresponds to a major goal of our group: to give a clear and rigorous picture, if not a theory, of the topics under investigation. Several aspects are persued at the same time: concept learning by generalization, developments of explanation‐based learning techniques, analogy reasoning, and automatic tuning of the description language. These different directions are related to or stimulated by different domains of tasks: learning of rule bases, games, computer‐aided teaching, learning in noisy environments, and so on. They are described in this article in the light of the main directions. The goal of a complete universal integrated system is still a far cry ahead, but as states a famous Chinese proverb: “The end lies in the way”.
Antoine Cornuéjols
Comput. Intell.1