Marc Sebban

dblp:s/MarcSebban · DBLP profile ↗
← Back
87ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-6851-169XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 10 first-author · 12 since 2021Databases, data management, data science and information retrieval · 29 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Bregman Proximal Viewpoint on Neural Operators
abstract
We present several advances on neural operators by viewing the action of operator layers as the minimizers of Bregman regularized optimization problems over Banach function spaces. The proposed framework allows interpreting the activation operators as Bregman proximity operators from dual to primal space. This novel viewpoint is general enough to recover classical neural operators as well as a new variant, coined Bregman neural operators, which includes the inverse activation operator and features the same expressivity of standard neural operators. Numerical experiments support the added benefits of the Bregman variant of Fourier neural operators for training deeper and more accurate models.
Abdel-Rahim Mezidi, Jordan Patracone, Saverio Salzo, Amaury Habrard, Massimiliano Pontil, Rémi Emonet, Marc Sebban
ICML7
2025 Physics-Informed Machine Learning for Modeling CO2 Capture from Scarce Data
abstract
Accurate modeling of complex industrial processes often relies on costly mechanistic simulations grounded in physical principles. In this paper, we investigate the subject of$\text{CO}_{2}$capture, a major environmental challenge, through the absorption column of an amine-based post-combustion process. The modeling of such unit at industrial scale faces two difficulties: (i) theoretical models, efficient at laboratory scale, might fail to fully reflect the complexity of the numerous intertwined phenomena occurring in the absorber, (ii) the cost and uncertainty of industrial observation data make purely data-driven approaches unfeasible. To tackle both this low data regime and inaccurate physical models, we envision this$\text{CO}_{2}$capture problem through the lens of Physics-informed Machine Learning (PiML). We present a hybrid (data+knowledge) model where the scarce observation data complement the physical model, while the latter ensures that the predictions remain physically consistent. Beyond the standard use of simulation data for learning and the embedding of physical laws as regularization, the originality of our PiML algorithm compared to other methods in the literature lies in a physical prior assumption about the network architecture and its countercurrent flow learning process inspired by the column's operation. Our experimental results showcase a significant improvement in accuracy and highlight the potential of our augmented model for generalizing across domains, especially when data is scarce.
Mickael Gault, Pierre Bachaud, Benoît Celse, Rémi Emonet, Marc Sebban
ICTAI5
2025 Provably Accurate Adaptive Sampling for Collocation Points in Physics-Informed Neural Networks
Antoine Caradot, Rémi Emonet, Amaury Habrard, Abdel-Rahim Mezidi, Marc Sebban
ECML/PKDD (5)5
2024 Physics-Informed Machine Learning for Better Understanding Laser-Matter Interaction
abstract
Physics-informed machine learning typically assumes that the underlying physical laws are known and abundant training data is available. These assumptions do not hold in the context of self-organization of matter, a phenomenon that leads to the emergence of patterns when a surface is irradiated with an ultrafast laser beam. Indeed, due to the constraints of the electronic data acquisition devices, the creation of large datasets is made impossible. Moreover, modeling this dynamic process is challenging as it involves coupling between electromagnetism, thermodynamics and fluid mechanics under far-from-equilibrium conditions that are not yet fully understood. This paper aims at taking a step forward towards a better understanding of this complex phenomenon. We specifically focus on the laser energy absorption of the surface, which is governed by the distinctive characteristics of Maxwell's equations in an inho-mogeneous lossy medium. This involves modelling physics at the nano scale and incurs high simulation costs that make any exploration impractical. To address this major issue, we investigate different physics-informed learning models. In this low data regime, our study reveals that learning a simple U-Net-based surrogate model surpasses (i) more sophisticated neural architectures and (ii) the FDTD-based solver in speed by several orders of magnitude. Interestingly, our study highlights a link between the formation of patterns and the magnitude of absorbed energy.
Fayad Ali Banna, Jean-Philippe Colombier, Rémi Emonet, Marc Sebban
ICTAI4
2024 Unsupervised Learning and Effective Complexity: Introducing JPG and Neural Sophistication
abstract
Measuring the complexity of arbitrary data has been of interest to many scientific domains, including machine learning and particularly unsupervised learning. In this paper, we cover relevant concepts including Kolmogorov complexity, entropy and minimum description length. We argue that these measures alone are failing to distinguish noise from meaningful complexity. We push for the concept sophistication which measures the complexity of the structured part of the data, ignoring unstructured noise. This concept is reified in two manners: using image compression algorithms and using autoencoders.
Erick Gomez Soto, Rémi Emonet, Marc Sebban
ICTAI3
2024 Generative shape deformation with optimal transport using learned transformations
abstract
Shape deformation is a fundamental problem in computer graphics and computer vision, with numerous applications in fields such as animation, medical imaging, robotics to cite a few. We propose a method for shape deformation based on applying learned transformations with optimal transport (OT). Our method combines the power of the latter with the flexibility of learned transformations to provide an efficient and effective solution for 2D and 3D shape deformation. We formulate the problem as an OT task, where the goal is to learn the optimal way to move the mass distribution of a shape to another. We then use the learned geometric transformations, to achieve shape deformation. Our method can be applied to a wide range of shapes and applications. Interestingly, we show that it requires a small amount of data to learn the transformations. We demonstrate the performance of our method on our own crafted dataset of 2D and 3D shapes and evaluate its effectiveness using various metrics. The promising results obtained suggest that our method can be applied in a wide range of real-world applications.
Jorge Azorín López, Marc Sebban, Nahuel E. Garcia-D'Urso, Amaury Habrard, Andrés Fuster Guilló
IJCNN2
2024 Approximation Error of Sobolev Regular Functions with Tanh Neural Networks: Theoretical Impact on PINNs
Benjamin Girault, Rémi Emonet, Amaury Habrard, Jordan Patracone, Marc Sebban
ECML/PKDD (4)5
2023 Is My Neural Net Driven by the MDL Principle?
Eduardo Brandao, Stefan Duffner, Rémi Emonet, Amaury Habrard, François Jacquenet, Marc Sebban
ECML/PKDD (2)6
2022 Optimal Tensor Transport
abstract
Optimal Transport (OT) has become a popular tool in machine learning to align finite datasets typically lying in the same vector space. To expand the range of possible applications, Co-Optimal Transport (Co-OT) jointly estimates two distinct transport plans, one for the rows (points) and one for the columns (features), to match two data matrices that might use different features. On the other hand, Gromov Wasserstein (GW) looks for a single transport plan from two pairwise intra-domain distance matrices. Both Co-OT and GW can be seen as specific extensions of OT to more complex data. In this paper, we propose a unified framework, called Optimal Tensor Transport (OTT), which takes the form of a generic formulation that encompasses OT, GW and Co-OT and can handle tensors of any order by learning possibly multiple transport plans. We derive theoretical results for the resulting new distance and present an efficient way for computing it. We further illustrate the interest of such a formulation in Domain Adaptation and Comparison-based Clustering.
Tanguy Kerdoncuff, Rémi Emonet, Michaël Perrot, Marc Sebban
AAAI4
2022 Fast Multiscale Diffusion On Graphs
abstract
Diffusing a graph signal at multiple scales requires to compute the action of the exponential of as many versions of the Laplacian matrix. Considering the truncated Chebyshev polynomial approximation of the exponential, we derive a tightened bound on the approximation error, allowing thus for a better estimate of the polynomial degree that reaches a prescribed error. We leverage the properties of these approximations to factorize the computation of the action of the diffusion operator over multiple scales, thus drastically reducing its computational cost.
Sibylle Marcotte, Amélie Barbe, Rémi Gribonval, Titouan Vayer, Marc Sebban, Pierre Borgnat, Paulo Gonçalves 0001
ICASSP5
2022 MetaAP: A meta-tree-based ranking algorithm optimizing the average precision from imbalanced data
Rémi Viola, Léo Gautheron, Amaury Habrard, Marc Sebban
Pattern Recognit. Lett.4
2021 Optimization of the Diffusion Time in Graph Diffused-Wasserstein Distances: Application to Domain Adaptation
abstract
The use of the heat kernel on graphs has recently given rise to a family of so-called Diffusion-Wasserstein distances which resort to Optimal Transport theory for comparing attributed graphs. In this paper, we address the open problem of optimizing the diffusion time used in these distances. Inspired from the notion of triplet-based constraints, we design a loss function that aims at bringing two graphs closer together while keeping an impostor away. After a thorough analysis of the properties of this function, we show on synthetic data that the resulting Diffusion-Wasserstein distances outperforms the Gromov and Fused-Gromov Wasserstein distances on unsupervised graph domain adaptation tasks.
Amélie Barbe, Paulo Gonçalves 0001, Marc Sebban, Pierre Borgnat, Rémi Gribonval, Titouan Vayer
ICTAI3
2021 Sampled Gromov Wasserstein
Tanguy Kerdoncuff, Rémi Emonet, Marc Sebban
Mach. Learn.3
2020 A Swiss Army Knife for Minimax Optimal Transport
abstract
The Optimal transport (OT) problem and its associated Wasserstein distance have recently become a topic of great interest in the machine learning community. However, the underlying optimization problem is known to have two major restrictions: (i) it largely depends on the choice of the cost function and (ii) its sample complexity scales exponentially with the dimension. In this paper, we propose a general formulation of a minimax OT problem that can tackle these restrictions by jointly optimizing the cost matrix and the transport plan, allowing us to define a robust distance between distributions. We propose to use a cutting-set method to solve this general problem and show its links and advantages compared to other existing minimax OT approaches. Additionally, we use this method to define a notion of stability allowing us to select the most robust cost matrix. Finally, we provide an experimental study highlighting the efficiency of our approach.
Sofien Dhouib, Ievgen Redko, Tanguy Kerdoncuff, Rémi Emonet, Marc Sebban
ICML5
2020 Metric Learning in Optimal Transport for Domain Adaptation
abstract
Domain Adaptation aims at benefiting from a labeled dataset drawn from a source distribution to learn a model from examples generated from a different but related target distribution. Creating a domain-invariant representation between the two source and target domains is the most widely technique used. A simple and robust way to perform this task consists in (i) representing the two domains by subspaces described by their respective eigenvectors and (ii) seeking a mapping function which aligns them. In this paper, we propose to use Optimal Transport (OT) and its associated Wassertein distance to perform this alignment. While the idea of using OT in domain adaptation is not new, the original contribution of this paper is two-fold: (i) we derive a generalization bound on the target error involving several Wassertein distances. This prompts us to optimize the ground metric of OT to reduce the target risk; (ii) from this theoretical analysis, we design an algorithm (MLOT) which optimizes a Mahalanobis distance leading to a transportation plan that adapts better. Extensive experiments demonstrate the effectiveness of this original approach.
Tanguy Kerdoncuff, Rémi Emonet, Marc Sebban
IJCAI3
2020 Learning from Few Positives: a Provably Accurate Metric Learning Algorithm to Deal with Imbalanced Data
abstract
Learning from imbalanced data, where the positive examples are very scarce, remains a challenging task from both a theoretical and algorithmic perspective. In this paper, we address this problem using a metric learning strategy. Unlike the state-of-the-art methods, our algorithm MLFP, for Metric Learning from Few Positives, learns a new representation that is used only when a test query is compared to a minority training example. From a geometric perspective, it artificially brings positive examples closer to the query without changing the distances to the negative (majority class) data. This strategy allows us to expand the decision boundaries around the positives, yielding a better F-Measure, a criterion which is suited to deal with imbalanced scenarios. Beyond the algorithmic contribution provided by MLFP, our paper presents generalization guarantees on the false positive and false negative rates. Extensive experiments conducted on several imbalanced datasets show the effectiveness of our method.
Rémi Viola, Rémi Emonet, Amaury Habrard, Guillaume Metzler, Marc Sebban
IJCAI5
2020 Graph Diffusion Wasserstein Distances
Amélie Barbe, Marc Sebban, Paulo Gonçalves 0001, Pierre Borgnat, Rémi Gribonval
ECML/PKDD (2)2
2020 Landmark-Based Ensemble Learning with Random Fourier Features and Gradient Boosting
Léo Gautheron, Pascal Germain, Amaury Habrard, Guillaume Metzler, Emilie Morvant, Marc Sebban, Valentina Zantedeschi
ECML/PKDD (3)6
2020 Metric Learning from Imbalanced Data with Generalization Guarantees
Léo Gautheron, Amaury Habrard, Emilie Morvant, Marc Sebban
Pattern Recognit. Lett.4
2019 From Cost-Sensitive to Tight F-measure Bounds
Kevin Bascol, Rémi Emonet, Élisa Fromont, Amaury Habrard, Guillaume Metzler, Marc Sebban
AISTATS6
2019 Metric Learning from Imbalanced Data
abstract
A key element of any machine learning algorithm is the use of a function that measures the dis/similarity between data points. Given a task, such a function can be optimized with a metric learning algorithm. Although this research field has received a lot of attention during the past decade, very few approaches have focused on learning a metric in an imbalanced scenario where the number of positive examples is much smaller than the negatives. Here, we address this challenging task by designing a new Mahalanobis metric learning algorithm (IML) which deals with class imbalance. The empirical study performed shows the efficiency of IML.
Léo Gautheron, Amaury Habrard, Emilie Morvant, Marc Sebban
ICTAI4
2019 An Adjusted Nearest Neighbor Algorithm Maximizing the F-Measure from Imbalanced Data
abstract
In this paper, we address the challenging problem of learning from imbalanced data using a Nearest-Neighbor (NN) algorithm. In this setting, the minority examples typically belong to the class of interest requiring the optimization of specific criteria, like the F-Measure. Based on simple geometrical ideas, we introduce an algorithm that reweights the distance between a query sample and any positive training example. This leads to a modification of the Voronoi regions and thus of the decision boundaries of the NN algorithm. We provide a theoretical justification about the weighting scheme needed to reduce the False Negative rate while controlling the number of False Positives. We perform an extensive experimental study on many public imbalanced datasets, but also on large scale non public data from the French Ministry of Economy and Finance on a tax fraud detection task, showing that our method is very effective and, interestingly, yields the best performance when combined with state of the art sampling methods.
Rémi Viola, Rémi Emonet, Amaury Habrard, Guillaume Metzler, Sébastien Riou, Marc Sebban
ICTAI6
2019 Differentially Private Optimal Transport: Application to Domain Adaptation
abstract
Optimal transport has received much attention during the past few years to deal with domain adaptation tasks. The goal is to transfer knowledge from a source domain to a target domain by finding a transportation of minimal cost moving the source distribution to the target one. In this paper, we address the challenging task of privacy preserving domain adaptation by optimal transport. Using the Johnson-Lindenstrauss transform together with some noise, we present the first differentially private optimal transport model and show how it can be directly applied on both unsupervised and semi-supervised domain adaptation scenarios. Our theoretically grounded method allows the optimization of the transportation plan and the Wasserstein distance between the two distributions while protecting the data of both domains.We perform an extensive series of experiments on various benchmarks (VisDA, Office-Home and Office-Caltech datasets) that demonstrates the efficiency of our method compared to non-private strategies.
Tien-Nam Le, Amaury Habrard, Marc Sebban
IJCAI3
2019 On the analysis of adaptability in multi-source domain adaptation
Ievgen Redko, Amaury Habrard, Marc Sebban
Mach. Learn.3
2019 Deep multi-Wasserstein unsupervised domain adaptation
Tien-Nam Le, Amaury Habrard, Marc Sebban
Pattern Recognit. Lett.3
2018 Online Non-linear Gradient Boosting in Multi-latent Spaces
Jordan Fréry, Amaury Habrard, Marc Sebban, Olivier Caelen, Liyun He-Guelton
IDA3
2018 Tree-Based Cost Sensitive Methods for Fraud Detection in Imbalanced Data
Guillaume Metzler, Xavier Badiche, Brahim Belkasmi, Élisa Fromont, Amaury Habrard, Marc Sebban
IDA6
2018 Fast and Provably Effective Multi-view Classification with Landmark-Based SVM
Valentina Zantedeschi, Rémi Emonet, Marc Sebban
ECML/PKDD (2)3
2018 Learning maximum excluding ellipsoids from imbalanced data with theoretical guarantees
Guillaume Metzler, Xavier Badiche, Brahim Belkasmi, Élisa Fromont, Amaury Habrard, Marc Sebban
Pattern Recognit. Lett.6
2017 Efficient Top Rank Optimization with Gradient Boosting for Supervised Anomaly Detection
Jordan Fréry, Amaury Habrard, Marc Sebban, Olivier Caelen, Liyun He-Guelton
ECML/PKDD (1)3
2017 Theoretical Analysis of Domain Adaptation with Optimal Transport
Ievgen Redko, Amaury Habrard, Marc Sebban
ECML/PKDD (2)3
2016 Metric Learning as Convex Combinations of Local Models with Generalization Guarantees
abstract
Over the past ten years, metric learning allowed the improvement of numerous machine learning approaches that manipulate distances or similarities. In this field, local metric learning has been shown to be very efficient, especially to take into account non linearities in the data and better capture the peculiarities of the application of interest. However, it is well known that local metric learning (i) can entail overfitting and (ii) face difficulties to compare two instances that are assigned to two different local models. In this paper, we address these two issues by introducing a novel metric learning algorithm that linearly combines local models (C2LM). Starting from a partition of the space in regions and a model (a score function) for each region, C2LM defines a metric between points as a weighted combination of the models. A weight vector is learned for each pair of regions, and a spatial regularization ensures that the weight vectors evolve smoothly and that nearby models are favored in the combination. The proposed approach has the particularity of working in a regression setting, of working implicitly at different scales, and of being generic enough so that it is applicable to similarities and distances. We prove theoretical guarantees of the approach using the framework of algorithmic robustness. We carry out experiments with datasets using both distances (perceptual color distances, using Mahalanobis-like distances) and similarities (semantic word similarities, using bilinear forms), showing that C2LM consistently improves regression accuracy even in the case where the amount of training data is small.
Valentina Zantedeschi, Rémi Emonet, Marc Sebban
CVPR3
2016 beta-risk: a New Surrogate Risk for Learning from Weakly Labeled Data
abstract
During the past few years, the machine learning community has paid attention to developping new methods for learning from weakly labeled data. This field covers different settings like semi-supervised learning, learning with label proportions, multi-instance learning, noise-tolerant learning, etc. This paper presents a generic framework to deal with these weakly labeled scenarios. We introduce the beta-risk as a generalized formulation of the standard empirical risk based on surrogate margin-based loss functions. This risk allows us to express the reliability on the labels and to derive different kinds of learning algorithms. We specifically focus on SVMs and propose a soft margin beta-svm algorithm which behaves better that the state of the art.
Valentina Zantedeschi, Rémi Emonet, Marc Sebban
NIPS3
2016 Learning discriminative tree edit similarities for linear classification - Application to melody recognition
Aurélien Bellet, José Francisco Bernabeu, Amaury Habrard, Marc Sebban
Neurocomputing4
2016 A new boosting algorithm for provably accurate unsupervised domain adaptation
Amaury Habrard, Jean-Philippe Peyrache, Marc Sebban
Knowl. Inf. Syst.3
2015 Landmarks-based kernelized subspace alignment for unsupervised domain adaptation
abstract
Domain adaptation (DA) has gained a lot of success in the recent years in computer vision to deal with situations where the learning process has to transfer knowledge from a source to a target domain. In this paper, we introduce a novel unsupervised DA approach based on both subspace alignment and selection of landmarks similarly distributed between the two domains. Those landmarks are selected so as to reduce the discrepancy between the domains and then are used to non linearly project the data in the same space where an efficient subspace alignment (in closed-form) is performed. We carry out a large experimental comparison in visual domain adaptation showing that our new method outperforms the most recent unsupervised DA approaches.
Rahaf Aljundi, Rémi Emonet, Damien Muselet, Marc Sebban
CVPR4
2015 Algorithmic Robustness for Semi-Supervised (ε, γ, τ) -Good Metric Learning
Maria-Irina Nicolae, Marc Sebban, Amaury Habrard, Éric Gaussier, Massih-Reza Amini
ICONIP (1)2
2015 Joint Semi-supervised Similarity Learning for Linear Classification
Maria-Irina Nicolae, Éric Gaussier, Amaury Habrard, Marc Sebban
ECML/PKDD (1)4
2014 Modeling Perceptual Color Differences by Local Metric Learning
Michaël Perrot, Amaury Habrard, Damien Muselet, Marc Sebban
ECCV (5)4
2014 Learning a priori constrained weighted majority votes
Aurélien Bellet, Amaury Habrard, Emilie Morvant, Marc Sebban
Mach. Learn.4
2013 Unsupervised Visual Domain Adaptation Using Subspace Alignment
abstract
In this paper, we introduce a new domain adaptation (DA) algorithm where the source and target domains are represented by subspaces described by eigenvectors. In this context, our method seeks a domain adaptation solution by learning a mapping function which aligns the source subspace with the target one. We show that the solution of the corresponding optimization problem can be obtained in a simple closed form, leading to an extremely fast algorithm. We use a theoretical result to tune the unique hyper parameter corresponding to the size of the subspaces. We run our method on various datasets and show that, despite its intrinsic simplicity, it outperforms state of the art DA methods.
Basura Fernando, Amaury Habrard, Marc Sebban, Tinne Tuytelaars
ICCV3
2013 Boosting for Unsupervised Domain Adaptation
Amaury Habrard, Jean-Philippe Peyrache, Marc Sebban
ECML/PKDD (2)3
2012 Discriminative feature fusion for image classification
abstract
Bag-of-words-based image classification approaches mostly rely on low level local shape features. However, it has been shown that combining multiple cues such as color, texture, or shape is a challenging and promising task which can improve the classification accuracy. Most of the state-of-the-art feature fusion methods usually aim to weight the cues without considering their statistical dependence in the application at hand. In this paper, we present a new logistic regression-based fusion method, called LRFF, which takes advantage of the different cues without being tied to any of them. We also design a new marginalized kernel by making use of the output of the regression model. We show that such kernels, surprisingly ignored so far by the computer vision community, are particularly well suited to achieve image classification tasks. We compare our approach with existing methods that combine color and shape on three datasets. The proposed learning-based feature fusion process clearly outperforms the state-of-the art fusion methods for image classification.
Basura Fernando, Élisa Fromont, Damien Muselet, Marc Sebban
CVPR4
2012 Similarity Learning for Provably Accurate Sparse Linear Classification
Aurélien Bellet, Amaury Habrard, Marc Sebban
ICML3
2012 Good edit similarity learning by loss minimization
Aurélien Bellet, Amaury Habrard, Marc Sebban
Mach. Learn.3
2012 Supervised learning of Gaussian mixture models for visual vocabulary generation
Basura Fernando, Élisa Fromont, Damien Muselet, Marc Sebban
Pattern Recognit.4
2011 An Experimental Study on Learning with Good Edit Similarity Functions
abstract
Similarity functions are essential to many learning algorithms. To allow their use in support vector machines (SVM), i.e., for the convergence of the learning algorithm to be guaranteed, they must be valid kernels. In the case of structured data, the similarities based on the popular edit distance often do not satisfy this requirement, which explains why they are typically used with k-nearest neighbor (k-NN). A common approach to use such edit similarities in SVM is to transform them into potentially (but not provably) valid kernels. Recently, a different theory of learning with (e,g,t) -good similarity functions was proposed, allowing the use of non-kernel similarity functions. Moreover, the resulting models are supposedly sparse, as opposed to standard SVM models that can be unnecessarily dense. In this paper, we study the relevance and applicability of this theory in the context of string edit similarities. We show that they are naturally good for a given string classification task and provide experimental evidence that the obtained models not only clearly outperform the k-NN approach, but are also competitive with standard SVM models learned with state-of-the-art edit kernels, while being much sparser.
Aurélien Bellet, Marc Sebban, Amaury Habrard
ICTAI2
2011 Using the H-Divergence to Prune Probabilistic Automata
abstract
A problem usually encountered in probabilistic automata learning is the difficulty to deal with large training samples and/or wide alphabets. This is partially due to the size of the resulting Probabilistic Prefix Tree (PPT) from which state merging-based learning algorithms are generally applied. In this paper, we propose a novel method to prune PPTs by making use of the H-divergence dH, recently introduced in the field of domain adaptation. dHis based on the classification error made by an hypothesis learned from unlabeled examples drawn according to two distributions to compare. Through a thorough comparison with state-of-the-art divergence measures, we provide experimental evidences that demonstrate the efficiency of our method based on this simple and intuitive criterion.
Marc Bernard, Baptiste Jeudy, Jean-Philippe Peyrache, Marc Sebban, Franck Thollard
ICTAI4
2011 Domain Adaptation with Good Edit Similarities: A Sparse Way to Deal with Scaling and Rotation Problems in Image Classification
abstract
In many real-life applications, the available source training information is either too small or not representative enough of the underlying target test problem. In the past few years, a new line of machine learning research has been developed to overcome such awkward situations, called Domain Adaptation (DA), giving rise to many adaptation algorithms and theoretical results in the form of generalization bounds. In this paper, a novel contribution is proposed in the form of a DA algorithm dealing with string-structured data, inspired from the DA support vector machine (SVM) technique introduced in [Bruzzone et al, PAMI 2010]. To ensure the convergence of SVM-based learning, the similarity functions involved in the process must be valid kernels, i.e. positive semi-definite (PSD) and symmetric. However, in the string-based context that we are considering in this paper, this condition is often not satisfied. Indeed, it has been proven that most string similarity functions based on the edit distance are not PSD. To overcome this drawback, we make use in this paper of the new theory of learning with good similarity functions introduced by Balcan et al., which (i) does not require the use of a valid kernel to learn well and (ii) allows us to induce sparser models. We take advantage of this theoretical framework to propose a new DA algorithm using good edit similarity functions. Using a suitable string-representation of handwritten digits, we show that are our new algorithm is very efficient to deal with the scaling and rotation problems usually encountered in image classification.
Amaury Habrard, Jean-Philippe Peyrache, Marc Sebban
ICTAI3
2011 Learning Good Edit Similarities with Generalization Guarantees
Aurélien Bellet, Amaury Habrard, Marc Sebban
ECML/PKDD (1)3
2010 Weighted Symbols-Based Edit Distance for String-Structured Image Classification
Cécile Barat, Christophe Ducottet, Élisa Fromont, Anne-Claire Legrand, Marc Sebban
ECML/PKDD (1)5
2010 Learning state machine-based string edit kernels
Aurélien Bellet, Marc Bernard, Thierry Murgue, Marc Sebban
Pattern Recognit.4
2009 Learning Constrained Edit State Machines
abstract
Learning the parameters of the edit distance has been increasingly studied during the past few years to improve the assessment of similarities between structured data, such as strings, trees or graphs. Often based on the optimization of the likelihood of pairs of data, the learned models usually take the form of probabilistic state machines, such as pair-Hidden Markov Models (pair-HMM), stochastic transducers, or probabilistic deterministic automata. Although the use of such models has lead to significant improvements of edit distance-based classification tasks, a new challenge has appeared on the horizon: How integrating background knowledge during the learning process? This is the subject matter of this paper in the case of (input,output) pairs of strings. We present a generalization of the pair-HMM in the form of a constrained state machine, where a transition between two states is driven by constraints fulfilled on the input string. Experimental results are provided on a task in molecular biology, aiming to detect transcription factor binding sites.
Laurent Boyer 0002, Olivier Gandrillon, Amaury Habrard, Mathilde Pellerin, Marc Sebban
ICTAI5
2009 Discovering Patterns in Flows: A Privacy Preserving Approach with the ACSM Prototype
Stéphanie Jacquemont, François Jacquenet, Marc Sebban
ECML/PKDD (2)3
2009 Boosting Classifiers Built from Different Subsets of Features
abstract
We focus on the adaptation of boosting to representation spaces composed of different subsets of features. Rather than imposing a single weak learner to handle data that could come from different sources (e.g., images and texts and sounds), we suggest the decomposition of the learning task into several dependent sub-problems of boosting, treated by different weak learners, that will optimally collaborate during the weight update stage. To achieve this task, we introduce a new weighting scheme for which we provide theoretical results. Experiments are carried out and show that ourmethod works significantly better than any combination of independent boosting procedures.
Jean-Christophe Janodet, Marc Sebban, Henri-Maxime Suchier
Fundam. Informaticae2
2009 Mining probabilistic automata: a statistical view of sequential pattern mining
Stéphanie Jacquemont, François Jacquenet, Marc Sebban
Mach. Learn.3
2009 A lower bound on the sample size needed to perform a significant frequent pattern mining task
Stéphanie Jacquemont, François Jacquenet, Marc Sebban
Pattern Recognit. Lett.3
2008 SEDiL: Software for Edit Distance Learning
Laurent Boyer 0002, Yann Esposito, Amaury Habrard, José Oncina, Marc Sebban
ECML/PKDD (2)5
2008 Learning probabilistic models of tree edit distance
Marc Bernard, Laurent Boyer 0002, Amaury Habrard, Marc Sebban
Pattern Recognit.4
2007 Learning Metrics Between Tree Structured Data: Application to Image Recognition
Laurent Boyer 0002, Amaury Habrard, Marc Sebban
ECML3
2007 Correct your text with Google
abstract
With the increasing amount of text files that are produced nowadays, spell checkers have become essential tools for everyday tasks of millions of end users. Among the years, several tools have been designed that show decent performances. Of course, grammatical checkers may improve corrections of texts, nevertheless, this requires large resources. We think that basic spell checking may be improved (a step towards) using the Web as a corpus and taking into account the context of words that are identified as potential misspellings. We propose to use the Google search engine and some machine learning techniques, in order to design a flexible and dynamic spell checker that may evolve among the time with new linguistic features.
Stéphanie Jacquemont, François Jacquenet, Marc Sebban
Web Intelligence3
2006 Learning Stochastic Tree Edit Distance
Marc Bernard, Amaury Habrard, Marc Sebban
ECML3
2006 Sequence Mining Without Sequences: A New Way for Privacy Preserving
abstract
During the last decade, sequential pattern mining has been the core of numerous researches. It is now possible to efficiently discover users' behavior in various domains such as purchases in supermarkets, Web site visits, etc. Nevertheless, classical algorithms do not respect individual's privacy, exploiting personal information (name, IP address, etc.). We provide an original solution to privacy preserving by using a probabilistic automaton instead of the original data. An application in car flow modeling is presented, showing the ability of our algorithm to discover frequent routes without any individual information. A comparison with SPAM is done showing that even if we sample from the automaton, our approach is more efficient
Stéphanie Jacquemont, François Jacquenet, Marc Sebban
ICTAI3
2006 Learning stochastic edit distance: Application in handwritten character recognition
José Oncina, Marc Sebban
Pattern Recognit.2
2005 Detecting Irrelevant Subtrees to Improve Probabilistic Learning from Tree-structured Data
Amaury Habrard, Marc Bernard, Marc Sebban
Fundam. Informaticae3
2004 Boosting grammatical inference with confidence oracles
abstract
In this paper we focus on the adaptation of boosting to grammatical inference. We aim at improving the performance of state merging algorithms in the presence of noisy data by using, in the update rule, additional information provided by an oracle. This strategy requires the construction of a new weighting scheme that takes into account the confidence in the labels of the examples. We prove that our new framework preserves the theoretical properties of boosting. Using the state merging algorithm RPNI*, we describe an experimental study on various datasets, showing a dramatic improvement of performances.
Jean-Christophe Janodet, Richard Nock, Marc Sebban, Henri-Maxime Suchier
ICML3
2004 Mining Decision Rules from Deterministic Finite Automata
abstract
This work presents a novel approach for knowledge discovery from sequential data. Instead of mining the examples in their sequential form, we suppose they have been processed by a machine learning algorithm that has generalized them into a deterministic finite automaton (DFA). Thus, we present a theoretical framework to extract decision rules from this DFA. Our method relies on statistical inference theory and contrary to usual support-based frequent pattern mining techniques. It does not depend on such a global threshold, but rather allows us to determine an adaptive relevance threshold. Various experiments show the advantage of mining DFA instead of mining sequences.
François Jacquenet, Marc Sebban, Georges Valétudie
ICTAI2
2003 Improvement of the State Merging Rule on Noisy Data in Probabilistic Grammatical Inference
Amaury Habrard, Marc Bernard, Marc Sebban
ECML3
2003 On Boosting Improvement: Error Reduction and Convergence Speed-Up
Marc Sebban, Henri-Maxime Suchier
ECML1
2003 On State Merging in Grammatical Inference: A Statistical Approach for Dealing with Noisy Data
Marc Sebban, Jean-Christophe Janodet
ICML1
2003 A Simple Locally Adaptive Nearest Neighbor Rule With Application To Pollution Forecasting
abstract
In this paper, we propose a thorough investigation of a nearest neighbor rule which we call the "Symmetric Nearest Neighbor (sNN) rule". Basically, it symmetrises the classical nearest neighbor relationship from which are computed the points voting for some instances. Experiments on 29 datasets, most of which are readily available, show that the method significantly outperforms the traditional Nearest Neighbors methods. Experiments on a domain of interest related to tropical pollution normalization also show the greater potential of this method. We finally discuss the reasons for the rule's efficiency, provide methods for speeding-up the classification time, and derive from the sNN rule a reliable and fast algorithm to fix the parameter k in the k-NN rule, a longstanding problem in this field.
Richard Nock, Marc Sebban, Didier Bernard
Int. J. Pattern Recognit. Artif. Intell.2
2002 Boosting Density Function Estimators
Franck Thollard, Marc Sebban, Philippe Ézéquel
ECML2
2002 A data-mining approach to spacer oligonucleotide typing of Mycobacterium tuberculosis
abstract
MOTIVATION: The Direct Repeat (DR) locus of Mycobacterium tuberculosis is a suitable model to study (i) molecular epidemiology and (ii) the evolutionary genetics of tuberculosis. This is achieved by a DNA analysis technique (genotyping), called sp acer oligo nucleotide typing (spoligotyping ). In this paper, we investigated data analysis methods to discover intelligible knowledge rules from spoligotyping, that has not yet been applied on such representation. This processing was achieved by applying the C4.5 induction algorithm and knowledge rules were produced. Finally, a Prototype Selection (PS) procedure was applied to eliminate noisy data. This both simplified decision rules, as well as the number of spacers to be tested to solve classification tasks. In the second part of this paper, the contribution of 25 new additional spacers and the knowledge rules inferred were studied from a machine learning point of view. From a statistical point of view, the correlations between spacers were analyzed and suggested that both negative and positive ones may be related to potential structural constraints within the DR locus that may shape its evolution directly or indirectly. RESULTS: By generating knowledge rules induced from decision trees, it was shown that not only the expert knowledge may be modeled but also improved and simplified to solve automatic classification tasks on unknown patterns. A practical consequence of this study may be a simplification of the spoligotyping technique, resulting in a reduction of the experimental constraints and an increase in the number of samples processed.
Marc Sebban, Igor Mokrousov, Nalin Rastogi, Christophe Sola
Bioinform.1
2002 Stopping Criterion for Boosting-Based Data Reduction Techniques: from Binary to Multiclass Problem
Marc Sebban, Richard Nock, Stéphane Lallich
J. Mach. Learn. Res.1
2002 A hybrid filter/wrapper approach of feature selection using information theory
Marc Sebban, Richard Nock
Pattern Recognit.1
2001 Boosting Neighborhood-Based Classifiers
Marc Sebban, Richard Nock, Stéphane Lallich
ICML1
2001 An improved bound on the finite-sample risk of the nearest neighbor rule
Richard Nock, Marc Sebban
Pattern Recognit. Lett.2
2001 A Bayesian boosting theorem
Richard Nock, Marc Sebban
Pattern Recognit. Lett.2
2000 Sharper Bounds for the Hardness of Prototype and Feature Selection
Richard Nock, Marc Sebban
ALT2
2000 Instance Pruning as an Information Preserving Problem
Marc Sebban, Richard Nock
ICML1
2000 Contribution of Dataset Reduction Techniques to Tree-Simplification and Knowledge Discovery
Marc Sebban, Richard Nock
PKDD1
2000 Combining Feature and Example Pruning by Uncertainty Minimization
Marc Sebban, Richard Nock
UAI1
1999 From Theoretical Learnability to Statistical Measures of the Learnable
Marc Sebban, Gilles Richard
IDA1
1999 Experiments on a Representation-Independent "Top-Down and Prune" Induction Scheme
Richard Nock, Marc Sebban, Pascal Jappy
PKDD2
1999 Contribution of Boosting in Wrapper Models
Marc Sebban, Richard Nock
PKDD1
1999 Selection and Statistical Validation of Features and Prototypes
Marc Sebban, Djamel A. Zighed, S. Di Palma
PKDD1
1996 A Comparison of Some Contextual Discretization Methods
Sabine Loudcher, Ricco Rakotomalala, Marc Sebban
Inf. Sci.3