Léon Bottou

dblp:30/1046 · DBLP profile ↗
← Back
79ranked-venue papers
15as first author
12since 2021 · last 2025
0000-0002-9894-8128ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 13 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 MagicPIG: LSH Sampling for Efficient LLM Generation
abstract
Large language models (LLMs) with long context windows have gained significant attention. However, the KV cache, stored to avoid re-computation, becomes a bottleneck. Various dynamic sparse or TopK-based attention approximation methods have been proposed to leverage the common insight that attention is sparse. In this paper, we first show that TopK attention itself suffers from quality degradation in certain downstream tasks because attention is not always as sparse as expected. Rather than selecting the keys and values with the highest attention scores, sampling with theoretical guarantees can provide a better estimation for attention output. To make the sampling-based approximation practical in LLM generation, we propose MagicPIG, a heterogeneous system based on Locality Sensitive Hashing (LSH). MagicPIG significantly reduces the workload of attention computation while preserving high accuracy for diverse tasks. MagicPIG stores the LSH hash tables and runs the attention computation on the CPU, which allows it to serve longer contexts and larger batch sizes with high approximation accuracy. MagicPIG can improve decoding throughput by up to $5\times$ across various GPU hardware and achieve 54ms decoding latency on a single RTX 4090 for Llama-3.1-8B-Instruct model with a context of 96k tokens.
Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye 0001, Niklas Nolte, Yuandong Tian, Matthijs Douze, Léon Bottou, Beidi Chen
ICLR9
2025 Memory Mosaics
abstract
Memory Mosaics are networks of associative memories working in concert to achieve a prediction task of interest. Like transformers, memory mosaics possess compositional capabilities and in-context learning capabilities. Unlike transformers, memory mosaics achieve these capabilities in comparatively transparent way (“predictive disentanglement”). We illustrate these capabilities on a toy example and also show that memory mosaics perform as well or better than transformers on medium-scale language modeling tasks.
Niklas Nolte, Ranajoy Sadhukhan, Beidi Chen, Léon Bottou
ICLR5
2025 Memory Mosaics at scale
abstract
Memory Mosaics, networks of associative memories, have demonstrated appealing compositional and in-context learning capabilities on medium-scale networks (GPT-2 scale) and synthetic small datasets. This work shows that these favorable properties remain when we scale memory mosaics to large language model sizes (llama-8B scale) and real-world datasets. To this end, we scale memory mosaics to 10B size, we train them on one trillion tokens, we introduce a couple architectural modifications (*memory mosaics v2*), we assess their capabilities across three evaluation dimensions: training-knowledge storage, new-knowledge storage, and in-context learning. Throughout the evaluation, memory mosaics v2 match transformers on the learning of training knowledge (first dimension) and significantly outperforms transformers on carrying out new tasks at inference time (second and third dimensions). These improvements cannot be easily replicated by simply increasing the training data for transformers. A memory mosaics v2 trained on one trillion tokens still perform better on these tasks than a transformer trained on eight trillion tokens.
Léon Bottou
NeurIPS2
2023 Active Self-Supervised Learning: A Few Low-Cost Relationships Are All You Need
abstract
Self-Supervised Learning (SSL) has emerged as the solution of choice to learn transferable representations from unlabeled data. However, SSL requires to build samples that are known to be semantically akin, i.e. positive views. Requiring such knowledge is the main limitation of SSL and is often tackled by ad-hoc strategies e.g. applying known data-augmentations to the same input. In this work, we formalize and generalize this principle through Positive Active Learning (PAL) where an oracle queries semantic relationships between samples. PAL achieves three main objectives. First, it unveils a theoretically grounded learning framework beyond SSL, based on similarity graphs, that can be extended to tackle supervised and semi-supervised learning depending on the employed oracle. Second, it provides a consistent algorithm to embed a priori knowledge, e.g. some observed labels, into any SSL losses without any change in the training pipeline. Third, it provides a proper active learning framework yielding low-cost solutions to annotate datasets, arguably bringing the gap between theory and practice of active learning that is based on simple-to-answer-by-non-experts queries of semantic relationships between inputs.
Vivien Cabannes, Léon Bottou, Yann LeCun, Randall Balestriero
ICCV2
2023 Model Ratatouille: Recycling Diverse Models for Out-of-Distribution Generalization
abstract
Foundation models are redefining how AI systems are built. Practitioners now follow a standard procedure to build their machine learning solutions: from a pre-trained foundation model, they fine-tune the weights on the target task of interest. So, the Internet is swarmed by a handful of foundation models fine-tuned on many diverse tasks: these individual fine-tunings exist in isolation without benefiting from each other. In our opinion, this is a missed opportunity, as these specialized models contain rich and diverse features. In this paper, we thus propose model ratatouille, a new strategy to recycle the multiple fine-tunings of the same foundation model on diverse auxiliary tasks. Specifically, we repurpose these auxiliary weights as initializations for multiple parallel fine-tunings on the target task; then, we average all fine-tuned weights to obtain the final model. This recycling strategy aims at maximizing the diversity in weights by leveraging the diversity in auxiliary tasks. Empirically, it improves the state of the art on the reference DomainBed benchmark for out-of-distribution generalization. Looking forward, this work contributes to the emerging paradigm of updatable machine learning where, akin to open-source software development, the community collaborates to reliably update machine learning models.
Alexandre Ramé, Kartik Ahuja, Matthieu Cord, Léon Bottou, David Lopez-Paz
ICML5
2023 Learning useful representations for shifting tasks and distributions
abstract
Does the dominant approach to learn representations (as a side effect of optimizing an expected cost for a single training distribution) remain a good approach when we are dealing with multiple distributions? Our thesis is that *such scenarios are better served by representations that are richer than those obtained with a single optimization episode.* We support this thesis with simple theoretical arguments and with experiments utilizing an apparently näive ensembling technique: concatenating the representations obtained from multiple training episodes using the same data, model, algorithm, and hyper-parameters, but different random seeds. These independently trained networks perform similarly. Yet, in a number of scenarios involving new distributions, the concatenated representation performs substantially better than an equivalently sized network trained with a single training run. This proves that the representations constructed by multiple training episodes are in fact different. Although their concatenation carries little additional information about the training task under the training distribution, it becomes substantially more informative when tasks or distributions change. Meanwhile, a single training episode is unlikely to yield such a redundant representation because the optimization process has no reason to accumulate features that do not incrementally improve the training performance.
Léon Bottou
ICML2
2023 Birth of a Transformer: A Memory Viewpoint
abstract
Large language models based on transformers have achieved great empirical successes. However, as they are deployed more widely, there is a growing need to better understand their internal mechanisms in order to make them more reliable. These models appear to store vast amounts of knowledge from their training data, and to adapt quickly to new information provided in their context or prompt. We study how transformers balance these two types of knowledge by considering a synthetic setup where tokens are generated from either global or context-specific bigram distributions. By a careful empirical analysis of the training process on a simplified two-layer transformer, we illustrate the fast learning of global bigrams and the slower development of an "induction head" mechanism for the in-context bigrams. We highlight the role of weight matrices as associative memories, provide theoretical insights on how gradients enable their learning during training, and study the role of data-distributional properties.
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou, Léon Bottou
NeurIPS5
2022 On the Relation between Distributionally Robust Optimization and Data Curation (Student Abstract)
abstract
Machine learning systems based on minimizing average error have been shown to perform inconsistently across notable subsets of the data, which is not exposed by a low average error for the entire dataset. In consequential social and economic applications, where data represent people, this can lead to discrimination of underrepresented gender and ethnic groups. Distributionally Robust Optimization (DRO) seemingly addresses this problem by minimizing the worst expected risk across subpopulations. We establish theoretical results that clarify the relation between DRO and the optimization of the same loss averaged on an adequately weighted training dataset. A practical implication of our results is that neither DRO nor curating the training set should be construed as a complete solution for bias mitigation.
Agnieszka Slowik, Léon Bottou, Sean B. Holden, Mateja Jamnik
AAAI2
2022 On Distributionally Robust Optimization and Data Rebalancing
abstract
Machine learning systems based on minimizing average error have been shown to perform inconsistently across notable subsets of the data, which is not exposed by a low average error for the entire dataset. Distributionally Robust Optimization (DRO) seemingly addresses this problem by minimizing the worst expected risk across subpopulations. We establish theoretical results that clarify the relation between DRO and the optimization of the same loss averaged on an adequately weighted training dataset. The results cover finite and infinite number of training distributions, as well as convex and non-convex loss functions. An implication of our results is that for each DRO problem there exists a data distribution such that learning this distribution is equivalent to solving the DRO problem. Yet, important problems that DRO seeks to address (for instance, adversarial robustness and fighting bias) cannot be reduced to finding the one ’unbiased’ dataset. Our discussion section addresses this important discrepancy.
Agnieszka Slowik, Léon Bottou
AISTATS2
2022 Rich Feature Construction for the Optimization-Generalization Dilemma
abstract
There often is a dilemma between ease of optimization and robust out-of-distribution (OoD) generalization. For instance, many OoD methods rely on penalty terms whose optimization is challenging. They are either too strong to optimize reliably or too weak to achieve their goals. We propose to initialize the networks with a rich representation containing a palette of potentially useful features, ready to be used by even simple models. On the one hand, a rich representation provides a good initialization for the optimizer. On the other hand, it also provides an inductive bias that helps OoD generalization. Such a representation is constructed with the Rich Feature Construction (RFC) algorithm, also called the Bonsai algorithm, which consists of a succession of training episodes. During discovery episodes, we craft a multi-objective optimization criterion and its associated datasets in a manner that prevents the network from using the features constructed in the previous iterations. During synthesis episodes, we use knowledge distillation to force the network to simultaneously represent all the previously discovered features. Initializing the networks with Bonsai representations consistently helps six OoD methods achieve top performance on ColoredMNIST benchmark. The same technique substantially outperforms comparable results on the Wilds Camelyon17 task, eliminates the high result variance that plagues other methods, and makes hyperparameter tuning and model selection more reliable.
David Lopez-Paz, Léon Bottou
ICML3
2022 The Effects of Regularization and Data Augmentation are Class Dependent
abstract
Regularization is a fundamental technique to prevent over-fitting and to improve generalization performances by constraining a model's complexity. Current Deep Networks heavily rely on regularizers such as Data-Augmentation (DA) or weight-decay, and employ structural risk minimization, i.e. cross-validation, to select the optimal regularization hyper-parameters. In this study, we demonstrate that techniques such as DA or weight decay produce a model with a reduced complexity that is unfair across classes. The optimal amount of DA or weight decay found from cross-validation over all classes leads to disastrous model performances on some classes e.g. on Imagenet with a resnet50, the ``barn spider'' classification test accuracy falls from $68\%$ to $46\%$ only by introducing random crop DA during training. Even more surprising, such performance drop also appears when introducing uninformative regularization techniques such as weight decay. Those results demonstrate that our search for ever increasing generalization performance ---averaged over all classes and samples--- has left us with models and regularizers that silently sacrifice performances on some classes. This scenario can become dangerous when deploying a model on downstream tasks e.g. an Imagenet pre-trained resnet50 deployed on INaturalist sees its performances fall from $70\%$ to $30\%$ on class \#8889 when introducing random crop DA during the Imagenet pre-training phase. Those results demonstrate that finding a correct measure of a model's complexity without class-dependent preference remains an open research question.
Randall Balestriero, Léon Bottou, Yann LeCun
NeurIPS2
2022 A scaling calculus for the design and initialization of ReLU networks
abstract
Abstract We propose a system for calculating a “scaling constant” for layers and weights of neural networks. We relate this scaling constant to two important quantities that relate to the optimizability of neural networks, and argue that a network that is “preconditioned” via scaling, in the sense that all weights have the same scaling constant, will be easier to train. This scaling calculus results in a number of consequences, among them the fact that the geometric mean of the fan-in and fan-out, rather than the fan-in, fan-out, or arithmetic mean, should be used for the initialization of the variance of weights in a neural network. Our system allows for the off-line design & engineering of ReLU (Rectified Linear Unit) neural networks, potentially replacing blind experimentation. We verify the effectiveness of our approach on a set of benchmark problems.
Aaron Defazio, Léon Bottou
Neural Comput. Appl.2
2020 Symplectic Recurrent Neural Networks
Zhengdao Chen, Martín Arjovsky, Léon Bottou
ICLR4
2019 First-Order Adversarial Vulnerability of Neural Networks and Input Dimension
abstract
Over the past few years, neural networks were proven vulnerable to adversarial images: targeted but imperceptible image perturbations lead to drastically different predictions. We show that adversarial vulnerability increases with the gradients of the training objective when viewed as a function of the inputs. Surprisingly, vulnerability does not depend on network topology: for many standard network architectures, we prove that at initialization, the L1-norm of these gradients grows as the square root of the input dimension, leaving the networks increasingly vulnerable with growing image size. We empirically show that this dimension-dependence persists after either usual or robust training, but gets attenuated with higher regularization.
Carl-Johann Simon-Gabriel, Yann Ollivier, Léon Bottou, Bernhard Schölkopf, David Lopez-Paz
ICML3
2019 AdaGrad stepsizes: sharp convergence over nonconvex landscapes
abstract
Adaptive gradient methods such as AdaGrad and its variants update the stepsize in stochastic gradient descent on the fly according to the gradients received along the way; such methods have gained widespread use in large-scale optimization for their ability to converge robustly, without the need to fine-tune parameters such as the stepsize schedule. Yet, the theoretical guarantees to date for AdaGrad are for online and convex optimization. We bridge this gap by providing strong theoretical guarantees for the convergence of AdaGrad over smooth, nonconvex landscapes. We show that the norm version of AdaGrad (AdaGrad-Norm) converges to a stationary point at the $\mathcal{O}(\log(N)/\sqrt{N})$ rate in the stochastic setting, and at the optimal $\mathcal{O}(1/N)$ rate in the batch (non-stochastic) setting – in this sense, our convergence guarantees are “sharp”. In particular, both our theoretical results and extensive numerical experiments imply that AdaGrad-Norm is robust to the unknown Lipschitz constant and level of stochastic noise on the gradient.
Rachel A. Ward, Xiaoxia Wu, Léon Bottou
ICML3
2019 On the Ineffectiveness of Variance Reduced Optimization for Deep Learning
abstract
The application of stochastic variance reduction to optimization has shown remarkable recent theoretical and practical success. The applicability of these techniques to the hard non-convex optimization problems encountered during training of modern deep neural networks is an open problem. We show that naive application of the SVRG technique and related approaches fail, and explore why.
Aaron Defazio, Léon Bottou
NeurIPS2
2019 Cold Case: The Lost MNIST Digits
abstract
Although the popular MNIST dataset \citep{mnist} is derived from the NIST database \citep{nist-sd19}, precise processing steps of this derivation have been lost to time. We propose a reconstruction that is accurate enough to serve as a replacement for the MNIST dataset, with insignificant changes in accuracy. We trace each MNIST digit to its NIST source and its rich metadata such as writer identifier, partition identifier, etc. We also reconstruct the complete MNIST test set with 60,000 samples instead of the usual 10,000. Since the balance 50,000 were never distributed, they enable us to investigate the impact of twenty-five years of MNIST experiments on the reported testing performances. Our results unambiguously confirm the trends observed by \citet{recht2018cifar,recht2019imagenet}: although the misclassification rates are slightly off, classifier ordering and model selection remain broadly reliable. We attribute this phenomenon to the pairing benefits of comparing classifiers on the same digits.
Chhavi Yadav, Léon Bottou
NeurIPS2
2018 SING: Symbol-to-Instrument Neural Generator
abstract
Recent progress in deep learning for audio synthesis opens the way to models that directly produce the waveform, shifting away from the traditional paradigm of relying on vocoders or MIDI synthesizers for speech or music generation. Despite their successes, current state-of-the-art neural audio synthesizers such as WaveNet and SampleRNN suffer from prohibitive training and inference times because they are based on autoregressive models that generate audio samples one at a time at a rate of 16kHz. In this work, we study the more computationally efficient alternative of generating the waveform frame-by-frame with large strides. We present a lightweight neural audio synthesizer for the original task of generating musical notes given desired instrument, pitch and velocity. Our model is trained end-to-end to generate notes from nearly 1000 instruments with a single decoder, thanks to a new loss function that minimizes the distances between the log spectrograms of the generated and target waveforms. On the generalization task of synthesizing notes for pairs of pitch and instrument not seen during training, SING produces audio with significantly improved perceptual quality compared to a state-of-the-art autoencoder based on WaveNet as measured by a Mean Opinion Score (MOS), and is about 32 times faster for training and 2, 500 times faster for inference.
Alexandre Défossez, Neil Zeghidour, Nicolas Usunier, Léon Bottou, Francis R. Bach
NeurIPS4
2018 An efficient distributed learning algorithm based on effective local functional approximations
abstract
Scalable machine learning over big data is an important problem that is receiving a lot of attention in recent years. On popular distributed environments such as Hadoop running on a cluster of commodity machines, communication costs are substantial and algorithms need to be designed suitably considering those costs. In this paper we give a novel approach to the distributed training of linear classifiers (involving smooth losses and $L_2$ regularization) that is designed to reduce the total communication costs. At each iteration, the nodes minimize locally formed approximate objective functions; then the resulting minimizers are combined to form a descent direction to move. Our approach gives a lot of freedom in the formation of the approximate objective function as well as in the choice of methods to solve them. The method is shown to have $O(\log(1/\epsilon))$ time convergence. The method can be viewed as an iterative parameter mixing method. A special instantiation yields a parallel stochastic gradient descent method with strong convergence. When communication times between nodes are large, our method is much faster than the Terascale method (Agarwal et al., 2011), which is a state of the art distributed solver based on the statistical query model (Chu et al., 2006) that computes function and gradient values in a distributed fashion. We also evaluate against other recent distributed methods and demonstrate superior performance of our method.
Dhruv Mahajan 0001, Nikunj Agrawal, S. Sathiya Keerthi, Sundararajan Sellamanickam, Léon Bottou
J. Mach. Learn. Res.5
2017 Discovering Causal Signals in Images
abstract
This paper establishes the existence of observable footprints that reveal the causal dispositions of the object categories appearing in collections of images. We achieve this goal in two steps. First, we take a learning approach to observational causal discovery, and build a classifier that achieves state-of-the-art performance on finding the causal direction between pairs of random variables, given samples from their joint distribution. Second, we use our causal direction classifier to effectively distinguish between features of objects and features of their contexts in collections of static images. Our experiments demonstrate the existence of a relation between the direction of causality and the difference between objects and their contexts, and by the same token, the existence of observable signals that reveal the causal dispositions of objects.
David Lopez-Paz, Robert Nishihara, Soumith Chintala, Bernhard Schölkopf, Léon Bottou
CVPR5
2017 Towards Principled Methods for Training Generative Adversarial Networks
Martín Arjovsky, Léon Bottou
ICLR2
2017 Wasserstein Generative Adversarial Networks
abstract
We introduce a new algorithm named WGAN, an alternative to traditional GAN training. In this new model, we show that we can improve the stability of learning, get rid of problems like mode collapse, and provide meaningful learning curves useful for debugging and hyperparameter searches. Furthermore, we show that the corresponding optimization problem is sound, and provide extensive theoretical work highlighting the deep connections to different distances between distributions.
Martín Arjovsky, Soumith Chintala, Léon Bottou
ICML3
2016 No Regret Bound for Extreme Bandits
abstract
Algorithms for hyperparameter optimization abound, all of which work well under different and often unverifiable assumptions. Motivated by the general challenge of sequentially choosing which algorithm to use, we study the more specific task of choosing among distributions to use for random hyperparameter optimization. This work is naturally framed in the extreme bandit setting, which deals with sequentially choosing which distribution from a collection to sample in order to minimize (maximize) the single best cost (reward). Whereas the distributions in the standard bandit setting are primarily characterized by their means, a number of subtleties arise when we care about the minimal cost as opposed to the average cost. For example, there may not be a well-defined “best” distribution as there is in the standard bandit setting. The best distribution depends on the rewards that have been obtained and on the remaining time horizon. Whereas in the standard bandit setting, it is sensible to compare policies with an oracle which plays the single best arm, in the extreme bandit setting, there are multiple sensible oracle models. We define a sensible notion of “extreme regret” in the extreme bandit setting, which parallels the concept of regret in the standard bandit setting. We then prove that no policy can asymptotically achieve no extreme regret.
Robert Nishihara, David Lopez-Paz, Léon Bottou
AISTATS3
2015 How big data changes statistical machine learning
abstract
Summary form only given. This presentation illustrates how big data forces change on algorithmic techniques and the goals of machine learning, bringing along challenges and opportunities. 1. The theoretical foundations of statistical machine learning traditionally assume that training data is scarce. If one assumes instead that data is abundant and that the bottleneck is the computation time, stochastic algorithms with poor optimization performance become very attractive learning algorithms. These algorithms quickly became the backbone of large-scale machine learning and are the object of very active research. 2. Increasing the training set size cannot improve average errors indefinitely. However this diminishing returns problem vanishes if we measure instead the diversity of conditions in which the trained system performs well. In other words, big data is not an opportunity to increase the average accuracy, but an opportunity to increase coverage. Machine learning research must broaden its statistical framework in order to embrace all the (changing) aspects of real big data problems. Transfer learning, causal inference, and deep learning are successful steps in this direction.
Léon Bottou
IEEE BigData1
2015 Is object localization for free? - Weakly-supervised learning with convolutional neural networks
abstract
Successful methods for visual object recognition typically rely on training datasets containing lots of richly annotated images. Detailed image annotation, e.g. by object bounding boxes, however, is both expensive and often subjective. We describe a weakly supervised convolutional neural network (CNN) for object classification that relies only on image-level labels, yet can learn from cluttered scenes containing multiple objects. We quantify its object classification and object location prediction performance on the Pascal VOC 2012 (20 object classes) and the much larger Microsoft COCO (80 object classes) datasets. We find that the network (i) outputs accurate image-level labels, (ii) predicts approximate locations (but not extents) of objects, and (iii) performs comparably to its fully-supervised counterparts using object bounding box annotation for training.
Maxime Oquab, Léon Bottou, Ivan Laptev, Josef Sivic
CVPR2
2015 A Lower Bound for the Optimization of Finite Sums
abstract
This paper presents a lower bound for optimizing a finite sum of n functions, where each function is L-smooth and the sum is μ-strongly convex. We show that no algorithm can reach an error εin minimizing all functions from this class in fewer than Ω(n + \sqrtn(κ-1)\log(1/ε)) iterations, where κ=L/μis a surrogate condition number. We then compare this lower bound to upper bounds for recently developed methods specializing to this setting. When the functions involved in this sum are not arbitrary, but based on i.i.d. random data, then we further contrast these complexity results with those for optimal first-order methods to directly optimize the sum. The conclusion we draw is that a lot of caution is necessary for an accurate comparison, and identify machine learning scenarios where the new methods help computationally.
Alekh Agarwal, Léon Bottou
ICML2
2014 Learning and Transferring Mid-level Image Representations Using Convolutional Neural Networks
abstract
Convolutional neural networks (CNN) have recently shown outstanding image classification performance in the large- scale visual recognition challenge (ILSVRC2012). The success of CNNs is attributed to their ability to learn rich mid-level image representations as opposed to hand-designed low-level features used in other image classification methods. Learning CNNs, however, amounts to estimating millions of parameters and requires a very large number of annotated image samples. This property currently prevents application of CNNs to problems with limited training data. In this work we show how image representations learned with CNNs on large-scale annotated datasets can be efficiently transferred to other visual recognition tasks with limited amount of training data. We design a method to reuse layers trained on the ImageNet dataset to compute mid-level image representation for images in the PASCAL VOC dataset. We show that despite differences in image statistics and tasks in the two datasets, the transferred representation leads to significantly improved results for object and action classification, outperforming the current state of the art on Pascal VOC 2007 and 2012 datasets. We also show promising results for object and action localization.
Maxime Oquab, Léon Bottou, Ivan Laptev, Josef Sivic
CVPR2
2014 Learning Image Embeddings using Convolutional Neural Networks for Improved Multi-Modal Semantics
abstract
We construct multi-modal concept representations by concatenating a skip-gram linguistic representation vector with a visual concept representation vector computed using the feature extraction layers of a deep convolutional neural network (CNN) trained on a large labeled object recognition dataset.This transfer learning approach brings a clear performance gain over features based on the traditional bag-of-visual-word approach.Experimental results are reported on the WordSim353 and MEN semantic relatedness evaluation tasks.We use visual features computed using either ImageNet or ESP Game images.
Douwe Kiela, Léon Bottou
EMNLP2
2014 Introduction to the special issue on learning semantics
Antoine Bordes, Léon Bottou, Ronan Collobert, Dan Roth 0001, Jason Weston, Luke Zettlemoyer
Mach. Learn.2
2014 From machine learning to machine reasoning - An essay
Léon Bottou
Mach. Learn.1
2013 Counterfactual reasoning and learning systems: the example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero Candela, Denis Xavier Charles, David Maxwell Chickering, Elon Portugaly, Dipankar Ray, Patrice Y. Simard, Ed Snelson
J. Mach. Learn. Res.1
2011 Natural Language Processing (Almost) from Scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, Pavel P. Kuksa
J. Mach. Learn. Res.3
2011 Nonconvex Online Support Vector Machines
abstract
In this paper, we propose a nonconvex online Support Vector Machine (SVM) algorithm (LASVM-NC) based on the Ramp Loss, which has the strong ability of suppressing the influence of outliers. Then, again in the online learning setting, we propose an outlier filtering mechanism (LASVM-I) based on approximating nonconvex behavior in convex optimization. These two algorithms are built upon another novel SVM algorithm (LASVM-G) that is capable of generating accurate intermediate models in its iterative steps by leveraging the duality gap. We present experimental results that demonstrate the merit of our frameworks in achieving significant robustness to outliers in noisy data classification where mislabeled training instances are in abundance. Experimental evaluation shows that the proposed approaches yield a more scalable online SVM algorithm with sparser models and less computational running time, both in the training and recognition phases, without sacrificing generalization performance.We also point out the relation between nonconvex optimization and min-margin active learning.
Seyda Ertekin, Léon Bottou, C. Lee Giles
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Batch and online learning algorithms for nonconvex neyman-pearson classification
abstract
We describe and evaluate two algorithms for Neyman-Pearson (NP) classification problem which has been recently shown to be of a particular importance for bipartite ranking problems. NP classification is a nonconvex problem involving a constraint on false negatives rate. We investigated batch algorithm based on DC programming and stochastic gradient method well suited for large-scale datasets. Empirical evidences illustrate the potential of the proposed methods.
Gilles Gasso, Aristidis Pappaioannou, Marina Spivak, Léon Bottou
ACM Trans. Intell. Syst. Technol.4
2010 Erratum: SGDQN is Less Careful than Expected
Antoine Bordes, Léon Bottou, Patrick Gallinari, Jonathan D. Chang, S. Alex Smith
J. Mach. Learn. Res.2
2009 SGD-QN: Careful Quasi-Newton Stochastic Gradient Descent
Antoine Bordes, Léon Bottou, Patrick Gallinari
J. Mach. Learn. Res.2
2008 Sequence Labelling SVMs Trained in One Pass
Antoine Bordes, Nicolas Usunier, Léon Bottou
ECML/PKDD (1)3
2007 Learning on the border: active learning in imbalanced data classification
abstract
This paper is concerned with the class imbalance problem which has been known to hinder the learning performance of classification algorithms. The problem occurs when there are significantly less number of observations of the target concept. Various real-world classification tasks, such as medical diagnosis, text categorization and fraud detection suffer from this phenomenon. The standard machine learning algorithms yield better prediction performance with balanced datasets. In this paper, we demonstrate that active learning is capable of solving the class imbalance problem by providing the learner more balanced classes. We also propose an efficient way of selecting informative instances from a smaller pool of samples for active learning which does not necessitate a search through the entire dataset. The proposed method yields an efficient querying system and allows active learning to be applied to very large datasets. Our experimental results show that with an early stopping criteria, active learning achieves a fast solution with competitive prediction performance in imbalanced data classification.
Seyda Ertekin, Jian Huang 0002, Léon Bottou, C. Lee Giles
CIKM3
2007 Solving multiclass support vector machines with LaRank
abstract
Optimization algorithms for large margin multiclass recognizers are often too costly to handle ambitious problems with structured outputs and exponential numbers of classes. Optimization algorithms that rely on the full gradient are not effective because, unlike the solution, the gradient is not sparse and is very large. The LaRank algorithm sidesteps this difficulty by relying on a randomized exploration inspired by the perceptron algorithm. We show that this approach is competitive with gradient based optimizers on simple multiclass problems. Furthermore, a single LaRank pass over the training examples delivers test error rates that are nearly as good as those of the final solution.
Antoine Bordes, Léon Bottou, Patrick Gallinari, Jason Weston
ICML2
2007 The Tradeoffs of Large Scale Learning
abstract
This contribution develops a theoretical framework that takes into account the effect of approximate optimization on learning algorithms. The analysis shows distinct tradeoffs for the case of small-scale and large-scale learning problems. Small-scale learning problems are subject to the usual approximation--estimation tradeoff. Large-scale learning problems are subject to a qualitatively different tradeoff involving the computational complexity of the underlying optimization algorithms in non-trivial ways.
Léon Bottou, Olivier Bousquet
NIPS1
2007 The Need for Open Source Software in Machine Learning
Sören Sonnenburg, Mikio L. Braun, Cheng Soon Ong, Samy Bengio, Léon Bottou, Geoff Holmes 0001, Yann LeCun, Klaus-Robert Müller, Fernando Pereira 0003, Carl E. Rasmussen, Gunnar Rätsch, Bernhard Schölkopf, Alexander J. Smola, Pascal Vincent, Jason Weston, Robert C. Williamson
J. Mach. Learn. Res.5
2006 Trading convexity for scalability
abstract
Convex learning algorithms, such as Support Vector Machines (SVMs), are often seen as highly desirable because they offer strong practical properties and are amenable to theoretical analysis. However, in this work we show how non-convexity can provide scalability advantages over convexity. We show how concave-convex programming can be applied to produce (i) faster SVMs where training errors are no longer support vectors, and (ii) much faster Transductive SVMs.
Ronan Collobert, Fabian H. Sinz, Jason Weston, Léon Bottou
ICML4
2006 Inference with the Universum
abstract
In this paper we study a new framework introduced by Vapnik (1998) and Vapnik (2006) that is an alternative capacity concept to the large margin approach. In the particular case of binary classification, we are given a set of labeled examples, and a collection of "non-examples" that do not belong to either class of interest. This collection, called the Universum, allows one to encode prior knowledge by representing meaningful concepts in the same domain as the problem at hand. We describe an algorithm to leverage the Universum by maximizing the number of observed contradictions, and show experimentally that this approach delivers accuracy improvements over using labeled data alone.
Jason Weston, Ronan Collobert, Fabian H. Sinz, Léon Bottou, Vladimir Vapnik
ICML4
2006 Large Scale Transductive SVMs
abstract
We show how the concave-convex procedure can be applied to transductive SVMs, which traditionally require solving a combinatorial search problem. This provides for the first time a highly scalable algorithm in the nonlinear case. Detailed experiments verify the utility of our approach. Software is available at http://www.kyb.tuebingen.mpg.de/bs/people/fabee/transduction.html.
Ronan Collobert, Fabian H. Sinz, Jason Weston, Léon Bottou
J. Mach. Learn. Res.4
2005 The Huller: A Simple and Efficient Online SVM
Antoine Bordes, Léon Bottou
ECML2
2005 Fast Kernel Classifiers with Online and Active Learning
abstract
Very high dimensional learning systems become theoretically possible when training examples are abundant. The computing cost then becomes the limiting factor. Any efficient learning algorithm should at least take a brief look at each example. But should all examples be given equal attention? This contribution proposes an empirical answer. We first present an online SVM algorithm based on this premise. LASVM yields competitive misclassification rates after a single pass over the training examples, outspeeding state-of-the-art SVM solvers. Then we show how active example selection can yield faster training, higher accuracies, and simpler models, using only a fraction of the training example labels.
Antoine Bordes, Seyda Ertekin, Jason Weston, Léon Bottou
J. Mach. Learn. Res.4
2005 Toward Automatic Phenotyping of Developing Embryos From Videos
abstract
We describe a trainable system for analyzing videos of developing C. elegans embryos. The system automatically detects, segments, and locates cells and nuclei in microscopic images. The system was designed as the central component of a fully automated phenotyping system. The system contains three modules 1) a convolutional network trained to classify each pixel into five categories: cell wall, cytoplasm, nucleus membrane, nucleus, outside medium; 2) an energy-based model, which cleans up the output of the convolutional network by learning local consistency constraints that must be satisfied by label images; 3) a set of elastic models of the embryo at various stages of development that are matched to the label images.
F. Ning, D. Delhomme, Yann LeCun, F. Piano, Léon Bottou, Paolo Emilio Barbano
IEEE Trans. Image Process.5
2004 Learning Methods for Generic Object Recognition with Invariance to Pose and Lighting
Yann LeCun, Fu Jie Huang, Léon Bottou
CVPR (2)3
2004 Breaking SVM Complexity with Cross-Training
abstract
We propose to selectively remove examples from the training set using probabilistic estimates related to editing algorithms (Devijver and Kittler, 1982). This heuristic procedure aims at creating a separable distribution of training examples with minimal impact on the position of the decision boundary. It breaks the linear dependency between the number of SVs and the number of training examples, and sharply reduces the complexity of SVMs during both the training and prediction stages.
Gökhan H. Bakir, Léon Bottou, Jason Weston
NIPS2
2004 Parallel Support Vector Machines: The Cascade SVM
abstract
We describe an algorithm for support vector machines (SVM) that can be parallelized efficiently and scales to very large problems with hundreds of thousands of training vectors. Instead of analyzing the whole training set in one optimization step, the data are split into subsets and optimized separately with multiple SVMs. The partial results are combined and filtered again in a ‘Cascade’ of SVMs, until the global optimum is reached. The Cascade SVM can be spread over multiple processors with minimal communication overhead and requires far less memory, since the kernel matrices are much smaller than for a regular SVM. Convergence to the global optimum is guaranteed with multiple passes through the Cascade, but already a single pass provides good generalization. A single pass is 5x – 10x faster than a regular SVM for problems of 100,000 vectors when implemented on a single processor. Parallel implementations on a cluster of 16 processors were tested with over 1 million vectors (2-class problems), converging in a day or two, while a regular SVM never converged in over a week.
Hans Peter Graf, Eric Cosatto, Léon Bottou, Igor Durdanovic, Vladimir Vapnik
NIPS3
2003 Large Scale Online Learning
abstract
We consider situations where training data is abundant and computing resources are comparatively scarce. We argue that suitably designed on- line learning algorithms asymptotically outperform any batch learning algorithm. Both theoretical and experimental evidences are presented.
Léon Bottou, Yann LeCun
NIPS1
2003 Geometric Clustering Using the Information Bottleneck Method
abstract
We argue that K–means and deterministic annealing algorithms for geo- metric clustering can be derived from the more general Information Bot- tleneck approach. If we cluster the identities of data points to preserve information about their location, the set of optimal solutions is massively degenerate. But if we treat the equations that define the optimal solution as an iterative algorithm, then a set of “smooth” initial conditions selects solutions with the desired geometrical properties. In addition to concep- tual unification, we argue that this approach can be more efficient and robust than classic algorithms.
Susanne Still, William Bialek, Léon Bottou
NIPS3
2003 Scalable video coding with managed drift
abstract
Traditional scalable video encoders sacrifice coding efficiency to reduce error propagation because they have avoided using enhancement-layer (EL) information to predict the base layer (BL) to prevent the error propagation termed "drift". Drift can produce very poor video quality if left unchecked. We propose a video coder with significantly better compression efficiency because it intentionally allows the drift produced by predicting the BL from the EL. Our drift management system balances the tradeoff between compression efficiency and error propagation. The proposed scalable coder uses a spatially adaptive procedure that optimally selects key encoder parameters: the quantizer and the prediction strategy. Our numerical results indicate the encoder is very powerful, and the selection procedure is effective. The video quality of our coder at low rates is only marginally worse than the drift-free case, while its overall compression efficiency is not much worse than a one-layer nonscalable encoder.
Amy R. Reibman, Léon Bottou, Andrea Basso 0001
IEEE Trans. Circuits Syst. Video Technol.2
2002 Electronic Document Publishing Using DjVu
Artem Mikheev, Luc Vincent, Michael Hawrylycz, Léon Bottou
Document Analysis Systems4
2001 Masked Wavelets: Applications to Image Compression
Steven Pigeon, Léon Bottou
Data Compression Conference2
2001 Managing Drift in DCT-Based Scalable Video Coding
abstract
When compressed video is transmitted over erasure-prone channels, errors will propagate whenever temporal or spatial prediction is used. Typical tools to combat this error propagation are packetization, re-synchronizing codewords, intra-coding, and scalability. In recent years, the concern over so-called "drift" has sent researchers toward structures for scalability that do not use enhancement-layer information to predict base-layer information and hence have no drift. In this paper, we propose alternative structures for scalability that use previous enhancement-layer information to predict the current base layer, while simultaneously managing the resulting possibility of drift. These structures allow better compression efficiency, while introducing only limited impairments in the quality of the reconstruction.
Amy R. Reibman, Léon Bottou
Data Compression Conference2
2001 Efficient Conversion of Digital Documents to Multilayer Raster Formats
abstract
How can we turn the description of a digital (i.e. electronically produced) document into something that is efficient for multi-layer raster formats? It is first shown that a foreground/background segmentation without overlapping foreground components can be more efficient for viewing or printing. Then, a new algorithm that prevents overlaps between foreground components while optimizing both the document quality and compression ratio is derived from the minimum description length (MDL) criterion. This algorithm makes the DjVu compression format significantly, more efficient on electronically produced documents. Comparisons with other formats are provided.
Léon Bottou, Patrick Haffner, Yann LeCun
ICDAR1
2001 DCT-based scalable video coding with drift
abstract
Scalable video coders have traditionally avoided using enhancement-layer (EL) information to predict the base layer (BL), so as to avoid so-called "drift". As a result, they are less efficient than a one-layer coder. Fine granularity scalable (FGS) coders avoid using EL information to predict the EL as well, suffering even further inefficiencies. In this paper, we explore a scalable video coder that allows drift, by predicting the BL from EL information. However, we show that through careful management of the amount of drift introduced, the video quality at low rates is only marginally worse than the drift-free case, while the overall compression efficiency is not much worse than a one-layer encoder.
Amy R. Reibman, Léon Bottou, Andrea Basso 0001
ICIP (2)2
2000 Vicinal Risk Minimization
abstract
The Vicinal Risk Minimization principle establishes a bridge between generative models and methods derived from the Structural Risk Mini(cid:173) mization Principle such as Support Vector Machines or Statistical Reg(cid:173) ularization. We explain how VRM provides a framework which inte(cid:173) grates a number of existing algorithms, such as Parzen windows, Support Vector Machines, Ridge Regression, Constrained Logistic Classifiers and Tangent-Prop. We then show how the approach implies new algorithm(cid:173) s for solving problems usually associated with generative models. New algorithms are described for dealing with pattern recognition problems with very different pattern distributions and dealing with unlabeled data. Preliminary empirical results are presented.
Olivier Chapelle, Jason Weston, Léon Bottou, Vladimir Vapnik
NIPS3
1999 DjVu: Analyzing and Compressing Scanned Documents for Internet Distribution
abstract
DjVu is an image compression technique specifically geared towards the compression of scanned documents in color at high resolution. Typical color magazine pages scanned at 300 dpi are compressed to between 40 and 80 kBytes, or 5 to 10 times smaller than with JPEG for a similar level of subjective quality. The foreground layer, which contains the text and drawings and requires high spatial resolution, is separated from the background layer, which contains pictures and backgrounds and requires less resolution. The foreground is compressed with a bi-tonal image compression technique that takes advantage of character shape similarities. The background is compressed with a new progressive, wavelet-based compression method. A real-time, memory-efficient version of the decoder is available as a plug-in for popular Web browsers.
Patrick Haffner, Léon Bottou, Paul G. Howard, Yann LeCun
ICDAR2
1999 Color Documents on the Web with DJVU
abstract
We present a new image compression technique called "DjVu" that is specifically geared towards the compression of scanned documents in color at high resolution. With DjVu, a magazine page in color at 300 dpi typically occupies between 40 KB and 80 KB, approximately 5 to 10 times better than JPEG for a similar level of readability. Using a combination of hidden Markov model techniques and MDL-driven heuristics, DjVu first classifies each pixel in the image as either foreground (text, drawings) or background (pictures, photos, paper texture). The pixel categories form a bitonal image which is compressed using a pattern matching technique that takes advantage of the similarities between character shapes. A progressive, wavelet-based compression technique, combined with a masking algorithm, is then used to compress the foreground and background images at lower resolutions while minimizing the number of bits spent on the pixels that are not visible in the foreground and background planes. Encoders, decoders, and real-time, memory efficient plug-ins for various web browsers are available for all the major platforms.
Patrick Haffner, Yann LeCun, Léon Bottou, Paul G. Howard, Pascal Vincent, Bill Riemers
ICIP (1)3
1998 The Z-Coder Adaptive Binary Coder
abstract
We present the Z-coder, a new adaptive data compression coder for coding binary data. The Z-coder is derived from the Golomb (1966) run-length coder, and retains most of the speed and simplicity of the earlier coder. The Z-coder can also be thought of as a multiplication-free approximate arithmetic coder, showing the close relationship between run-length coding and arithmetic coding. The Z-coder improves upon existing arithmetic coders by its speed and its principled design. We present a derivation of the Z-coder as well as details of the construction of its adaptive probability estimation table.
Léon Bottou, Paul G. Howard, Yoshua Bengio
Data Compression Conference1
1998 Lossy Compression of Partially Masked Still Images
abstract
Summary form only given. This article describes a component of the DejaVu color document image compression. The DejaVu scheme consists in separating the image into (a) a foreground layer containing text and drawings, and (b) a background layer containing background color and the pictures. These two components are then compressed using the most appropriate techniques. This article outlines a wavelet method for encoding the background layer without wasting bits on pixels masked by the foreground layer. Our method handles arbitrarily complex masks with reasonable computational requirements.
Léon Bottou, Steven Pigeon
Data Compression Conference1
1998 Boxlets: A Fast Convolution Algorithm for Signal Processing and Neural Networks
Patrice Y. Simard, Léon Bottou, Patrick Haffner, Yann LeCun
NIPS2
1998 Gradient-based learning applied to document recognition
abstract
Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network architecture, gradient-based learning algorithms can be used to synthesize a complex decision surface that can classify high-dimensional patterns, such as handwritten characters, with minimal preprocessing. This paper reviews various methods applied to handwritten character recognition and compares them on a standard handwritten digit recognition task. Convolutional neural networks, which are specifically designed to deal with the variability of 2D shapes, are shown to outperform all other techniques. Real-life document recognition systems are composed of multiple modules including field extraction, segmentation recognition, and language modeling. A new learning paradigm, called graph transformer networks (GTN), allows such multimodule systems to be trained globally using gradient-based methods so as to minimize an overall performance measure. Two systems for online handwriting recognition are described. Experiments demonstrate the advantage of global training, and the flexibility of graph transformer networks. A graph transformer network for reading a bank cheque is also described. It uses convolutional neural network character recognizers combined with global training techniques to provide record accuracy on business and personal cheques. It is deployed commercially and reads several million cheques per day.
Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner
Proc. IEEE2
1998 Image and video coding-emerging standards and beyond
abstract
Discusses coding standards for still images and motion video. We first briefly discuss standards already in use, including: Group 3 and Group 4 for bilevel fax images; JPEG for still color images; and H.261, H.263, MPEG-1, and MPEG-2 for motion video. We then cover newly emerging standards such as JBIG1 and JBIG2 for bilevel fax images, JPEG-2000 for still color images, and H.263+ and MPEG-4 for motion video. Finally, we describe some directions beyond the standards such as hybrid coding of graphics/photo images, MPEG-7 for multimedia metadata, and possible new technologies.
Barry G. Haskell, Paul G. Howard, Yann LeCun, Atul Puri, Jörn Ostermann, M. Reha Civanlar, Lawrence R. Rabiner, Léon Bottou, Patrick Haffner
IEEE Trans. Circuits Syst. Video Technol.8
1997 Global Training of Document Processing Systems Using Graph Transformer Networks
abstract
We propose a new machine learning paradigm called Graph Transformer Networks that extends the applicability of gradient-based learning algorithms to systems composed of modules that take graphs as inputs and produce graphs as output. Training is performed by computing gradients of a global objective function with respect to all the parameters in the system using a kind of back-propagation procedure. A complete check reading system based on these concepts is described. The system uses convolutional neural network character recognizers, combined with global training techniques to provide record accuracy on business and personal checks. It is presently deployed commercially and reads million of checks per month.
Léon Bottou, Yoshua Bengio, Yann LeCun
CVPR1
1997 Reading checks with multilayer graph transformer networks
abstract
We propose a new machine learning paradigm called multilayer graph transformer network that extends the applicability of gradient-based learning algorithms to systems composed of modules that take graphs as input and produce graphs as output. A complete check reading system based on this concept is described. The system combines convolutional neural network character recognizers with graph-based stochastic models trained cooperatively at the document level. It is deployed commercially and reads million of business and personal checks per month with record accuracy.
Yann LeCun, Léon Bottou, Yoshua Bengio
ICASSP2
1994 Comparison of classifier methods: a case study in handwritten digit recognition
abstract
This paper compares the performance of several classifier algorithms on a standard database of handwritten digits. We consider not only raw accuracy, but also training time, recognition time, and memory requirements. When available, we report measurements of the fraction of patterns that must be rejected so that the remaining patterns have misclassification rates less than a given threshold.
Léon Bottou, Corinna Cortes, John S. Denker, Harris Drucker, Isabelle Guyon, Lawrence D. Jackel, Yann LeCun, Urs A. Müller, Patrice Y. Simard, Vladimir Vapnik
ICPR (2)1
1994 Convergence Properties of the K-Means Algorithms
abstract
This paper studies the convergence properties of the well known K-Means clustering algorithm. The K-Means algorithm can be de(cid:173) scribed either as a gradient descent algorithm or by slightly extend(cid:173) ing the mathematics of the EM algorithm to this hard threshold case. We show that the K-Means algorithm actually minimizes the quantization error using the very fast Newton algorithm.
Léon Bottou, Yoshua Bengio
NIPS1
1993 Signature Verification Using A "Siamese" Time Delay Neural Network
abstract
This paper describes the development of an algorithm for verification of signatures written on a touch-sensitive pad. The signature verification algorithm is based on an artificial neural network. The novel network presented here, called a “Siamese” time delay neural network, consists of two identical networks joined at their output. During training the network learns to measure the similarity between pairs of signatures. When used for verification, only one half of the Siamese network is evaluated. The output of this half network is the feature vector for the input signature. Verification consists of comparing this feature vector with a stored feature vector for the signer. Signatures closer than a chosen threshold to this stored representation are accepted, all other signatures are rejected as forgeries. System performance is illustrated with experiments performed in the laboratory.
Jane Bromley, James W. Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Roopak Shah
Int. J. Pattern Recognit. Artif. Intell.3
1993 Local Algorithms for Pattern Recognition and Dependencies Estimation
abstract
In previous publications (Bottou and Vapnik 1992; Vapnik 1992) we described local learning algorithms, which result in performance improvements for real problems. We present here the theoretical framework on which these algorithms are based. First, we present a new statement of certain learning problems, namely the local risk minimization. We review the basic results of the uniform convergence theory of learning, and extend these results to local risk minimization. We also extend the structural risk minimization principle for both pattern recognition problems and regression problems. This extended induction principle is the basis for a new class of algorithms.
Vladimir Vapnik, Léon Bottou
Neural Comput.2
1992 Capacity control in linear classifiers for pattern recognition
abstract
Achieving good performance in statistical pattern recognition requires matching the capacity of the classifier to the amount of training data. If the classifier has too many adjustable parameters (large capacity), it is likely to learn the training data without difficulty, but will probably not generalize properly to patterns that do not belong to the training set. Conversely, if the capacity of the classifier is not large enough, it might not be able to learn the task at all. In between, there is an optimal classifier capacity which ensures the best expected generalization for a given amount of training data. The method of structural risk minimization (SRM) refers to tuning the capacity of the classifier to the available amount of training data. This paper illustrates the method of SRM with several examples of algorithms. Experiments confirm theoretical predictions of performance improvement in application to handwritten digit recognition.>
Isabelle Guyon, Vladimir Vapnik, Bernhard E. Boser, Léon Bottou, Sara A. Solla
ICPR (2)4
1992 Computer aided cleaning of large databases for character recognition
abstract
A method for computer-aided cleaning of undesirable patterns in large training databases has been developed. The method uses the trainable classifier itself, to point out patterns that are suspicious, and should be checked by the human supervisor. While suspicious patterns that are meaningless or mislabeled are considered garbage, and removed from the database, the remaining patterns, like ambiguous or atypical, represent valid patterns that are hard to learn and should be kept in the database. By using the method of pattern cleaning, combined with an emphasizing scheme applied on the patterns that are hard to learn, the error rate on the test set has been reduced by half, in the case of the database of handwritten lowercase characters entered on a touch terminal. The classifier is based on a time delay neural network (TDNN).>
Nada Matic, Isabelle Guyon, Léon Bottou, John S. Denker, Vladimir Vapnik
ICPR (2)3
1992 Local Learning Algorithms
abstract
Very rarely are training data evenly distributed in the input space. Local learning algorithms attempt to locally adjust the capacity of the training system to the properties of the training set in each area of the input space. The family of local learning algorithms contains known methods, like the k-nearest neighbors method (kNN) or the radial basis function networks (RBF), as well as new algorithms. A single analysis models some aspects of these algorithms. In particular, it suggests that neither kNN or RBF, nor nonlocal classifiers, achieve the best compromise between locality and capacity. A careful control of these parameters in a simple local learning algorithm has provided a performance breakthrough for an optical character recognition problem. Both the error rate and the rejection performance have been significantly improved.
Léon Bottou, Vladimir Vapnik
Neural Comput.1
1991 Structural Risk Minimization for Character Recognition
Isabelle Guyon, Vladimir Vapnik, Bernhard E. Boser, Léon Bottou, Sara A. Solla
NIPS4
1990 A Framework for the Cooperation of Learning Algorithms
Léon Bottou, Patrick Gallinari
NIPS1
1990 Speaker-independent isolated digit recognition: Multilayer perceptrons vs. Dynamic time warping
Léon Bottou, Françoise Fogelman-Soulié, Pascal Blanchet, Jean-Sylvain Liénard
Neural Networks1
1989 Experiments with time delay networks and dynamic time warping for speaker independent isolated digits recognition
Léon Bottou, Françoise Fogelman-Soulié, Pascal Blanchet, Jean-Sylvain Liénard
EUROSPEECH1