VLDB 2026 Research / reviewers in the wild / expert
Marc'Aurelio Ranzato
dblp:28/1732
· DBLP profile ↗
53ranked-venue papers
11as first author
8since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 11 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
42 papers |
Machine translation · 17% Deep learning architectures and training · 16% Transfer learning and domain adaptation · 13% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 86% Image and video processing · 14% |
Topics — the 30 heaviest of 83, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
continual learning |
1.8 | 4 | 2023 | Nevis'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision Research · J. Mach. Learn. Res. 2023 Efficient Continual Learning with Modular Networks and Task-Driven Priors · ICLR 2021 Efficient Lifelong Learning with A-GEM · ICLR (Poster) 2019 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
1.2 | 2 | 2023 | Nevis'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision Research · J. Mach. Learn. Res. 2023 Towards Learning Universal Hyperparameter Optimizers with Transformers · NeurIPS 2022 |
Natural language and speech › Machine translation
neural machine translation |
1.2 | 3 | 2021 | Discriminative Reranking for Neural Machine Translation · ACL/IJCNLP (1) 2021 Analyzing Uncertainty in Neural Machine Translation · ICML 2018 Phrase-Based & Neural Unsupervised Machine Translation · EMNLP 2018 |
Machine learning › Generative modeling
energy-based model |
1.1 | 4 | 2021 | Residual Energy-Based Models for Text · J. Mach. Learn. Res. 2021 Residual Energy-Based Models for Text Generation · ICLR 2020 On Autoencoders and Score Matching for Energy Based Models · ICML 2011 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 2 | 2022 | Unified Scaling Laws for Routed Language Models · ICML 2022 Hard Mixtures of Experts for Large Scale Weakly Supervised Vision · CVPR 2017 |
Natural language and speech › Machine translation
unsupervised machine translation |
0.7 | 2 | 2018 | Unsupervised Machine Translation Using Monolingual Corpora Only · ICLR (Poster) 2018 Phrase-Based & Neural Unsupervised Machine Translation · EMNLP 2018 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.6 | 1 | 2022 | Towards Learning Universal Hyperparameter Optimizers with Transformers · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
scaling laws |
0.6 | 1 | 2022 | Unified Scaling Laws for Routed Language Models · ICML 2022 |
Machine learning › Deep learning architectures and training
modular network |
0.5 | 1 | 2021 | Efficient Continual Learning with Modular Networks and Task-Driven Priors · ICLR 2021 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.4 | 5 | 2012 | Building high-level features using large scale unsupervised learning · ICML 2012 Learning invariant features through topographic filter maps · CVPR 2009 Sparse Feature Learning for Deep Belief Networks · NIPS 2007 |
Natural language and speech › Machine translation
machine translation evaluation |
0.4 | 1 | 2020 | On The Evaluation of Machine Translation SystemsTrained With Back-Translation · ACL 2020 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
neural sequence generation |
0.4 | 1 | 2020 | Revisiting Self-Training for Neural Sequence Generation · ICLR 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
0.4 | 1 | 2020 | Revisiting Self-Training for Neural Sequence Generation · ICLR 2020 |
Machine learning › Learning paradigms
semi-supervised learning |
0.4 | 1 | 2020 | Revisiting Self-Training for Neural Sequence Generation · ICLR 2020 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2020 | Residual Energy-Based Models for Text Generation · ICLR 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 2 | 2017 | Hard Mixtures of Experts for Large Scale Weakly Supervised Vision · CVPR 2017 Large Scale Distributed Deep Networks · NIPS 2012 |
Computer vision › Face, body and person analysis
face recognition |
0.4 | 2 | 2015 | Web-scale training for face identification · CVPR 2015 DeepFace: Closing the Gap to Human-Level Performance in Face Verification · CVPR 2014 |
Machine learning › Deep learning architectures and training › deep generative model
deep belief network |
0.4 | 4 | 2013 | Modeling Natural Images Using Gated MRFs · IEEE Trans. Pattern Anal. Mach. Intell. 2013 On deep generative models with applications to recognition · CVPR 2011 Sparse Feature Learning for Deep Belief Networks · NIPS 2007 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
compositional zero-shot learning |
0.4 | 1 | 2019 | Task-Driven Modular Networks for Zero-Shot Compositional Learning · ICCV 2019 |
Natural language and speech › Language models and text generation
controllable text generation |
0.4 | 1 | 2019 | Multiple-Attribute Text Rewriting · ICLR (Poster) 2019 |
Natural language and speech › Machine translation › controllable machine translation
diverse machine translation |
0.4 | 1 | 2019 | Mixture Models for Diverse Machine Translation: Tricks of the Trade · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.4 | 1 | 2019 | Mixture Models for Diverse Machine Translation: Tricks of the Trade · ICML 2019 |
Machine learning › Learning paradigms
lifelong learning |
0.4 | 1 | 2019 | Efficient Lifelong Learning with A-GEM · ICLR (Poster) 2019 |
Natural language and speech › Machine translation
low-resource machine translation |
0.4 | 1 | 2019 | The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English · EMNLP/IJCNLP (1) 2019 |
Machine learning › Deep learning architectures and training
memory-augmented neural networks |
0.4 | 1 | 2019 | Large Memory Layers with Product Keys · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.4 | 1 | 2019 | Mixture Models for Diverse Machine Translation: Tricks of the Trade · ICML 2019 |
Natural language and speech › Language models and text generation › text generation
text rewriting |
0.4 | 1 | 2019 | Multiple-Attribute Text Rewriting · ICLR (Poster) 2019 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.4 | 1 | 2019 | Task-Driven Modular Networks for Zero-Shot Compositional Learning · ICCV 2019 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding |
0.3 | 1 | 2018 | Word translation without parallel data · ICLR (Poster) 2018 |
Machine learning › Trustworthy machine learning › calibration
model calibration |
0.3 | 1 | 2018 | Analyzing Uncertainty in Neural Machine Translation · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
back-translation · 1.5transformer · 1.0gradient projection · 0.7episodic memory · 0.7supervised learning · 0.7AutoML · 0.7uncertainty estimation · 0.6power-law scaling · 0.6gaussian process · 0.6effective parameter count · 0.6encoder-decoder · 0.3adversarial training · 0.3asynchronous SGD · 0.1Sandblaster L-BFGS · 0.1Downpour SGD · 0.1weight sharing · 0.1latent variable modeling · 0.1gated MRF · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Nevis'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision ResearchabstractA shared goal of several machine learning communities like continual learning, meta-learning and transfer learning, is to design algorithms and models that efficiently and robustly adapt to unseen tasks. An even more ambitious goal is to build models that never stop adapting, and that become increasingly more efficient through time by suitably transferring the accrued knowledge. Beyond the study of the actual learning algorithm and model architecture, there are several hurdles towards our quest to build such models, such as the choice of learning protocol, metric of success and data needed to validate research hypotheses. In this work, we introduce the Never-Ending VIsual-classification Stream (NEVIS'22), a benchmark consisting of a stream of over 100 visual classification tasks, sorted chronologically and extracted from papers sampled uniformly from computer vision proceedings spanning the last three decades. The resulting stream reflects what the research community thought was meaningful at any point in time, and it serves as an ideal test bed to assess how well models can adapt to new tasks, and do so better and more efficiently as time goes by. Despite being limited to classification, the resulting stream has a rich diversity of tasks from OCR, to texture analysis, scene recognition, and so forth. The diversity is also reflected in the wide range of dataset sizes, spanning over four orders of magnitude. Overall, NEVIS'22 poses an unprecedented challenge for current sequential learning approaches due to the scale and diversity of tasks, yet with a low entry barrier as it is limited to a single modality and well understood supervised learning problems. Moreover, we provide a reference implementation including strong baselines and an evaluation protocol to compare methods in terms of their trade-off between accuracy and compute. We hope that NEVIS'22 can be useful to researchers working on continual learning, meta-learning, AutoML and more generally sequential learning, and help these communities join forces towards more robust models that efficiently adapt to a never ending stream of data. Jörg Bornschein, Alexandre Galashov, Ross Hemsley, Amal Rannen Triki, Yutian Chen 0001, Arslan Chaudhry, Xu Owen He, Arthur Douillard, Massimo Caccia, Qixuan Feng, Sylvestre-Alvise Rebuffi, Kitty Stacpoole, Diego de Las Casas, Will Hawkins, Angeliki Lazaridou, Yee Whye Teh, Andrei A. Rusu, Razvan Pascanu, Marc'Aurelio Ranzato |
J. Mach. Learn. Res. | 20 |
| 2022 | Unified Scaling Laws for Routed Language ModelsabstractThe performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters. Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan |
ICML | 22 |
| 2022 | Towards Learning Universal Hyperparameter Optimizers with TransformersabstractMeta-learning hyperparameter optimization (HPO) algorithms from prior experiments is a promising approach to improve optimization efficiency over objective functions from a similar distribution. However, existing methods are restricted to learning from experiments sharing the same set of hyperparameters. In this paper, we introduce the OptFormer, the first text-based Transformer HPO framework that provides a universal end-to-end interface for jointly learning policy and function prediction when trained on vast tuning data from the wild, such as Google’s Vizier database, one of the world’s largest HPO datasets. Our extensive experiments demonstrate that the OptFormer can simultaneously imitate at least 7 different HPO algorithms, which can be further improved via its function uncertainty estimates. Compared to a Gaussian Process, the OptFormer also learns a robust prior distribution for hyperparameter response functions, and can thereby provide more accurate and better calibrated predictions. This work paves the path to future extensions for training a Transformer-based model as a general HPO optimizer. Yutian Chen 0001, Xingyou Song, Chansoo Lee, Qiuyi Zhang 0001, David Dohan, Kazuya Kawakami, Greg Kochanski, Arnaud Doucet, Marc'Aurelio Ranzato, Sagi Perel, Nando de Freitas |
NeurIPS | 10 |
| 2022 | The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine TranslationabstractAbstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the Flores-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are fully aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond. Naman Goyal 0001, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzmán, Angela Fan |
Trans. Assoc. Comput. Linguistics | 8 |
| 2021 | Discriminative Reranking for Neural Machine TranslationabstractAnn Lee, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ann Lee 0001, Michael Auli, Marc'Aurelio Ranzato |
ACL/IJCNLP (1) | 3 |
| 2021 | The Source-Target Domain Mismatch Problem in Machine TranslationabstractJiajun Shen, Peng-Jen Chen, Matthew Le, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Peng-Jen Chen, Matt Le 0001, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc'Aurelio Ranzato |
EACL | 8 |
| 2021 | Efficient Continual Learning with Modular Networks and Task-Driven Priors
Tom Veniat, Ludovic Denoyer, Marc'Aurelio Ranzato |
ICLR | 3 |
| 2021 | Residual Energy-Based Models for TextabstractCurrent large-scale auto-regressive language models display impressive fluency and can generate convincing text. In this work we start by asking the question: Can the generations of these models be reliably distinguished from real text by statistical discriminators? We find experimentally that the answer is affirmative when we have access to the training data for the model, and guardedly affirmative even if we do not. This suggests that the auto-regressive models can be improved by incorporating the (globally normalized) discriminators into the generative process. We give a formalism for this using the Energy-Based Model framework, and show that it indeed improves the results of the generative models, measured both in terms of perplexity and in terms of human evaluation. Anton Bakhtin, Yuntian Deng, Sam Gross, Myle Ott, Marc'Aurelio Ranzato, Arthur Szlam |
J. Mach. Learn. Res. | 5 |
| 2020 | On The Evaluation of Machine Translation SystemsTrained With Back-TranslationabstractBack-translation is a widely used data augmentation technique which leverages target monolingual data.However, its effectiveness has been challenged since automatic metrics such as BLEU only show significant improvements for test examples where the source itself is a translation, or translationese.This is believed to be due to translationese inputs better matching the back-translated training data.In this work, we show that this conjecture is not empirically supported and that backtranslation improves translation quality of both naturally occurring text as well as translationese according to professional human translators.We provide empirical evidence to support the view that back-translation is preferred by humans because it produces more fluent outputs.BLEU cannot capture human preferences because references are translationese when source sentences are natural text.We recommend complementing BLEU with a language model score to measure fluency. Sergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael Auli |
ACL | 3 |
| 2020 | Residual Energy-Based Models for Text Generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, Marc'Aurelio Ranzato |
ICLR | 5 |
| 2020 | Revisiting Self-Training for Neural Sequence Generation
Junxian He, Jiatao Gu, Marc'Aurelio Ranzato |
ICLR | 4 |
| 2019 | The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-EnglishabstractFrancisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino 0001, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc'Aurelio Ranzato |
EMNLP/IJCNLP (1) | 8 |
| 2019 | Task-Driven Modular Networks for Zero-Shot Compositional LearningabstractOne of the hallmarks of human intelligence is the ability to compose learned knowledge into novel concepts which can be recognized without a single training example. In contrast, current state-of-the-art methods require hundreds of training examples for each possible category to build reliable and accurate classifiers. To alleviate this striking difference in efficiency, we propose a task-driven modular architecture for compositional reasoning and sample efficient learning. Our architecture consists of a set of neural network modules, which are small fully connected layers operating in semantic concept space. These modules are configured through a gating function conditioned on the task to produce features representing the compatibility between the input image and the concept under consideration. This enables us to express tasks as a combination of sub-tasks and to generalize to unseen categories by reweighting a set of small modules. Furthermore, the network can be trained efficiently as it is fully differentiable and its modules operate on small sub-spaces. We focus our study on the problem of compositional zero-shot classification of object-attribute categories. We show in our experiments that current evaluation metrics are flawed as they only consider unseen object-attribute pairs. When extending the evaluation to the generalized setting which accounts also for pairs seen during training, we discover that naive baseline methods perform similarly or better than current approaches. However, our modular network is able to outperform all existing approaches on two widely-used benchmark datasets. Senthil Purushwalkam, Maximilian Nickel, Abhinav Gupta 0001, Marc'Aurelio Ranzato |
ICCV | 4 |
| 2019 | Efficient Lifelong Learning with A-GEM
Arslan Chaudhry, Marc'Aurelio Ranzato, Marcus Rohrbach, Mohamed Elhoseiny 0001 |
ICLR (Poster) | 2 |
| 2019 | Multiple-Attribute Text Rewriting
Guillaume Lample, Sandeep Subramanian, Eric Michael Smith, Ludovic Denoyer, Marc'Aurelio Ranzato, Y-Lan Boureau |
ICLR (Poster) | 5 |
| 2019 | Mixture Models for Diverse Machine Translation: Tricks of the TradeabstractMixture models trained via EM are among the simplest, most widely used and well understood latent variable models in the machine learning literature. Surprisingly, these models have been hardly explored in text generation applications such as machine translation. In principle, they provide a latent variable to control generation and produce a diverse set of hypotheses. In practice, however, mixture models are prone to degeneracies—often only one component gets trained or the latent variable is simply ignored. We find that disabling dropout noise in responsibility computation is critical to successful training. In addition, the design choices of parameterization, prior distribution, hard versus soft EM and online versus offline assignment can dramatically affect model performance. We develop an evaluation protocol to assess both quality and diversity of generations against multiple references, and provide an extensive empirical study of several mixture model variants. Our analysis shows that certain types of mixture models are more robust and offer the best trade-off between translation quality and diversity compared to variational models and diverse decoding approaches.\footnote{Code to reproduce the results in this paper is available at \url{https://github.com/pytorch/fairseq}} Tianxiao Shen, Myle Ott, Michael Auli, Marc'Aurelio Ranzato |
ICML | 4 |
| 2019 | Large Memory Layers with Product KeysabstractThis paper introduces a structured memory which can be easily integrated into a neural network. The memory is very large by design and significantly increases the capacity of the architecture, by up to a billion parameters with a negligible computational overhead. Its design and access pattern is based on product keys, which enable fast and exact nearest neighbor search. The ability to increase the number of parameters while keeping the same computational budget lets the overall system strike a better trade-off between prediction accuracy and computation efficiency both at training and test time. This memory layer allows us to tackle very large scale language modeling tasks. In our experiments we consider a dataset with up to 30 billion words, and we plug our memory layer in a state-of-the-art transformer-based architecture. In particular, we found that a memory augmented model with only 12 layers outperforms a baseline transformer model with 24 layers, while being twice faster at inference time. We release our code for reproducibility purposes. Guillaume Lample, Alexandre Sablayrolles, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou |
NeurIPS | 3 |
| 2018 | Phrase-Based & Neural Unsupervised Machine TranslationabstractMachine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences, which hinders their applicability to the majority of language pairs.This work investigates how to learn to translate when having access to only large monolingual corpora in each language.We propose two model variants, a neural and a phrase-based model.Both versions leverage a careful initialization of the parameters, the denoising effect of language models and automatic generation of parallel data by iterative back-translation.These models are significantly better than methods from the literature, while being simpler and having fewer hyper-parameters.On the widely used WMT'14 English-French and WMT'16 German-English benchmarks, our models respectively obtain 28.1 and 25.2 BLEU points without using a single parallel sentence, outperforming the state of the art by more than 11 BLEU points.On low-resource languages like English-Urdu and English-Romanian, our methods achieve even better results than semisupervised and supervised approaches leveraging the paucity of available bitexts.Our code for NMT and PBSMT is publicly available. Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, Marc'Aurelio Ranzato |
EMNLP | 5 |
| 2018 | Unsupervised Machine Translation Using Monolingual Corpora Only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, Marc'Aurelio Ranzato |
ICLR (Poster) | 4 |
| 2018 | Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou |
ICLR (Poster) | 3 |
| 2018 | Analyzing Uncertainty in Neural Machine TranslationabstractMachine translation is a popular test bed for research in neural sequence-to-sequence models but despite much recent research, there is still a lack of understanding of these models. Practitioners report performance degradation with large beams, the under-estimation of rare words and a lack of diversity in the final translations. Our study relates some of these issues to the inherent uncertainty of the task, due to the existence of multiple valid translations for a single source sentence, and to the extrinsic uncertainty caused by noisy training data. We propose tools and metrics to assess how uncertainty in the data is captured by the model distribution and how it affects search strategies that generate translations. Our results show that search works remarkably well but that the models tend to spread too much probability mass over the hypothesis space. Next, we propose tools to assess model calibration and show how to easily fix some shortcomings of current models. We release both code and multiple human reference translations for two popular benchmarks. Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato |
ICML | 4 |
| 2018 | Classical Structured Prediction Losses for Sequence to Sequence LearningabstractSergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc’Aurelio Ranzato. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato |
NAACL-HLT | 5 |
| 2017 | Hard Mixtures of Experts for Large Scale Weakly Supervised VisionabstractTraining convolutional networks (CNNs) that fit on a single GPU with minibatch stochastic gradient descent has become effective in practice. However, there is still no effective method for training large networks that do not fit in the memory of a few GPU cards, or for parallelizing CNN training. In this work we show that a simple hard mixture of experts model can be efficiently trained to good effect on large scale hashtag (multilabel) prediction tasks. Mixture of experts models are not new [7, 3], but in the past, researchers have had to devise sophisticated methods to deal with data fragmentation. We show empirically that modern weakly supervised data sets are large enough to support naive partitioning schemes where each data point is assigned to a single expert. Because the experts are independent, training them in parallel is easy, and evaluation is cheap for the size of the model. Furthermore, we show that we can use a single decoding layer for all the experts, allowing a unified feature embedding space. We demonstrate that it is feasible (and in fact relatively painless) to train far larger models than could be practically trained with standard CNN architectures, and that the extra capacity can be well used on current datasets. Sam Gross, Marc'Aurelio Ranzato, Arthur Szlam |
CVPR | 2 |
| 2017 | Dialogue Learning With Human-in-the-Loop
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston |
ICLR (Poster) | 4 |
| 2017 | Learning through Dialogue Interactions by Asking Questions
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston |
ICLR (Poster) | 4 |
| 2017 | Fader Networks: Manipulating Images by Sliding AttributesabstractThis paper introduces a new encoder-decoder architecture that is trained to reconstruct images by disentangling the salient information of the image and the values of attributes directly in the latent space. As a result, after training, our model can generate different realistic versions of an input image by varying the attribute values. By using continuous attribute values, we can choose how much a specific attribute is perceivable in the generated image. This property could allow for applications where users can modify an image using sliding knobs, like faders on a mixing console, to change the facial expression of a portrait, or to update the color of some objects. Compared to the state-of-the-art which mostly relies on training adversarial networks in pixel space by altering attribute values at train time, our approach results in much simpler training schemes and nicely scales to multiple attributes. We present evidence that our model can significantly change the perceived value of the attributes while preserving the naturalness of images. Guillaume Lample, Neil Zeghidour, Nicolas Usunier, Antoine Bordes, Ludovic Denoyer, Marc'Aurelio Ranzato |
NIPS | 6 |
| 2017 | Gradient Episodic Memory for Continual LearningabstractOne major obstacle towards AI is the poor ability of models to solve new problems quicker, and without forgetting previously acquired knowledge. To better understand this issue, we study the problem of continual learning, where the model observes, once and one by one, examples concerning a sequence of tasks. First, we propose a set of metrics to evaluate models learning over a continuum of data. These metrics characterize models not only by their test accuracy, but also in terms of their ability to transfer knowledge across tasks. Second, we propose a model for continual learning, called Gradient Episodic Memory (GEM) that alleviates forgetting, while allowing beneficial transfer of knowledge to previous tasks. Our experiments on variants of the MNIST and CIFAR-100 datasets demonstrate the strong performance of GEM when compared to the state-of-the-art. David Lopez-Paz, Marc'Aurelio Ranzato |
NIPS | 2 |
| 2015 | Web-scale training for face identificationabstractScaling machine learning methods to very large datasets has attracted considerable attention in recent years, thanks to easy access to ubiquitous sensing and data from the web. We study face recognition and show that three distinct properties have surprising effects on the transferability of deep convolutional networks (CNN): (1) The bottleneck of the network serves as an important transfer learning regularizer, and (2) in contrast to the common wisdom, performance saturation may exist in CNN's (as the number of training samples grows); we propose a solution for alleviating this by replacing the naive random subsampling of the training set with a bootstrapping process. Moreover, (3) we find a link between the representation norm and the ability to discriminate in a target domain, which sheds lights on how such networks represent faces. Based on these discoveries, we are able to improve face recognition accuracy on the widely used LFW benchmark, both in the verification (1:1) and identification (1:N) protocols, and directly compare, for the first time, with the state of the art Commercially-Off-The-Shelf system and show a sizable leap in performance. Yaniv Taigman, Ming Yang 0007, Marc'Aurelio Ranzato, Lior Wolf |
CVPR | 3 |
| 2015 | Guest Editorial: Deep Learning
Marc'Aurelio Ranzato, Geoffrey E. Hinton, Yann LeCun |
Int. J. Comput. Vis. | 1 |
| 2014 | DeepFace: Closing the Gap to Human-Level Performance in Face VerificationabstractIn modern face recognition, the conventional pipeline consists of four stages: detect => align => represent => classify. We revisit both the alignment step and the representation step by employing explicit 3D face modeling in order to apply a piecewise affine transformation, and derive a face representation from a nine-layer deep neural network. This deep network involves more than 120 million parameters using several locally connected layers without weight sharing, rather than the standard convolutional layers. Thus we trained it on the largest facial dataset to-date, an identity labeled dataset of four million facial images belonging to more than 4, 000 identities. The learned representations coupling the accurate model-based alignment with the large facial database generalize remarkably well to faces in unconstrained environments, even with a simple classifier. Our method reaches an accuracy of 97.35% on the Labeled Faces in the Wild (LFW) dataset, reducing the error of the current state of the art by more than 27%, closely approaching human-level performance. Yaniv Taigman, Ming Yang 0007, Marc'Aurelio Ranzato, Lior Wolf |
CVPR | 3 |
| 2014 | PANDA: Pose Aligned Networks for Deep Attribute ModelingabstractWe propose a method for inferring human attributes (such as gender, hair style, clothes style, expression, action) from images of people under large variation of viewpoint, pose, appearance, articulation and occlusion. Convolutional Neural Nets (CNN) have been shown to perform very well on large scale object recognition problems. In the context of attribute classification, however, the signal is often subtle and it may cover only a small part of the image, while the image is dominated by the effects of pose and viewpoint. Discounting for pose variation would require training on very large labeled datasets which are not presently available. Part-based models, such as poselets [4] and DPM [12] have been shown to perform well for this problem but they are limited by shallow low-level features. We propose a new method which combines part-based models and deep learning by training pose-normalized CNNs. We show substantial improvement vs. state-of-the-art methods on challenging attribute classification tasks in unconstrained settings. Experiments confirm that our method outperforms both the best part-based methods on this problem and conventional CNNs trained on the full bounding box of the person. Ning Zhang 0014, Manohar Paluri, Marc'Aurelio Ranzato, Trevor Darrell, Lubomir D. Bourdev |
CVPR | 3 |
| 2013 | Multilingual acoustic models using distributed deep neural networksabstractToday's speech recognition technology is mature enough to be useful for many practical applications. In this context, it is of paramount importance to train accurate acoustic models for many languages within given resource constraints such as data, processing power, and time. Multilingual training has the potential to solve the data issue and close the performance gap between resource-rich and resource-scarce languages. Neural networks lend themselves naturally to parameter sharing across languages, and distributed implementations have made it feasible to train large networks. In this paper, we present experimental results for cross- and multi-lingual network training of eleven Romance languages on 10k hours of data in total. The average relative gains over the monolingual baselines are 4%/2% (data-scarce/data-rich languages) for cross- and 7%/2% for multi-lingual training. However, the additional gain from jointly training the languages on all data comes at an increased training time of roughly four weeks, compared to two weeks (monolingual) and one week (crosslingual). Georg Heigold, Vincent Vanhoucke, Andrew W. Senior, Patrick Nguyen, Marc'Aurelio Ranzato, Matthieu Devin, Jeffrey Dean |
ICASSP | 5 |
| 2013 | An empirical study of learning rates in deep neural networks for speech recognitionabstractRecent deep neural network systems for large vocabulary speech recognition are trained with minibatch stochastic gradient descent but use a variety of learning rate scheduling schemes. We investigate several of these schemes, particularly AdaGrad. Based on our analysis of its limitations, we propose a new variant `AdaDec' that decouples long-term learning-rate scheduling from per-parameter learning rate variation. AdaDec was found to result in higher frame accuracies than other methods. Overall, careful choice of learning rate schemes leads to faster convergence and lower word error rates. Andrew W. Senior, Georg Heigold, Marc'Aurelio Ranzato |
ICASSP | 3 |
| 2013 | On rectified linear units for speech processingabstractDeep neural networks have recently become the gold standard for acoustic modeling in speech recognition systems. The key computational unit of a deep network is a linear projection followed by a point-wise non-linearity, which is typically a logistic function. In this work, we show that we can improve generalization and make training of deep networks faster and simpler by substituting the logistic units with rectified linear units. These units are linear when their input is positive and zero otherwise. In a supervised setting, we can successfully train very deep nets from random initialization on a large vocabulary speech recognition task achieving lower word error rates than using a logistic network with the same topology. Similarly in an unsupervised setting, we show how we can learn sparse features that can be useful for discriminative tasks. All our experiments are executed in a distributed environment using several hundred machines and several hundred hours of speech data. Matthew D. Zeiler, Marc'Aurelio Ranzato, Rajat Monga, Mark Z. Mao, Quoc V. Le, Patrick Nguyen, Andrew W. Senior, Vincent Vanhoucke, Jeffrey Dean, Geoffrey E. Hinton |
ICASSP | 2 |
| 2013 | Predicting Parameters in Deep LearningabstractWe demonstrate that there is significant redundancy in the parameterization of several deep learning models. Given only a few weight values for each feature it is possible to accurately predict the remaining values. Moreover, we show that not only can the parameter values be predicted, but many of them need not be learned at all. We train several different architectures by learning only a small number of weights and predicting the rest. In the best case we are able to predict more than 95% of the weights of a network without any drop in accuracy. Misha Denil, Babak Shakibi, Laurent Dinh, Marc'Aurelio Ranzato, Nando de Freitas |
NIPS | 4 |
| 2013 | DeViSE: A Deep Visual-Semantic Embedding ModelabstractModern visual recognition systems are often limited in their ability to scale to large numbers of object categories. This limitation is in part due to the increasing difficulty of acquiring sufficient training data in the form of labeled images as the number of object categories grows. One remedy is to leverage data from other sources -- such as text data -- both to train visual models and to constrain their predictions. In this paper we present a new deep visual-semantic embedding model trained to identify visual objects using both labeled image data as well as semantic information gleaned from unannotated text. We demonstrate that this model matches state-of-the-art performance on the 1000-class ImageNet object recognition challenge while making more semantically reasonable errors, and also show that the semantic information can be exploited to make predictions about tens of thousands of image labels not observed during training. Semantic knowledge improves such zero-shot predictions by up to 65%, achieving hit rates of up to 10% across thousands of novel labels never seen by the visual model. Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc'Aurelio Ranzato, Tomás Mikolov |
NIPS | 6 |
| 2013 | Modeling Natural Images Using Gated MRFsabstractThis paper describes a Markov Random Field for real-valued image modeling that has two sets of latent variables. One set is used to gate the interactions between all pairs of pixels, while the second set determines the mean intensities of each pixel. This is a powerful model with a conditional distribution over the input that is Gaussian, with both mean and covariance determined by the configuration of latent variables, which is unlike previous models that were restricted to using Gaussians with either a fixed mean or a diagonal covariance matrix. Thanks to the increased flexibility, this gated MRF can generate more realistic samples after training on an unconstrained distribution of high-resolution natural images. Furthermore, the latent variables of the model can be inferred efficiently and can be used as very effective descriptors in recognition tasks. Both generation and discrimination drastically improve as layers of binary latent variables are added to the model, yielding a hierarchical model called a Deep Belief Network. Marc'Aurelio Ranzato, Volodymyr Mnih, Joshua M. Susskind, Geoffrey E. Hinton |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Building high-level features using large scale unsupervised learning
Quoc V. Le, Marc'Aurelio Ranzato, Rajat Monga, Matthieu Devin, Gregory S. Corrado, Kai Chen 0010, Jeffrey Dean, Andrew Y. Ng |
ICML | 2 |
| 2012 | Large Scale Distributed Deep NetworksabstractRecent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm. Jeffrey Dean, Gregory S. Corrado, Rajat Monga, Kai Chen 0010, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Andrew Y. Ng |
NIPS | 8 |
| 2011 | On deep generative models with applications to recognitionabstractThe most popular way to use probabilistic models in vision is first to extract some descriptors of small image patches or object parts using well-engineered features, and then to use statistical learning tools to model the dependencies among these features and eventual labels. Learning probabilistic models directly on the raw pixel values has proved to be much more difficult and is typically only used for regularizing discriminative methods. In this work, we use one of the best, pixel-level, generative models of natural images-a gated MRF-as the lowest level of a deep belief network (DBN) that has several hidden layers. We show that the resulting DBN is very good at coping with occlusion when predicting expression categories from face images, and it can produce features that perform comparably to SIFT descriptors for discriminating different types of scene. The generative ability of the model also makes it easy to see what information is captured and what is lost at each level of representation. Marc'Aurelio Ranzato, Joshua M. Susskind, Volodymyr Mnih, Geoffrey E. Hinton |
CVPR | 1 |
| 2011 | On Autoencoders and Score Matching for Energy Based Models
Kevin Swersky, Marc'Aurelio Ranzato, David Buchman, Benjamin M. Marlin, Nando de Freitas |
ICML | 2 |
| 2010 | Modeling pixel means and covariances using factorized third-order boltzmann machinesabstractLearning a generative model of natural images is a useful way of extracting features that capture interesting regularities. Previous work on learning such models has focused on methods in which the latent features are used to determine the mean and variance of each pixel independently, or on methods in which the hidden units determine the covariance matrix of a zero-mean Gaussian distribution. In this work, we propose a probabilistic model that combines these two approaches into a single framework. We represent each image using one set of binary latent features that model the image-specific covariance and a separate set that model the mean. We show that this approach provides a probabilistic framework for the widely used simple-cell complex-cell architecture, it produces very realistic samples of natural images and it extracts features that yield state-of-the-art recognition accuracy on the challenging CIFAR 10 dataset. Marc'Aurelio Ranzato, Geoffrey E. Hinton |
CVPR | 1 |
| 2010 | Phone Recognition with the Mean-Covariance Restricted Boltzmann MachineabstractStraightforward application of Deep Belief Nets (DBNs) to acoustic modeling produces a rich distributed representation of speech data that is useful for recognition and yields impressive results on the speaker-independent TIMIT phone recognition task. However, the first-layer Gaussian-Bernoulli Restricted Boltzmann Machine (GRBM) has an important limitation, shared with mixtures of diagonal-covariance Gaussians: GRBMs treat different components of the acoustic input vector as conditionally independent given the hidden state. The mean-covariance restricted Boltzmann machine (mcRBM), first introduced for modeling natural images, is a much more representationally efficient and powerful way of modeling the covariance structure of speech data. Every configuration of the precision units of the mcRBM specifies a different precision matrix for the conditional distribution over the acoustic space. In this work, we use the mcRBM to learn features of speech data that serve as input into a standard DBN. The mcRBM features combined with DBNs allow us to achieve a phone error rate of 20.5\%, which is superior to all published results on speaker-independent TIMIT to date. George E. Dahl, Marc'Aurelio Ranzato, Abdel-rahman Mohamed, Geoffrey E. Hinton |
NIPS | 2 |
| 2010 | Generating more realistic images using gated MRF'sabstractProbabilistic models of natural images are usually evaluated by measuring performance on rather indirect tasks, such as denoising and inpainting. A more direct way to evaluate a generative model is to draw samples from it and to check whether statistical properties of the samples match the statistics of natural images. This method is seldom used with high-resolution images, because current models produce samples that are very different from natural images, as assessed by even simple visual inspection. We investigate the reasons for this failure and we show that by augmenting existing models so that there are two sets of latent variables, one set modelling pixel intensities and the other set modelling image-specific pixel covariances, we are able to generate high-resolution images that look much more realistic than before. The overall model can be interpreted as a gated MRF where both pair-wise dependencies and mean intensities of pixels are modulated by the states of latent variables. Finally, we confirm that if we disallow weight-sharing between receptive fields that overlap each other, the gated MRF learns more efficient internal representations, as demonstrated in several recognition tasks. Marc'Aurelio Ranzato, Volodymyr Mnih, Geoffrey E. Hinton |
NIPS | 1 |
| 2009 | Learning invariant features through topographic filter mapsabstractSeveral recently-proposed architectures for high-performance object recognition are composed of two main stages: a feature extraction stage that extracts locally-invariant feature vectors from regularly spaced image patches, and a somewhat generic supervised classifier. The first stage is often composed of three main modules: (1) a bank of filters (often oriented edge detectors); (2) a non-linear transform, such as a point-wise squashing functions, quantization, or normalization; (3) a spatial pooling operation which combines the outputs of similar filters over neighboring regions. We propose a method that automatically learns such feature extractors in an unsupervised fashion by simultaneously learning the filters and the pooling units that combine multiple filter outputs together. The method automatically generates topographic maps of similar filters that extract features of orientations, scales, and positions. These similar filters are pooled together, producing locally-invariant outputs. The learned feature descriptors give comparable results as SIFT on image recognition tasks for which SIFT is well suited, and better results than SIFT on tasks for which SIFT is less well suited. Koray Kavukcuoglu, Marc'Aurelio Ranzato, Rob Fergus, Yann LeCun |
CVPR | 2 |
| 2009 | What is the best multi-stage architecture for object recognition?abstractIn many recent object recognition systems, feature extraction stages are generally composed of a filter bank, a non-linear transformation, and some sort of feature pooling layer. Most systems use only one stage of feature extraction in which the filters are hard-wired, or two stages where the filters in one or both stages are learned in supervised or unsupervised mode. This paper addresses three questions: 1. How does the non-linearities that follow the filter banks influence the recognition accuracy? 2. does learning the filter banks in an unsupervised or supervised manner improve the performance over random filters or hardwired filters? 3. Is there any advantage to using an architecture with two stages of feature extraction, rather than one? We show that using non-linearities that include rectification and local contrast normalization is the single most important ingredient for good accuracy on object recognition benchmarks. We show that two stages of feature extraction yield better accuracy than one. Most surprisingly, we show that a two-stage system with random filters can yield almost 63% recognition rate on Caltech-101, provided that the proper non-linearities and pooling layers are used. Finally, we show that with supervised refinement, the system achieves state-of-the-art performance on NORB dataset (5.6%) and unsupervised pre-training followed by supervised refinement produces good accuracy on Caltech-101 (> 65%), and the lowest known error rate on the undistorted, unprocessed MNIST dataset (0.53%). Kevin Jarrett, Koray Kavukcuoglu, Marc'Aurelio Ranzato, Yann LeCun |
ICCV | 3 |
| 2008 | Semi-supervised learning of compact document representations with deep networksabstractFinding good representations of text documents is crucial in information retrieval and classification systems. Today the most popular document representation is based on a vector of word counts in the document. This representation neither captures dependencies between related words, nor handles synonyms or polysemous words. In this paper, we propose an algorithm to learn text document representations based on semi-supervised autoencoders that are stacked to form a deep network. The model can be trained efficiently on partially labeled corpora, producing very compact representations of documents, while retaining as much class information and joint word statistics as possible. We show that it is advantageous to exploit even a few labeled samples during training. Marc'Aurelio Ranzato, Martin Szummer |
ICML | 1 |
| 2007 | Unsupervised Learning of Invariant Feature Hierarchies with Applications to Object RecognitionabstractWe present an unsupervised method for learning a hierarchy of sparse feature detectors that are invariant to small shifts and distortions. The resulting feature extractor consists of multiple convolution filters, followed by a feature-pooling layer that computes the max of each filter output within adjacent windows, and a point-wise sigmoid non-linearity. A second level of larger and more invariant features is obtained by training the same algorithm on patches of features from the first level. Training a supervised classifier on these features yields 0.64% error on MNIST, and 54% average recognition rate on Caltech 101 with 30 training samples per category. While the resulting architecture is similar to convolutional networks, the layer-wise unsupervised training procedure alleviates the over-parameterization problems that plague purely supervised learning procedures, and yields good performance with very few labeled training samples. Marc'Aurelio Ranzato, Fu Jie Huang, Y-Lan Boureau, Yann LeCun |
CVPR | 1 |
| 2007 | Energy-Based Models in Document Recognition and Computer VisionabstractThe machine learning and pattern recognition communities are facing two challenges: solving the normalization problem, and solving the deep learning problem. The normalization problem is related to the difficulty of training probabilistic models over large spaces while keeping them properly normalized. In recent years, the ML and natural language communities have devoted considerable efforts to circumventing this problem by developing "un-normalized" learning models for tasks in which the output is highly structured (e.g. English sentences). This class of models was in fact originally developed during the 90's in the handwriting recognition community, and includes graph transformer networks, conditional random fields, hidden Markov SVMs, and maximum margin Markov networks. We describe these models within the unifying framework of "energy-based models" (EBM). The deep learning problem is related to the issue of training all the levels of a recognition system (e.g. segmentation, feature extraction, recognition, etc) in an integrated fashion. We first consider " traditional" methods for deep learning, such as convolutional networks and back-propagation, and show that, although they produce very low error rates for handwriting and object recognition, they require many training samples. We show that using unsupervised learning to initialize the layers of a deep network dramatically reduces the required number of training samples, particularly for such tasks as the recognition of everyday objects at the category level. Yann LeCun, Sumit Chopra, Marc'Aurelio Ranzato, Fu Jie Huang |
ICDAR | 3 |
| 2007 | A Sparse and Locally Shift Invariant Feature Extractor Applied to Document ImagesabstractWe describe an unsupervised learning algorithm for extracting sparse and locally shift-invariant features. We also devise a principled procedure for learning hierarchies of invariant features. Each feature detector is composed of a set of trainable convolutional filters followed by a max-pooling layer over non-overlapping windows, and a point-wise sigmoid non-linearity. A second stage of more invariant features is fed with patches provided by the first stage feature extractor, and is trained in the same way. The method is used to pre-train the first four layers of a deep convolutional network which achieves state-of-the-art performance on the MNIST dataset of handwritten digits. The final testing error rate is equal to 0.42%. Preliminary experiments on compression of bitonal document images show very promising results in terms of compression ratio and reconstruction error. Marc'Aurelio Ranzato, Yann LeCun |
ICDAR | 1 |
| 2007 | Sparse Feature Learning for Deep Belief NetworksabstractUnsupervised learning algorithms aim to discover the structure hidden in the data, and to learn representations that are more suitable as input to a supervised machine than the raw input. Many unsupervised methods are based on reconstructing the input from the representation, while constraining the representation to have certain desirable properties (e.g. low dimension, sparsity, etc). Others are based on approximating density by stochastically reconstructing the input from the representation. We describe a novel and efficient algorithm to learn sparse representations, and compare it theoretically and experimentally with a similar machines trained probabilistically, namely a Restricted Boltzmann Machine. We propose a simple criterion to compare and select different unsupervised machines based on the trade-off between the reconstruction error and the information content of the representation. We demonstrate this method by extracting features from a dataset of handwritten numerals, and from a dataset of natural image patches. We show that by stacking multiple levels of such machines and by training sequentially, high-order dependencies between the input variables can be captured. Marc'Aurelio Ranzato, Y-Lan Boureau, Yann LeCun |
NIPS | 1 |
| 2007 | Automatic recognition of biological particles in microscopic images
Marc'Aurelio Ranzato, P. E. Taylor, James M. House, R. C. Flagan, Yann LeCun, Pietro Perona |
Pattern Recognit. Lett. | 1 |
| 2006 | Efficient Learning of Sparse Representations with an Energy-Based ModelabstractWe describe a novel unsupervised method for learning sparse, overcomplete features. The model uses a linear encoder, and a linear decoder preceded by a sparsifying non-linearity that turns a code vector into a quasi-binary sparse code vector. Given an input, the optimal code minimizes the distance between the output of the decoder and the input patch while being as similar as possible to the encoder output. Learning proceeds in a two-phase EM-like fashion: (1) compute the minimum-energy code vector, (2) adjust the parameters of the encoder and decoder so as to decrease the energy. The model produces "stroke detectors" when trained on handwritten numerals, and Gabor-like filters when trained on natural image patches. Inference and learning are very fast, requiring no preprocessing, and no expensive sampling. Using the proposed unsupervised method to initialize the first layer of a convolutional network, we achieved an error rate slightly lower than the best reported result on the MNIST dataset. Finally, an extension of the method is described to learn topographical filter maps. Marc'Aurelio Ranzato, Christopher S. Poultney, Sumit Chopra, Yann LeCun |
NIPS | 1 |