Michaël Mathieu

dblp:47/9509 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 27% Representation and self-supervised learning · 26% Generative modeling · 26%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
energy-based model
0.312017
Energy-based Generative Adversarial Networks · ICLR (Poster) 2017
Machine learning › Generative modeling
generative adversarial network
0.312017
Energy-based Generative Adversarial Networks · ICLR (Poster) 2017
Machine learning › Deep learning architectures and training › autoencoder
adversarial autoencoder
0.212016
Disentangling factors of variation in deep representation using adversarial training · NIPS 2016
Machine learning › Deep learning architectures and training
autoencoder
0.212016
Disentangling factors of variation in deep representation using adversarial training · NIPS 2016
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.212016
Disentangling factors of variation in deep representation using adversarial training · NIPS 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent generative model
0.212015
Learning to Linearize Under Uncertainty · NIPS 2015
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.212015
Learning to Linearize Under Uncertainty · NIPS 2015
Computer vision › Video understanding and tracking
video prediction
0.212015
Learning to Linearize Under Uncertainty · NIPS 2015
Machine learning › Deep learning architectures and training
convolutional neural network
0.112010
Learning Convolutional Feature Hierarchies for Visual Recognition · NIPS 2010
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
convolutional sparse coding
0.112010
Learning Convolutional Feature Hierarchies for Visual Recognition · NIPS 2010
Computer vision › Image recognition and object detection
visual recognition
0.012010
Learning Convolutional Feature Hierarchies for Visual Recognition · NIPS 2010

Methods — techniques the papers use, named apart from their topics

adversarial training · 0.5energy-based model · 0.3convolutional autoencoder · 0.2latent variable modeling · 0.2unsupervised learning · 0.1sparse coding · 0.1feed-forward encoder · 0.1
YearPublicationVenuePosition
2017 Energy-based Generative Adversarial Networks
Junbo Jake Zhao, Michaël Mathieu, Yann LeCun
ICLR (Poster)2
2016 Disentangling factors of variation in deep representation using adversarial training
abstract
We propose a deep generative model for learning to distill the hidden factors of variation within a set of labeled observations into two complementary codes. One code describes the factors of variation relevant to solving a specified task. The other code describes the remaining factors of variation that are irrelevant to solving this task. The only available source of supervision during the training process comes from our ability to distinguish among different observations belonging to the same category. Concrete examples include multiple images of the same object from different viewpoints, or multiple speech samples from the same speaker. In both of these instances, the factors of variation irrelevant to classification are implicitly expressed by intra-class variabilities, such as the relative position of an object in an image, or the linguistic content of an utterance. Most existing approaches for solving this problem rely heavily on having access to pairs of observations only sharing a single factor of variation, e.g. different objects observed in the exact same conditions. This assumption is often not encountered in realistic settings where data acquisition is not controlled and labels for the uninformative components are not available. In this work, we propose to overcome this limitation by augmenting deep convolutional autoencoders with a form of adversarial training. Both factors of variation are implicitly captured in the organization of the learned embedding space, and can be used for solving single-image analogies. Experimental results on synthetic and real datasets show that the proposed method is capable of disentangling the influences of style and content factors using a flexible representation, as well as generalizing to unseen styles or content classes.
Michaël Mathieu, Junbo Jake Zhao, Pablo Sprechmann, Aditya Ramesh, Yann LeCun
NIPS1
2015 The Loss Surfaces of Multilayer Networks
abstract
We study the connection between the highly non-convex loss function of a simple model of the fully-connected feed-forward neural network and the Hamiltonian of the spherical spin-glass model under the assumptions of: i) variable independence, ii) redundancy in network parametrization, and iii) uniformity. These assumptions enable us to explain the complexity of the fully decoupled neural network through the prism of the results from random matrix theory. We show that for large-size decoupled networks the lowest critical values of the random loss function form a layered structure and they are located in a well-defined band lower-bounded by the global minimum. The number of local minima outside that band diminishes exponentially with the size of the network. We empirically verify that the mathematical model exhibits similar behavior as the computer simulations, despite the presence of high dependencies in real networks. We conjecture that both simulated annealing and SGD converge to the band of low critical points, and that all critical points found there are local minima of high quality measured by the test error. This emphasizes a major difference between large- and small-size networks where for the latter poor quality local minima have non-zero probability of being recovered. Finally, we prove that recovering the global minimum becomes harder as the network size increases and that it is in practice irrelevant as global minimum often leads to overfitting.
Anna Choromanska, Mikael Henaff, Michaël Mathieu, Gérard Ben Arous, Yann LeCun
AISTATS3
2015 Learning to Linearize Under Uncertainty
abstract
Training deep feature hierarchies to solve supervised learning tasks has achieving state of the art performance on many problems in computer vision. However, a principled way in which to train such hierarchies in the unsupervised setting has remained elusive. In this work we suggest a new architecture and loss for training deep feature hierarchies that linearize the transformations observed in unlabelednatural video sequences. This is done by training a generative model to predict video frames. We also address the problem of inherent uncertainty in prediction by introducing a latent variables that are non-deterministic functions of the input into the network architecture.
Ross Goroshin, Michaël Mathieu, Yann LeCun
NIPS2
2010 Learning Convolutional Feature Hierarchies for Visual Recognition
abstract
We propose an unsupervised method for learning multi-stage hierarchies of sparse convolutional features. While sparse coding has become an increasingly popular method for learning visual features, it is most often trained at the patch level. Applying the resulting filters convolutionally results in highly redundant codes because overlapping patches are encoded in isolation. By training convolutionally over large image windows, our method reduces the redudancy between feature vectors at neighboring locations and improves the efficiency of the overall representation. In addition to a linear decoder that reconstructs the image from sparse features, our method trains an efficient feed-forward encoder that predicts quasi-sparse features from the input. While patch-based training rarely produces anything but oriented edge detectors, we show that convolutional training produces highly diverse filters, including center-surround filters, corner detectors, cross detectors, and oriented grating detectors. We show that using these filters in multi-stage convolutional network architecture improves performance on a number of visual recognition and detection tasks.
Koray Kavukcuoglu, Pierre Sermanet, Y-Lan Boureau, Karol Gregor, Michaël Mathieu, Yann LeCun
NIPS5