Yaroslav Ganin

dblp:147/5336 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 47% Face, body and person analysis · 18% Transfer learning and domain adaptation · 13%
Computer graphics and multimedia
5 papers
Geometric modeling and processing · 45% Visual content generation and editing · 42% Image and video processing · 13%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
0.922021
Computer-Aided Design as Language · NeurIPS 2021
PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020
Computer vision › Face, body and person analysis
gaze analysis
0.622018
Photorealistic Monocular Gaze Redirection Using Machine Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018
DeepWarp: Photorealistic Image Resynthesis for Gaze Manipulation · ECCV (2) 2016
Machine learning › Generative modeling › diffusion model
conditional generation
0.512021
Computer-Aided Design as Language · NeurIPS 2021
Machine learning › Generative modeling › image generation
sketch generation
0.512021
Computer-Aided Design as Language · NeurIPS 2021
Machine learning › Generative modeling › 3d generative model
mesh generative model
0.412020
PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020
Geometric modeling and processing › mesh generation
autoregressive mesh generation
0.412020
PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020
Geometric modeling and processing
mesh generation
0.412020
PolyGen: An Autoregressive Generative Model of 3D Meshes · ICML 2020
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.312018
Synthesizing Programs for Images using Reinforced Adversarial Learning · ICML 2018
Computer vision › Face, body and person analysis › gaze estimation
gaze redirection
0.312018
Photorealistic Monocular Gaze Redirection Using Machine Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › 3D vision
inverse rendering
0.312018
Synthesizing Programs for Images using Reinforced Adversarial Learning · ICML 2018
Visual content generation and editing › face editing
gaze redirection
0.312018
Photorealistic Monocular Gaze Redirection Using Machine Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Machine learning › Generative modeling
adversarial inference
0.312017
GibbsNet: Iterative Adversarial Inference for Deep Graphical Models · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.312017
GibbsNet: Iterative Adversarial Inference for Deep Graphical Models · NIPS 2017
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation
0.212016
Domain-Adversarial Training of Neural Networks · J. Mach. Learn. Res. 2016
Machine learning › Transfer learning and domain adaptation
domain-invariant representation learning
0.212016
Domain-Adversarial Training of Neural Networks · J. Mach. Learn. Res. 2016
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation
0.212015
Unsupervised Domain Adaptation by Backpropagation · ICML 2015
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.212015
Unsupervised Domain Adaptation by Backpropagation · ICML 2015
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
gibbs sampling
0.112017
GibbsNet: Iterative Adversarial Inference for Deep Graphical Models · NIPS 2017
Computer vision › Face, body and person analysis
person re-identification
0.112016
Domain-Adversarial Training of Neural Networks · J. Mach. Learn. Res. 2016

Methods — techniques the papers use, named apart from their topics

language modeling · 1.0data serialization · 1.0transformer · 0.9autoregressive modeling · 0.9reinforcement learning · 0.7discriminator reward · 0.7deep neural network · 0.7gradient reversal layer · 0.5backpropagation · 0.5knowledge distillation · 0.3decision forests · 0.3decision forest · 0.3deep warping · 0.2
YearPublicationVenuePosition
2021 Computer-Aided Design as Language
abstract
Computer-Aided Design (CAD) applications are used in manufacturing to model everything from coffee mugs to sports cars. These programs are complex and require years of training and experience to master. A component of all CAD models particularly difficult to make are the highly structured 2D sketches that lie at the heart of every 3D construction. In this work, we propose a machine learning model capable of automatically generating such sketches. Through this, we pave the way for developing intelligent tools that would help engineers create better designs with less effort. The core of our method is a combination of a general-purpose language modeling technique alongside an off-the-shelf data serialization protocol. Additionally, we explore several extensions allowing us to gain finer control over the generation process. We show that our approach has enough flexibility to accommodate the complexity of the domain and performs well for both unconditional synthesis and image-to-sketch translation.
Yaroslav Ganin, Sergey Bartunov, Ethan Keller, Stefano Saliceti
NeurIPS1
2020 PolyGen: An Autoregressive Generative Model of 3D Meshes
abstract
Polygon meshes are an efficient representation of 3D geometry, and are of central importance in computer graphics, robotics and games development. Existing learning-based approaches for object synthesis have avoided the challenges of working with 3D meshes, instead using alternative object representations that are more compatible with neural architectures and training approaches. We present PolyGen, a generative model of 3D objects which models the mesh directly, predicting vertices and faces sequentially using a Transformer-based architecture. Our model can condition on a range of inputs, including object classes, voxels, and images, and because the model is probabilistic it can produce samples that capture uncertainty in ambiguous scenarios. We show that the model is capable of producing high-quality, usable meshes, and establish log-likelihood benchmarks for the mesh-modelling task. We also evaluate the conditional models on surface reconstruction metrics against alternative methods, and demonstrate competitive performance despite not training directly on this task.
Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. Battaglia
ICML2
2018 Synthesizing Programs for Images using Reinforced Adversarial Learning
abstract
Advances in deep generative networks have led to impressive results in recent years. Nevertheless, such models can often waste their capacity on the minutiae of datasets, presumably due to weak inductive biases in their decoders. This is where graphics engines may come in handy since they abstract away low-level details and represent images as high-level programs. Current methods that combine deep learning and renderers are limited by hand-crafted likelihood or distance functions, a need for large amounts of supervision, or difficulties in scaling their inference algorithms to richer datasets. To mitigate these issues, we present SPIRAL, an adversarially trained agent that generates a program which is executed by a graphics engine to interpret and sample images. The goal of this agent is to fool a discriminator network that distinguishes between real and rendered data, trained with a distributed reinforcement learning setup without any supervision. A surprising finding is that using the discriminator’s output as a reward signal is the key to allow the agent to make meaningful progress at matching the desired output rendering. To the best of our knowledge, this is the first demonstration of an end-to-end, unsupervised and adversarial inverse graphics agent on challenging real world (MNIST, Omniglot, CelebA) and synthetic 3D datasets. A video of the agent can be found at https://youtu.be/iSyvwAwa7vk.
Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, S. M. Ali Eslami, Oriol Vinyals
ICML1
2018 Photorealistic Monocular Gaze Redirection Using Machine Learning
abstract
We propose a general approach to the gaze redirection problem in images that utilizes machine learning. The idea is to learn to re-synthesize images by training on pairs of images with known disparities between gaze directions. We show that such learning-based re-synthesis can achieve convincing gaze redirection based on monocular input, and that the learned systems generalize well to people and imaging conditions unseen during training. We describe and compare three instantiations of our idea. The first system is based on efficient decision forest predictors and redirects the gaze by a fixed angle in real-time (on a single CPU), being particularly suitable for the videoconferencing gaze correction. The second system is based on a deep architecture and allows gaze redirection by a range of angles. The second system achieves higher photorealism, while being several times slower. The third system is based on real-time decision forests at test time, while using the supervision from a "teacher" deep network during training. The third system approaches the quality of a teacher network in our experiments, and thus provides a highly realistic real-time monocular solution to the gaze correction problem. We present in-depth assessment and comparisons of the proposed systems based on quantitative measurements and a user study.
Daniil Kononenko, Yaroslav Ganin, Diana Sungatullina, Victor S. Lempitsky
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Multi-Region bilinear convolutional neural networks for person re-identification
abstract
In this work we propose a new architecture for person re-identification. As the task of re-identification is inherently associated with embedding learning and non-rigid appearance description, our architecture is based on the deep bilinear convolutional network (Bilinear-CNN) that has been proposed recently for fine-grained classification of highly non-rigid objects. While the last stages of the original Bilinear-CNN architecture completely removes the geometric information from consideration by performing orderless pooling, we observe that a better embedding can be learned by performing bilinear pooling in a more local way, where each pooling is confined to a predefined region. Our architecture thus represents a compromise between traditional convolutional networks and bilinear CNNs and strikes a balance between rigid matching and completely ignoring spatial information. We perform the experimental validation of the new architecture on the three popular benchmark datasets (Market-1501, CUHK01, CUHK03), comparing it to baselines that include Bilinear-CNN as well as prior art. The new architecture outperforms the baseline on all three datasets, while performing better than state-of-the-art on two out of three. The code and the pretrained models of the approach will be made available at the time of publication.
Evgeniya Ustinova, Yaroslav Ganin, Victor S. Lempitsky
AVSS2
2017 GibbsNet: Iterative Adversarial Inference for Deep Graphical Models
abstract
Directed latent variable models that formulate the joint distribution as $p(x,z) = p(z) p(x \mid z)$ have the advantage of fast and exact sampling. However, these models have the weakness of needing to specify $p(z)$, often with a simple fixed prior that limits the expressiveness of the model. Undirected latent variable models discard the requirement that $p(z)$ be specified with a prior, yet sampling from them generally requires an iterative procedure such as blocked Gibbs-sampling that may require many steps to draw samples from the joint distribution $p(x, z)$. We propose a novel approach to learning the joint distribution between the data and a latent code which uses an adversarially learned iterative procedure to gradually refine the joint distribution, $p(x, z)$, to better match with the data distribution on each step. GibbsNet is the best of both worlds both in theory and in practice. Achieving the speed and simplicity of a directed latent variable model, it is guaranteed (assuming the adversarial game reaches the virtual training criteria global minimum) to produce samples from $p(x, z)$ with only a few sampling iterations. Achieving the expressiveness and flexibility of an undirected latent variable model, GibbsNet does away with the need for an explicit $p(z)$ and has the ability to do attribute prediction, class-conditional generation, and joint image-attribute modeling in a single model which is not trained for any of these specific tasks. We show empirically that GibbsNet is able to learn a more complex $p(z)$ and show that this leads to improved inpainting and iterative refinement of $p(x, z)$ for dozens of steps and stable generation without collapse for thousands of steps, despite being trained on only a few steps.
Alex Lamb, R. Devon Hjelm, Yaroslav Ganin, Joseph Paul Cohen, Aaron C. Courville, Yoshua Bengio
NIPS3
2016 DeepWarp: Photorealistic Image Resynthesis for Gaze Manipulation
Yaroslav Ganin, Daniil Kononenko, Diana Sungatullina, Victor S. Lempitsky
ECCV (2)1
2016 Domain-Adversarial Training of Neural Networks
abstract
We introduce a new representation learning approach for domain adaptation, in which data at training and test time come from similar but different distributions. Our approach is directly inspired by the theory on domain adaptation suggesting that, for effective domain transfer to be achieved, predictions must be made based on features that cannot discriminate between the training (source) and test (target) domains. The approach implements this idea in the context of neural network architectures that are trained on labeled data from the source domain and unlabeled data from the target domain (no labeled target-domain data is necessary). As the training progresses, the approach promotes the emergence of features that are (i) discriminative for the main learning task on the source domain and (ii) indiscriminate with respect to the shift between the domains. We show that this adaptation behaviour can be achieved in almost any feed-forward model by augmenting it with few standard layers and a new gradient reversal layer. The resulting augmented architecture can be trained using standard backpropagation and stochastic gradient descent, and can thus be implemented with little effort using any of the deep learning packages. We demonstrate the success of our approach for two distinct classification problems (document sentiment analysis and image classification), where state-of-the-art domain adaptation performance on standard benchmarks is achieved. We also validate the approach for descriptor learning task in the context of person re-identification application.
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, Victor S. Lempitsky
J. Mach. Learn. Res.1
2015 Unsupervised Domain Adaptation by Backpropagation
abstract
Top-performing deep architectures are trained on massive amounts of labeled data. In the absence of labeled data for a certain task, domain adaptation often provides an attractive option given that labeled data of similar nature but from a different domain (e.g. synthetic images) are available. Here, we propose a new approach to domain adaptation in deep architectures that can be trained on large amount of labeled data from the source domain and large amount of unlabeled data from the target domain (no labeled target-domain data is necessary). As the training progresses, the approach promotes the emergence of "deep" features that are (i) discriminative for the main learning task on the source domain and (ii) invariant with respect to the shift between the domains. We show that this adaptation behaviour can be achieved in almost any feed-forward model by augmenting it with few standard layers and a simple new gradient reversal layer. The resulting augmented architecture can be trained using standard backpropagation. Overall, the approach can be implemented with little effort using any of the deep-learning packages. The method performs very well in a series of image classification experiments, achieving adaptation effect in the presence of big domain shifts and outperforming previous state-of-the-art on Office datasets.
Yaroslav Ganin, Victor S. Lempitsky
ICML1
2014 N^4 -Fields: Neural Network Nearest Neighbor Fields for Image Transforms
Yaroslav Ganin, Victor S. Lempitsky
ACCV (2)1