Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mikolaj Binkowski

dblp:198/0887 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 27% Language models and text generation · 21% Learning theory · 12%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
1.242021
High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020
On gradient regularizers for MMD GANs · NeurIPS 2018
Demystifying MMD GANs · ICLR (Poster) 2018
Machine learning › Learning theory › probability metric › integral probability metric
maximum mean discrepancy
0.722018
On gradient regularizers for MMD GANs · NeurIPS 2018
Demystifying MMD GANs · ICLR (Poster) 2018
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Machine learning › Generative modeling
diffusion model
0.612022
Step-unrolled Denoising Autoencoders for Text Generation · ICLR 2022
Natural language and speech › Language models and text generation
in-context learning
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Natural language and speech › Language models and text generation
multimodal language model
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Natural language and speech › Language models and text generation
text generation
0.612022
Step-unrolled Denoising Autoencoders for Text Generation · ICLR 2022
Computer vision › Vision and language
vision-language model
0.612022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.512021
End-to-end Adversarial Text-to-Speech · ICLR 2021
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.412020
High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.412019
Batch Weight for Domain Adaptation With Mass Shift · ICCV 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.312018
Autoregressive Convolutional Neural Networks for Asynchronous Time Series · ICML 2018
Machine learning › Learning theory › probability metric
integral probability metric
0.312018
Demystifying MMD GANs · ICLR (Poster) 2018
Machine learning › Generative modeling › generative adversarial network
MMD-GAN
0.312018
Demystifying MMD GANs · ICLR (Poster) 2018
Computer vision › Vision and language
visual question answering
0.212022
Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.112019
Batch Weight for Domain Adaptation With Mass Shift · ICCV 2019

Methods — techniques the papers use, named apart from their topics

adversarial training · 0.9step unrolling · 0.6pretrained vision encoder · 0.6pre-trained language model · 0.6interleaved multimodal pretraining · 0.6denoising autoencoder · 0.6adversarial network · 0.4sample reweighting · 0.4cycle consistency · 0.4generative adversarial network · 0.3
YearPublicationVenuePosition
2022 Step-unrolled Denoising Autoencoders for Text Generation
Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, Aäron van den Oord
ICLR3
2022 Flamingo: a Visual Language Model for Few-Shot Learning
abstract
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan
NeurIPS23
2021 End-to-end Adversarial Text-to-Speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, Karen Simonyan
ICLR3
2020 High Fidelity Speech Synthesis with Adversarial Networks
Mikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, Karen Simonyan
ICLR1
2019 Batch Weight for Domain Adaptation With Mass Shift
abstract
Unsupervised domain transfer is the task of transferring or translating samples from a source distribution to a different target distribution. Current solutions unsupervised domain transfer often operate on data on which the modes of the distribution are well-matched, for instance have the same frequencies of classes between source and target distributions. However, these models do not perform well when the modes are not well-matched, as would be the case when samples are drawn independently from two different, but related, domains. This mode imbalance is problematic as generative adversarial networks (GANs), a successful approach in this setting, are sensitive to mode frequency, which results in a mismatch of semantics between source samples and generated samples of the target distribution. We propose a principled method of re-weighting training samples to correct for such mass shift between the transferred distributions, which we call batch weight. We also provide rigorous probabilistic setting for domain transfer and new simplified objective for training transfer networks, an alternative to complex, multi-component loss functions used in the current state-of-the art image-to-image translation models. The new objective stems from the discrimination of joint distributions and enforces cycle-consistency in an abstract, high-level, rather than pixel-wise, sense. Lastly, we experimentally show the effectiveness of the proposed methods in several image-to-image translation tasks.
Mikolaj Binkowski, R. Devon Hjelm, Aaron C. Courville
ICCV1
2018 Demystifying MMD GANs
Mikolaj Binkowski, Danica J. Sutherland, Michael Arbel, Arthur Gretton
ICLR (Poster)1
2018 Autoregressive Convolutional Neural Networks for Asynchronous Time Series
abstract
We propose Significance-Offset Convolutional Neural Network, a deep convolutional network architecture for regression of multivariate asynchronous time series. The model is inspired by standard autoregressive (AR) models and gating mechanisms used in recurrent neural networks. It involves an AR-like weighting system, where the final predictor is obtained as a weighted sum of adjusted regressors, while the weights are data-dependent functions learnt through a convolutional network. The architecture was designed for applications on asynchronous time series and is evaluated on such datasets: a hedge fund proprietary dataset of over 2 million quotes for a credit derivative index, an artificially generated noisy autoregressive series and UCI household electricity consumption dataset. The proposed architecture achieves promising results as compared to convolutional and recurrent neural networks.
Mikolaj Binkowski, Gautier Marti, Philippe Donnat
ICML1
2018 On gradient regularizers for MMD GANs
abstract
We propose a principled method for gradient-based regularization of the critic of GAN-like models trained by adversarially optimizing the kernel of a Maximum Mean Discrepancy (MMD). We show that controlling the gradient of the critic is vital to having a sensible loss function, and devise a method to enforce exact, analytical gradient constraints at no additional cost compared to existing approximate techniques based on additive regularizers. The new loss function is provably continuous, and experiments show that it stabilizes and accelerates training, giving image generation models that outperform state-of-the art methods on $160 \times 160$ CelebA and $64 \times 64$ unconditional ImageNet.
Michael Arbel, Danica J. Sutherland, Mikolaj Binkowski, Arthur Gretton
NeurIPS3