EDBT 2026 Demo / reviewers in the wild / expert
Kevin Swersky
dblp:35/9381
· DBLP profile ↗
37ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
28 papers |
Generative modeling · 17% Optimization for machine learning · 14% Graph learning · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 30 heaviest of 79, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
1.6 | 5 | 2024 | Pre-trained Gaussian Processes for Bayesian Optimization · J. Mach. Learn. Res. 2024 Taking the Human Out of the Loop: A Review of Bayesian Optimization · Proc. IEEE 2016 Scalable Bayesian Optimization Using Deep Neural Networks · ICML 2015 |
Machine learning › Generative modeling
energy-based model |
1.4 | 5 | 2021 | No MCMC for me: Amortized sampling for fast and stable training of energy-based models · ICLR 2021 Your classifier is secretly an energy based model and you should treat it like one · ICLR 2020 Oops I Took A Gradient: Scalable Sampling for Discrete Distributions · ICML 2021 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
1.1 | 3 | 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020 Meta-Learning for Semi-Supervised Few-Shot Classification · ICLR (Poster) 2018 Prototypical Networks for Few-shot Learning · NIPS 2017 |
Machine learning › Graph learning
graph neural network |
1.0 | 2 | 2022 | Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks · ICDM 2022 Graph Normalizing Flows · NeurIPS 2019 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.9 | 3 | 2024 | Pre-trained Gaussian Processes for Bayesian Optimization · J. Mach. Learn. Res. 2024 Scalable Bayesian Optimization Using Deep Neural Networks · ICML 2015 Multi-Task Bayesian Optimization · NIPS 2013 |
Memory systems › memory access optimization
neural prefetching |
0.8 | 2 | 2021 | A hierarchical neural model of data prefetching · ASPLOS 2021 Learning Memory Access Patterns · ICML 2018 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.8 | 2 | 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples · ICLR 2020 Meta-Learning for Semi-Supervised Few-Shot Classification · ICLR (Poster) 2018 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Directly Fine-Tuning Diffusion Models on Differentiable Rewards · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process prior |
0.8 | 1 | 2024 | Pre-trained Gaussian Processes for Bayesian Optimization · J. Mach. Learn. Res. 2024 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
reward fine-tuning |
0.8 | 1 | 2024 | Directly Fine-Tuning Diffusion Models on Differentiable Rewards · ICLR 2024 |
Image and video processing › super-resolution › image super-resolution
arbitrary-scale super-resolution |
0.7 | 1 | 2023 | CUF: Continuous Upsampling Filters · CVPR 2023 |
Image and video processing › super-resolution
image super-resolution |
0.7 | 1 | 2023 | CUF: Continuous Upsampling Filters · CVPR 2023 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.6 | 1 | 2022 | Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks · ICDM 2022 |
Machine learning › Graph learning › graph neural network
heterophily |
0.6 | 1 | 2022 | Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks · ICDM 2022 |
Machine learning › Graph learning › graph neural network › deep graph neural network
over-smoothing |
0.6 | 1 | 2022 | Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural Networks · ICDM 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.5 | 2 | 2019 | Flexibly Fair Representation Learning by Disentanglement · ICML 2019 Learning Fair Representations · ICML (3) 2013 |
Machine learning › Trustworthy machine learning › fairness
fair representation learning |
0.5 | 2 | 2019 | Flexibly Fair Representation Learning by Disentanglement · ICML 2019 Learning Fair Representations · ICML (3) 2013 |
Machine learning › Generative modeling
amortized sampling |
0.5 | 1 | 2021 | No MCMC for me: Amortized sampling for fast and stable training of energy-based models · ICLR 2021 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
discrete sampling |
0.5 | 1 | 2021 | Oops I Took A Gradient: Scalable Sampling for Discrete Distributions · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.5 | 1 | 2021 | Oops I Took A Gradient: Scalable Sampling for Discrete Distributions · ICML 2021 |
Machine learning › Reinforcement learning
sample efficiency |
0.5 | 1 | 2021 | No MCMC for me: Amortized sampling for fast and stable training of energy-based models · ICLR 2021 |
Memory systems
cache management |
0.5 | 1 | 2021 | A hierarchical neural model of data prefetching · ASPLOS 2021 |
Memory systems › cache › prefetching
data prefetching |
0.5 | 1 | 2021 | A hierarchical neural model of data prefetching · ASPLOS 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.4 | 1 | 2020 | Big Self-Supervised Models are Strong Semi-Supervised Learners · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2020 | Big Self-Supervised Models are Strong Semi-Supervised Learners · NeurIPS 2020 |
Natural language and speech › Language models and text generation › compositional generalization
length generalization |
0.4 | 1 | 2020 | Neural Execution Engines: Learning to Execute Subroutines · NeurIPS 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge incorporation › knowledge-infused learning › neuro-symbolic learning
neural algorithmic reasoning |
0.4 | 1 | 2020 | Neural Execution Engines: Learning to Execute Subroutines · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
numerical embedding |
0.4 | 1 | 2020 | Neural Execution Engines: Learning to Execute Subroutines · NeurIPS 2020 |
Machine learning › Learning paradigms
semi-supervised learning |
0.4 | 1 | 2020 | Big Self-Supervised Models are Strong Semi-Supervised Learners · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.4 | 1 | 2020 | Your classifier is secretly an energy based model and you should treat it like one · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
gaussian process · 1.9meta-learning · 1.5equilibrium selection · 0.9constrained matching · 0.9transfer learning · 0.8reward backpropagation · 0.8reinforcement learning · 0.8KL divergence · 0.8neural field · 0.7implicit neural representation · 0.7spectral graph theory · 0.6offline optimization · 0.6machine learning · 0.6edge correction · 0.6markov chain monte carlo · 0.5hierarchical neural network · 0.5amortized inference · 0.5imitation learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do generative video models understand physical principles?abstractAI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics—or, alternatively, are they merely sophisticated pixel predictors that achieve visual realism without understanding the physical principles of reality? We address this question by developing Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics. We find that across a range of current models (Sora, Runway, Pika, Lumiere, Stable Video Diffusion, and VideoPoet), physical understanding is severely limited, and unrelated to visual realism. At the same time, some test cases can already be successfully solved. This indicates that acquiring certain physical principles from observation alone may be possible, but significant challenges remain. While we expect rapid advances ahead, our work demonstrates that visual realism does not imply physical understanding. Our project page is at Physics-IQ website; code at Physics-IQ benchmark. Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, Robert Geirhos |
WACV | 3 |
| 2024 | Directly Fine-Tuning Diffusion Models on Differentiable RewardsabstractWe present Direct Reward Fine-Tuning (DRaFT), a simple and effective method for fine-tuning diffusion models to maximize differentiable reward functions, such as scores from human preference models. We first show that it is possible to backpropagate the reward function gradient through the full sampling procedure, and that doing so achieves strong performance on a variety of rewards, outperforming reinforcement learning-based approaches. We then propose more efficient variants of DRaFT: DRaFT-K, which truncates backpropagation to only the last K steps of sampling, and DRaFT-LV, which obtains lower-variance gradient estimates for the case when K=1. We show that our methods work well for a variety of reward functions and can be used to substantially improve the aesthetic quality of images generated by Stable Diffusion 1.4. Finally, we draw connections between our approach and prior work, providing a unifying perspective on the design space of gradient-based fine-tuning algorithms. Kevin Clark, Paul Vicol, Kevin Swersky, David J. Fleet |
ICLR | 3 |
| 2024 | Pre-trained Gaussian Processes for Bayesian OptimizationabstractBayesian optimization (BO) has become a popular strategy for global optimization of expensive real-world functions. Contrary to a common expectation that BO is suited to optimizing black-box functions, it actually requires domain knowledge about those functions to deploy BO successfully. Such domain knowledge often manifests in Gaussian process (GP) priors that specify initial beliefs on functions. However, even with expert knowledge, it is non-trivial to quantitatively define a prior. This is especially true for hyperparameter tuning problems on complex machine learning models, where landscapes of tuning objectives are often difficult to comprehend. We seek an alternative practice for setting these functional priors. In particular, we consider the scenario where we have data from similar functions that allow us to pre-train a tighter distribution a priori. We detail what pre-training entails for GPs using a KL divergence based loss function, and propose a new pre-training based BO framework named HyperBO. Theoretically, we show bounded posterior predictions and near-zero regrets for HyperBO without assuming the "ground truth" GP prior is known. To verify our approach in realistic setups, we collect a large multi-task hyperparameter tuning dataset by training tens of thousands of configurations of near-state-of-the-art deep learning models on popular image and text datasets, as well as a protein sequence dataset. Our results show that on average, HyperBO is able to locate good hyperparameters at least 3 times more efficiently than the best competing methods on both our new tuning dataset and existing multi-task BO benchmarks. George E. Dahl, Kevin Swersky, Chansoo Lee, Zachary Nado, Justin Gilmer, Jasper Snoek, Zoubin Ghahramani |
J. Mach. Learn. Res. | 3 |
| 2023 | CUF: Continuous Upsampling FiltersabstractNeural fields have rapidly been adopted for representing 3D signals, but their application to more classical 2D image-processing has been relatively limited. In this paper, we consider one of the most important operations in image processing: upsampling. In deep learning, learnable upsampling layers have extensively been used for single image super-resolution. We propose to parameterize upsampling kernels as neural fields. This parameterization leads to a compact architecture that obtains a 40-fold reduction in the number of parameters when compared with competing arbitrary-scale super-resolution architectures. When upsampling images of size 256×256 we show that our architecture is 2x-10x more efficient than competing arbitrary-scale super-resolution architectures, and more efficient than sub-pixel convolutions when instantiated to a single-scale model. In the general setting, these gains grow polynomially with the square of the target scale. We validate our method on standard benchmarks showing such efficiency gains can be achieved without sacrifices in super-resolution performance. https://cuf-paper.github.io Cristina Nader Vasconcelos, A. Cengiz Öztireli, Mark J. Matthews, Milad Hashemi, Kevin Swersky, Andrea Tagliasacchi |
CVPR | 5 |
| 2022 | Two Sides of the Same Coin: Heterophily and Oversmoothing in Graph Convolutional Neural NetworksabstractIn node classification tasks, graph convolutional neural networks (GCNs) have demonstrated competitive performance over traditional methods on diverse graph data. However, it is known that the performance of GCNs degrades with increasing number of layers (oversmoothing problem) and recent studies have also shown that GCNs may perform worse in heterophilous graphs, where neighboring nodes tend to belong to different classes (heterophily problem). These two problems are usually viewed as unrelated, and thus are studied independently, often at the graph filter level from a spectral perspective.We are the first to take a unified perspective to jointly explain the oversmoothing and heterophily problems at the node level. Specifically, we profile the nodes via two quantitative metrics: the relative degree of a node (compared to its neighbors) and the node-level heterophily. Our theory shows that the interplay of these two profiling metrics defines three cases of node behaviors, which explain the oversmoothing and heterophily problems jointly and can predict the performance of GCNs. Based on insights from our theory, we show theoretically and empirically the effectiveness of two strategies: structure-based edge correction, which learns corrected edge weights from structural properties (i.e., degrees), and feature-based edge correction, which learns signed edge weights from node features. Compared to other approaches, which tend to handle well either heterophily or oversmoothing, we show that our model, GGCN, which incorporates the two strategies performs well in both problems. We provide a longer version of this paper in [1] and codes on https://github.com/YujunYan/Heterophily_and_oversmoothing. Yujun Yan, Milad Hashemi, Kevin Swersky, Yaoqing Yang 0002, Danai Koutra |
ICDM | 3 |
| 2022 | Data-Driven Offline Optimization for Architecting Hardware Accelerators
Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, Sergey Levine |
ICLR | 4 |
| 2021 | A hierarchical neural model of data prefetchingabstractThis paper presents Voyager, a novel neural network for data prefetching. Unlike previous neural models for prefetching, which are limited to learning delta correlations, our model can also learn address correlations, which are important for prefetching irregular sequences of memory accesses. The key to our solution is its hierarchical structure that separates addresses into pages and offsets and that introduces a mechanism for learning important relations among pages and offsets. Voyager provides significant prediction benefits over current data prefetchers. For a set of irregular programs from the SPEC 2006 and GAP benchmark suites, Voyager sees an average IPC improvement of 41.6% over a system with no prefetcher, compared with 21.7% and 28.2%, respectively, for idealized Domino and ISB prefetchers. We also find that for two commercial workloads for which current data prefetchers see very little benefit, Voyager dramatically improves both accuracy and coverage. At present, slow training and prediction preclude neural models from being practically used in hardware, but Voyager’s overheads are significantly lower—in every dimension—than those of previous neural models. For example, computation cost is reduced by 15- 20×, and storage overhead is reduced by 110-200×. Thus, Voyager represents a significant step towards a practical neural prefetcher. Akanksha Jain, Kevin Swersky, Milad Hashemi, Parthasarathy Ranganathan, Calvin Lin |
ASPLOS | 3 |
| 2021 | No MCMC for me: Amortized sampling for fast and stable training of energy-based models
Will Grathwohl, Jacob Kelly, Milad Hashemi, Mohammad Norouzi 0002, Kevin Swersky, David Duvenaud |
ICLR | 5 |
| 2021 | Oops I Took A Gradient: Scalable Sampling for Discrete DistributionsabstractWe propose a general and scalable approximate sampling strategy for probabilistic models with discrete variables. Our approach uses gradients of the likelihood function with respect to its discrete inputs to propose updates in a Metropolis-Hastings sampler. We show empirically that this approach outperforms generic samplers in a number of difficult settings including Ising models, Potts models, restricted Boltzmann machines, and factorial hidden Markov models. We also demonstrate our improved sampler for training deep energy-based models on high dimensional discrete image data. This approach outperforms variational auto-encoders and existing energy-based models. Finally, we give bounds showing that our approach is near-optimal in the class of samplers which propose local updates. Will Grathwohl, Kevin Swersky, Milad Hashemi, David Duvenaud, Chris J. Maddison |
ICML | 2 |
| 2020 | Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi 0002, Kevin Swersky |
ICLR | 6 |
| 2020 | Learning Execution through Neural Code fusion
Kevin Swersky, Daniel Tarlow, Parthasarathy Ranganathan, Milad Hashemi |
ICLR | 2 |
| 2020 | Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, Hugo Larochelle |
ICLR | 9 |
| 2020 | An Imitation Learning Approach for Cache ReplacementabstractProgram execution speed critically depends on increasing cache hits, as cache hits are orders of magnitude faster than misses. To increase cache hits, we focus on the problem of cache replacement: choosing which cache line to evict upon inserting a new line. This is challenging because it requires planning far ahead and currently there is no known practical solution. As a result, current replacement policies typically resort to heuristics designed for specific common access patterns, which fail on more diverse and complex access patterns. In contrast, we propose an imitation learning approach to automatically learn cache access patterns by leveraging Belady’s, an oracle policy that computes the optimal eviction decision given the future cache accesses. While directly applying Belady’s is infeasible since the future is unknown, we train a policy conditioned only on past accesses that accurately approximates Belady’s even on diverse and complex access patterns, and call this approach Parrot. When evaluated on 13 of the most memory-intensive SPEC applications, Parrot increases cache miss rates by 20% over the current state of the art. In addition, on a large-scale web search benchmark, Parrot increases cache hit rates by 61% over a conventional LRU policy. We release a Gym environment to facilitate research in this area, as data is plentiful, and further advancements can have significant real-world impact. Evan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan, Junwhan Ahn |
ICML | 3 |
| 2020 | Optimizing Long-term Social Welfare in Recommender Systems: A Constrained Matching ApproachabstractMost recommender systems (RS) research assumes that a user’s utility can be maximized independently of the utility of the other agents (e.g., other users, content providers). In realistic settings, this is often not true – the dynamics of an RS ecosystem couple the long-term utility of all agents. In this work, we explore settings in which content providers cannot remain viable unless they receive a certain level of user engagement. We formulate this problem as one of equilibrium selection in the induced dynamical system, and show that it can be solved as an optimal constrained matching problem. Our model ensures the system reaches an equilibrium with maximal social welfare supported by a sufficiently diverse set of viable providers. We demonstrate that even in a simple, stylized dynamical RS model, the standard myopic approach to recommendation - always matching a user to the best provider - performs poorly. We develop several scalable techniques to solve the matching problem, and also draw connections to various notions of user regret and fairness, arguing that these outcomes are fairer in a utilitarian sense. Martin Mladenov, Elliot Creager, Omer Ben-Porat, Kevin Swersky, Richard S. Zemel, Craig Boutilier |
ICML | 4 |
| 2020 | Big Self-Supervised Models are Strong Semi-Supervised LearnersabstractOne paradigm for learning from few labeled examples while making best use of a large amount of unlabeled data is unsupervised pretraining followed by supervised fine-tuning. Although this paradigm uses unlabeled data in a task-agnostic way, in contrast to common approaches to semi-supervised learning for computer vision, we show that it is surprisingly effective for semi-supervised learning on ImageNet. A key ingredient of our approach is the use of big (deep and wide) networks during pretraining and fine-tuning. We find that, the fewer the labels, the more this approach (task-agnostic use of unlabeled data) benefits from a bigger network. After fine-tuning, the big network can be further improved and distilled into a much smaller one with little loss in classification accuracy by using the unlabeled examples for a second time, but in a task-specific way. The proposed semi-supervised learning algorithm can be summarized in three steps: unsupervised pretraining of a big ResNet model using SimCLRv2, supervised fine-tuning on a few labeled examples, and distillation with unlabeled examples for refining and transferring the task-specific knowledge. This procedure achieves 73.9% ImageNet top-1 accuracy with just 1% of the labels ($\le$13 labeled images per class) using ResNet-50, a 10X improvement in label efficiency over the previous state-of-the-art. With 10% of labels, ResNet-50 trained with our method achieves 77.5% top-1 accuracy, outperforming standard supervised training with all of the labels. Ting Chen 0007, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 0002, Geoffrey E. Hinton |
NeurIPS | 3 |
| 2020 | Neural Execution Engines: Learning to Execute SubroutinesabstractA significant effort has been made to train neural networks that replicate algorithmic reasoning, but they often fail to learn the abstract concepts underlying these algorithms. This is evidenced by their inability to generalize to data distributions that are outside of their restricted training sets, namely larger inputs and unseen data. We study these generalization issues at the level of numerical subroutines that comprise common algorithms like sorting, shortest paths, and minimum spanning trees. First, we observe that transformer-based sequence-to-sequence models can learn subroutines like sorting a list of numbers, but their performance rapidly degrades as the length of lists grows beyond those found in the training set. We demonstrate that this is due to attention weights that lose fidelity with longer sequences, particularly when the input numbers are numerically similar. To address the issue, we propose a learned conditional masking mechanism, which enables the model to strongly generalize far outside of its training range with near-perfect accuracy on a variety of algorithms. Second, to generalize to unseen data, we show that encoding numbers with a binary representation leads to embeddings with rich structure once trained on downstream tasks like addition or multiplication. This allows the embedding to handle missing data by faithfully interpolating numbers not seen during training. Yujun Yan, Kevin Swersky, Danai Koutra, Parthasarathy Ranganathan, Milad Hashemi |
NeurIPS | 2 |
| 2020 | Amortized Bayesian Optimization over Discrete SpacesabstractBayesian optimization is a principled approach for globally optimizing expensive, black-box functions by using a surrogate model of the objective. However, each step of Bayesian optimization involves solving an inner optimization problem, in which we maximize an acquisition function derived from the surrogate model to decide where to query next. This inner problem can be challenging to solve, particularly in discrete spaces, such as protein sequences or molecular graphs, where gradient-based optimization cannot be used. Our key insight is that we can train a generative model to generate candidates that maximize the acquisition function. This is faster than standard model-free local search methods, since we can amortize the cost of learning the model across multiple rounds of Bayesian optimization. We therefore call this Amortized Bayesian Optimization. On several challenging discrete design problems, we show this method generally outperforms other methods at optimizing the inner acquisition function, resulting in more efficient optimization of the outer black-box objective. Kevin Swersky, Yulia Rubanova, David Dohan, Kevin Murphy 0002 |
UAI | 1 |
| 2019 | Flexibly Fair Representation Learning by DisentanglementabstractWe consider the problem of learning representations that achieve group and subgroup fairness with respect to multiple sensitive attributes. Taking inspiration from the disentangled representation learning literature, we propose an algorithm for learning compact representations of datasets that are useful for reconstruction and prediction, but are also flexibly fair, meaning they can be easily modified at test time to achieve subgroup demographic parity with respect to multiple sensitive attributes and their conjunctions. We show empirically that the resulting encoder—which does not require the sensitive attributes for inference—allows for the adaptation of a single representation to a variety of fair classification tasks with new target labels and subgroup definitions. Elliot Creager, David Madras, Jörn-Henrik Jacobsen, Marissa A. Weis, Kevin Swersky, Toniann Pitassi, Richard S. Zemel |
ICML | 5 |
| 2019 | Graph Normalizing FlowsabstractWe introduce graph normalizing flows: a new, reversible graph neural network model for prediction and generation. On supervised tasks, graph normalizing flows perform similarly to message passing neural networks, but at a significantly reduced memory footprint, allowing them to scale to larger graphs. In the unsupervised case, we combine graph normalizing flows with a novel graph auto-encoder to create a generative model of graph structures. Our model is permutation-invariant, generating entire graphs with a single feed-forward pass, and achieves competitive results with the state-of-the art auto-regressive models, while being better suited to parallel computing architectures. Jenny Liu, Aviral Kumar, Jimmy Ba, Jamie Kiros, Kevin Swersky |
NeurIPS | 5 |
| 2018 | Learning Hard Alignments with Variational InferenceabstractThere has recently been significant interest in hard attention models for tasks such as object recognition, visual captioning and speech recognition. Hard attention offers benefits over soft attention such as decreased computational cost, but training hard attention models can be difficult because of the discrete latent variables they introduce. Previous work used REINFORCE to approach these issues, however, it suffers from high-variance gradient estimates, resulting in slow convergence. In this paper, we tackle the problem of learning hard attention for a sequential task using variational inference methods, specifically the recently introduced Variational Inference for Monte Carlo Objectives (VIMCO) and Neural Variational Inference (NVIL). Furthermore, we propose a novel baseline that adapts VIMCO to this setting. We demonstrate our method on a phoneme recognition task in clean and noisy environments and show that our method outperforms REINFORCE, with the difference being greater for a more complicated task. Dieterich Lawson, Chung-Cheng Chiu, George Tucker, Colin Raffel, Kevin Swersky, Navdeep Jaitly |
ICASSP | 5 |
| 2018 | Meta-Learning for Semi-Supervised Few-Shot Classification
Mengye Ren, Eleni Triantafillou, Sachin Ravi, Jake Snell, Kevin Swersky, Josh Tenenbaum, Hugo Larochelle, Richard S. Zemel |
ICLR (Poster) | 5 |
| 2018 | Learning Memory Access PatternsabstractThe explosion in workload complexity and the recent slow-down in Moore’s law scaling call for new approaches towards efficient computing. Researchers are now beginning to use recent advances in machine learning in software optimizations; augmenting or replacing traditional heuristics and data structures. However, the space of machine learning for computer hardware architecture is only lightly explored. In this paper, we demonstrate the potential of deep learning to address the von Neumann bottleneck of memory performance. We focus on the critical problem of learning memory access patterns, with the goal of constructing accurate and efficient memory prefetchers. We relate contemporary prefetching strategies to n-gram models in natural language processing, and show how recurrent neural networks can serve as a drop-in replacement. On a suite of challenging benchmark datasets, we find that neural networks consistently demonstrate superior performance in terms of precision and recall. This work represents the first step towards practical neural-network based prefetching, and opens a wide range of exciting directions for machine learning in computer architecture research. Milad Hashemi, Kevin Swersky, Jamie A. Smith, Grant Ayers, Heiner Litz, Jichuan Chang, Christoforos E. Kozyrakis, Parthasarathy Ranganathan |
ICML | 2 |
| 2017 | Prototypical Networks for Few-shot LearningabstractWe propose Prototypical Networks for the problem of few-shot classification, where a classifier must generalize to new classes not seen in the training set, given only a small number of examples of each new class. Prototypical Networks learn a metric space in which classification can be performed by computing distances to prototype representations of each class. Compared to recent approaches for few-shot learning, they reflect a simpler inductive bias that is beneficial in this limited-data regime, and achieve excellent results. We provide an analysis showing that some simple design decisions can yield substantial improvements over recent approaches involving complicated architectural choices and meta-learning. We further extend Prototypical Networks to zero-shot learning and achieve state-of-the-art results on the CU-Birds dataset. Jake Snell, Kevin Swersky, Richard S. Zemel |
NIPS | 2 |
| 2016 | Taking the Human Out of the Loop: A Review of Bayesian OptimizationabstractBig Data applications are typically associated with systems involving large numbers of users, massive complex software systems, and large-scale heterogeneous computing and storage architectures. The construction of such systems involves many distributed design choices. The end products (e.g., recommendation systems, medical analysis tools, real-time game engines, speech recognizers) thus involve many tunable configuration parameters. These parameters are often specified and hard-coded into the software by various developers or teams. If optimized jointly, these parameters can result in significant improvements. Bayesian optimization is a powerful tool for the joint optimization of design choices that is gaining great popularity in recent years. It promises greater automation so as to increase both product quality and human productivity. This review paper introduces Bayesian optimization, highlights some of its methodological aspects, and showcases a wide range of applications. Bobak Shahriari, Kevin Swersky, Ziyu Wang 0001, Ryan P. Adams, Nando de Freitas |
Proc. IEEE | 2 |
| 2015 | Predicting Deep Zero-Shot Convolutional Neural Networks Using Textual DescriptionsabstractOne of the main challenges in Zero-Shot Learning of visual categories is gathering semantic attributes to accompany images. Recent work has shown that learning from textual descriptions, such as Wikipedia articles, avoids the problem of having to explicitly define these attributes. We present a new model that can classify unseen categories from their textual description. Specifically, we use text features to predict the output weights of both the convolutional and the fully connected layers in a deep convolutional neural network (CNN). We take advantage of the architecture of CNNs and learn features at different layers, rather than just learning an embedding space for both modalities, as is common with existing approaches. The proposed model also allows us to automatically generate a list of pseudo-attributes for each visual category consisting of words from Wikipedia articles. We train our models end-to-end using the Caltech-UCSD bird and flower datasets and evaluate both ROC and Precision-Recall curves. Our empirical results show that the proposed model significantly outperforms previous methods. Jimmy Ba, Kevin Swersky, Sanja Fidler, Ruslan Salakhutdinov |
ICCV | 2 |
| 2015 | Generative Moment Matching NetworksabstractWe consider the problem of learning deep generative models from data. We formulate a method that generates an independent sample via a single feedforward pass through a multilayer preceptron, as in the recently proposed generative adversarial networks (Goodfellow et al., 2014). Training a generative adversarial network, however, requires careful optimization of a difficult minimax program. Instead, we utilize a technique from statistical hypothesis testing known as maximum mean discrepancy (MMD), which leads to a simple objective that can be interpreted as matching all orders of statistics between a dataset and samples from the model, and can be trained by backpropagation. We further boost the performance of this approach by combining our generative network with an auto-encoder network, using MMD to learn to generate codes that can then be decoded to produce samples. We show that the combination of these techniques yields excellent generative models compared to baseline approaches as measured on MNIST and the Toronto Face Database. Yujia Li 0001, Kevin Swersky, Richard S. Zemel |
ICML | 2 |
| 2015 | Scalable Bayesian Optimization Using Deep Neural NetworksabstractBayesian optimization is an effective methodology for the global optimization of functions with expensive evaluations. It relies on querying a distribution over functions defined by a relatively cheap surrogate model. An accurate model for this distribution over functions is critical to the effectiveness of the approach, and is typically fit using Gaussian processes (GPs). However, since GPs scale cubically with the number of observations, it has been challenging to handle objectives whose optimization requires many evaluations, and as such, massively parallelizing the optimization. In this work, we explore the use of neural networks as an alternative to GPs to model distributions over functions. We show that performing adaptive basis function regression with a neural network as the parametric form performs competitively with state-of-the-art GP-based approaches, but scales linearly with the number of data rather than cubically. This allows us to achieve a previously intractable degree of parallelism, which we apply to large scale hyperparameter optimization, rapidly finding competitive models on benchmark object recognition tasks using convolutional networks, and image caption generation using neural language models. Jasper Snoek, Oren Rippel, Kevin Swersky, Jamie Kiros, Nadathur Satish, Narayanan Sundaram, Md. Mostofa Ali Patwary, Prabhat, Ryan P. Adams |
ICML | 3 |
| 2014 | Input Warping for Bayesian Optimization of Non-Stationary FunctionsabstractBayesian optimization has proven to be a highly effective methodology for the global optimization of unknown, expensive and multimodal functions. The ability to accurately model distributions over functions is critical to the effectiveness of Bayesian optimization. Although Gaussian processes provide a flexible prior over functions, there are various classes of functions that remain difficult to model. One of the most frequently occurring of these is the class of non-stationary functions. The optimization of the hyperparameters of machine learning algorithms is a problem domain in which parameters are often manually transformed a priori, for example by optimizing in "log-space", to mitigate the effects of spatially-varying length scale. We develop a methodology for automatically learning a wide family of bijective transformations or warpings of the input space using the Beta cumulative distribution function. We further extend the warping framework to multi-task Bayesian optimization so that multiple tasks can be warped into a jointly stationary space. On a set of challenging benchmark optimization tasks, we observe that the inclusion of warping greatly improves on the state-of-the-art, producing better results faster and more reliably. Jasper Snoek, Kevin Swersky, Richard S. Zemel, Ryan P. Adams |
ICML | 2 |
| 2013 | Stochastic k-Neighborhood Selection for Supervised and Unsupervised LearningabstractNeighborhood Components Analysis (NCA) is a popular method for learning a distance metric to be used within a k-nearest neighbors (kNN) classifier. A key assumption built into the model is that each point stochastically selects a single neighbor, which makes the model well-justified only for kNN with k=1. However, kNN classifiers with k>1 are more robust and usually preferred in practice. Here we present kNCA, which generalizes NCA by learning distance metrics that are appropriate for kNN with arbitrary k. The main technical contribution is showing how to efficiently compute and optimize the expected accuracy of a kNN classifier. We apply similar ideas in an unsupervised setting to yield kSNE and ktSNE, generalizations of Stochastic Neighbor Embedding (SNE, tSNE) that operate on neighborhoods of size k, which provide an axis of control over embeddings that allow for more homogeneous and interpretable regions. Empirically, we show that kNCA often improves classification accuracy over state of the art methods, produces qualitative differences in the embeddings as k is varied, and is more robust with respect to label noise. Daniel Tarlow, Kevin Swersky, Laurent Charlin, Ilya Sutskever, Richard S. Zemel |
ICML (3) | 2 |
| 2013 | Learning Fair RepresentationsabstractWe propose a learning algorithm for fair classification that achieves both group fairness (the proportion of members in a protected group receiving positive classification is identical to the proportion in the population as a whole), and individual fairness (similar individuals should be treated similarly). We formulate fairness as an optimization problem of finding a good representation of the data with two competing goals: to encode the data as well as possible, while simultaneously obfuscating any information about membership in the protected group. We show positive results of our algorithm relative to other known techniques, on three datasets. Moreover, we demonstrate several advantages to our approach. First, our intermediate representation can be used for other classification tasks (i.e., transfer learning is possible); secondly, we take a step toward learning a distance metric which can find important dimensions of the data for classification. Richard S. Zemel, Kevin Swersky, Toniann Pitassi, Cynthia Dwork |
ICML (3) | 3 |
| 2013 | Multi-Task Bayesian OptimizationabstractBayesian optimization has recently been proposed as a framework for automatically tuning the hyperparameters of machine learning models and has been shown to yield state-of-the-art performance with impressive ease and efficiency. In this paper, we explore whether it is possible to transfer the knowledge gained from previous optimizations to new tasks in order to find optimal hyperparameter settings more efficiently. Our approach is based on extending multi-task Gaussian processes to the framework of Bayesian optimization. We show that this method significantly speeds up the optimization process when compared to the standard single-task approach. We further propose a straightforward extension of our algorithm in order to jointly minimize the average error across multiple tasks and demonstrate how this can be used to greatly speed up $k$-fold cross-validation. Lastly, our most significant contribution is an adaptation of a recently proposed acquisition function, entropy search, to the cost-sensitive and multi-task settings. We demonstrate the utility of this new acquisition function by utilizing a small dataset in order to explore hyperparameter settings for a large dataset. Our algorithm dynamically chooses which dataset to query in order to yield the most information per unit cost. Kevin Swersky, Jasper Snoek, Ryan P. Adams |
NIPS | 1 |
| 2012 | Prediction and Fault Detection of Environmental Signals with Uncharacterised FaultsabstractMany signals of interest are corrupted by faults of anunknown type. We propose an approach that uses Gaus-sian processes and a general “fault bucket” to capturea priori uncharacterised faults, along with an approxi-mate method for marginalising the potential faultinessof all observations. This gives rise to an efficient, flexible algorithm for the detection and automatic correction of faults. Our method is deployed in the domain of water monitoring and management, where it is able to solve several fault detection, correction, and prediction problems. The method works well despite the fact that the data is plagued with numerous difficulties, including missing observations, multiple discontinuities, nonlinearity and many unanticipated types of fault. Michael A. Osborne, Roman Garnett, Kevin Swersky, Nando de Freitas |
AAAI | 3 |
| 2012 | Estimating the Hessian by Back-propagating Curvature
James Martens, Ilya Sutskever, Kevin Swersky |
ICML | 3 |
| 2012 | Probabilistic n-Choose-k Models for Classification and RankingabstractIn categorical data there is often structure in the number of variables that take on each label. For example, the total number of objects in an image and the number of highly relevant documents per query in web search both tend to follow a structured distribution. In this paper, we study a probabilistic model that explicitly includes a prior distribution over such counts, along with a count-conditional likelihood that defines probabilities over all subsets of a given size. When labels are binary and the prior over counts is a Poisson-Binomial distribution, a standard logistic regression model is recovered, but for other count distributions, such priors induce global dependencies and combinatorics that appear to complicate learning and inference. However, we demonstrate that simple, efficient learning procedures can be derived for more general forms of this model. We illustrate the utility of the formulation by exploring applications to multi-object classification, learning to rank, and top-K classification. Kevin Swersky, Daniel Tarlow, Ryan P. Adams, Richard S. Zemel, Brendan J. Frey |
NIPS | 1 |
| 2012 | Cardinality Restricted Boltzmann MachinesabstractThe Restricted Boltzmann Machine (RBM) is a popular density model that is also good for extracting features. A main source of tractability in RBM models is the model's assumption that given an input, hidden units activate independently from one another. Sparsity and competition in the hidden representation is believed to be beneficial, and while an RBM with competition among its hidden units would acquire some of the attractive properties of sparse coding, such constraints are not added due to the widespread belief that the resulting model would become intractable. In this work, we show how a dynamic programming algorithm developed in 1981 can be used to implement exact sparsity in the RBM's hidden units. We then expand on this and show how to pass derivatives through a layer of exact sparsity, which makes it possible to fine-tune a deep belief network (DBN) consisting of RBMs with sparse hidden layers. We show that sparsity in the RBM's hidden layer improves the performance of both the pre-trained representations and of the fine-tuned model. Kevin Swersky, Daniel Tarlow, Ilya Sutskever, Ruslan Salakhutdinov, Richard S. Zemel, Ryan P. Adams |
NIPS | 1 |
| 2012 | Fast Exact Inference for Recursive Cardinality Models
Daniel Tarlow, Kevin Swersky, Richard S. Zemel, Ryan P. Adams, Brendan J. Frey |
UAI | 2 |
| 2011 | On Autoencoders and Score Matching for Energy Based Models
Kevin Swersky, Marc'Aurelio Ranzato, David Buchman, Benjamin M. Marlin, Nando de Freitas |
ICML | 1 |