Brian McWilliams

dblp:44/905 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
6since 2021 · last 2024
0009-0002-7433-1702ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Deep learning architectures and training · 23% Optimization for machine learning · 21% Representation and self-supervised learning · 21%
Computer graphics and multimedia
6 papers
Rendering · 44% Audio and music processing · 30% Image and video processing · 26%
Theoretical computer science
3 papers
Algorithms and data structures · 47% Algorithmic game theory and mechanism design · 38% Mathematical optimization · 14%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium
1.732023
The Symmetric Generalized Eigenvalue Problem as a Nash Equilibrium · ICLR 2023
EigenGame Unloaded: When playing games is better than optimizing · ICLR 2022
EigenGame: PCA as a Nash Equilibrium · ICLR 2021
Algorithms and data structures › numerical linear algebra
matrix factorization
1.122022
EigenGame Unloaded: When playing games is better than optimizing · ICLR 2022
EigenGame: PCA as a Nash Equilibrium · ICLR 2021
Algorithms and data structures › numerical linear algebra › dimensionality reduction
principal component analysis
1.122022
EigenGame Unloaded: When playing games is better than optimizing · ICLR 2022
EigenGame: PCA as a Nash Equilibrium · ICLR 2021
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
MusicRL: Aligning Music Generation to Human Preferences · ICML 2024
Audio and music processing
music generation
0.812024
MusicRL: Aligning Music Generation to Human Preferences · ICML 2024
Audio and music processing › music generation
text-to-music generation
0.812024
MusicRL: Aligning Music Generation to Human Preferences · ICML 2024
Natural language and speech › Language models and text generation
pre-trained language model
0.712023
TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations at Twitter · KDD 2023
Mathematical optimization › numerical analysis › eigenvalue problem
generalized eigenvalue problem
0.712023
The Symmetric Generalized Eigenvalue Problem as a Nash Equilibrium · ICLR 2023
Rendering
monte carlo rendering
0.622018
Denoising with kernel prediction and asymmetric loss functions · ACM Trans. Graph. 2018
Kernel-predicting convolutional networks for denoising Monte Carlo renderings · ACM Trans. Graph. 2017
Machine learning › Representation and self-supervised learning
causal representation learning
0.512021
Representation Learning via Invariant Causal Mechanisms · ICLR 2021
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.512021
Representation Learning via Invariant Causal Mechanisms · ICLR 2021
Machine learning › Optimization for machine learning
stochastic gradient descent
0.522016
Scalable Adaptive Stochastic Optimization Using Random Projections · NIPS 2016
Variance Reduced Stochastic Gradient Descent with Neighbors · NIPS 2015
Rendering
monte carlo integration
0.412019
Neural Importance Sampling · ACM Trans. Graph. 2019
Image and video processing › image restoration
image denoising
0.312018
Denoising with kernel prediction and asymmetric loss functions · ACM Trans. Graph. 2018
Image and video processing
video frame interpolation
0.312018
PhaseNet for Video Frame Interpolation · CVPR 2018
Machine learning › Optimization for machine learning
convergence guarantees
0.312017
Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks · ICML 2017
Machine learning › Deep learning architectures and training
rectifier network
0.312017
Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks · ICML 2017
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.312017
The Shattered Gradients Problem: If resnets are the answer, then what is the question? · ICML 2017
Machine learning › Deep learning architectures and training
skip connections
0.312017
The Shattered Gradients Problem: If resnets are the answer, then what is the question? · ICML 2017
Machine learning › Deep learning architectures and training › training dynamics
vanishing and exploding gradients
0.312017
The Shattered Gradients Problem: If resnets are the answer, then what is the question? · ICML 2017
Rendering › light transport
atmospheric rendering
0.312017
Deep scattering: rendering atmospheric clouds with radiance-predicting neural networks · ACM Trans. Graph. 2017
Rendering › participating media rendering
cloud rendering
0.312017
Deep scattering: rendering atmospheric clouds with radiance-predicting neural networks · ACM Trans. Graph. 2017
Image and video processing › image restoration
denoising
0.312017
Kernel-predicting convolutional networks for denoising Monte Carlo renderings · ACM Trans. Graph. 2017
Rendering
neural rendering
0.312017
Deep scattering: rendering atmospheric clouds with radiance-predicting neural networks · ACM Trans. Graph. 2017
Machine learning › Optimization for machine learning › adaptive optimization
adaptive stochastic optimization
0.212016
Scalable Adaptive Stochastic Optimization Using Random Projections · NIPS 2016
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.212016
Scalable Adaptive Stochastic Optimization Using Random Projections · NIPS 2016
Computer vision › Video understanding and tracking
video object segmentation
0.212016
A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation · CVPR 2016
Machine learning › Optimization for machine learning › online optimization
streaming optimization
0.212015
Variance Reduced Stochastic Gradient Descent with Neighbors · NIPS 2015
Machine learning › Optimization for machine learning
variance reduction
0.212015
Variance Reduced Stochastic Gradient Descent with Neighbors · NIPS 2015
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
least squares regression
0.212014
Fast and Robust Least Squares Estimation in Corrupted Linear Models · NIPS 2014

Methods — techniques the papers use, named apart from their topics

game theory · 1.7reinforcement learning from human feedback · 1.5autoregressive model · 1.5deep neural network · 0.7social objective · 0.7heterogeneous information network · 0.7eigenvalue computation · 0.7kernel prediction · 0.6convolutional neural network · 0.6invariant causal mechanisms · 0.5piecewise-polynomial coupling transforms · 0.4one-blob encoding · 0.4NICE · 0.4phase-based motion representation · 0.3neural network decoder · 0.3initialization · 0.3gradient descent · 0.3batch normalization · 0.3
YearPublicationVenuePosition
2024 MusicRL: Aligning Music Generation to Human Preferences
abstract
We propose MusicRL, the first music generation system finetuned from human feedback. Appreciation of text-to-music models is particularly subjective since the concept of musicality as well as the specific intention behind a caption are user-dependent (e.g. a caption such as “upbeat workout music” can map to a retro guitar solo or a technopop beat). Not only this makes supervised training of such models challenging, but it also calls for integrating continuous human feedback in their post-deployment finetuning. MusicRL is a pretrained autoregressive MusicLM model of discrete audio tokens finetuned with reinforcement learning to maximize sequence-level rewards. We design reward functions related specifically to text-adherence and audio quality with the help from selected raters, and use those to finetune MusicLM into MusicRL-R. We deploy MusicLM to users and collect a substantial dataset comprising 300,000 pairwise preferences. Using Reinforcement Learning from Human Feedback (RLHF), we train MusicRL-U, the first text-to-music model that incorporates human feedback at scale. Human evaluations show that both MusicRL-R and MusicRL-U are preferred to the baseline. Ultimately, MusicRL-RU combines the two approaches and results in the best model according to human raters. Ablation studies shed light on the musical attributes influencing human preferences, indicating that text adherence and quality only account for a part of it. This underscores the prevalence of subjectivity in musical appreciation and calls for further involvement of human listeners in the finetuning of music generation models. Samples can be found at google-research.github.io/seanet/musiclm/rlhf/.
Geoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent, Matej Kastelic, Zalan Borsos, Brian McWilliams, Victor Ungureanu, Olivier Bachem, Olivier Pietquin, Matthieu Geist, Léonard Hussenot, Neil Zeghidour, Andrea Agostinelli
ICML7
2023 The Symmetric Generalized Eigenvalue Problem as a Nash Equilibrium
Ian Gemp, Charlie Chen, Brian McWilliams
ICLR3
2023 TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations at Twitter
abstract
Pre-trained language models (PLMs) are fundamental for natural language processing applications. Most existing PLMs are not tailored to the noisy user-generated text on social media, and the pre-training does not factor in the valuable social engagement logs available in a social network. We present TwHIN-BERT, a multilingual language model productionized at Twitter, trained on in-domain data from the popular social network. TwHIN-BERT differs from prior pre-trained language models as it is trained with not only text-based self-supervision but also with a social objective based on the rich social engagements within a Twitter heterogeneous information network (TwHIN). Our model is trained on 7 billion tweets covering over 100 distinct languages, providing a valuable representation to model short, noisy, user-generated text. We evaluate our model on various multilingual social recommendation and semantic understanding tasks and demonstrate significant metric improvement over established pre-trained language models. We open-source TwHIN-BERT and our curated hashtag prediction and social engagement benchmark datasets to the research community.
Xinyang Zhang 0002, Yury Malkov, Omar Florez, Se Rim Park, Brian McWilliams, Jiawei Han 0001, Ahmed El-Kishky
KDD5
2022 EigenGame Unloaded: When playing games is better than optimizing
Ian Gemp, Brian McWilliams, Claire Vernade, Thore Graepel
ICLR2
2021 EigenGame: PCA as a Nash Equilibrium
Ian Gemp, Brian McWilliams, Claire Vernade, Thore Graepel
ICLR2
2021 Representation Learning via Invariant Causal Mechanisms
Jovana Mitrovic, Brian McWilliams, Jacob C. Walker, Lars Buesing, Charles Blundell
ICLR2
2019 Neural Importance Sampling
abstract
We propose to use deep neural networks for generating samples in Monte Carlo integration. Our work is based on non-linear independent components estimation (NICE), which we extend in numerous ways to improve performance and enable its application to integration problems. First, we introduce piecewise-polynomial coupling transforms that greatly increase the modeling power of individual coupling layers. Second, we propose to preprocess the inputs of neural networks using one-blob encoding, which stimulates localization of computation and improves inference. Third, we derive a gradient-descent-based optimization for the Kullback-Leibler and the χ 2 divergence for the specific application of Monte Carlo integration with unnormalized stochastic estimates of the target distribution. Our approach enables fast and accurate inference and efficient sample generation independently of the dimensionality of the integration domain. We show its benefits on generating natural images and in two applications to light-transport simulation: first, we demonstrate learning of joint path-sampling densities in the primary sample space and importance sampling of multi-dimensional path prefixes thereof. Second, we use our technique to extract conditional directional densities driven by the product of incident illumination and the BSDF in the rendering equation, and we leverage the densities for path guiding. In all applications, our approach yields on-par or higher performance than competing techniques at equal sample count.
Thomas Müller 0013, Brian McWilliams, Fabrice Rousselle, Markus Gross 0001, Jan Novák
ACM Trans. Graph.2
2018 PhaseNet for Video Frame Interpolation
abstract
Most approaches for video frame interpolation require accurate dense correspondences to synthesize an in-between frame. Therefore, they do not perform well in challenging scenarios with e.g. lighting changes or motion blur. Recent deep learning approaches that rely on kernels to represent motion can only alleviate these problems to some extent. In those cases, methods that use a per-pixel phase-based motion representation have been shown to work well. However, they are only applicable for a limited amount of motion. We propose a new approach, PhaseNet, that is designed to robustly handle challenging scenarios while also coping with larger motion. Our approach consists of a neural network decoder that directly estimates the phase decomposition of the intermediate frame. We show that this is superior to the hand-crafted heuristics previously used in phase-based methods and also compares favorably to recent deep learning based approaches for video frame interpolation on challenging datasets.
Simone Schaub-Meyer, Abdelaziz Djelouah, Brian McWilliams, Alexander Sorkine-Hornung, Markus Gross 0001, Christopher Schroers
CVPR3
2018 Denoising with kernel prediction and asymmetric loss functions
abstract
We present a modular convolutional architecture for denoising rendered images. We expand on the capabilities of kernel-predicting networks by combining them with a number of task-specific modules, and optimizing the assembly using an asymmetric loss. The source-aware encoder---the first module in the assembly---extracts low-level features and embeds them into a common feature space, enabling quick adaptation of a trained network to novel data. The spatial and temporal modules extract abstract, high-level features for kernel-based reconstruction, which is performed at three different spatial scales to reduce low-frequency artifacts. The complete network is trained using a class of asymmetric loss functions that are designed to preserve details and provide the user with a direct control over the variance-bias trade-off during inference. We also propose an error-predicting module for inferring reconstruction error maps that can be used to drive adaptive sampling. Finally, we present a theoretical analysis of convergence rates of kernel-predicting architectures, shedding light on why kernel prediction performs better than synthesizing the colors directly, complementing the empirical evidence presented in this and previous works. We demonstrate that our networks attain results that compare favorably to state-of-the-art methods in terms of detail preservation, low-frequency noise removal, and temporal stability on a variety of production and academic datasets.
Thijs Vogels, Fabrice Rousselle, Brian McWilliams, Gerhard Röthlin, Alex Harvill, David Adler, Mark Meyer, Jan Novák
ACM Trans. Graph.3
2017 The Shattered Gradients Problem: If resnets are the answer, then what is the question?
abstract
A long-standing obstacle to progress in deep learning is the problem of vanishing and exploding gradients. Although, the problem has largely been overcome via carefully constructed initializations and batch normalization, architectures incorporating skip-connections such as highway and resnets perform much better than standard feedforward architectures despite well-chosen initialization and batch normalization. In this paper, we identify the shattered gradients problem. Specifically, we show that the correlation between gradients in standard feedforward networks decays exponentially with depth resulting in gradients that resemble white noise whereas, in contrast, the gradients in architectures with skip-connections are far more resistant to shattering, decaying sublinearly. Detailed empirical evidence is presented in support of the analysis, on both fully-connected networks and convnets. Finally, we present a new “looks linear” (LL) initialization that prevents shattering, with preliminary experiments showing the new initialization allows to train very deep networks without the addition of skip-connections.
David Balduzzi, Marcus Frean, Lennox Leary, John P. Lewis, Kurt Wan-Duo Ma, Brian McWilliams
ICML6
2017 Neural Taylor Approximations: Convergence and Exploration in Rectifier Networks
abstract
Modern convolutional networks, incorporating rectifiers and max-pooling, are neither smooth nor convex; standard guarantees therefore do not apply. Nevertheless, methods from convex optimization such as gradient descent and Adam are widely used as building blocks for deep learning algorithms. This paper provides the first convergence guarantee applicable to modern convnets, which furthermore matches a lower bound for convex nonsmooth functions. The key technical tool is the neural Taylor approximation – a straightforward application of Taylor expansions to neural networks – and the associated Taylor loss. Experiments on a range of optimizers, layers, and tasks provide evidence that the analysis accurately captures the dynamics of neural optimization. The second half of the paper applies the Taylor approximation to isolate the main difficulty in training rectifier nets – that gradients are shattered – and investigates the hypothesis that, by exploring the space of activation configurations more thoroughly, adaptive optimizers such as RMSProp and Adam are able to converge to better solutions.
David Balduzzi, Brian McWilliams, Tony Butler-Yeoman
ICML2
2017 Kernel-predicting convolutional networks for denoising Monte Carlo renderings
abstract
Regression-based algorithms have shown to be good at denoising Monte Carlo (MC) renderings by leveraging its inexpensive by-products (e.g., feature buffers). However, when using higher-order models to handle complex cases, these techniques often overfit to noise in the input. For this reason, supervised learning methods have been proposed that train on a large collection of reference examples, but they use explicit filters that limit their denoising ability. To address these problems, we propose a novel, supervised learning approach that allows the filtering kernel to be more complex and general by leveraging a deep convolutional neural network (CNN) architecture. In one embodiment of our framework, the CNN directly predicts the final denoised pixel value as a highly non-linear combination of the input features. In a second approach, we introduce a novel, kernel-prediction network which uses the CNN to estimate the local weighting kernels used to compute each denoised pixel from its neighbors. We train and evaluate our networks on production data and observe improvements over state-of-the-art MC denoisers, showing that our methods generalize well to a variety of scenes. We conclude by analyzing various components of our architecture and identify areas of further research in deep learning for MC denoising.
Steve Bako, Thijs Vogels, Brian McWilliams, Mark Meyer, Jan Novák, Alex Harvill, Pradeep Sen, Tony DeRose, Fabrice Rousselle
ACM Trans. Graph.3
2017 Deep scattering: rendering atmospheric clouds with radiance-predicting neural networks
abstract
We present a technique for efficiently synthesizing images of atmospheric clouds using a combination of Monte Carlo integration and neural networks. The intricacies of Lorenz-Mie scattering and the high albedo of cloud-forming aerosols make rendering of clouds---e.g. the characteristic silverlining and the "whiteness" of the inner body---challenging for methods based solely on Monte Carlo integration or diffusion theory. We approach the problem differently. Instead of simulating all light transport during rendering, we pre-learn the spatial and directional distribution of radiant flux from tens of cloud exemplars. To render a new scene, we sample visible points of the cloud and, for each, extract a hierarchical 3D descriptor of the cloud geometry with respect to the shading location and the light source. The descriptor is input to a deep neural network that predicts the radiance function for each shading configuration. We make the key observation that progressively feeding the hierarchical descriptor into the network enhances the network's ability to learn faster and predict with higher accuracy while using fewer coefficients. We also employ a block design with residual connections to further improve performance. A GPU implementation of our method synthesizes images of clouds that are nearly indistinguishable from the reference solution within seconds to minutes. Our method thus represents a viable solution for applications such as cloud design and, thanks to its temporal stability, for high-quality production of animated content.
Simon Kallweit, Thomas Müller 0013, Brian McWilliams, Markus Gross 0001, Jan Novák
ACM Trans. Graph.3
2016 DUAL-LOCO: Distributing Statistical Estimation Using Random Projections
abstract
We present DUAL-LOCO, a communication-efficient algorithm for distributed statistical estimation. DUAL-LOCO assumes that the data is distributed across workers according to the features rather than the samples. It requires only a single round of communication where low-dimensional random projections are used to approximate the dependencies between features available to different workers. We show that DUAL-LOCO has bounded approximation error which only depends weakly on the number of workers. We compare DUAL-LOCO against a state-of-the-art distributed optimization method on a variety of real world datasets and show that it obtains better speedups while retaining good accuracy. In particular, DUAL-LOCO allows for fast cross validation as only part of the algorithm depends on the regularization parameter.
Christina Heinze-Deml, Brian McWilliams, Nicolai Meinshausen
AISTATS2
2016 A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation
abstract
Over the years, datasets and benchmarks have proven their fundamental importance in computer vision research, enabling targeted progress and objective comparisons in many fields. At the same time, legacy datasets may impend the evolution of a field due to saturated algorithm performance and the lack of contemporary, high quality data. In this work we present a new benchmark dataset and evaluation methodology for the area of video object segmentation. The dataset, named DAVIS (Densely Annotated VIdeo Segmentation), consists of fifty high quality, Full HD video sequences, spanning multiple occurrences of common video object segmentation challenges such as occlusions, motionblur and appearance changes. Each video is accompanied by densely annotated, pixel-accurate and per-frame ground truth segmentation. In addition, we provide a comprehensive analysis of several state-of-the-art segmentation approaches using three complementary metrics that measure the spatial extent of the segmentation, the accuracy of the silhouette contours and the temporal coherence. The results uncover strengths and weaknesses of current approaches, opening up promising directions for future works.
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross 0001, Alexander Sorkine-Hornung
CVPR3
2016 Scalable Adaptive Stochastic Optimization Using Random Projections
abstract
Adaptive stochastic gradient methods such as AdaGrad have gained popularity in particular for training deep neural networks. The most commonly used and studied variant maintains a diagonal matrix approximation to second order information by accumulating past gradients which are used to tune the step size adaptively. In certain situations the full-matrix variant of AdaGrad is expected to attain better performance, however in high dimensions it is computationally impractical. We present Ada-LR and RadaGrad two computationally efficient approximations to full-matrix AdaGrad based on randomized dimensionality reduction. They are able to capture dependencies between features and achieve similar performance to full-matrix AdaGrad but at a much smaller computational cost. We show that the regret of Ada-LR is close to the regret of full-matrix AdaGrad which can have an up-to exponentially smaller dependence on the dimension than the diagonal variant. Empirically, we show that Ada-LR and RadaGrad perform similarly to full-matrix AdaGrad. On the task of training convolutional neural networks as well as recurrent neural networks, RadaGrad achieves faster convergence than diagonal AdaGrad.
Gabriel Krummenacher, Brian McWilliams, Yannic Kilcher, Joachim M. Buhmann, Nicolai Meinshausen
NIPS2
2015 Variance Reduced Stochastic Gradient Descent with Neighbors
abstract
Stochastic Gradient Descent (SGD) is a workhorse in machine learning, yet it is also known to be slow relative to steepest descent. Recently, variance reduction techniques such as SVRG and SAGA have been proposed to overcome this weakness. With asymptotically vanishing variance, a constant step size can be maintained, resulting in geometric convergence rates. However, these methods are either based on occasional computations of full gradients at pivot points (SVRG), or on keeping per data point corrections in memory (SAGA). This has the disadvantage that one cannot employ these methods in a streaming setting and that speed-ups relative to SGD may need a certain number of epochs in order to materialize. This paper investigates a new class of algorithms that can exploit neighborhood structure in the training data to share and re-use information about past stochastic gradients across data points. While not meant to be offering advantages in an asymptotic setting, there are significant benefits in the transient optimization phase, in particular in a streaming or single-epoch setting. We investigate this family of algorithms in a thorough analysis and show supporting experimental results. As a side-product we provide a simple and unified proof technique for a broad class of variance reduction algorithms.
Thomas Hofmann 0001, Aurélien Lucchi, Simon Lacoste-Julien, Brian McWilliams
NIPS4
2014 Fast and Robust Least Squares Estimation in Corrupted Linear Models
Brian McWilliams, Gabriel Krummenacher, Mario Lucic, Joachim M. Buhmann
NIPS1
2014 Subspace clustering of high-dimensional data: a predictive approach
Brian McWilliams, Giovanni Montana
Data Min. Knowl. Discov.1
2013 Correlated random features for fast semi-supervised learning
abstract
This paper presents Correlated Nystrom Views (XNV), a fast semi-supervised algorithm for regression and classification. The algorithm draws on two main ideas. First, it generates two views consisting of computationally inexpensive random features. Second, multiview regression, using Canonical Correlation Analysis (CCA) on unlabeled data, biases the regression towards useful features. It has been shown that CCA regression can substantially reduce variance with a minimal increase in bias if the views contains accurate estimators. Recent theoretical and empirical work shows that regression with random features closely approximates kernel regression, implying that the accuracy requirement holds for random views. We show that XNV consistently outperforms a state-of-the-art algorithm for semi-supervised learning: substantially improving predictive performance and reducing the variability of performance on a wide variety of real-world datasets, whilst also reducing runtime by orders of magnitude.
Brian McWilliams, David Balduzzi, Joachim M. Buhmann
NIPS1
1996 The value brokers: how to measure client/server payback
abstract
Discusses the measurement of client server payback. suggests that companies should not be using return on investment because financial yardsticks cannot measure directly the long‐term competitive advantage given by increased productivity and communication related to information systems.
Brian McWilliams
Inf. Manag. Comput. Secur.1