Brian Cheung

dblp:34/7493 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
7since 2021 · last 2025
0009-0000-7771-1618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Representation and self-supervised learning · 36% Transfer learning and domain adaptation · 13% Efficient and distributed learning · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 24 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.022023
Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks · ICML 2023
Superposition of many models into one · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.912025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Machine learning › Learning theory
inductive bias
0.912025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
knowledge transfer
0.912025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation alignment
0.812024
Position: The Platonic Representation Hypothesis · ICML 2024
Machine learning › Representation and self-supervised learning › representation analysis
representation similarity
0.712023
System Identification of Neural Systems: If We Got It Right, Would We Know? · ICML 2023
Machine learning › Optimization for machine learning › gradient estimation
straight-through estimator
0.712023
Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks · ICML 2023
Machine learning › Representation and self-supervised learning
vector quantization
0.712023
Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks · ICML 2023
Bioinformatics and computational biology
computational neuroscience
0.712023
System Identification of Neural Systems: If We Got It Right, Would We Know? · ICML 2023
Bioinformatics and computational biology › computational neuroscience › neural response modeling
neural system identification
0.712023
System Identification of Neural Systems: If We Got It Right, Would We Know? · ICML 2023
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.612022
Equivariant Self-Supervised Learning: Encouraging Equivariance in Representations · ICLR 2022
Machine learning › Reinforcement learning
safe reinforcement learning
0.412020
Cautious Adaptation For Reinforcement Learning in Safety-Critical Settings · ICML 2020
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
synthetic-to-real domain adaptation
0.412020
Cautious Adaptation For Reinforcement Learning in Safety-Critical Settings · ICML 2020
Machine learning › Optimization for machine learning
learned optimizer
0.412019
Meta-Learning Update Rules for Unsupervised Representation Learning · ICLR 2019
Machine learning › Reinforcement learning › meta-reinforcement learning
learned update rules
0.412019
Meta-Learning Update Rules for Unsupervised Representation Learning · ICLR 2019
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412019
Meta-Learning Update Rules for Unsupervised Representation Learning · ICLR 2019
Machine learning › Efficient and distributed learning
parameter sharing
0.412019
Superposition of many models into one · NeurIPS 2019
Machine learning › Trustworthy machine learning › neural network interpretability
superposition
0.412019
Superposition of many models into one · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.412019
Meta-Learning Update Rules for Unsupervised Representation Learning · ICLR 2019
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.312018
Adversarial Examples that Fool both Computer Vision and Time-Limited Humans · NeurIPS 2018
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial transferability
0.312018
Adversarial Examples that Fool both Computer Vision and Time-Limited Humans · NeurIPS 2018
Machine learning › Trustworthy machine learning
robustness
0.312018
Adversarial Examples that Fool both Computer Vision and Time-Limited Humans · NeurIPS 2018
Machine learning › Deep learning architectures and training › attention mechanism
attention learning
0.312017
Emergence of foveal image sampling from learning to attend in visual scenes · ICLR (Poster) 2017
Computer vision › Image recognition and object detection
object recognition
0.312025
Training the Untrainable: Introducing Inductive Bias via Representational Alignment · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

linear encoding model · 1.3centered kernel alignment · 1.3neural distance function · 0.9layerwise representational similarity · 0.9knowledge distillation · 0.9representational similarity analysis · 0.8commitment loss · 0.7alternating optimization · 0.7affine re-parameterization · 0.7contrastive learning · 0.6
YearPublicationVenuePosition
2025 Training the Untrainable: Introducing Inductive Bias via Representational Alignment
abstract
We demonstrate that architectures which traditionally are considered to be ill-suited for a task can be trained using inductive biases from another architecture. We call a network untrainable when it overfits, underfits, or converges to poor results even when tuning their hyperparameters. For example, fully connected networks overfit on object recognition while deep convolutional networks without residual connections underfit. The traditional answer is to change the architecture to impose some inductive bias, although the nature of that bias is unknown. We introduce guidance, where a guide network steers a target network using a neural distance function. The target minimizes its task loss plus a layerwise representational similarity against the frozen guide. If the guide is trained, this transfers over the architectural prior and knowledge of the guide to the target. If the guide is untrained, this transfers over only part of the architectural prior of the guide. We show that guidance prevents FCN overfitting on ImageNet, narrows the vanilla RNN–Transformer gap, boosts plain CNNs toward ResNet accuracy, and aids Transformers on RNN-favored tasks. We further identify that guidance-driven initialization alone can mitigate FCN overfitting. Our method provides a mathematical tool to investigate priors and architectures, and in the long term, could automate architecture design.
Vighnesh Subramaniam, David Mayo, Colin Conwell, Tomaso A. Poggio, Boris Katz, Brian Cheung, Andrei Barbu
NeurIPS6
2024 Position: The Platonic Representation Hypothesis
abstract
We argue that representations in AI models, particularly deep networks, are converging. First, we survey many examples of convergence in the literature: over time and across multiple domains, the ways by which different neural networks represent data are becoming more aligned. Next, we demonstrate convergence across data modalities: as vision models and language models get larger, they measure distance between datapoints in a more and more alike way. We hypothesize that this convergence is driving toward a shared statistical model of reality, akin to Plato’s concept of an ideal reality. We term such a representation the platonic representation and discuss several possible selective pressures toward it. Finally, we discuss the implications of these trends, their limitations, and counterexamples to our analysis.
Minyoung Huh, Brian Cheung, Tongzhou Wang 0001, Phillip Isola
ICML2
2023 System Identification of Neural Systems: If We Got It Right, Would We Know?
abstract
Artificial neural networks are being proposed as models of parts of the brain. The networks are compared to recordings of biological neurons, and good performance in reproducing neural responses is considered to support the model’s validity. A key question is how much this system identification approach tells us about brain computation. Does it validate one model architecture over another? We evaluate the most commonly used comparison techniques, such as a linear encoding model and centered kernel alignment, to correctly identify a model by replacing brain recordings with known ground truth models. System identification performance is quite variable; it also depends significantly on factors independent of the ground truth architecture, such as stimuli images. In addition, we show the limitations of using functional similarity scores in identifying higher-level architectural motifs.
Yena Han, Tomaso A. Poggio, Brian Cheung
ICML3
2023 Straightening Out the Straight-Through Estimator: Overcoming Optimization Challenges in Vector Quantized Networks
abstract
This work examines the challenges of training neural networks using vector quantization using straight-through estimation. We find that the main cause of training instability is the discrepancy between the model embedding and the code-vector distribution. We identify the factors that contribute to this issue, including the codebook gradient sparsity and the asymmetric nature of the commitment loss, which leads to misaligned code-vector assignments. We propose to address this issue via affine re-parameterization of the code vectors. Additionally, we introduce an alternating optimization to reduce the gradient error introduced by the straight-through estimation. Moreover, we propose an improvement to the commitment loss to ensure better alignment between the codebook representation and the model embedding. These optimization methods improve the mathematical approximation of the straight-through estimation and, ultimately, the model performance. We demonstrate the effectiveness of our methods on several common model architectures, such as AlexNet, ResNet, and ViT, across various tasks, including image classification and generative modeling.
Minyoung Huh, Brian Cheung, Pulkit Agrawal 0001, Phillip Isola
ICML2
2023 Compact and Optimal Deep Learning with Recurrent Parameter Generators
abstract
Deep learning has achieved tremendous success by training increasingly large models, which are then compressed for practical deployment. We propose a drastically different approach to compact and optimal deep learning: We decouple the Degrees of freedom (DoF) and the actual number of parameters of a model, optimize a small DoF with predefined random linear constraints for a large model of an arbitrary architecture, in one-stage end-to-end learning.Specifically, we create a recurrent parameter generator (RPG), which repeatedly fetches parameters from a ring and unpacks them onto a large model with random permutation and sign flipping to promote parameter decorrelation. We show that gradient descent can automatically find the best model under constraints with in fact faster convergence.Our extensive experimentation reveals a log-linear relationship between model DoF and accuracy. Our RPG demonstrates remarkable DoF reduction, and can be further pruned and quantized for additional run-time performance gain. For example, in terms of top-1 accuracy on ImageNet, RPG achieves 96% of ResNet18’s performance with only 18% DoF (the equivalent of one convolutional layer) and 52% of ResNet34’s performance with only 0.25% DoF! Our work shows significant potential of constrained neural opti-mization in compact and optimal deep learning.
Yubei Chen, Stella X. Yu, Brian Cheung, Yann LeCun
WACV4
2022 Equivariant Self-Supervised Learning: Encouraging Equivariance in Representations
Rumen Dangovski, Li Jing 0001, Charlotte Loh, Seungwook Han, Akash Srivastava, Brian Cheung, Pulkit Agrawal 0001, Marin Soljacic
ICLR6
2021 Adapting Deep Learning Models to New Meteorological Contexts Using Transfer Learning
abstract
Meteorological applications such as precipitation nowcasting, synthetic radar generation, statistical downscaling and others have benefited from deep learning (DL) approaches, however several challenges remain for widespread adaptation of these complex models in operational systems. One of these challenges is adequate generalizability; deep learning models trained from datasets collected in specific contexts should not be expected to perform as well when applied to different contexts required by large operational systems. One obvious mitigation for this is to collect massive amounts of training data that cover all expected meteorological contexts, however this is not only costly and difficult to manage, but is also not possible in many parts of the globe where certain sensing platforms are sparse. In this paper, we describe an application of transfer learning to perform domain transfer for deep learning models. We demonstrate a transfer learning algorithm called weight superposition to adapt a Convolutional Neural Network trained in a source context to a new target context. Weight superposition is a method for storing multiple models within a single set of parameters thus greatly simplifying model maintenance and training. This approach also addresses the issue of catastrophic forgetting where a model, once adapted to a new context, performs poorly in the original context. We apply weight superposition to the problem of synthetic weather radar generation and show that in scenarios where the target context has less data, a model adapted with weight superposition is better at maintaining performance when compared to simpler methods. Conversely, the simple adapted model performs better on the source context when the source and target contexts have comparable amounts of data.
Pooya Khorrami, Olga Simek, Brian Cheung, Mark Veillette, Rumen Dangovski, Ileana Rugina, Marin Soljacic, Pulkit Agrawal 0001
IEEE BigData3
2020 Cautious Adaptation For Reinforcement Learning in Safety-Critical Settings
abstract
Reinforcement learning (RL) in real-world safety-critical target settings like urban driving is hazardous, imperiling the RL agent, other agents, and the environment. To overcome this difficulty, we propose a "safety-critical adaptation" task setting: an agent first trains in non-safety-critical "source" environments such as in a simulator, before it adapts to the target environment where failures carry heavy costs. We propose a solution approach, CARL, that builds on the intuition that prior experience in diverse environments equips an agent to estimate risk, which in turn enables relative safety through risk-averse, cautious adaptation. CARL first employs model-based RL to train a probabilistic model to capture uncertainty about transition dynamics and catastrophic states across varied source environments. Then, when exploring a new safety-critical environment with unknown dynamics, the CARL agent plans to avoid actions that could lead to catastrophic states. In experiments on car driving, cartpole balancing, and half-cheetah locomotion, CARL successfully acquires cautious exploration behaviors, yielding higher rewards with fewer failures than strong RL adaptation baselines.
Jesse Zhang, Brian Cheung, Chelsea Finn, Sergey Levine, Dinesh Jayaraman
ICML2
2019 Meta-Learning Update Rules for Unsupervised Representation Learning
Luke Metz, Niru Maheswaranathan, Brian Cheung, Jascha Sohl-Dickstein
ICLR3
2019 Superposition of many models into one
abstract
We present a method for storing multiple models within a single set of parameters. Models can coexist in superposition and still be retrieved individually. In experiments with neural networks, we show that a surprisingly large number of models can be effectively stored within a single parameter instance. Furthermore, each of these models can undergo thousands of training steps without significantly interfering with other models within the superposition. This approach may be viewed as the online complement of compression: rather than reducing the size of a network after training, we make use of the unrealized capacity of a network during training.
Brian Cheung, Alexander Terekhov, Yubei Chen, Pulkit Agrawal 0001, Bruno A. Olshausen
NeurIPS1
2018 Probabilistic Multi Hypothesis Tracker for an Event Based Sensor
abstract
The Event-Based Sensor (EBS) is a new class of imaging sensor where each pixel independently reports “events” in response to changes in log intensity, rather than outputting image frames containing the absolute intensity at each pixel. Positive and negative events are emitted from the sensor when the change in log intensity exceeds certain controllable thresholds internal to the device. For objects moving through the field of view, a change in intensity can be related to motion. The sensor records events independently and asynchronously for each pixel with a very high temporal resolution, allowing the detection of objects moving very quickly through the field of view. Recently this type of sensor has been applied to the detection of orbiting space objects using a ground-based telescope. This paper describes a method to treat the data generated by the EBS as a classical detect-then-track problem by collating the events spatially and temporally to form target measurements. An efficient multi-target tracking algorithm, the probabilistic multi-hypothesis tracker (PMHT) is then applied to the EBS measurements to produce tracks. This method is demonstrated by automatically generating tracks on orbiting space objects from data collected by the EBS.
Brian Cheung, Mark Rutten, Samuel J. Davey, Greg Cohen
FUSION1
2018 Adversarial Examples that Fool both Computer Vision and Time-Limited Humans
abstract
Machine learning models are vulnerable to adversarial examples: small changes to images can cause computer vision models to make mistakes such as identifying a school bus as an ostrich. However, it is still an open question whether humans are prone to similar mistakes. Here, we address this question by leveraging recent techniques that transfer adversarial examples from computer vision models with known parameters and architecture to other models with unknown parameters and architecture, and by matching the initial processing of the human visual system. We find that adversarial examples that strongly transfer across computer vision models influence the classifications made by time-limited human observers.
Gamaleldin F. Elsayed, Shreya Shankar, Brian Cheung, Nicolas Papernot, Alexey Kurakin, Ian J. Goodfellow, Jascha Sohl-Dickstein
NeurIPS3
2017 Emergence of foveal image sampling from learning to attend in visual scenes
Brian Cheung, Eric Weiss, Bruno A. Olshausen
ICLR (Poster)1
2012 Convolutional Neural Networks Applied to Human Face Classification
abstract
Convolutional neural network models have covered a broad scope of computer vision applications, achieving competitive performance with minimal domain knowledge. In this work, we apply such a model to a task designed to deter automated systems. We trained a convolutional neural network to distinguish between images of human faces from computer generated avatars as part of the ICMLA 2012 Face Recognition Challenge. The network achieved a classification accuracy of 99\% on the \textit{Avatar CAPTCHA} dataset. Furthermore, we demonstrated the potential of utilizing support vector machines on the same problem and achieved equally competitive performance.
Brian Cheung
ICMLA (2)1
2010 PMHT for tracking with timing uncertainty
Brian Cheung, Samuel J. Davey, Douglas A. Gray 0001
FUSION1
2010 Comparison of the PMHT path planning algorithm with the Genetic Algorithm for multiple platforms
Brian Cheung, Samuel J. Davey, Douglas A. Gray 0001
FUSION1
2009 Combining PMHT with classifications to perform SLAM
Brian Cheung, Samuel J. Davey, Douglas A. Gray 0001
FUSION1
2009 Track-Before-Detect for sensors with complex measurements
Samuel J. Davey, Brian Cheung, Mark Rutten
FUSION2
2008 A comparison of detection performance for several Track-Before-Detect algorithms
Samuel J. Davey, Mark Rutten, Brian Cheung
FUSION3