Ronald G. Junkins

dblp:359/2513 · also Ronald Guenther Junkins · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 26% Optimization for machine learning · 26% Deep learning architectures and training · 13%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
backpropagation
0.912025
Restructuring Vector Quantization with the Rotation Trick · ICLR 2025
Machine learning › Optimization for machine learning
gradient estimation
0.912025
Restructuring Vector Quantization with the Rotation Trick · ICLR 2025
Machine learning › Optimization for machine learning › gradient estimation
straight-through estimator
0.912025
Restructuring Vector Quantization with the Rotation Trick · ICLR 2025
Machine learning › Generative modeling
variational autoencoder
0.912025
Restructuring Vector Quantization with the Rotation Trick · ICLR 2025
Machine learning › Generative modeling › variational autoencoder
vector-quantized variational autoencoder
0.912025
Restructuring Vector Quantization with the Rotation Trick · ICLR 2025
Natural language and speech › Language models and text generation
in-context learning
0.812024
Context-Aware Meta-Learning · ICLR 2024
Machine learning › Transfer learning and domain adaptation
meta-learning
0.812024
Context-Aware Meta-Learning · ICLR 2024

Methods — techniques the papers use, named apart from their topics

vector quantization · 0.9straight-through estimator · 0.9rotation trick · 0.9sequence modeling · 0.8meta-learning · 0.8in-context learning · 0.8
YearPublicationVenuePosition
2025 Restructuring Vector Quantization with the Rotation Trick
abstract
Vector Quantized Variational AutoEncoders (VQ-VAEs) are designed to compress a continuous input to a discrete latent space and reconstruct it with minimal distortion. They operate by maintaining a set of vectors---often referred to as the codebook---and quantizing each encoder output to the nearest vector in the codebook. However, as vector quantization is non-differentiable, the gradient to the encoder flows _around_ the vector quantization layer rather than _through_ it in a straight-through approximation. This approximation may be undesirable as all information from the vector quantization operation is lost. In this work, we propose a way to propagate gradients through the vector quantization layer of VQ-VAEs. We smoothly transform each encoder output into its corresponding codebook vector via a rotation and rescaling linear transformation that is treated as a constant during backpropagation. As a result, the relative magnitude and angle between encoder output and codebook vector becomes encoded into the gradient as it propagates through the vector quantization layer and back to the encoder. Across 11 different VQ-VAE training paradigms, we find this restructuring improves reconstruction metrics, codebook utilization, and quantization error.
Christopher Fifty, Ronald G. Junkins, Dennis Duan, Aniketh Iyengar, Jerry W. Liu, Ehsan Amid, Sebastian Thrun, Christopher Ré
ICLR2
2024 Context-Aware Meta-Learning
abstract
Large Language Models like ChatGPT demonstrate a remarkable capacity to learn new concepts during inference without any fine-tuning. However, visual models trained to detect new objects during inference have been unable to replicate this ability, and instead either perform poorly or require meta-training and/or fine-tuning on similar objects. In this work, we propose a meta-learning algorithm that emulates Large Language Models by learning new visual concepts during inference without fine-tuning. Our approach leverages a frozen pre-trained feature extractor, and analogous to in-context learning, recasts meta-learning as sequence modeling over datapoints with known labels and a test datapoint with an unknown label. On 8 out of 11 meta-learning benchmarks, our approach---without meta-training or fine-tuning---exceeds or matches the state-of-the-art algorithm, P>M>F, which is meta-trained on these benchmarks.
Christopher Fifty, Dennis Duan, Ronald G. Junkins, Ehsan Amid, Jure Leskovec, Christopher Ré, Sebastian Thrun
ICLR3