Karan Goel

dblp:175/1290 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
10since 2021 · last 2023
0009-0009-0687-5805ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Deep learning architectures and training · 55% Trustworthy machine learning · 12% Reinforcement learning · 9%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 21 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
state space model
3.052023
Effectively Modeling Time Series with Simple Discrete State Spaces · ICLR 2023
S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces · NeurIPS 2022
On the Parameterization and Initialization of Diagonal State Space Models · NeurIPS 2022
Machine learning › Deep learning architectures and training
sequence modeling
2.342023
Effectively Modeling Time Series with Simple Discrete State Spaces · ICLR 2023
On the Parameterization and Initialization of Diagonal State Space Models · NeurIPS 2022
Efficiently Modeling Long Sequences with Structured State Spaces · ICLR 2022
Machine learning › Time series and sequential data
time series modeling
0.712023
Effectively Modeling Time Series with Simple Discrete State Spaces · ICLR 2023
Computer vision › Image recognition and object detection
image classification
0.612022
S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces · NeurIPS 2022
Machine learning › Deep learning architectures and training › sequence modeling
long sequence modeling
0.612022
Efficiently Modeling Long Sequences with Structured State Spaces · ICLR 2022
Machine learning › Deep learning architectures and training › state space model
structured state space model
0.612022
Efficiently Modeling Long Sequences with Structured State Spaces · ICLR 2022
Computer vision › Video understanding and tracking
video classification
0.612022
S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces · NeurIPS 2022
Machine learning › Deep learning architectures and training
data augmentation
0.512021
Model Patching: Closing the Subgroup Performance Gap with Data Augmentation · ICLR 2021
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.512021
Mandoline: Model Evaluation under Distribution Shift · ICML 2021
Machine learning › Trustworthy machine learning
fairness
0.512021
Model Patching: Closing the Subgroup Performance Gap with Data Augmentation · ICLR 2021
Machine learning › Transfer learning and domain adaptation › instance weighting
importance weighting
0.512021
Mandoline: Model Evaluation under Distribution Shift · ICML 2021
Machine learning › Trustworthy machine learning › robustness evaluation
model evaluation under distribution shift
0.512021
Mandoline: Model Evaluation under Distribution Shift · ICML 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.512021
Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers · NeurIPS 2021
Machine learning and data management › data management for machine learning
feature store
0.512021
Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems · Proc. VLDB Endow. 2021
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412019
Learning Procedural Abstractions and Evaluating Discrete Latent Temporal Structure · ICLR (Poster) 2019
Machine learning › Reinforcement learning › online decision making
optimal stopping
0.312017
Sample Efficient Policy Search for Optimal Stopping Domains · IJCAI 2017
Machine learning › Reinforcement learning
policy search
0.312017
Sample Efficient Policy Search for Optimal Stopping Domains · IJCAI 2017
Machine learning › Reinforcement learning › sample efficiency
sample-efficient policy learning
0.312017
Sample Efficient Policy Search for Optimal Stopping Domains · IJCAI 2017
Machine learning › Deep learning architectures and training › sequence modeling
long-range dependency modeling
0.212022
On the Parameterization and Initialization of Diagonal State Space Models · NeurIPS 2022
Machine learning › Deep learning architectures and training
neural differential equations
0.112021
Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers · NeurIPS 2021
Machine learning › Trustworthy machine learning
robustness
0.112021
Mandoline: Model Evaluation under Distribution Shift · ICML 2021

Methods — techniques the papers use, named apart from their topics

state space model · 1.1diffusion model · 1.1discrete state space · 0.7structured state space model · 0.6parameterization · 0.6diagonal initialization · 0.6continuous-signal modeling · 0.6band-limiting · 0.6HiPPO matrix · 0.6self-supervised pretrained embeddings · 0.5model patching · 0.5
YearPublicationVenuePosition
2023 Effectively Modeling Time Series with Simple Discrete State Spaces
Khaled Saab 0002, Michael Poli, Tri Dao, Karan Goel, Christopher Ré
ICLR5
2022 Data Management Opportunities for Foundation Models
Laurel J. Orr, Karan Goel, Christopher Ré
CIDR2
2022 Efficiently Modeling Long Sequences with Structured State Spaces
Albert Gu, Karan Goel, Christopher Ré
ICLR2
2022 It's Raw! Audio Generation with State-Space Models
abstract
Developing architectures suitable for modeling raw audio is a challenging problem due to the high sampling rates of audio waveforms. Standard sequence modeling approaches like RNNs and CNNs have previously been tailored to fit the demands of audio, but the resultant architectures make undesirable computational tradeoffs and struggle to model waveforms effectively. We propose SaShiMi, a new multi-scale architecture for waveform modeling built around the recently introduced S4 model for long sequence modeling. We identify that S4 can be unstable during autoregressive generation, and provide a simple improvement to its parameterization by drawing connections to Hurwitz matrices. SaShiMi yields state-of-the-art performance for unconditional waveform generation in the autoregressive setting. Additionally, SaShiMi improves non-autoregressive generation performance when used as the backbone architecture for a diffusion model. Compared to prior architectures in the autoregressive generation setting, SaShiMi generates piano and speech waveforms which humans find more musical and coherent respectively, e.g. 2{\texttimes} better mean opinion scores than WaveNet on an unconditional speech generation task. On a music generation task, SaShiMi outperforms WaveNet on density estimation and speed at both training and inference even when using 3{\texttimes} fewer parameters
Karan Goel, Albert Gu, Chris Donahue, Christopher Ré
ICML1
2022 On the Parameterization and Initialization of Diagonal State Space Models
abstract
State space models (SSM) have recently been shown to be very effective as a deep learning layer as a promising alternative to sequence models such as RNNs, CNNs, or Transformers. The first version to show this potential was the S4 model, which is particularly effective on tasks involving long-range dependencies by using a prescribed state matrix called the HiPPO matrix. While this has an interpretable mathematical mechanism for modeling long dependencies, it also requires a custom representation and algorithm that makes the model difficult to understand and implement. On the other hand, a recent variant of S4 called DSS showed that restricting the state matrix to be fully diagonal can still preserve the performance of the original model when using a specific initialization based on approximating S4's matrix. This work seeks to systematically understand how to parameterize and initialize diagonal state space models. While it follows from classical results that almost all SSMs have an equivalent diagonal form, we show that the initialization is critical for performance. First, we explain why DSS works mathematically, as the diagonal approximation to S4 surprisingly recovers the same dynamics in the limit of infinite state dimension. We then systematically describe various design choices in parameterizing and computing diagonal SSMs, and perform a controlled empirical study ablating the effects of these choices. Our final model S4D is a simple diagonal version of S4 whose kernel computation requires just 3 lines of code and performs comparably to S4 in almost all settings, with state-of-the-art results in image, audio, and medical time-series domains, and 85\% average on the Long Range Arena benchmark.
Albert Gu, Karan Goel, Ankit Gupta 0001, Christopher Ré
NeurIPS2
2022 S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces
abstract
Visual data such as images and videos are typically modeled as discretizations of inherently continuous, multidimensional signals. Existing continuous-signal models attempt to exploit this fact by modeling the underlying signals of visual (e.g., image) data directly. However, these models have not yet been able to achieve competitive performance on practical vision tasks such as large-scale image and video classification. Building on a recent line of work on deep state space models (SSMs), we propose \method, a new multidimensional SSM layer that extends the continuous-signal modeling ability of SSMs to multidimensional data including images and videos. We show that S4ND can model large-scale visual data in $1$D, $2$D, and $3$D as continuous multidimensional signals and demonstrates strong performance by simply swapping Conv2D and self-attention layers with \method\ layers in existing state-of-the-art models. On ImageNet-1k, \method\ exceeds the performance of a Vision Transformer baseline by $1.5\%$ when training with a $1$D sequence of patches, and matches ConvNeXt when modeling images in $2$D. For videos, S4ND improves on an inflated $3$D ConvNeXt in activity classification on HMDB-51 by $4\%$. S4ND implicitly learns global, continuous convolutional kernels that are resolution invariant by construction, providing an inductive bias that enables generalization across multiple resolutions. By developing a simple bandlimiting modification to S4 to overcome aliasing, S4ND achieves strong zero-shot (unseen at training time) resolution performance, outperforming a baseline Conv2D by $40\%$ on CIFAR-10 when trained on $8 \times 8$ and tested on $32 \times 32$ images. When trained with progressive resizing, S4ND comes within $\sim 1\%$ of a high-resolution model while training $22\%$ faster.
Eric Nguyen, Karan Goel, Albert Gu, Gordon W. Downs, Preey Shah, Tri Dao, Stephen A. Baccus, Christopher Ré
NeurIPS2
2021 Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
Karan Goel, Albert Gu, Yixuan Li 0001, Christopher Ré
ICLR1
2021 Mandoline: Model Evaluation under Distribution Shift
abstract
Machine learning models are often deployed in different settings than they were trained and validated on, posing a challenge to practitioners who wish to predict how well the deployed model will perform on a target distribution. If an unlabeled sample from the target distribution is available, along with a labeled sample from a possibly different source distribution, standard approaches such as importance weighting can be applied to estimate performance on the target. However, importance weighting struggles when the source and target distributions have non-overlapping support or are high-dimensional. Taking inspiration from fields such as epidemiology and polling, we develop Mandoline, a new evaluation framework that mitigates these issues. Our key insight is that practitioners may have prior knowledge about the ways in which the distribution shifts, which we can use to better guide the importance weighting procedure. Specifically, users write simple "slicing functions" {–} noisy, potentially correlated binary functions intended to capture possible axes of distribution shift {–} to compute reweighted performance estimates. We further describe a density ratio estimation framework for the slices and show how its estimation error scales with slice quality and dataset size. Empirical validation on NLP and vision tasks shows that Mandoline can estimate performance on the target distribution up to 3x more accurately compared to standard baselines.
Mayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms, Kayvon Fatahalian, Christopher Ré
ICML2
2021 Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers
abstract
Recurrent neural networks (RNNs), temporal convolutions, and neural differential equations (NDEs) are popular families of deep learning models for time-series data, each with unique strengths and tradeoffs in modeling power and computational efficiency. We introduce a simple sequence model inspired by control systems that generalizes these approaches while addressing their shortcomings. The Linear State-Space Layer (LSSL) maps a sequence $u \mapsto y$ by simply simulating a linear continuous-time state-space representation $\dot{x} = Ax + Bu, y = Cx + Du$. Theoretically, we show that LSSL models are closely related to the three aforementioned families of models and inherit their strengths. For example, they generalize convolutions to continuous-time, explain common RNN heuristics, and share features of NDEs such as time-scale adaptation. We then incorporate and generalize recent theory on continuous-time memorization to introduce a trainable subset of structured matrices $A$ that endow LSSLs with long-range memory. Empirically, stacking LSSL layers into a simple deep neural network obtains state-of-the-art results across time series benchmarks for long dependencies in sequential image classification, real-world healthcare regression tasks, and speech. On a difficult speech classification task with length-16000 sequences, LSSL outperforms prior approaches by 24 accuracy points, and even outperforms baselines that use hand-crafted features on 100x shorter sequences.
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab 0002, Tri Dao, Atri Rudra, Christopher Ré
NeurIPS3
2021 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems
abstract
The industrial machine learning pipeline requires iterating on model features, training and deploying models, and monitoring deployed models at scale. Feature stores were developed to manage and standardize the engineer's workflow in this end-to-end pipeline, focusing on traditional tabular feature data. In recent years, however, model development has shifted towards using self-supervised pretrained embeddings as model features. Managing these embeddings and the downstream systems that use them introduces new challenges with respect to managing embedding training data, measuring embedding quality, and monitoring downstream models that use embeddings. These challenges are largely unaddressed in standard feature stores. Our goal in this tutorial is to introduce the feature store system and discuss the challenges and current solutions to managing these new embedding-centric pipelines.
Laurel J. Orr, Atindriyo Sanyal, Karan Goel, Megan Leszczynski
Proc. VLDB Endow.4
2019 Learning Procedural Abstractions and Evaluating Discrete Latent Temporal Structure
Karan Goel, Emma Brunskill
ICLR (Poster)1
2017 Octopus: A Framework for Cost-Quality-Time Optimization in Crowdsourcing
abstract
We present Octopus, an AI agent to jointly balance three conflicting task objectives on a micro-crowdsourcing marketplace – the quality of work, total cost incurred, and time to completion. Previous control agents have mostly focused on cost-quality, or cost-time tradeoffs, but not on directly controlling all three in concert. A naive formulation of three-objective optimization is intractable; Octopus takes a hierarchical POMDP approach, with three different components responsible for setting the pay per task, selecting the next task, and controlling task-level quality. We demonstrate that Octopus significantly outperforms existing state-of-the-art approaches on real experiments. We also deploy Octopus on Amazon Mechanical Turk, showing its ability to manage tasks in a real-world, dynamic setting.
Karan Goel, Shreya Rajpal, Mausam
HCOMP1
2017 Sample Efficient Policy Search for Optimal Stopping Domains
abstract
Optimal stopping problems consider the question of deciding when to stop an observation-generating process in order to maximize a return. We examine the problem of simultaneously learning and planning in such domains, when data is collected directly from the environment. We propose GFSE, a simple and flexible model-free policy search method that reuses data for sample efficiency by leveraging problem structure. We bound the sample complexity of our approach to guarantee uniform convergence of policy value estimates, tightening existing PAC bounds to achieve logarithmic dependence on horizon length for our setting. We also examine the benefit of our method against prevalent model-based and model-free approaches on 3 domains taken from diverse fields.
Karan Goel, Christoph Dann, Emma Brunskill
IJCAI1