VLDB 2026 Research / reviewers in the wild / expert
Stephen A. Baccus
dblp:116/5417 · also Stephen Baccus
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0001-8692-5685ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 32% Deep learning architectures and training · 28% Trustworthy machine learning · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
computational neuroscience |
1.3 | 3 | 2023 | Information Geometry of the Retinal Representation Manifold · NeurIPS 2023 From deep learning to mechanistic understanding in neuroscience: the structure of retinal prediction · NeurIPS 2019 Deep Learning Models of the Retinal Response to Natural Scenes · NIPS 2016 |
Natural language and speech › Language models and text generation › efficient language model
efficient language model architectures |
0.7 | 1 | 2023 | Hyena Hierarchy: Towards Larger Convolutional Language Models · ICML 2023 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.7 | 1 | 2023 | Hyena Hierarchy: Towards Larger Convolutional Language Models · ICML 2023 |
Bioinformatics and computational biology › genomics › machine learning for genomics
genomic foundation models |
0.7 | 1 | 2023 | HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution · NeurIPS 2023 |
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling |
0.7 | 1 | 2023 | HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution · NeurIPS 2023 |
Computer vision › Image recognition and object detection
image classification |
0.6 | 1 | 2022 | S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
state space model |
0.6 | 1 | 2022 | S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces · NeurIPS 2022 |
Computer vision › Video understanding and tracking
video classification |
0.6 | 1 | 2022 | S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › interpretability
attribution methods |
0.4 | 1 | 2019 | From deep learning to mechanistic understanding in neuroscience: the structure of retinal prediction · NeurIPS 2019 |
Machine learning › Trustworthy machine learning
neural network interpretability |
0.4 | 1 | 2019 | From deep learning to mechanistic understanding in neuroscience: the structure of retinal prediction · NeurIPS 2019 |
Natural language and speech › Language models and text generation
in-context learning |
0.2 | 1 | 2023 | HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution · NeurIPS 2023 |
Information theory › information measures
fisher information |
0.2 | 1 | 2023 | Information Geometry of the Retinal Representation Manifold · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2016 | Deep Learning Models of the Retinal Response to Natural Scenes · NIPS 2016 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 1.6stochastic encoding model · 1.3pre-training · 1.3information geometry · 1.3implicit convolution · 1.3dimensionality reduction · 0.8attribution methods · 0.8state space model · 0.7single-nucleotide tokenization · 0.7single nucleotide tokenization · 0.7implicit long convolutions · 0.7data-controlled gating · 0.7continuous-signal modeling · 0.6band-limiting · 0.6linear-nonlinear model · 0.2generalized linear model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Hyena Hierarchy: Towards Larger Convolutional Language ModelsabstractRecent advances in deep learning have relied heavily on the use of large Transformers due to their ability to learn at scale. However, the core building block of Transformers, the attention operator, exhibits quadratic cost in sequence length, limiting the amount of context accessible. Existing subquadratic methods based on low-rank and sparse approximations need to be combined with dense attention layers to match Transformers at scale, indicating a gap in capability. In this work, we propose Hyena, a subquadratic drop-in replacement for attention constructed by interleaving implicitly parametrized long convolutions and data-controlled gating. In challenging reasoning tasks on sequences of thousands to hundreds of thousands of tokens, Hyena improves accuracy by more than 50 points over operators relying on state-space models, transfer functions, and other implicit and explicit methods, matching attention-based models. We set a new state-of-the-art for dense-attention-free architectures on language modeling in standard datasets WikiText103 and The Pile, reaching Transformer quality with a 20% reduction in training compute required at sequence length 2k. Hyena operators are 2x faster than highly optimized attention at sequence length 8k, with speedups of 100x at 64k. Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen A. Baccus, Yoshua Bengio, Stefano Ermon, Christopher Ré |
ICML | 6 |
| 2023 | Information Geometry of the Retinal Representation ManifoldabstractThe ability for the brain to discriminate among visual stimuli is constrained by their retinal representations. Previous studies of visual discriminability have been limited to either low-dimensional artificial stimuli or pure theoretical considerations without a realistic encoding model. Here we propose a novel framework for understanding stimulus discriminability achieved by retinal representations of naturalistic stimuli with the method of information geometry. To model the joint probability distribution of neural responses conditioned on the stimulus, we created a stochastic encoding model of a population of salamander retinal ganglion cells based on a three-layer convolutional neural network model. This model not only accurately captured the mean response to natural scenes but also a variety of second-order statistics. With the model and the proposed theory, we computed the Fisher information metric over stimuli to study the most discriminable stimulus directions. We found that the most discriminable stimulus varied substantially across stimuli, allowing an examination of the relationship between the most discriminable stimulus and the current stimulus. By examining responses generated by the most discriminable stimuli we further found that the most discriminative response mode is often aligned with the most stochastic mode. This finding carries the important implication that under natural scenes, retinal noise correlations are information-limiting rather than increasing information transmission as has been previously speculated. We additionally observed that sensitivity saturates less in the population than for single cells and that as a function of firing rate, Fisher information varies less than sensitivity. We conclude that under natural scenes, population coding benefits from complementary coding and helps to equalize the information carried by different firing rates, which may facilitate decoding of the stimulus under principles of information maximization. Xuehao Ding, Dongsoo Lee, Joshua Melander, George Sivulka, Surya Ganguli, Stephen A. Baccus |
NeurIPS | 6 |
| 2023 | HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionabstractGenomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation models in genomics to learn generalizable features from unlabeled genome data that can then be fine-tuned for downstream tasks such as identifying regulatory elements. Due to the quadratic scaling of attention, previous Transformer-based genomic models have used 512 to 4k tokens as context (<0.001% of the human genome), significantly limiting the modeling of long-range interactions in DNA. In addition, these methods rely on tokenizers or fixed k-mers to aggregate meaningful DNA units, losing single nucleotide resolution (i.e. DNA "characters") where subtle genetic variations can completely alter protein function via single nucleotide polymorphisms (SNPs). Recently, Hyena, a large language model based on implicit convolutions was shown to match attention in quality while allowing longer context lengths and lower time complexity. Leveraging Hyena’s new long-range capabilities, we present HyenaDNA, a genomic foundation model pretrained on the human reference genome with context lengths of up to 1 million tokens at the single nucleotide-level – an up to 500x increase over previous dense attention-based models. HyenaDNA scales sub-quadratically in sequence length (training up to 160x faster than Transformer), uses single nucleotide tokens, and has full global context at each layer. We explore what longer context enables - including the first use of in-context learning in genomics for simple adaptation to novel tasks without updating pretrained model weights. On fine-tuned benchmarks from the Nucleotide Transformer, HyenaDNA reaches state-of-the-art (SotA) on 12 of 18 datasets using a model with orders of magnitude less parameters and pretraining data.1 On the GenomicBenchmarks, HyenaDNA surpasses SotA on 7 of 8 datasets on average by +10 accuracy points. Code at https://github.com/HazyResearch/hyena-dna. Eric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas, Michael Wornow, Callum Birch-Sykes, Stefano Massaroli, Aman Patel, Clayton M. Rabideau, Yoshua Bengio, Stefano Ermon, Christopher Ré, Stephen A. Baccus |
NeurIPS | 13 |
| 2022 | S4ND: Modeling Images and Videos as Multidimensional Signals with State SpacesabstractVisual data such as images and videos are typically modeled as discretizations of inherently continuous, multidimensional signals. Existing continuous-signal models attempt to exploit this fact by modeling the underlying signals of visual (e.g., image) data directly. However, these models have not yet been able to achieve competitive performance on practical vision tasks such as large-scale image and video classification. Building on a recent line of work on deep state space models (SSMs), we propose \method, a new multidimensional SSM layer that extends the continuous-signal modeling ability of SSMs to multidimensional data including images and videos. We show that S4ND can model large-scale visual data in $1$D, $2$D, and $3$D as continuous multidimensional signals and demonstrates strong performance by simply swapping Conv2D and self-attention layers with \method\ layers in existing state-of-the-art models. On ImageNet-1k, \method\ exceeds the performance of a Vision Transformer baseline by $1.5\%$ when training with a $1$D sequence of patches, and matches ConvNeXt when modeling images in $2$D. For videos, S4ND improves on an inflated $3$D ConvNeXt in activity classification on HMDB-51 by $4\%$. S4ND implicitly learns global, continuous convolutional kernels that are resolution invariant by construction, providing an inductive bias that enables generalization across multiple resolutions. By developing a simple bandlimiting modification to S4 to overcome aliasing, S4ND achieves strong zero-shot (unseen at training time) resolution performance, outperforming a baseline Conv2D by $40\%$ on CIFAR-10 when trained on $8 \times 8$ and tested on $32 \times 32$ images. When trained with progressive resizing, S4ND comes within $\sim 1\%$ of a high-resolution model while training $22\%$ faster. Eric Nguyen, Karan Goel, Albert Gu, Gordon W. Downs, Preey Shah, Tri Dao, Stephen A. Baccus, Christopher Ré |
NeurIPS | 7 |
| 2019 | From deep learning to mechanistic understanding in neuroscience: the structure of retinal predictionabstractRecently, deep feedforward neural networks have achieved considerable success in modeling biological sensory processing, in terms of reproducing the input-output map of sensory neurons. However, such models raise profound questions about the very nature of explanation in neuroscience. Are we simply replacing one complex system (a biological circuit) with another (a deep network), without understanding either? Moreover, beyond neural representations, are the deep network's computational mechanisms for generating neural responses the same as those in the brain? Without a systematic approach to extracting and understanding computational mechanisms from deep neural network models, it can be difficult both to assess the degree of utility of deep learning approaches in neuroscience, and to extract experimentally testable hypotheses from deep networks. We develop such a systematic approach by combining dimensionality reduction and modern attribution methods for determining the relative importance of interneurons for specific visual computations. We apply this approach to deep network models of the retina, revealing a conceptual understanding of how the retina acts as a predictive feature extractor that signals deviations from expectations for diverse spatiotemporal stimuli. For each stimulus, our extracted computational mechanisms are consistent with prior scientific literature, and in one case yields a new mechanistic hypothesis. Thus overall, this work not only yields insights into the computational mechanisms underlying the striking predictive capabilities of the retina, but also places the framework of deep networks as neuroscientific models on firmer theoretical foundations, by providing a new roadmap to go beyond comparing neural representations to extracting and understand computational mechanisms. Hidenori Tanaka, Aran Nayebi, Niru Maheswaranathan, Lane McIntosh, Stephen A. Baccus, Surya Ganguli |
NeurIPS | 5 |
| 2018 | Inferring hidden structure in multilayered neural circuitsabstractA central challenge in sensory neuroscience involves understanding how neural circuits shape computations across cascaded cell layers. Here we attempt to reconstruct the response properties of experimentally unobserved neurons in the interior of a multilayered neural circuit, using cascaded linear-nonlinear (LN-LN) models. We combine non-smooth regularization with proximal consensus algorithms to overcome difficulties in fitting such models that arise from the high dimensionality of their parameter space. We apply this framework to retinal ganglion cell processing, learning LN-LN models of retinal circuitry consisting of thousands of parameters, using 40 minutes of responses to white noise. Our models demonstrate a 53% improvement in predicting ganglion cell spikes over classical linear-nonlinear (LN) models. Internal nonlinear subunits of the model match properties of retinal bipolar cells in both receptive field structure and number. Subunits have consistently high thresholds, supressing all but a small fraction of inputs, leading to sparse activity patterns in which only one subunit drives ganglion cell spiking at any time. From the model's parameters, we predict that the removal of visual redundancies through stimulus decorrelation across space, a central tenet of efficient coding theory, originates primarily from bipolar cell synapses. Furthermore, the composite nonlinear computation performed by retinal circuitry corresponds to a boolean OR function applied to bipolar cell feature detectors. Our methods are statistically and computationally efficient, enabling us to rapidly learn hierarchical non-linear models as well as efficiently compute widely used descriptive statistics such as the spike triggered average (STA) and covariance (STC) for high dimensional stimuli. This general computational framework may aid in extracting principles of nonlinear hierarchical sensory processing across diverse modalities from limited data. Niru Maheswaranathan, David B. Kastner, Stephen A. Baccus, Surya Ganguli |
PLoS Comput. Biol. | 3 |
| 2018 | Adaptive feature detection from differential processing in parallel retinal pathwaysabstractTo transmit information efficiently in a changing environment, the retina adapts to visual contrast by adjusting its gain, latency and mean response. Additionally, the temporal frequency selectivity, or bandwidth changes to encode the absolute intensity when the stimulus environment is noisy, and intensity differences when noise is low. We show that the On pathway of On-Off retinal amacrine and ganglion cells is required to change temporal bandwidth but not other adaptive properties. This remarkably specific adaptive mechanism arises from differential effects of contrast on the On and Off pathways. We analyzed a biophysical model fit only to a cell's membrane potential, and verified pharmacologically that it accurately revealed the two pathways. We conclude that changes in bandwidth arise mostly from differences in synaptic threshold in the two pathways, rather than synaptic release dynamics as has previously been proposed to underlie contrast adaptation. Different efficient codes are selected by different thresholds in two independently adapting neural pathways. Yusuf Ozuysal, David B. Kastner, Stephen A. Baccus |
PLoS Comput. Biol. | 3 |
| 2016 | Deep Learning Models of the Retinal Response to Natural ScenesabstractA central challenge in sensory neuroscience is to understand neural computations and circuit mechanisms that underlie the encoding of ethologically relevant, natural stimuli. In multilayered neural circuits, nonlinear processes such as synaptic transmission and spiking dynamics present a significant obstacle to the creation of accurate computational models of responses to natural stimuli. Here we demonstrate that deep convolutional neural networks (CNNs) capture retinal responses to natural scenes nearly to within the variability of a cell's response, and are markedly more accurate than linear-nonlinear (LN) models and Generalized Linear Models (GLMs). Moreover, we find two additional surprising properties of CNNs: they are less susceptible to overfitting than their LN counterparts when trained on small amounts of data, and generalize better when tested on stimuli drawn from a different distribution (e.g. between natural scenes and white noise). An examination of the learned CNNs reveals several properties. First, a richer set of feature maps is necessary for predicting the responses to natural scenes compared to white noise. Second, temporally precise responses to slowly varying inputs originate from feedforward inhibition, similar to known retinal mechanisms. Third, the injection of latent noise sources in intermediate layers enables our model to capture the sub-Poisson spiking variability observed in retinal ganglion cells. Fourth, augmenting our CNNs with recurrent lateral connections enables them to capture contrast adaptation as an emergent property of accurately describing retinal responses to natural scenes. These methods can be readily generalized to other sensory modalities and stimulus ensembles. Overall, this work demonstrates that CNNs not only accurately capture sensory circuit responses to natural scenes, but also can yield information about the circuit's internal structure and function. Lane McIntosh, Niru Maheswaranathan, Aran Nayebi, Surya Ganguli, Stephen A. Baccus |
NIPS | 5 |