David E. Carlson

dblp:30/11107 · also David Edwin Carlson · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0003-1005-6385ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Probabilistic and Bayesian machine learning · 47% 3D vision · 14% Generative modeling · 13%
Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 92% Computational social science and digital humanities · 8%
Theoretical computer science
2 papers
Information theory · 62% Mathematical optimization · 38%

Topics — the 30 heaviest of 69, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
computational neuroscience
1.552021
Directed Spectrum Measures Improve Latent Network Models Of Neural Populations · NeurIPS 2021
YASS: Yet Another Spike Sorter · NIPS 2017
Cross-Spectral Factor Analysis · NIPS 2017
Machine learning › Trustworthy machine learning
uncertainty estimation
1.222023
Estimating Causal Effects using a Multi-task Deep Ensemble · ICML 2023
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.912025
MOTTO: A Mixture-of-Experts Framework for Multi-Treatment, Multi-Outcome Treatment Effect Estimation · KDD (2) 2025
Computer vision › 3D vision
pose estimation
0.912025
Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation
0.912025
MOTTO: A Mixture-of-Experts Framework for Multi-Treatment, Multi-Outcome Treatment Effect Estimation · KDD (2) 2025
Machine learning › Generative modeling
generative adversarial network
0.722019
StoryGAN: A Sequential Conditional GAN for Story Visualization · CVPR 2019
Video Generation From Text · AAAI 2018
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal effect estimation
0.712023
Estimating Causal Effects using a Multi-task Deep Ensemble · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.632017
YASS: Yet Another Spike Sorter · NIPS 2017
Real-Time Inference for a Gamma Process Model of Neural Spiking · NIPS 2013
On the Analysis of Multi-Channel Neural Spike Data · NIPS 2011
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
neural signal analysis
0.622017
Targeting EEG/LFP Synchrony with Neural Nets · NIPS 2017
Cross-Spectral Factor Analysis · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.522017
Stochastic Bouncy Particle Sampler · ICML 2017
Partition Functions from Rao-Blackwellized Tempered Sampling · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian prediction
predictive distribution
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
0.512021
Estimating Uncertainty Intervals from Collaborating Networks · J. Mach. Learn. Res. 2021
Bioinformatics and computational biology › computational neuroscience
neural population analysis
0.512021
Directed Spectrum Measures Improve Latent Network Models Of Neural Populations · NeurIPS 2021
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
spike sorting
0.522017
YASS: Yet Another Spike Sorter · NIPS 2017
Real-Time Inference for a Gamma Process Model of Neural Spiking · NIPS 2013
Machine learning › Graph learning
dynamic graph learning
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Graph learning
graph representation learning
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Graph learning
network embedding
0.412020
Dynamic Embedding on Textual Networks via a Gaussian Process · AAAI 2020
Machine learning › Generative modeling › cross-modal generation
story visualization
0.412019
StoryGAN: A Sequential Conditional GAN for Story Visualization · CVPR 2019
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.412019
StoryGAN: A Sequential Conditional GAN for Story Visualization · CVPR 2019
Information theory › information measures
mutual information
0.422014
A Bregman Matrix and the Gradient of Mutual Information for Vector Poisson and Gaussian Channels · IEEE Trans. Inf. Theory 2014
Designed Measurements for Vector Count Data · NIPS 2013
Machine learning › Transfer learning and domain adaptation
domain generalization
0.312018
Extracting Relationships by Multi-Domain Matching · NeurIPS 2018
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.312018
Extracting Relationships by Multi-Domain Matching · NeurIPS 2018
Machine learning › Generative modeling › video generation
text-to-video generation
0.312018
Video Generation From Text · AAAI 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Video Generation From Text · AAAI 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.312017
Stochastic Bouncy Particle Sampler · ICML 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
Targeting EEG/LFP Synchrony with Neural Nets · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.312017
YASS: Yet Another Spike Sorter · NIPS 2017

Methods — techniques the papers use, named apart from their topics

mixture of experts · 1.7gaussian process · 1.0shape carving · 0.9rotation-invariant embedding · 0.93d gaussian splatting · 0.9multi-task gaussian process · 0.7deep ensembles · 0.7neural network detection · 0.6matching pursuit deconvolution · 0.6coreset · 0.6neural network · 0.5linear factor model · 0.5directed spectrum estimation · 0.5cumulative distribution function estimation · 0.5asymptotic consistency analysis · 0.5poisson signal model · 0.5gradient-based optimization · 0.5bregman divergence · 0.4
YearPublicationVenuePosition
2025 MOTTO: A Mixture-of-Experts Framework for Multi-Treatment, Multi-Outcome Treatment Effect Estimation
abstract
Multi-treatment multi-outcome treatment effect estimation plays a vital role in today's industry-level applications. For example, in social media ads, practitioners simultaneously deploy multiple interventions to users' experience and track multi-faceted metrics (e.g., ad performance, engagement, churn). However, existing methods for estimating treatment effects struggle to simultaneously address the complex interplays and ensure robust counterfactual balancing across treatment-outcome pairs.
Yiling Liu, Wei Shi 0011, Ziyang Jiang, Zhigang Hua, David E. Carlson
KDD (2)6
2025 Pose Splatter: A 3D Gaussian Splatting Model for Quantifying Animal Pose and Appearance
abstract
Accurate and scalable quantification of animal pose and appearance is crucial for studying behavior. Current 3D pose estimation techniques, such as keypoint- and mesh-based techniques, often face challenges including limited representational detail, labor-intensive annotation requirements, and expensive per-frame optimization. These limitations hinder the study of subtle movements and can make large-scale analyses impractical. We propose *Pose Splatter*, a novel framework leveraging shape carving and 3D Gaussian splatting to model the complete pose and appearance of laboratory animals without prior knowledge of animal geometry, per-frame optimization, or manual annotations. We also propose a rotation-invariant visual embedding technique for encoding pose and appearance, designed to be a plug-in replacement for 3D keypoint data in downstream behavioral analyses. Experiments on datasets of mice, rats, and zebra finches show *Pose Splatter* learns accurate 3D animal geometries. Notably, *Pose Splatter* represents subtle variations in pose, provides better low-dimensional pose embeddings over state-of-the-art as evaluated by humans, and generalizes to unseen data. By eliminating annotation and per-frame optimization bottlenecks, *Pose Splatter* enables analysis of large-scale, longitudinal behavior needed to map genotype, neural activity, and behavior at high resolutions.
Jack Goffinet, Youngjo Min, Carlo Tomasi, David E. Carlson
NeurIPS4
2025 Efficient few-shot medical image segmentation via self-supervised variational autoencoder
Yanjie Zhou, Fengjun Xi, David E. Carlson, Liyun Tu
Medical Image Anal.6
2025 Scale-free and unbiased transformer with tokenization for cell type annotation from single-cell RNA-seq data
Ziyang Jiang, Liyun Tu, David E. Carlson
Pattern Recognit.5
2023 Estimating Causal Effects using a Multi-task Deep Ensemble
abstract
A number of methods have been proposed for causal effect estimation, yet few have demonstrated efficacy in handling data with complex structures, such as images. To fill this gap, we propose Causal Multi-task Deep Ensemble (CMDE), a novel framework that learns both shared and group-specific information from the study population. We provide proofs demonstrating equivalency of CDME to a multi-task Gaussian process (GP) with a coregionalization kernel a priori. Compared to multi-task GP, CMDE efficiently handles high-dimensional and multi-modal covariates and provides pointwise uncertainty estimates of causal effects. We evaluate our method across various types of datasets and tasks and find that CMDE outperforms state-of-the-art methods on a majority of these tasks.
Ziyang Jiang, Zhuoran Hou, Yiling Liu, Yiman Ren, David E. Carlson
ICML6
2022 Learning to Weight Filter Groups for Robust Classification
abstract
In many real-world tasks, a canonical “big data” problem is created by combining data from several individual groups or domains. Because test data will likely come from a new group of data, we want to utilize the grouped structure of our training data to enforce generalization between groups of data, not just individual samples. This can be viewed as a multiple-domain generalization problem. Specifically, the goal is to encourage generalization between previously seen labeled source data from multiple domains and unlabeled target domain data. To address this challenge, we introduce Domain-Specific Filter Group (DSFG), where each training domain has a unique filter group and each test data point is predicted by a weighted sum over the outputs of different domain filters. A separate neural network learns to estimate the appropriate filter group weights through a meta-learning strategy. Empirically, experiments on three benchmark datasets demonstrate improved performance compared to current state-of-the-art approaches.
Siyang Yuan, Yitong Li 0001, Dong Wang 0037, Ke Bai 0001, Lawrence Carin, David E. Carlson
WACV6
2021 Directed Spectrum Measures Improve Latent Network Models Of Neural Populations
abstract
Systems neuroscience aims to understand how networks of neurons distributed throughout the brain mediate computational tasks. One popular approach to identify those networks is to first calculate measures of neural activity (e.g. power spectra) from multiple brain regions, and then apply a linear factor model to those measures. Critically, despite the established role of directed communication between brain regions in neural computation, measures of directed communication have been rarely utilized in network estimation because they are incompatible with the implicit assumptions of the linear factor model approach. Here, we develop a novel spectral measure of directed communication called the Directed Spectrum (DS). We prove that it is compatible with the implicit assumptions of linear factor models, and we provide a method to estimate the DS. We demonstrate that latent linear factor models of DS measures better capture underlying brain networks in both simulated and real neural recording data compared to available alternatives. Thus, linear factor models of the Directed Spectrum offer neuroscientists a simple and effective way to explicitly model directed communication in networks of neural populations.
Neil Gallagher, Kafui Dzirasa, David E. Carlson
NeurIPS3
2021 Estimating Uncertainty Intervals from Collaborating Networks
abstract
Effective decision making requires understanding the uncertainty inherent in a prediction. In regression, this uncertainty can be estimated by a variety of methods; however, many of these methods are laborious to tune, generate overconfident uncertainty intervals, or lack sharpness (give imprecise intervals). We address these challenges by proposing a novel method to capture predictive distributions in regression by defining two neural networks with two distinct loss functions. Specifically, one network approximates the cumulative distribution function, and the second network approximates its inverse. We refer to this method as Collaborating Networks (CN). Theoretical analysis demonstrates that a fixed point of the optimization is at the idealized solution, and that the method is asymptotically consistent to the ground truth distribution. Empirically, learning is straightforward and robust. We benchmark CN against several common approaches on two synthetic and six real-world datasets, including forecasting A1c values in diabetic patients from electronic health records, where uncertainty is critical. In the synthetic data, the proposed approach essentially matches ground truth. In the real-world datasets, CN improves results on many performance metrics, including log-likelihood estimates, mean absolute errors, coverage estimates, and prediction interval widths.
Tianhui Zhou, Yitong Li 0001, David E. Carlson
J. Mach. Learn. Res.4
2020 Dynamic Embedding on Textual Networks via a Gaussian Process
abstract
Textual network embedding aims to learn low-dimensional representations of text-annotated nodes in a graph. Prior work in this area has typically focused on fixed graph structures; however, real-world networks are often dynamic. We address this challenge with a novel end-to-end node-embedding model, called Dynamic Embedding for Textual Networks with a Gaussian Process (DetGP). After training, DetGP can be applied efficiently to dynamic graphs without re-training or backpropagation. The learned representation of each node is a combination of textual and structural embeddings. Because the structure is allowed to be dynamic, our method uses the Gaussian process to take advantage of its non-parametric properties. To use both local and global graph structures, diffusion is used to model multiple hops between neighbors. The relative importance of global versus local structure for the embeddings is learned automatically. With the non-parametric nature of the Gaussian process, updating the embeddings for a changed graph structure requires only a forward pass through the learned model. Considering link prediction and node classification, experiments demonstrate the empirical effectiveness of our method compared to baseline approaches. We further show that DetGP can be straightforwardly and efficiently applied to dynamic textual networks.
Pengyu Cheng, Yitong Li 0001, Xinyuan Zhang 0001, Liqun Chen 0001, David E. Carlson, Lawrence Carin
AAAI5
2019 On Target Shift in Adversarial Domain Adaptation
abstract
Discrepancy between training and testing domains is a fundamental problem in the generalization of machine learning techniques. Recently, several approaches have been proposed to learn domain invariant feature representations through adversarial deep learning. However, label shift, where the percentage of data in each class is different between domains, has received less attention. Label shift naturally arises in many contexts, especially in behavioral studies where the behaviors are freely chosen. In this work, we propose a method called Domain Adversarial nets for Target Shift (DATS) to address label shift while learning a domain invariant representation. This is accomplished by using distribution matching to estimate label proportions in a blind test set. We extend this framework to handle multiple domains by developing a scheme to upweight source domains most similar to the target domain. Empirical results show that this framework performs well under large label shift in synthetic and real experiments, demonstrating the practical importance.
Yitong Li 0001, Michael Murias, Samantha Major, Geraldine Dawson, David E. Carlson
AISTATS5
2019 StoryGAN: A Sequential Conditional GAN for Story Visualization
abstract
In this work, we propose a new task called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization focuses less on the continuity in generated images (frames), but more on the global consistency across dynamic scenes and characters -- a challenge that has not been addressed by any single-image or video generation methods. Therefore, we propose a new story-to-image-sequence generation model, StoryGAN, based on the sequential conditional GAN framework. Our model is unique in that it consists of a deep Context Encoder that dynamically tracks the story flow, and two discriminators at the story and image levels, to enhance the image quality and the consistency of the generated sequences. To evaluate the model, we modified existing datasets to create the CLEVR-SV and Pororo-SV datasets. Empirically, StoryGAN outperformed state-of-the-art models in image quality, contextual consistency metrics, and human evaluation.
Yitong Li 0001, Zhe Gan, Yelong Shen, Jingjing Liu 0001, Yu Cheng 0001, Yuexin Wu, Lawrence Carin, David E. Carlson, Jianfeng Gao 0001
CVPR8
2018 Video Generation From Text
abstract
Generating videos from text has proven to be a significant challenge for existing generative models. We tackle this problem by training a conditional generative model to extract both static and dynamic information from text. This is manifested in a hybrid framework, employing a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN). The static features, called "gist," are used to sketch text-conditioned background color and object layout structure. Dynamic features are considered by transforming input text into an image filter. To obtain a large amount of data for training the deep-learning model, we develop a method to automatically create a matched text-video corpus from publicly available online videos. Experimental results show that the proposed framework generates plausible and diverse short-duration smooth videos, while accurately reflecting the input text information. It significantly outperforms baseline models that directly adapt text-to-image generation procedures to produce videos. Performance is evaluated both visually and by adapting the inception score used to evaluate image generation in GANs.
Yitong Li 0001, Martin Renqiang Min, Dinghan Shen, David E. Carlson, Lawrence Carin
AAAI4
2018 Extracting Relationships by Multi-Domain Matching
abstract
In many biological and medical contexts, we construct a large labeled corpus by aggregating many sources to use in target prediction tasks. Unfortunately, many of the sources may be irrelevant to our target task, so ignoring the structure of the dataset is detrimental. This work proposes a novel approach, the Multiple Domain Matching Network (MDMN), to exploit this structure. MDMN embeds all data into a shared feature space while learning which domains share strong statistical relationships. These relationships are often insightful in their own right, and they allow domains to share strength without interference from irrelevant data. This methodology builds on existing distribution-matching approaches by assuming that source domains are varied and outcomes multi-factorial. Therefore, each domain should only match a relevant subset. Theoretical analysis shows that the proposed approach can have a tighter generalization bound than existing multiple-domain adaptation approaches. Empirically, we show that the proposed methodology handles higher numbers of source domains (up to 21 empirically), and provides state-of-the-art performance on image, text, and multi-channel time series classification, including clinically relevant data of a novel treatment of Autism Spectrum Disorder.
Yitong Li 0001, Michael Murias, Geraldine Dawson, David E. Carlson
NeurIPS4
2017 Stochastic Bouncy Particle Sampler
abstract
We introduce a stochastic version of the non-reversible, rejection-free Bouncy Particle Sampler (BPS), a Markov process whose sample trajectories are piecewise linear, to efficiently sample Bayesian posteriors in big datasets. We prove that in the BPS no bias is introduced by noisy evaluations of the log-likelihood gradient. On the other hand, we argue that efficiency considerations favor a small, controllable bias, in exchange for faster mixing. We introduce a simple method that controls this trade-off. We illustrate these ideas in several examples which outperform previous approaches.
Ari Pakman, Dar Gilboa, David E. Carlson, Liam Paninski
ICML3
2017 Cross-Spectral Factor Analysis
abstract
In neuropsychiatric disorders such as schizophrenia or depression, there is often a disruption in the way that regions of the brain synchronize with one another. To facilitate understanding of network-level synchronization between brain regions, we introduce a novel model of multisite low-frequency neural recordings, such as local field potentials (LFPs) and electroencephalograms (EEGs). The proposed model, named Cross-Spectral Factor Analysis (CSFA), breaks the observed signal into factors defined by unique spatio-spectral properties. These properties are granted to the factors via a Gaussian process formulation in a multiple kernel learning framework. In this way, the LFP signals can be mapped to a lower dimensional space in a way that retains information of relevance to neuroscientists. Critically, the factors are interpretable. The proposed approach empirically allows similar performance in classifying mouse genotype and behavioral context when compared to commonly used approaches that lack the interpretability of CSFA. We also introduce a semi-supervised approach, termed discriminative CSFA (dCSFA). CSFA and dCSFA provide useful tools for understanding neural dynamics, particularly by aiding in the design of causal follow-up experiments.
Neil Gallagher, Kyle R. Ulrich, Austin Talbot, Kafui Dzirasa, Lawrence Carin, David E. Carlson
NIPS6
2017 YASS: Yet Another Spike Sorter
abstract
Spike sorting is a critical first step in extracting neural signals from large-scale electrophysiological data. This manuscript describes an efficient, reliable pipeline for spike sorting on dense multi-electrode arrays (MEAs), where neural signals appear across many electrodes and spike sorting currently represents a major computational bottleneck. We present several new techniques that make dense MEA spike sorting more robust and scalable. Our pipeline is based on an efficient multi-stage ''triage-then-cluster-then-pursuit'' approach that initially extracts only clean, high-quality waveforms from the electrophysiological time series by temporarily skipping noisy or ''collided'' events (representing two neurons firing synchronously). This is accomplished by developing a neural network detection method followed by efficient outlier triaging. The clean waveforms are then used to infer the set of neural spike waveform templates through nonparametric Bayesian clustering. Our clustering approach adapts a ''coreset'' approach for data reduction and uses efficient inference methods in a Dirichlet process mixture model framework to dramatically improve the scalability and reliability of the entire pipeline. The ''triaged'' waveforms are then finally recovered with matching-pursuit deconvolution techniques. The proposed methods improve on the state-of-the-art in terms of accuracy and stability on both real and biophysically-realistic simulated MEA data. Furthermore, the proposed pipeline is efficient, learning templates and clustering faster than real-time for a 500-electrode dataset, largely on a single CPU core.
Jin Hyung Lee, David E. Carlson, Hooshmand Shokri Razaghi, Weichi Yao, Georges Goetz, Espen Hagen, Eleanor Batty, E. J. Chichilnisky, Gaute T. Einevoll, Liam Paninski
NIPS2
2017 Targeting EEG/LFP Synchrony with Neural Nets
abstract
We consider the analysis of Electroencephalography (EEG) and Local Field Potential (LFP) datasets, which are “big” in terms of the size of recorded data but rarely have sufficient labels required to train complex models (e.g., conventional deep learning methods). Furthermore, in many scientific applications, the goal is to be able to understand the underlying features related to the classification, which prohibits the blind application of deep networks. This motivates the development of a new model based on {\em parameterized} convolutional filters guided by previous neuroscience research; the filters learn relevant frequency bands while targeting synchrony, which are frequency-specific power and phase correlations between electrodes. This results in a highly expressive convolutional neural network with only a few hundred parameters, applicable to smaller datasets. The proposed approach is demonstrated to yield competitive (often state-of-the-art) predictive performance during our empirical tests while yielding interpretable features. Furthermore, a Gaussian process adapter is developed to combine analysis over distinct electrode layouts, allowing the joint processing of multiple datasets to address overfitting and improve generalizability. Finally, it is demonstrated that the proposed framework effectively tracks neural dynamics on children in a clinical trial on Autism Spectrum Disorder.
Yitong Li 0001, Michael Murias, Samantha Major, Geraldine Dawson, Kafui Dzirasa, Lawrence Carin, David E. Carlson
NIPS7
2016 Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
abstract
Effective training of deep neural networks suffers from two main issues. The first is that the parameter space of these models exhibit pathological curvature. Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD). These methods improve convergence by adapting to the local geometry of parameter space. A second issue is overfitting, which is typically addressed by early stopping. However, recent work has demonstrated that Bayesian model averaging mitigates this problem. The posterior can be sampled by using Stochastic Gradient Langevin Dynamics (SGLD). However, the rapidly changing curvature renders default SGLD methods inefficient. Here, we propose combining adaptive preconditioners with SGLD. In support of this idea, we give theoretical properties on asymptotic convergence and predictive risk. We also provide empirical results for Logistic Regression, Feedforward Neural Nets, and Convolutional Neural Nets, demonstrating that our preconditioned SGLD method gives state-of-the-art performance on these models.
Chunyuan Li, Changyou Chen, David E. Carlson, Lawrence Carin
AAAI3
2016 Bridging the Gap between Stochastic Gradient MCMC and Stochastic Optimization
abstract
Stochastic gradient Markov chain Monte Carlo (SG-MCMC) methods are Bayesian analogs to popular stochastic optimization methods; however, this connection is not well studied. We explore this relationship by applying simulated annealing to an SG-MCMC algorithm. Furthermore, we extend recent SG-MCMC methods with two key components: i) adaptive preconditioners (as in ADAgrad or RMSprop), and ii) adaptive element-wise momentum weights. The zero-temperature limit gives a novel stochastic optimization method with adaptive element-wise momentum weights, while conventional optimization methods only have a shared, static momentum weight. Under certain assumptions, our theoretical analysis suggests the proposed simulated annealing approach converges close to the global optima. Experiments on several deep neural network models show state-of-the-art results compared to related stochastic optimization algorithms.
Changyou Chen, David E. Carlson, Zhe Gan, Chunyuan Li, Lawrence Carin
AISTATS2
2016 Parallel Majorization Minimization with Dynamically Restricted Domains for Nonconvex Optimization
abstract
We propose an optimization framework for nonconvex problems based on majorization-minimization that is particularity well-suited for parallel computing. It reduces the optimization of a high dimensional nonconvex objective function to successive optimizations of locally tight and convex upper bounds which are additively separable into low dimensional objectives. The original problem is then broken into simpler and parallel tasks, while guaranteeing the monotonic reduction of the original objective function and convergence to a local minimum. This framework also allows one to restrict the upper bound to a local dynamic convex domain, so that the bound is better matched to the local curvature of the objective function, resulting in accelerated convergence. We test the proposed framework on a nonconvex support vector machine based on a sigmoid loss function and on nonconvex penalized logistic regression.
Yan Kaganovsky, Ikenna Odinaka, David E. Carlson, Lawrence Carin
AISTATS3
2016 Learning Sigmoid Belief Networks via Monte Carlo Expectation Maximization
abstract
Belief networks are commonly used generative models of data, but require expensive posterior estimation to train and test the model. Learning typically proceeds by posterior sampling, variational approximations, or recognition networks, combined with stochastic optimization. We propose using an online Monte Carlo expectation-maximization (MCEM) algorithm to learn the maximum a posteriori (MAP) estimator of the generative model or optimize the variational lower bound of a recognition network. The E-step in this algorithm requires posterior samples, which are already generated in current learning schema. For the M-step, we augment with Polya-Gamma (PG) random variables to give an analytic updating scheme. We show relationships to standard learning approaches by deriving stochastic gradient ascent in the MCEM framework. We apply the proposed methods to both binary and count data. Experimental results show that MCEM improves the convergence speed and often improves hold-out performance over existing learning methods. Our approach is readily generalized to other recognition networks.
Zhao Song 0001, Ricardo Henao, David E. Carlson, Lawrence Carin
AISTATS3
2016 Partition Functions from Rao-Blackwellized Tempered Sampling
abstract
Partition functions of probability distributions are important quantities for model evaluation and comparisons. We present a new method to compute partition functions of complex and multimodal distributions. Such distributions are often sampled using simulated tempering, which augments the target space with an auxiliary inverse temperature variable. Our method exploits the multinomial probability law of the inverse temperatures, and provides estimates of the partition function in terms of a simple quotient of Rao-Blackwellized marginal inverse temperature probability estimates, which are updated while sampling. We show that the method has interesting connections with several alternative popular methods, and offers some significant advantages. In particular, we empirically find that the new method provides more accurate estimates than Annealed Importance Sampling when calculating partition functions of large Restricted Boltzmann Machines (RBM); moreover, the method is sufficiently accurate to track training and validation log-likelihoods during learning of RBMs, at minimal computational cost.
David E. Carlson, Patrick Stinson, Ari Pakman, Liam Paninski
ICML1
2016 Neuroprosthetic Decoder Training as Imitation Learning
abstract
Neuroprosthetic brain-computer interfaces function via an algorithm which decodes neural activity of the user into movements of an end effector, such as a cursor or robotic arm. In practice, the decoder is often learned by updating its parameters while the user performs a task. When the user's intention is not directly observable, recent methods have demonstrated value in training the decoder against a surrogate for the user's intended movement. Here we show that training a decoder in this way is a novel variant of an imitation learning problem, where an oracle or expert is employed for supervised training in lieu of direct observations, which are not available. Specifically, we describe how a generic imitation learning meta-algorithm, dataset aggregation (DAgger), can be adapted to train a generic brain-computer interface. By deriving existing learning algorithms for brain-computer interfaces in this framework, we provide a novel analysis of regret (an important metric of learning efficacy) for brain-computer interfaces. This analysis allows us to characterize the space of algorithmic variants and bounds on their regret rates. Existing approaches for decoder learning have been performed in the cursor control setting, but the available design principles for these decoders are such that it has been impossible to scale them to naturalistic settings. Leveraging our findings, we then offer an algorithm that combines imitation learning with optimal control, which should allow for training of arbitrary effectors for which optimal control can generate goal-oriented control. We demonstrate this novel and general BCI algorithm with simulated neuroprosthetic control of a 26 degree-of-freedom model of an arm, a sophisticated and realistic end effector.
Josh Merel, David E. Carlson, Liam Paninski, John P. Cunningham
PLoS Comput. Biol.2
2015 Stochastic Spectral Descent for Restricted Boltzmann Machines
abstract
Restricted Boltzmann Machines (RBMs) are widely used as building blocks for deep learning models. Learning typically proceeds by using stochastic gradient descent, and the gradients are estimated with sampling methods. However, the gradient estimation is a computational bottleneck, so better use of the gradients will speed up the descent algorithm. To this end, we first derive upper bounds on the RBM cost function, then show that descent methods can have natural ad- vantages by operating in the L∞and Shatten-∞norm. We introduce a new method called “Stochastic Spectral Descent” that updates parameters in the normed space. Empirical results show dramatic improvements over stochastic gradient descent, and have only have a fractional increase on the per-iteration cost.
David E. Carlson, Volkan Cevher, Lawrence Carin
AISTATS1
2015 Learning Deep Sigmoid Belief Networks with Data Augmentation
abstract
Deep directed generative models are developed. The multi-layered model is designed by stacking sigmoid belief networks, with sparsity-encouraging priors placed on the model parameters. Learning and inference of layer-wise model parameters are implemented in a Bayesian setting. By exploring the idea of data augmentation and introducing auxiliary Polya-Gamma variables, simple and efficient Gibbs sampling and mean-field variational Bayes (VB) inference are implemented. To address large-scale datasets, an online version of VB is also developed. Experimental results are presented for three publicly available datasets: MNIST, Caltech 101 Silhouettes and OCR letters.
Zhe Gan, Ricardo Henao, David E. Carlson, Lawrence Carin
AISTATS3
2015 Scalable Deep Poisson Factor Analysis for Topic Modeling
abstract
A new framework for topic modeling is developed, based on deep graphical models, where interactions between topics are inferred through deep latent binary hierarchies. The proposed multi-layer model employs a deep sigmoid belief network or restricted Boltzmann machine, the bottom binary layer of which selects topics for use in a Poisson factor analysis model. Under this setting, topics live on the bottom layer of the model, while the deep specification serves as a flexible prior for revealing topic structure. Scalable inference algorithms are derived by applying Bayesian conditional density filtering algorithm, in addition to extending recently proposed work on stochastic gradient thermostats. Experimental results on several corpora show that the proposed approach readily handles very large collections of text documents, infers structured topic representations, and obtains superior test perplexities when compared with related models.
Zhe Gan, Changyou Chen, Ricardo Henao, David E. Carlson, Lawrence Carin
ICML4
2015 Preconditioned Spectral Descent for Deep Learning
abstract
Deep learning presents notorious computational challenges. These challenges include, but are not limited to, the non-convexity of learning objectives and estimating the quantities needed for optimization algorithms, such as gradients. While we do not address the non-convexity, we present an optimization solution that ex- ploits the so far unused “geometry” in the objective function in order to best make use of the estimated gradients. Previous work attempted similar goals with preconditioned methods in the Euclidean space, such as L-BFGS, RMSprop, and ADA-grad. In stark contrast, our approach combines a non-Euclidean gradient method with preconditioning. We provide evidence that this combination more accurately captures the geometry of the objective function compared to prior work. We theoretically formalize our arguments and derive novel preconditioned non-Euclidean algorithms. The results are promising in both computational time and quality when applied to Restricted Boltzmann Machines, Feedforward Neural Nets, and Convolutional Neural Nets.
David E. Carlson, Edo Collins, Ya-Ping Hsieh, Lawrence Carin, Volkan Cevher
NIPS1
2015 Deep Temporal Sigmoid Belief Networks for Sequence Modeling
abstract
Deep dynamic generative models are developed to learn sequential dependencies in time-series data. The multi-layered model is designed by constructing a hierarchy of temporal sigmoid belief networks (TSBNs), defined as a sequential stack of sigmoid belief networks (SBNs). Each SBN has a contextual hidden state, inherited from the previous SBNs in the sequence, and is used to regulate its hidden bias. Scalable learning and inference algorithms are derived by introducing a recognition model that yields fast sampling from the variational posterior. This recognition model is trained jointly with the generative model, by maximizing its variational lower bound on the log-likelihood. Experimental results on bouncing balls, polyphonic music, motion capture, and text streams show that the proposed approach achieves state-of-the-art predictive performance, and has the capacity to synthesize various sequences.
Zhe Gan, Chunyuan Li, Ricardo Henao, David E. Carlson, Lawrence Carin
NIPS4
2015 GP Kernels for Cross-Spectrum Analysis
abstract
Multi-output Gaussian processes provide a convenient framework for multi-task problems. An illustrative and motivating example of a multi-task problem is multi-region electrophysiological time-series data, where experimentalists are interested in both power and phase coherence between channels. Recently, Wilson and Adams (2013) proposed the spectral mixture (SM) kernel to model the spectral density of a single task in a Gaussian process framework. In this paper, we develop a novel covariance kernel for multiple outputs, called the cross-spectral mixture (CSM) kernel. This new, flexible kernel represents both the power and phase relationship between multiple observation channels. We demonstrate the expressive capabilities of the CSM kernel through implementation of a Bayesian hidden Markov model, where the emission distribution is a multi-output Gaussian process with a CSM covariance kernel. Results are presented for measured multi-region electrophysiological data.
Kyle R. Ulrich, David E. Carlson, Kafui Dzirasa, Lawrence Carin
NIPS2
2014 Latent Gaussian Models for Topic Modeling
abstract
A new approach is proposed for topic modeling, in which the latent matrix factorization employs Gaussian priors, rather than the Dirichlet-class priors widely used in such models. The use of a latent-Gaussian model permits simple and efficient approximate Bayesian posterior inference, via the Laplace approximation. On multiple datasets, the proposed approach is demonstrated to yield results as accurate as state-of-the-art approaches based on Dirichlet constructions, at a small fraction of the computation. The framework is general enough to jointly model text and binary data, here demonstrated to produce accurate and fast results for joint analysis of voting rolls and the associated legislative text. Further, it is demonstrated how the technique may be scaled up to massive data, with encouraging performance relative to alternative methods.
Changwei Hu, Eunsu Ryu, David E. Carlson, Yingjian Wang 0004, Lawrence Carin
AISTATS3
2014 On the relations of LFPs & Neural Spike Trains
David E. Carlson, Jana Schaich Borg, Kafui Dzirasa, Lawrence Carin
NIPS1
2014 Analysis of Brain States from Multi-Region LFP Time-Series
Kyle R. Ulrich, David E. Carlson, Wenzhao Lian, Jana Schaich Borg, Kafui Dzirasa, Lawrence Carin
NIPS2
2014 A Bregman Matrix and the Gradient of Mutual Information for Vector Poisson and Gaussian Channels
abstract
A generalization of Bregman divergence is developed and utilized to unify vector Poisson and Gaussian channel models, from the perspective of the gradient of mutual information. The gradient is with respect to the measurement matrix in a compressive-sensing setting, and mutual information is considered for signal recovery and classification. Existing gradient-of-mutual-information results for scalar Poisson models are recovered as special cases, as are known results for the vector Gaussian model. The Bregman-divergence generalization yields a Bregman matrix, and this matrix induces numerous matrix-valued metrics. The metrics associated with the Bregman matrix are detailed, as are its other properties. The Bregman matrix is also utilized to connect the relative entropy and mismatched minimum mean squared error. Two applications are considered: 1) compressive sensing with a Poisson measurement model and 2) compressive topic modeling for analysis of a document corpora (word-count data). In both of these settings, we use the developed theory to optimize the compressive measurement matrix, for signal recovery and classification.
Liming Wang 0004, David E. Carlson, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin
IEEE Trans. Inf. Theory2
2013 Real-Time Inference for a Gamma Process Model of Neural Spiking
abstract
With simultaneous measurements from ever increasing populations of neurons, there is a growing need for sophisticated tools to recover signals from individual neurons. In electrophysiology experiments, this classically proceeds in a two-step process: (i) threshold the waveforms to detect putative spikes and (ii) cluster the waveforms into single units (neurons). We extend previous Bayesian nonparamet- ric models of neural spiking to jointly detect and cluster neurons using a Gamma process model. Importantly, we develop an online approximate inference scheme enabling real-time analysis, with performance exceeding the previous state-of-the- art. Via exploratory data analysis—using data with partial ground truth as well as two novel data sets—we find several features of our model collectively contribute to our improved performance including: (i) accounting for colored noise, (ii) de- tecting overlapping spikes, (iii) tracking waveform dynamics, and (iv) using mul- tiple channels. We hope to enable novel experiments simultaneously measuring many thousands of neurons and possibly adapting stimuli dynamically to probe ever deeper into the mysteries of the brain.
David E. Carlson, Vinayak A. Rao, Joshua T. Vogelstein, Lawrence Carin
NIPS1
2013 Designed Measurements for Vector Count Data
abstract
We consider design of linear projection measurements for a vector Poisson signal model. The projections are performed on the vector Poisson rate, $X\in\mathbb{R}_+^n$, and the observed data are a vector of counts, $Y\in\mathbb{Z}_+^m$. The projection matrix is designed by maximizing mutual information between $Y$ and $X$, $I(Y;X)$. When there is a latent class label $C\in\{1,\dots,L\}$ associated with $X$, we consider the mutual information with respect to $Y$ and $C$, $I(Y;C)$. New analytic expressions for the gradient of $I(Y;X)$ and $I(Y;C)$ are presented, with gradient performed with respect to the measurement matrix. Connections are made to the more widely studied Gaussian measurement model. Example results are presented for compressive topic modeling of a document corpora (word counting), and hyperspectral compressive sensing for chemical classification (photon counting).
Liming Wang 0004, David E. Carlson, Miguel R. D. Rodrigues, David Wilcox, A. Robert Calderbank, Lawrence Carin
NIPS2
2011 On the Analysis of Multi-Channel Neural Spike Data
abstract
Nonparametric Bayesian methods are developed for analysis of multi-channel spike-train data, with the feature learning and spike sorting performed jointly. The feature learning and sorting are performed simultaneously across all channels. Dictionary learning is implemented via the beta-Bernoulli process, with spike sorting performed via the dynamic hierarchical Dirichlet process (dHDP), with these two models coupled. The dHDP is augmented to eliminate refractoryperiod violations, it allows the “appearance” and “disappearance” of neurons over time, and it models smooth variation in the spike statistics.
Bo Chen 0001, David E. Carlson, Lawrence Carin
NIPS2
1980 Bit-Oriented Data Link Control Procedures
abstract
The rapid growth of data communications in recent years, coupled with a movement from batch-oriented to transaction-oriented (interactive) type of operation, generated the need for a new, more efficient, more reliable, more flexible form of data link control procedure. A bit-oriented approach to data link control that has become the generally accepted standard around the world is discussed in this paper. The basic elements and structure of the procedure are described, some typical examples of operation reviewed, and a "crystal-balling" of the future is offered.
David E. Carlson
IEEE Trans. Commun.1