Joseph G. Makin

dblp:136/4647 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0002-0053-7006ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 48% Generative modeling · 40% Deep learning architectures and training · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
energy-based model
0.912025
Exponential-Family Harmoniums with Neural Sufficient Statistics · AAAI 2025
Machine learning › Probabilistic and Bayesian machine learning › boltzmann machine
restricted boltzmann machine
0.912025
Exponential-Family Harmoniums with Neural Sufficient Statistics · AAAI 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.212014
Sensory Integration and Density Estimation · NIPS 2014
Bioinformatics and computational biology › computational neuroscience › sensory processing
multisensory integration
0.112014
Sensory Integration and Density Estimation · NIPS 2014
Bioinformatics and computational biology › computational neuroscience
sensory processing
0.112014
Sensory Integration and Density Estimation · NIPS 2014

Methods — techniques the papers use, named apart from their topics

langevin dynamics · 0.9gibbs sampling · 0.9deep neural network · 0.9density estimation · 0.4
YearPublicationVenuePosition
2025 Exponential-Family Harmoniums with Neural Sufficient Statistics
abstract
Exponential-family harmoniums (EFHs) generalize the restricted Boltzmann machine beyond Bernoulli random variables to other exponential families. Here we show how to extend the EFH beyond standard exponential families (Poisson, Gaussian, etc.), by allowing the sufficient statistics for the hidden units to be arbitrary functions of the observed data, parameterized by deep neural networks. This rules out the standard sampling scheme, block Gibbs sampling, so we replace it with a form of Langevin dynamics within Gibbs, inspired by a recent method for training Gaussian restricted Boltzmann machines (GRBMs). With Gibbs-Langevin, the GRBM can successfully model small datasets like MNIST and CelebA-32, but struggles with CIFAR-10, and cannot scale to larger images because it lacks convolutions. In contrast, our neural-network EFHs (NN-EFHs) generate high-quality samples from CIFAR-10 and scale well to CelebA-HQ. On these datasets, the NN-EFH achieves FID scores that are 25--50% lower than a standard energy-based model with a similar neural-network architecture and the same number of parameters; and competitive with noise-conditional score networks, which utilize more complex neural networks (U-nets) and require considerably more sampling steps.
Azwar Abdulsalam, Joseph G. Makin
AAAI2
2025 Deep neural networks explain spiking activity in auditory cortex
abstract
For static stimuli or at gross (∼1-s) time scales, artificial neural networks (ANNs) that have been trained on challenging engineering tasks, like image classification and automatic speech recognition, are now the best predictors of neural responses in primate visual and auditory cortex. It is, however, unknown whether this success can be extended to spiking activity at fine time scales, which are particularly relevant to audition. Here we address this question with ANNs trained on speech audio, and acute multi-electrode recordings from the auditory cortex of squirrel monkeys. We show that layers of trained ANNs can predict the spike counts of multi-units responding to speech audio and to monkey vocalizations at bin widths of 50 ms and below. For some multi-units, the ANNs explain close to all of the explainable variance-much more than traditional spectrotemporal receptive fields, and more than untrained networks. Non-primary neurons tend to be more predictable by deeper layers of the ANNs, but there is much variation by neuron, which would be invisible to coarser recording modalities.
Joshua D. Downer, Brian J. Malone, Joseph G. Makin
PLoS Comput. Biol.4
2015 Learning to Estimate Dynamical State with Probabilistic Population Codes
abstract
Tracking moving objects, including one's own body, is a fundamental ability of higher organisms, playing a central role in many perceptual and motor tasks. While it is unknown how the brain learns to follow and predict the dynamics of objects, it is known that this process of state estimation can be learned purely from the statistics of noisy observations. When the dynamics are simply linear with additive Gaussian noise, the optimal solution is the well known Kalman filter (KF), the parameters of which can be learned via latent-variable density estimation (the EM algorithm). The brain does not, however, directly manipulate matrices and vectors, but instead appears to represent probability distributions with the firing rates of population of neurons, "probabilistic population codes." We show that a recurrent neural network-a modified form of an exponential family harmonium (EFH)-that takes a linear probabilistic population code as input can learn, without supervision, to estimate the state of a linear dynamical system. After observing a series of population responses (spike counts) to the position of a moving object, the network learns to represent the velocity of the object and forms nearly optimal predictions about the position at the next time-step. This result builds on our previous work showing that a similar network can learn to perform multisensory integration and coordinate transformations for static stimuli. The receptive fields of the trained network also make qualitative predictions about the developing and learning brain: tuning gradually emerges for higher-order dynamical states not explicitly present in the inputs, appearing as delayed tuning for the lower-order states.
Joseph G. Makin, Benjamin K. Dichter, Philip N. Sabes
PLoS Comput. Biol.1
2014 Sensory Integration and Density Estimation
Joseph G. Makin, Philip N. Sabes
NIPS1
2013 Learning Multisensory Integration and Coordinate Transformation via Density Estimation
abstract
Sensory processing in the brain includes three key operations: multisensory integration-the task of combining cues into a single estimate of a common underlying stimulus; coordinate transformations-the change of reference frame for a stimulus (e.g., retinotopic to body-centered) effected through knowledge about an intervening variable (e.g., gaze position); and the incorporation of prior information. Statistically optimal sensory processing requires that each of these operations maintains the correct posterior distribution over the stimulus. Elements of this optimality have been demonstrated in many behavioral contexts in humans and other animals, suggesting that the neural computations are indeed optimal. That the relationships between sensory modalities are complex and plastic further suggests that these computations are learned-but how? We provide a principled answer, by treating the acquisition of these mappings as a case of density estimation, a well-studied problem in machine learning and statistics, in which the distribution of observed data is modeled in terms of a set of fixed parameters and a set of latent variables. In our case, the observed data are unisensory-population activities, the fixed parameters are synaptic connections, and the latent variables are multisensory-population activities. In particular, we train a restricted Boltzmann machine with the biologically plausible contrastive-divergence rule to learn a range of neural computations not previously demonstrated under a single approach: optimal integration; encoding of priors; hierarchical integration of cues; learning when not to integrate; and coordinate transformation. The model makes testable predictions about the nature of multisensory representations.
Joseph G. Makin, Matthew R. Fellows, Philip N. Sabes
PLoS Comput. Biol.1