Jonathan Kadmon

dblp:191/6697 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-3970-5684ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Optimization for machine learning · 55% Deep learning architectures and training · 27% Graph learning · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 75% Computational science and engineering · 25%
Theoretical computer science
1 paper
Algorithms and data structures · 25% Mathematical optimization · 25% Information theory · 25%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
convergence analysis
0.912025
Curl Descent : Non-Gradient Learning Dynamics with Sign-Diverse Plasticity · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.912025
Curl Descent : Non-Gradient Learning Dynamics with Sign-Diverse Plasticity · NeurIPS 2025
Machine learning › Deep learning architectures and training › biologically plausible learning
synaptic plasticity
0.912025
Curl Descent : Non-Gradient Learning Dynamics with Sign-Diverse Plasticity · NeurIPS 2025
Bioinformatics and computational biology
computational neuroscience
0.412020
Predictive coding in balanced neural networks with noise, chaos and delays · NeurIPS 2020
Computational science and engineering › statistical physics
mean-field theory
0.412020
Predictive coding in balanced neural networks with noise, chaos and delays · NeurIPS 2020
Bioinformatics and computational biology › computational neuroscience
neural coding
0.412020
Predictive coding in balanced neural networks with noise, chaos and delays · NeurIPS 2020
Bioinformatics and computational biology › computational neuroscience › neural coding
predictive coding
0.412020
Predictive coding in balanced neural networks with noise, chaos and delays · NeurIPS 2020
Machine learning › Graph learning › graph neural network › message passing
approximate message passing
0.312018
Statistical mechanics of low-rank tensor decomposition · NeurIPS 2018
Mathematical optimization › tensor optimization
low-rank tensor decomposition
0.312018
Statistical mechanics of low-rank tensor decomposition · NeurIPS 2018
Computational complexity
phase transition
0.312018
Statistical mechanics of low-rank tensor decomposition · NeurIPS 2018
Information theory › statistical inference
statistical mechanics of inference
0.312018
Statistical mechanics of low-rank tensor decomposition · NeurIPS 2018
Algorithms and data structures › numerical linear algebra › matrix and tensor decomposition
tensor decomposition
0.312018
Statistical mechanics of low-rank tensor decomposition · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

student-teacher framework · 0.9hebbian plasticity · 0.9bayesian approximate message passing · 0.7alternating least squares · 0.7recurrent spiking network model · 0.4mean field theory · 0.4dynamic mean-field theory · 0.3dynamic mean field theory · 0.3statistical mechanics · 0.2recursion relations · 0.2
YearPublicationVenuePosition
2025 Curl Descent : Non-Gradient Learning Dynamics with Sign-Diverse Plasticity
abstract
Gradient-based algorithms are a cornerstone of artificial neural network training, yet it remains unclear whether biological neural networks use similar gradient-based strategies during learning. Experiments often discover a diversity of synaptic plasticity rules, but whether these amount to an approximation to gradient descent is unclear. Here we investigate a previously overlooked possibility: that learning dynamics may include fundamentally non-gradient "curl"-like components while still being able to effectively optimize a loss function. Curl terms naturally emerge in networks with excitatory-inhibitory connectivity or Hebbian/anti-Hebbian plasticity, resulting in learning dynamics that cannot be framed as gradient descent on any objective. To investigate the impact of these curl terms, we analyze feedforward networks within an analytically tractable student-teacher framework, systematically introducing non-gradient dynamics through rule-flipped neurons. Small curl terms preserve the stability of the original solution manifold, resulting in learning dynamics similar to gradient descent. Beyond a critical value, strong curl terms destabilize the solution manifold. Depending on the network architecture, this loss of stability can lead to chaotic learning dynamics that destroy performance. In other cases, the curl terms can counterintuitively speed up learning compared to gradient descent by allowing the weight dynamics to escape saddles by temporarily ascending the loss. Our results identify specific architectures capable of supporting robust learning via diverse learning rules, providing an important counterpoint to normative theories of gradient-based learning in neural networks.
Hugo Ninou, Jonathan Kadmon, N. Alex Cayco-Gajic
NeurIPS2
2025 Neural mechanisms of flexible perceptual inference
abstract
What seems obvious in one context can take on an entirely different meaning if that context shifts. While context-dependent inference has been widely studied, a fundamental question remains: how does the brain simultaneously infer both the meaning of sensory input and the underlying context itself, especially when the context is changing? Here, we study flexible perceptual inference-the ability to adapt rapidly to implicit contextual shifts without trial and error. We introduce a novel change-detection task in dynamic environments that requires tracking latent state and context. We find that mice exhibit first-trial behavioral adaptation to latent context shifts driven by inference rather than reward feedback. By deriving the Bayes-optimal policy under a partially observable Markov decision process, we show that rapid adaptation emerges from sequential updates of an internal belief state. In addition, we show that artificial neural networks trained via reinforcement learning achieve near-optimal performance, implementing Bayesian inference-like mechanisms within their recurrent dynamics. These networks develop flexible internal representations that enable adaptive inference in real-time. Our findings establish flexible perceptual inference as a core principle of cognitive flexibility, offering computational and neural-mechanistic insights into adaptive behavior in uncertain environments.
John Schwarcz, Haneen Rajabi, Gabrielle Marmur, Robert Reiner, Eran Lottem, Jonathan Kadmon
PLoS Comput. Biol.7
2022 Optimal noise level for coding with tightly balanced networks of spiking neurons in the presence of transmission delays
abstract
Neural circuits consist of many noisy, slow components, with individual neurons subject to ion channel noise, axonal propagation delays, and unreliable and slow synaptic transmission. This raises a fundamental question: how can reliable computation emerge from such unreliable components? A classic strategy is to simply average over a population of N weakly-coupled neurons to achieve errors that scale as [Formula: see text]. But more interestingly, recent work has introduced networks of leaky integrate-and-fire (LIF) neurons that achieve coding errors that scale superclassically as 1/N by combining the principles of predictive coding and fast and tight inhibitory-excitatory balance. However, spike transmission delays preclude such fast inhibition, and computational studies have observed that such delays can cause pathological synchronization that in turn destroys superclassical coding performance. Intriguingly, it has also been observed in simulations that noise can actually improve coding performance, and that there exists some optimal level of noise that minimizes coding error. However, we lack a quantitative theory that describes this fascinating interplay between delays, noise and neural coding performance in spiking networks. In this work, we elucidate the mechanisms underpinning this beneficial role of noise by deriving analytical expressions for coding error as a function of spike propagation delay and noise levels in predictive coding tight-balance networks of LIF neurons. Furthermore, we compute the minimal coding error and the associated optimal noise level, finding that they grow as power-laws with the delay. Our analysis reveals quantitatively how optimal levels of noise can rescue neural coding performance in spiking neural networks with delays by preventing the build up of pathological synchrony without overwhelming the overall spiking dynamics. This analysis can serve as a foundation for the further study of precise computation in the presence of noise and delays in efficient spiking neural circuits.
Jonathan Timcheck, Jonathan Kadmon, Kwabena Boahen 0001, Surya Ganguli
PLoS Comput. Biol.2
2020 Predictive coding in balanced neural networks with noise, chaos and delays
abstract
Biological neural networks face a formidable task: performing reliable computations in the face of intrinsic stochasticity in individual neurons, imprecisely specified synaptic connectivity, and nonnegligible delays in synaptic transmission. A common approach to combatting such biological heterogeneity involves averaging over large redundant networks of N neurons resulting in coding errors that decrease classically as the square root of N. Recent work demonstrated a novel mechanism whereby recurrent spiking networks could efficiently encode dynamic stimuli achieving a superclassical scaling in which coding errors decrease as 1/N. This specific mechanism involved two key ideas: predictive coding, and a tight balance, or cancellation between strong feedforward inputs and strong recurrent feedback. However, the theoretical principles governing the efficacy of balanced predictive coding and its robustness to noise, synaptic weight heterogeneity and communication delays remain poorly understood. To discover such principles, we introduce an analytically tractable model of balanced predictive coding, in which the degree of balance and the degree of weight disorder can be dissociated unlike in previous balanced network models, and we develop a mean-field theory of coding accuracy. Overall, our work provides and solves a general theoretical framework for dissecting the differential contributions neural noise, synaptic disorder, chaos, synaptic delays, and balance to the fidelity of predictive neural codes, reveals the fundamental role that balance plays in achieving superclassical scaling, and unifies previously disparate models in theoretical neuroscience.
Jonathan Kadmon, Jonathan Timcheck, Surya Ganguli
NeurIPS1
2018 Statistical mechanics of low-rank tensor decomposition
abstract
Often, large, high dimensional datasets collected across multiple modalities can be organized as a higher order tensor. Low-rank tensor decomposition then arises as a powerful and widely used tool to discover simple low dimensional structures underlying such data. However, we currently lack a theoretical understanding of the algorithmic behavior of low-rank tensor decompositions. We derive Bayesian approximate message passing (AMP) algorithms for recovering arbitrarily shaped low-rank tensors buried within noise, and we employ dynamic mean field theory to precisely characterize their performance. Our theory reveals the existence of phase transitions between easy, hard and impossible inference regimes, and displays an excellent match with simulations. Moreover, it reveals several qualitative surprises compared to the behavior of symmetric, cubic tensor decomposition. Finally, we compare our AMP algorithm to the most commonly used algorithm, alternating least squares (ALS), and demonstrate that AMP significantly outperforms ALS in the presence of noise.
Jonathan Kadmon, Surya Ganguli
NeurIPS1
2016 Optimal Architectures in a Solvable Model of Deep Networks
abstract
Deep neural networks have received a considerable attention due to the success of their training for real world machine learning applications. They are also of great interest to the understanding of sensory processing in cortical sensory hierarchies. The purpose of this work is to advance our theoretical understanding of the computational benefits of these architectures. Using a simple model of clustered noisy inputs and a simple learning rule, we provide analytically derived recursion relations describing the propagation of the signals along the deep network. By analysis of these equations, and defining performance measures, we show that these model networks have optimal depths. We further explore the dependence of the optimal architecture on the system parameters.
Jonathan Kadmon, Haim Sompolinsky
NIPS1