Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ezra Edelman

dblp:369/3333 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 47% Deep learning architectures and training · 21% Probabilistic and Bayesian machine learning · 11%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model inference
inference-time computation
0.912025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Natural language and speech › Language models and text generation
in-context learning
0.812024
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains · NeurIPS 2024
Machine learning › Trustworthy machine learning › language model interpretability
induction head
0.812024
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
markov chain
0.812024
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains · NeurIPS 2024
Machine learning › Learning theory
phase transition
0.812024
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains · NeurIPS 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains · NeurIPS 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains · NeurIPS 2024
Graph algorithms and graph theory
graph algorithms
0.312025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Graph algorithms and graph theory
graph connectivity
0.312025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

majority voting · 1.7chain-of-thought · 1.7theoretical analysis · 0.8empirical study · 0.8
YearPublicationVenuePosition
2025 Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones
abstract
Inference-time computation has emerged as a promising scaling axis for improving large language model reasoning. However, despite yielding impressive performance, the optimal allocation of inference-time computation remains poorly understood. A central question is whether to prioritize sequential scaling (e.g., longer chains of thought) or parallel scaling (e.g., majority voting across multiple short chains of thought). In this work, we seek to illuminate the landscape of test-time scaling by demonstrating the existence of reasoning settings where sequential scaling offers an exponential advantage over parallel scaling. These settings are based on graph connectivity problems in challenging distributions of graphs. We validate our theoretical findings with comprehensive experiments across a range of language models, including models trained from scratch for graph connectivity with different chain of thought strategies as well as large reasoning models.
Parsa Mirtaheri, Ezra Edelman, Samy Jelassi, Eran Malach, Enric Boix-Adserà
NeurIPS2
2024 The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
abstract
Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this in-context learning capability emerges. In our setting, each example is sampled from a Markov chain drawn from a prior distribution over Markov chains. Transformers trained on this task form \emph{statistical induction heads} which compute accurate next-token probabilities given the bigram statistics of the context. During the course of training, models pass through multiple phases: after an initial stage in which predictions are uniform, they learn to sub-optimally predict using in-context single-token statistics (unigrams); then, there is a rapid phase transition to the correct in-context bigram solution. We conduct an empirical and theoretical investigation of this multi-phase process, showing how successful learning results from the interaction between the transformer's layers, and uncovering evidence that the presence of the simpler unigram solution may delay formation of the final bigram solution. We examine how learning is affected by varying the prior distribution over Markov chains, and consider the generalization of our in-context learning of Markov chains (ICL-MC) task to $n$-grams for $n > 2$.
Ezra Edelman, Nikolaos Tsilivis 0002, Benjamin L. Edelman, Eran Malach, Surbhi Goel
NeurIPS1