Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sean McLeish

dblp:374/9044 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 49% Deep learning architectures and training · 38% Optimization for machine learning · 13%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › hyperparameter optimization
hyperparameter sensitivity
0.912025
Gemstones: A Model Suite for Multi-Faceted Scaling Laws · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model reasoning › inference-time reasoning
latent reasoning
0.912025
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach · NeurIPS 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach · NeurIPS 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Gemstones: A Model Suite for Multi-Faceted Scaling Laws · NeurIPS 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach · NeurIPS 2025
Natural language and speech › Language models and text generation › compositional generalization
length generalization
0.812024
Transformers Can Do Arithmetic with the Right Embeddings · NeurIPS 2024
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning
0.812024
Transformers Can Do Arithmetic with the Right Embeddings · NeurIPS 2024
Machine learning › Deep learning architectures and training
positional encoding
0.812024
Transformers Can Do Arithmetic with the Right Embeddings · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

speculative decoding · 0.9scaling law fitting · 0.9recurrent depth · 0.9recurrent layer · 0.8positional embedding · 0.8input injection · 0.8
YearPublicationVenuePosition
2025 Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
abstract
We study a novel language model architecture that is capable of scaling test-time computation by implicitly reasoning in latent space. Our model works by iterating a recurrent block, thereby unrolling to arbitrary depth at test-time. This stands in contrast to mainstream reasoning models that scale up compute by producing more tokens. Unlike approaches based on chain-of-thought, our approach does not require any specialized training data, can work with small context windows, and can capture types of reasoning that are not easily represented in words. We train a proof-of-concept model from scratch with 3.5 billion parameters and 800 billion tokens. We show that this model can effortlessly use varying levels of compute, significantly improving with additional compute especially on reasoning tasks, such as math and coding. Further, this architecture naturally reduces compute costs via zero-shot per-token adaptive compute, KV-cache sharing and speculative decoding.
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein
NeurIPS2
2025 Gemstones: A Model Suite for Multi-Faceted Scaling Laws
abstract
Scaling laws are typically fit using a family of models with a narrow range of frozen hyperparameter choices. In this work we study scaling laws using multiple architectural shapes and hyperparameter choices, highlighting their impact on resulting prescriptions. As a primary artifact of our research, we release the Gemstones: an open-source scaling law dataset, consisting of over 4000 checkpoints from transformers with up to 2 billion parameters and diverse architectural shapes; including ablations over learning rate and cooldown. Our checkpoints enable more complex studies of scaling, such as analyzing the relationship between width and depth. By examining our model suite, we find that the prescriptions of scaling laws can be highly sensitive to the experimental design process and the specific model checkpoints used during fitting.
Sean McLeish, John Kirchenbauer, David Yu Miller, Abhinav Bhatele, Micah Goldblum, Ashwinee Panda, Tom Goldstein
NeurIPS1
2024 Transformers Can Do Arithmetic with the Right Embeddings
abstract
The poor performance of transformers on arithmetic tasks seems to stem in large part from their inability to keep track of the exact position of each digit inside of a large span of digits. We mend this problem by adding an embedding to each digit that encodes its position relative to the start of the number. In addition to the boost these embeddings provide on their own, we show that this fix enables architectural modifications such as input injection and recurrent layers to improve performance even further. With positions resolved, we can study the logical extrapolation ability of transformers. Can they solve arithmetic problems that are larger and more complex than those in their training data? We find that training on only 20 digit numbers with a single GPU for one day, we can reach state-of-the-art performance, achieving up to 99% accuracy on 100 digit addition problems. Finally, we show that these gains in numeracy also unlock improvements on other multi-step reasoning tasks including sorting and multiplication.
Sean McLeish, Arpit Bansal, Alex Stein, Neel Jain, John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, Tom Goldstein
NeurIPS1