David Hall 0006

dblp:133/2070 · also David Leo Wright Hall · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
3since 2021 · last 2025
0000-0002-9596-6136ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Optimization for machine learning · 30% Information extraction and text analysis · 24% Language models and text generation · 15%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 100%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model training
language model pretraining
1.122025
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View · ICLR 2025
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training · ICLR 2024
Machine learning › Trustworthy machine learning
interpretability
0.912025
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning · ICLR 2025
Machine learning › Optimization for machine learning
learning rate schedule
0.912025
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning · ICLR 2025
Machine learning › Optimization for machine learning › learning rate schedule
warmup-stable-decay schedule
0.912025
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View · ICLR 2025
Machine learning › Optimization for machine learning
second-order optimization
0.812024
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training · ICLR 2024
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.742014
Less Grammar, More Features · ACL (1) 2014
Sparser, Better, Faster GPU Parsing · ACL (1) 2014
Parser Showdown at the Wall Street Corral: An Empirical Investigation of Error Types in Parser Output · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis › syntactic parsing
constituency parsing
0.422014
Less Grammar, More Features · ACL (1) 2014
A Multi-Teraflop Constituency Parser using GPUs · EMNLP 2013
Machine learning › Deep learning architectures and training
loss landscape
0.312025
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View · ICLR 2025
GPUs and heterogeneous computing
GPU computing
0.222014
Sparser, Better, Faster GPU Parsing · ACL (1) 2014
A Multi-Teraflop Constituency Parser using GPUs · EMNLP 2013
Natural language and speech › Information extraction and text analysis › multilingual NLP
cognate identification
0.222011
Large-Scale Cognate Recovery · EMNLP 2011
Finding Cognate Groups Using Phylogenies · ACL 2010
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.212014
Less Grammar, More Features · ACL (1) 2014
Natural language and speech › Information extraction and text analysis
topic model
0.222009
Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora · EMNLP 2009
Studying the History of Ideas Using Topic Models · EMNLP 2008
Natural language and speech › Information extraction and text analysis
coreference resolution
0.212013
Decentralized Entity-Level Modeling for Coreference Resolution · ACL (1) 2013
Natural language and speech › Language models and text generation › grammar formalisms
probabilistic context-free grammar
0.112012
Training Factored PCFGs with Expectation Propagation · EMNLP-CoNLL 2012
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
bounded-suboptimal search
0.112011
Optimal Graph Search with Iterated Graph Cuts · AAAI 2011
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
graph search
0.112011
Optimal Graph Search with Iterated Graph Cuts · AAAI 2011
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search
0.112011
Optimal Graph Search with Iterated Graph Cuts · AAAI 2011
Natural language and speech › Information extraction and text analysis
historical linguistics
0.112010
Finding Cognate Groups Using Phylogenies · ACL 2010
Bioinformatics and computational biology
phylogenetics
0.112010
Finding Cognate Groups Using Phylogenies · ACL 2010
Machine learning › Learning paradigms
multi-label classification
0.112009
Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora · EMNLP 2009
Natural language and speech › Information extraction and text analysis › topic model
supervised topic model
0.112009
Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora · EMNLP 2009
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
0.012012
Training Factored PCFGs with Expectation Propagation · EMNLP-CoNLL 2012
Computational social science and digital humanities
historical linguistics
0.012011
Large-Scale Cognate Recovery · EMNLP 2011
Algorithms and data structures › search algorithms › heuristic search
a* search
0.012011
Optimal Graph Search with Iterated Graph Cuts · AAAI 2011
Graph algorithms and graph theory
shortest path
0.012011
Optimal Graph Search with Iterated Graph Cuts · AAAI 2011
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation
0.012009
Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora · EMNLP 2009

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 0.9spectral analysis · 0.9neural ordinary differential equation · 0.9lyapunov exponent · 0.9bi-gram dataset · 0.9stochastic second-order optimization · 0.8gradient clipping · 0.8diagonal hessian estimation · 0.8minimum bayes risk · 0.4coarse-to-fine pruning · 0.4grammar compilation · 0.2cache-sharing · 0.2CKY chart evaluation · 0.2heuristic search · 0.1graph cuts · 0.1cognate recovery · 0.1phylogeny · 0.1topic model · 0.1
YearPublicationVenuePosition
2025 Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
abstract
Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we introduce a novel approach to modeling transformer architectures using highly flexible non-autonomous neural ordinary differential equations (ODEs). Our proposed model parameterizes all weights of attention and feed-forward blocks through neural networks, expressing these weights as functions of a continuous layer index. Through spectral analysis of the model's dynamics, we uncover an increase in eigenvalue magnitude that challenges the weight-sharing assumption prevalent in existing theoretical studies. We also leverage the Lyapunov exponent to examine token-level sensitivity, enhancing model interpretability. Our neural ODE transformer demonstrates performance comparable to or better than vanilla transformers across various configurations and datasets, while offering flexible fine-tuning capabilities that can adapt to different architectural constraints.
Anh Tong, Thanh Nguyen-Tang, Dongeun Lee 0001, Toan M. Tran, David Hall 0006, Cheongwoong Kang, Jaesik Choi
ICLR6
2025 Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View
abstract
Training language models currently requires pre-determining a fixed compute budget because the typical cosine learning rate schedule depends on the total number of steps. In contrast, the Warmup-Stable-Decay (WSD) schedule uses a constant learning rate to produce a main branch of iterates that can in principle continue indefinitely without a pre-specified compute budget. Then, given any compute budget, one can branch out from the main branch at a proper time with a rapidly decaying learning rate to produce a strong model. Empirically, WSD generates an intriguing, non-traditional loss curve: the loss remains elevated during the stable phase but sharply declines during the decay phase. Towards explaining this phenomenon, we conjecture that pretraining loss exhibits a river valley landscape, which resembles a deep valley with a river at its bottom. Under this assumption, we show that during the stable phase, the iterate undergoes large oscillations due to the high learning rate, yet it progresses swiftly along the river. During the decay phase, the rapidly dropping learning rate minimizes the iterate’s oscillations, moving it closer to the river and revealing true optimization progress. Therefore, the sustained high learning rate phase and fast decaying phase are responsible for progress in the river and the mountain directions, respectively, and are both critical. Our analysis predicts phenomenons consistent with empirical observations and shows that this landscape can naturally emerge from pretraining on a simple bi-gram dataset. Inspired by the theory, we introduce WSD-S, a variant of WSD that reuses previous checkpoints’ decay phases and keeps only one main branch, where we resume from a decayed checkpoint. WSD-S empirically outperforms WSD and Cyclic-Cosine in obtaining multiple pretrained language model checkpoints across various compute budgets in a single run for parameters scaling from 0.1B to 1.2B.
Kaiyue Wen, Zhiyuan Li 0005, Jason S. Wang, David Hall 0006, Percy Liang, Tengyu Ma 0001
ICLR4
2024 Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
abstract
Given the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training. Adam and its variants have been state-of-the-art for years, and more sophisticated second-order (Hessian-based) optimizers often incur too much per-step overhead. In this paper, we propose Sophia, a simple scalable second-order optimizer that uses a light-weight estimate of the diagonal Hessian as the pre-conditioner. The update is the moving average of the gradients divided by the moving average of the estimated Hessian, followed by element-wise clipping. The clipping controls the worst-case update size and tames the negative impact of non-convexity and rapid change of Hessian along the trajectory. Sophia only estimates the diagonal Hessian every handful of iterations, which has negligible average per-step time and memory overhead. On language modeling with GPT models of sizes ranging from 125M to 1.5B, Sophia achieves a 2x speed-up compared to Adam in the number of steps, total compute, and wall-clock time, achieving the same perplexity with 50\% fewer steps, less total compute, and reduced wall-clock time.
Zhiyuan Li 0005, David Hall 0006, Percy Liang, Tengyu Ma 0001
ICLR3
2020 Task-Oriented Dialogue as Dataflow Synthesis
abstract
We describe an approach to task-oriented dialogue in which dialogue state is represented as a dataflow graph. A dialogue agent maps each user utterance to a program that extends this graph. Programs include metacomputation operators for reference and revision that reuse dataflow fragments from previous turns. Our graph-based state enables the expression and manipulation of complex user intents, and explicit metacomputation makes these intents easier for learned models to predict. We introduce a new dataset, SMCalFlow, featuring complex dialogues about events, weather, places, and people. Experiments show that dataflow graphs and metacomputation substantially improve representability and predictability in these natural dialogues. Additional experiments on the MultiWOZ dataset show that our dataflow representation enables an otherwise off-the-shelf sequence-to-sequence model to match the best existing task-specific state tracking model. The SMCalFlow dataset, code for replicating experiments, and a public leaderboard are available at https://www.microsoft.com/en-us/research/project/dataflow-based-dialogue-semantic-machines .
Jacob Andreas, John Bufe, David Burkett, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang 0002, Alan Guo, David Hall 0006, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Daniel Klein 0001, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christopher H. Lin, Ilya Lintsbakh, Andy McGovern, Aleksandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth 0001, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su 0001, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Andrew Wolfe, Abby Wray, Yuchen Zhang 0002, Alexander Zotov
Trans. Assoc. Comput. Linguistics13
2014 Sparser, Better, Faster GPU Parsing
abstract
Due to their origin in computer graphics, graphics processing units (GPUs) are highly optimized for dense problems, where the exact same operation is applied repeatedly to all data points.Natural language processing algorithms, on the other hand, are traditionally constructed in ways that exploit structural sparsity.Recently, Canny et al. (2013) presented an approach to GPU parsing that sacrifices traditional sparsity in exchange for raw computational power, obtaining a system that can compute Viterbi parses for a high-quality grammar at about 164 sentences per second on a mid-range GPU.In this work, we reintroduce sparsity to GPU parsing by adapting a coarse-to-fine pruning approach to the constraints of a GPU.The resulting system is capable of computing over 404 Viterbi parses per second-more than a 2x speedup-on the same hardware.Moreover, our approach allows us to efficiently implement less GPU-friendly minimum Bayes risk inference, improving throughput for this more accurate algorithm from only 32 sentences per second unpruned to over 190 sentences per second using pruning-nearly a 6x speedup.
David Hall 0006, Taylor Berg-Kirkpatrick, Daniel Klein 0001
ACL (1)1
2014 Less Grammar, More Features
abstract
We present a parser that relies primar-ily on extracting information directly from surface spans rather than on propagat-ing information through enriched gram-mar structure. For example, instead of cre-ating separate grammar symbols to mark the definiteness of an NP, our parser might instead capture the same information from the first word of the NP. Moving context out of the grammar and onto surface fea-tures can greatly simplify the structural component of the parser: because so many deep syntactic cues have surface reflexes, our system can still parse accurately with context-free backbones as minimal as X-bar grammars. Keeping the structural backbone simple and moving features to the surface also allows easy adaptation to new languages and even to new tasks. On the SPMRL 2013 multilingual con-stituency parsing shared task (Seddah et al., 2013), our system outperforms the top single parser system of Björkelund et al. (2013) on a range of languages. In addi-tion, despite being designed for syntactic analysis, our system also achieves state-of-the-art numbers on the structural senti-ment task of Socher et al. (2013). Finally, we show that, in both syntactic parsing and sentiment analysis, many broad linguistic trends can be captured via surface features. 1
David Hall 0006, Greg Durrett, Daniel Klein 0001
ACL (1)1
2013 Decentralized Entity-Level Modeling for Coreference Resolution
Greg Durrett, David Hall 0006, Daniel Klein 0001
ACL (1)2
2013 A Multi-Teraflop Constituency Parser using GPUs
abstract
Constituency parsing with rich grammars remains a computational challenge.Graphics Processing Units (GPUs) have previously been used to accelerate CKY chart evaluation, but gains over CPU parsers were modest.In this paper, we describe a collection of new techniques that enable chart evaluation at close to the GPU's practical maximum speed (a Teraflop), or around a half-trillion rule evaluations per second.Net parser performance on a 4-GPU system is over 1 thousand length-30 sentences/second (1 trillion rules/sec), and 400 general sentences/second for the Berkeley Parser Grammar.The techniques we introduce include grammar compilation, recursive symbol blocking, and cache-sharing.
John F. Canny, David Hall 0006, Daniel Klein 0001
EMNLP2
2012 Training Factored PCFGs with Expectation Propagation
David Hall 0006, Daniel Klein 0001
EMNLP-CoNLL1
2012 Parser Showdown at the Wall Street Corral: An Empirical Investigation of Error Types in Parser Output
Jonathan K. Kummerfeld, David Hall 0006, James R. Curran, Daniel Klein 0001
EMNLP-CoNLL2
2011 Optimal Graph Search with Iterated Graph Cuts
abstract
Informed search algorithms such as A* use heuristics to focus exploration on states with low total path cost. To the extent that heuristics underestimate forward costs, a wider cost radius of suboptimal states will be explored. For many weighted graphs, however, a small distance in terms of cost may encompass a large fraction of the unweighted graph. We present a new informed search algorithm, Iterative Monotonically Bounded A* (IMBA*), which first proves that no optimal paths exist in a bounded cut of the graph before considering larger cuts. We prove that IMBA* has the same optimality and completeness guarantees as A* and, in a non-uniform pathfinding application, we empirically demonstrate substantial speed improvements over classic A*.
David Burkett, David Hall 0006, Daniel Klein 0001
AAAI2
2011 Large-Scale Cognate Recovery
David Hall 0006, Daniel Klein 0001
EMNLP1
2010 Finding Cognate Groups Using Phylogenies
David Hall 0006, Daniel Klein 0001
ACL1
2009 Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora
Daniel Ramage, David Hall 0006, Ramesh Nallapati, Christopher D. Manning
EMNLP2
2008 Studying the History of Ideas Using Topic Models
David Hall 0006, Daniel Jurafsky, Christopher D. Manning
EMNLP1