EDBT 2026 Demo / reviewers in the wild / expert
Zifan Carl Guo
dblp:332/9532
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 33% Deep learning architectures and training · 33% Trustworthy machine learning · 18% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.9 | 1 | 2025 | (How) Do Language Models Track State? · ICML 2025 |
Machine learning › Deep learning architectures and training › sequence modeling
state tracking |
0.9 | 1 | 2025 | (How) Do Language Models Track State? · ICML 2025 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model |
0.9 | 1 | 2025 | (How) Do Language Models Track State? · ICML 2025 |
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training |
0.8 | 1 | 2024 | Algorithmic progress in language models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.8 | 1 | 2024 | Algorithmic progress in language models · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
scaling laws |
0.8 | 1 | 2024 | Algorithmic progress in language models · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
permutation composition · 0.9intermediate training tasks · 0.9associative scan · 0.9scaling law estimation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | (How) Do Language Models Track State?abstractTransformer language models (LMs) exhibit behaviors—from storytelling to code generation—that seem to require tracking the unobserved state of an evolving world. How do they do this? We study state tracking in LMs trained or fine-tuned to compose permutations (i.e., to compute the order of a set of objects after a sequence of swaps). Despite the simple algebraic structure of this problem, many other tasks (e.g., simulation of finite automata and evaluation of boolean expressions) can be reduced to permutation composition, making it a natural model for state tracking in general. We show that LMs consistently learn one of two state tracking mechanisms for this task. The first closely resembles the “associative scan” construction used in recent theoretical work by Liu et al. (2023) and Merrill et al. (2024). The second uses an easy-to-compute feature (permutation parity) to partially prune the space of outputs, and then refines this with an associative scan. LMs that learn the former algorithm tend to generalize better and converge faster, and we show how to steer LMs toward one or the other with intermediate training tasks that encourage or suppress the heuristics. Our results demonstrate that transformer LMs, whether pre-trained or fine-tuned, can learn to implement efficient and interpretable state-tracking mechanisms, and the emergence of these mechanisms can be predicted and controlled. Code and data are available at https://github.com/belindal/state-tracking Belinda Z. Li, Zifan Carl Guo, Jacob Andreas |
ICML | 2 |
| 2024 | Algorithmic progress in language modelsabstractWe investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 90\% confidence interval of around 2 to 22 months, substantially faster than hardware gains per Moore's Law. We estimate augmented scaling laws, which enable us to quantify algorithmic progress and determine the relative contributions of scaling models versus innovations in training algorithms. Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period. Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms. Anson Ho, Tamay Besiroglu, Ege Erdil, Zifan Carl Guo, David Owen 0001, Robi Rahman, David Atkinson, Neil Thompson, Jaime Sevilla |
NeurIPS | 4 |