VLDB 2026 Research / reviewers in the wild / expert
William L. Tong
dblp:315/0406 · also William Lingxiao Tong
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Learning theory · 31% Probabilistic and Bayesian machine learning · 23% Language models and text generation · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › neural network theory › neural network analysis
architecture comparison |
0.9 | 1 | 2025 | MLPs Learn In-Context on Regression and Classification Tasks · ICLR 2025 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.9 | 1 | 2025 | No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions · ICML 2025 |
Machine learning › Learning theory
generalization bounds |
0.9 | 1 | 2025 | No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions · ICML 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | MLPs Learn In-Context on Regression and Classification Tasks · ICLR 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.7 | 1 | 2023 | Neural Circuits for Fast Poisson Compressed Sensing in the Olfactory Bulb · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling |
0.7 | 1 | 2023 | Neural Circuits for Fast Poisson Compressed Sensing in the Olfactory Bulb · NeurIPS 2023 |
Bioinformatics and computational biology
computational neuroscience |
0.7 | 1 | 2023 | Neural Circuits for Fast Poisson Compressed Sensing in the Olfactory Bulb · NeurIPS 2023 |
Bioinformatics and computational biology › computational neuroscience › neural coding
olfactory coding |
0.7 | 1 | 2023 | Neural Circuits for Fast Poisson Compressed Sensing in the Olfactory Bulb · NeurIPS 2023 |
Information theory › signal processing
compressed sensing |
0.7 | 1 | 2023 | Neural Circuits for Fast Poisson Compressed Sensing in the Olfactory Bulb · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
normative modeling · 2.0ridge regression · 0.9random features · 0.9multi-layer perceptron · 0.9deterministic equivalent · 0.9MLP-Mixer · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MLPs Learn In-Context on Regression and Classification TasksabstractIn-context learning (ICL), the remarkable ability to solve a task from only input exemplars, is often assumed to be a unique hallmark of Transformer models. By examining commonly employed synthetic ICL tasks, we demonstrate that multi-layer perceptrons (MLPs) can also learn in-context. Moreover, MLPs, and the closely related MLP-Mixer models, learn in-context comparably with Transformers under the same compute budget in this setting. We further show that MLPs outperform Transformers on a series of classical tasks from psychology designed to test relational reasoning, which are closely related to in-context classification. These results underscore a need for studying in-context learning beyond attention-based architectures, while also challenging prior arguments against MLPs' ability to solve relational tasks. Altogether, our results highlight the unexpected competence of MLPs in a synthetic setting, and support the growing interest in all-MLP alternatives to Transformer architectures. It remains unclear how MLPs perform against Transformers at scale on real-world tasks, and where a performance gap may originate. We encourage further exploration of these architectures in more complex settings to better understand the potential comparative advantage of attention-based schemes. William L. Tong, Cengiz Pehlevan |
ICLR | 1 |
| 2025 | No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality ConditionsabstractGiven a fixed budget for total model size, one must choose between training a single large model or combining the predictions of multiple smaller models.
We investigate this trade-off for ensembles of random-feature ridge regression models in both the overparameterized and underparameterized regimes.
Using deterministic equivalent risk estimates, we prove that when a fixed number of parameters is distributed among $K$ independently trained models, the ridge-optimized test risk increases with $K$.
Consequently, a single large model achieves optimal performance. We then ask when ensembles can achieve *near*-optimal performance.
In the overparameterized regime, we show that, to leading order, the test error depends on ensemble size and model size only through the total feature count, so that overparameterized ensembles consistently achieve near-optimal performance.
To understand underparameterized ensembles, we derive scaling laws for the test risk as a function of total parameter count when the ensemble size and parameters per ensemble member are jointly scaled according to a ``growth exponent'' $\ell$.
While the optimal error scaling is always achieved by increasing model size with a fixed ensemble size, our analysis identifies conditions on the kernel and task eigenstructure under which near-optimal scaling laws can be obtained by joint scaling of ensemble size and model size. Benjamin S. Ruben, William L. Tong, Hamza Tahir Chaudhry, Cengiz Pehlevan |
ICML | 2 |
| 2025 | Adaptive algorithms for shaping behaviorabstractDogs and laboratory mice are commonly trained to perform complex tasks by guiding them through a curriculum of simpler tasks ('shaping'). What are the principles behind effective shaping strategies? Here, we propose a teacher-student framework for shaping behavior, where an autonomous teacher agent decides its student's task based on the student's transcript of successes and failures on previously assigned tasks. Using algorithms for Monte Carlo planning under uncertainty, we show that near-optimal shaping algorithms achieve a careful balance between reinforcement and extinction. Near-optimal algorithms track learning rate to adaptively alternate between simpler and harder tasks. Based on this intuition, we derive an adaptive shaping heuristic with minimal parameters, which we show is near-optimal on a sequence learning task and robustly trains deep reinforcement learning agents on navigation tasks that involve sparse, delayed rewards. Extensions to continuous curricula are explored. Our work provides a starting point towards a general computational framework for shaping behavior that applies to both animals and artificial agents. William L. Tong, Venkatesh N. Murthy, Gautam Reddy |
PLoS Comput. Biol. | 1 |
| 2023 | Neural Circuits for Fast Poisson Compressed Sensing in the Olfactory BulbabstractWithin a single sniff, the mammalian olfactory system can decode the identity and concentration of odorants wafted on turbulent plumes of air. Yet, it must do so given access only to the noisy, dimensionally-reduced representation of the odor world provided by olfactory receptor neurons. As a result, the olfactory system must solve a compressed sensing problem, relying on the fact that only a handful of the millions of possible odorants are present in a given scene. Inspired by this principle, past works have proposed normative compressed sensing models for olfactory decoding. However, these models have not captured the unique anatomy and physiology of the olfactory bulb, nor have they shown that sensing can be achieved within the 100-millisecond timescale of a single sniff. Here, we propose a rate-based Poisson compressed sensing circuit model for the olfactory bulb. This model maps onto the neuron classes of the olfactory bulb, and recapitulates salient features of their connectivity and physiology. For circuit sizes comparable to the human olfactory bulb, we show that this model can accurately detect tens of odors within the timescale of a single sniff. We also show that this model can perform Bayesian posterior sampling for accurate uncertainty estimation. Fast inference is possible only if the geometry of the neural code is chosen to match receptor properties, yielding a distributed neural code that is not axis-aligned to individual odor identities. Our results illustrate how normative modeling can help us map function onto specific neural circuits to generate new hypotheses. Jacob A. Zavatone-Veth, Paul Masset, William L. Tong, Joseph D. Zak, Venkatesh Murthy, Cengiz Pehlevan |
NeurIPS | 3 |