VLDB 2026 Research / reviewers in the wild / expert
Ashish Rao
dblp:260/0227
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
1 paper |
Algorithms and data structures · 50% Mathematical optimization · 50% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization
least squares |
0.9 | 1 | 2025 | Towards Learning High-Precision Least Squares Algorithms with Sequence Models · ICLR 2025 |
Algorithms and data structures
numerical algorithms |
0.9 | 1 | 2025 | Towards Learning High-Precision Least Squares Algorithms with Sequence Models · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.7polynomial architectures · 1.7linear attention · 1.7high-precision training · 1.7gated convolution · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Learning High-Precision Least Squares Algorithms with Sequence ModelsabstractThis paper investigates whether sequence models can learn to perform numerical algorithms, e.g. gradient descent, on the fundamental problem of least squares. Our goal is to inherit two properties of standard algorithms from numerical analysis: (1) machine precision, i.e. we want to obtain solutions that are accurate to near floating point error, and (2) numerical generality, i.e. we want them to apply broadly across problem instances. We find that prior approaches using Transformers fail to meet these criteria, and identify limitations present in existing architectures and training procedures. First, we show that softmax Transformers struggle to perform high-precision multiplications, which prevents them from precisely learning numerical algorithms. Second, we identify an alternate class of architectures, comprised entirely of polynomials, that can efficiently represent high-precision gradient descent iterates. Finally, we investigate precision bottlenecks during training and address them via a high-precision training recipe that reduces stochastic gradient noise. Our recipe enables us to train two polynomial architectures, gated convolutions and linear attention, to perform gradient descent iterates on least squares problems. For the first time, we demonstrate the ability to train to near machine precision. Applied iteratively, our models obtain $100,000\times$ lower MSE than standard Transformers trained end-to-end and they incur a $10,000\times$ smaller generalization gap on out-of-distribution problems. We make progress towards end-to-end learning of numerical algorithms for least squares. Jerry W. Liu, Jessica Grogan, Owen Dugan, Ashish Rao, Simran Arora, Atri Rudra, Christopher Ré |
ICLR | 4 |