VLDB 2026 Research / reviewers in the wild / expert
Ammar Vora
dblp:421/0507
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Deep learning architectures and training · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Attention-Level Speculation · ICML 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Attention-Level Speculation · ICML 2025 |
Machine learning › Deep learning architectures and training › transformer › efficient transformer
self-attention approximation |
0.9 | 1 | 2025 | Attention-Level Speculation · ICML 2025 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
0.9 | 1 | 2025 | Attention-Level Speculation · ICML 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit |
0.3 | 1 | 2025 | Attention-Level Speculation · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
speculative execution · 1.7attention approximation · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attention-Level SpeculationabstractAs Large Language Models (LLMs) grow in size and context length, efficient inference strategies are essential to maintain low-latency token generation. Unfortunately, conventional tensor and data parallelism face diminishing returns when scaling across multiple devices. We propose a novel form—attention-level speculative parallelism (ALSpec)—that predicts self-attention outputs to execute subsequent operations early on separate devices. Our approach overlaps attention and non-attention computations, reducing the attention latency overhead at 128K context length by up to 5x and improving end-to-end decode latency by up to 1.65x, all without sacrificing quality. We establish the fundamental pillars for speculative execution and provide an execution paradigm that simplifies implementation. We show that existing attention-approximation methods perform well on simple information retrieval tasks, but they fail in advanced reasoning and math. Combined with speculative execution, we can approximate up to 90% of self-attention without harming model correctness. Demonstrated on Tenstorrent’s NPU devices, we scale up LLM inference beyond current techniques, paving the way for faster inference in transformer models. Jack Cai, Ammar Vora, Randolph Zhang, Mark O'Connor, Mark C. Jeffrey |
ICML | 2 |