VLDB 2026 Research / reviewers in the wild / expert
Yinwei Dai
dblp:293/8320
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-9291-2060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remembrall: Leaning into Memory for Accurate Video Analytics on System-on-Chip GPUs
Murali Ramanujam, Yinwei Dai, Kyle Jamieson, Ravi Netravali |
NSDI | 2 |
| 2025 | SpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningabstractRecent advances in inference-time compute have significantly improved performance on complex tasks by generating long chains of thought (CoTs) using Large Reasoning Models (LRMs). However, this improved accuracy comes at the cost of high inference latency due to the length of generated reasoning sequences and the autoregressive nature of decoding. Our key insight in tackling these overheads is that LRM inference, and the reasoning that it embeds, is highly tolerant of approximations: complex tasks are typically broken down into simpler steps, each of which brings utility based on the semantic insight it provides for downstream steps rather than the exact tokens it generates. Accordingly, we introduce SpecReason, a system that automatically accelerates LRM inference by using a lightweight model to (speculatively) carry out simpler intermediate reasoning steps and reserving the costly base model only to efficiently assess (and potentially correct) the speculated outputs. Importantly, SpecReason's focus on exploiting the semantic flexibility of thinking tokens in preserving final-answer accuracy is complementary to prior speculation techniques, most notably speculative decoding, which demands token-level equivalence at each step. Across a variety of cross-domain reasoning benchmarks, SpecReason achieves 1.4-3.0$\times$ speedup over vanilla LRM inference while improving accuracy by 0.4-9.0%. Compared to speculative decoding without SpecReason, their combination yields an additional 8.8-58.0% latency reduction. We open-source SpecReason at \url{https://anonymous.4open.science/r/specreason/}. Rui Pan 0003, Yinwei Dai, Zhihao Zhang 0001, Gabriele Oliaro, Ravi Netravali |
NeurIPS | 2 |
| 2024 | Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML ServingabstractMachine learning (ML) inference platforms are tasked with balancing two competing goals: ensuring high throughput given many requests, and delivering low-latency responses to support interactive applications. Unfortunately, existing platform knobs (e.g., batch sizes) fail to ease this fundamental tension, and instead only enable users to harshly trade off one property for the other. This paper explores an alternate strategy to taming throughput-latency tradeoffs by changing the granularity at which inference is performed. We present Apparate, a system that automatically applies and manages early exits (EEs) in ML models, whereby certain inputs can exit with results at intermediate layers. To cope with the time-varying overhead and accuracy challenges that EEs bring, Apparate repurposes exits to provide continual feedback that powers several novel runtime monitoring and adaptation strategies. Apparate lowers median response latencies by 40.5--91.5% and 10.0--24.2% for diverse CV and NLP classification workloads, and median time-per-token latencies by 22.6--77.9% for generative scenarios, without affecting throughputs or violating tight accuracy constraints. Yinwei Dai, Rui Pan 0003, Anand Padmanabha Iyer, Kai Li 0001, Ravi Netravali |
SOSP | 1 |
| 2024 | Improving DNN Inference Throughput Using Practical, Per-Input Compute AdaptationabstractMachine learning inference platforms continue to face high request rates and strict latency constraints. Existing solutions largely focus on compressing models to substantially lower compute costs (and time) with mild accuracy degradations. This paper explores an alternate (but complementary) technique that trades off accuracy and resource costs on a perinput granularity: early exit models, which selectively allow certain inputs to exit a model from an intermediate layer. Though intuitive, early exits face fundamental deployment challenges, largely owing to the effects that exiting inputs have on batch size (and resource utilization) throughout model execution. We present E3, the first system that makes early exit models practical for realistic inference deployments. Our key insight is to split and replicate blocks of layers in models in a manner that maintains a constant batch size throughout execution, all the while accounting for resource requirements and communication overheads. Evaluations with NLP and vision models show that E3 can deliver up to 1.74× improvement in goodput (for a fixed cost) or 1.78× reduction in cost (for a fixed goodput). Additionally, E3's goodput wins generalize to autoregressive LLMs (2.8--3.8×) and compressed models (1.67×). Anand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan 0003, Swapnil Gandhi, Ravi Netravali |
SOSP | 3 |
| 2023 | Auxo: Efficient Federated Learning via Scalable Client ClusteringabstractFederated learning (FL) is an emerging machine learning (ML) paradigm that enables heterogeneous edge devices to collaboratively train ML models without revealing their raw data to a logically centralized server. However, beyond the heterogeneous device capacity, FL participants often exhibit differences in their data distributions, which are not independent and identically distributed (Non-IID). Many existing works present point solutions to address issues like slow convergence, low final accuracy, and bias in FL, all stemming from client heterogeneity. Fan Lai 0001, Yinwei Dai, Aditya Akella, Harsha V. Madhyastha, Mosharaf Chowdhury |
SoCC | 3 |
| 2023 | ModelKeeper: Accelerating DNN Training via Automated Training Warmup
Fan Lai 0001, Yinwei Dai, Harsha V. Madhyastha, Mosharaf Chowdhury |
NSDI | 2 |
| 2022 | FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleabstractWe present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging from image classification and object detection to language modeling and speech recognition. Each dataset comes with a unified evaluation protocol using real-world data splits and evaluation metrics. To reproduce realistic FL behavior, FedScale contains a scalable and extensible runtime. It provides high-level APIs to implement FL algorithms, deploy them at scale across diverse hardware and software backends, and evaluate them at scale, all with minimal developer efforts. We combine the two to perform systematic benchmarking experiments and highlight potential opportunities for heterogeneity-aware co-optimizations in FL. FedScale is open-source and actively maintained by contributors from different institutions at http://fedscale.ai. We welcome feedback and contributions from the community. Fan Lai 0001, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
ICML | 2 |