VLDB 2026 Research / reviewers in the wild / expert
Sabri Eyuboglu
dblp:298/7563 · also Evan Sabri Eyuboglu
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-8412-0266ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Efficient and distributed learning · 32% Deep learning architectures and training · 23% Language models and text generation · 23% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 36% Data integration and cleaning · 32% Distributed and cloud data management · 32% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 19 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | Cost-efficient Collaboration between On-device and Cloud Language Models · ICML 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | Cost-efficient Collaboration between On-device and Cloud Language Models · ICML 2025 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation |
0.9 | 1 | 2025 | Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters · ICLR 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters · ICLR 2025 |
Machine learning › Deep learning architectures and training
associative recall |
0.8 | 1 | 2024 | Zoology: Measuring and Improving Recall in Efficient Language Models · ICLR 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.8 | 1 | 2024 | Simple linear attention language models balance the recall-throughput tradeoff · ICML 2024 |
Natural language and speech › Language models and text generation
efficient language model |
0.8 | 1 | 2024 | Zoology: Measuring and Improving Recall in Efficient Language Models · ICLR 2024 |
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention |
0.8 | 1 | 2024 | Simple linear attention language models balance the recall-throughput tradeoff · ICML 2024 |
Machine learning › Trustworthy machine learning
Data-centric AI |
0.7 | 1 | 2023 | DataPerf: Benchmarks for Data-Centric AI Development · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
efficient architecture design |
0.7 | 1 | 2023 | Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | HAPI Explorer: Comprehension, Discovery, and Explanation on History of ML APIs · AAAI 2023 |
Machine learning › Efficient and distributed learning › model compression
structured matrices |
0.7 | 1 | 2023 | Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture · NeurIPS 2023 |
Distributed and cloud data management
data lake |
0.7 | 1 | 2023 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes · Proc. VLDB Endow. 2023 |
Visualization and visual analytics
interactive data exploration |
0.7 | 1 | 2023 | HAPI Explorer: Comprehension, Discovery, and Explanation on History of ML APIs · AAAI 2023 |
Machine learning › Trustworthy machine learning
fairness |
0.6 | 1 | 2022 | HAPI: A Large-scale Longitudinal Dataset of Commercial ML API Predictions · NeurIPS 2022 |
Computer vision › Vision and language › cross-modal alignment
visual-semantic embedding |
0.6 | 1 | 2022 | Domino: Discovering Systematic Errors with Cross-Modal Embeddings · ICLR 2022 |
Machine learning and data management › machine learning systems
machine learning as a service |
0.6 | 1 | 2022 | HAPI: A Large-scale Longitudinal Dataset of Commercial ML API Predictions · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
weak supervision · 1.3large language model · 1.3in-context learning · 1.3task decomposition · 0.9rank adapter · 0.9low-rank matrix decomposition · 0.9local-remote collaboration · 0.9adaptive masking · 0.9sparse attention · 0.8multi-query associative recall · 0.8linear attention · 0.8gated convolution · 0.8IO-aware algorithms · 0.8visual analytics · 0.7natural language explanation · 0.7longitudinal analysis · 0.6benchmark evaluation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA AdaptersabstractLarge Language Models (LLMs) are computationally intensive, particularly during inference. Neuron-adaptive techniques, which selectively activate neurons in Multi-Layer Perceptron (MLP) layers, offer some speedups but suffer from limitations in modern Transformers. These include reliance on sparse activations, incompatibility with attention layers, and the use of costly neuron masking techniques. To address these issues, we propose the Adaptive Rank Allocation framework and introduce the Rank and Neuron Allocator (RaNA) adapter. RaNA adapters leverage rank adapters, which operate on linear layers by applying both low-rank matrix decompositions and adaptive masking to efficiently allocate compute without depending on activation sparsity. This enables RaNA to be generally applied to MLPs and linear components of attention modules, while eliminating the need for expensive maskers found in neuron-adaptive methods. Notably, when compared to neuron adapters, RaNA improves perplexity by up to 7 points and increases accuracy by up to 8 percentage-points when reducing FLOPs by $\sim$44\% in state-of-the-art Transformer architectures. These results position RaNA as a robust solution for improving inference efficiency in modern Transformer architectures. Roberto Garcia, Jerry W. Liu, Daniel Sorvisto, Sabri Eyuboglu |
ICLR | 4 |
| 2025 | Cost-efficient Collaboration between On-device and Cloud Language ModelsabstractWe investigate an emerging setup in which a small, on-device language model (LM) with access to local data collaborates with a frontier, cloud-hosted LM to solve real-world tasks involving financial, medical, and scientific reasoning over long documents. Can a local-remote collaboration reduce cloud inference costs while preserving quality? First, we consider a naïve collaboration protocol, coined MINION, where the local and remote models simply chat back and forth. Because only the local model ingests the full context, this protocol reduces cloud costs by 30.4x, but recovers only 87% of the performance of the frontier model. We identify two key limitations of this protocol: the local model struggles to (1) follow the remote model’s multi-step instructions and (2) reason over long contexts. Motivated by these observations, we propose MINIONS, a protocol in which the remote model decomposes the task into easier subtasks over shorter chunks of the document, that are executed locally in parallel. MINIONS reduces costs by 5.7$\times$ on average while recovering 97.9% of the remote-only performance. Our analysis reveals several key design choices that influence the trade-off between cost and performance in local-remote systems. Avanika Narayan, Dan Biderman, Sabri Eyuboglu, Avner May, Scott W. Linderman, James Zou 0001, Christopher Ré |
ICML | 3 |
| 2024 | Zoology: Measuring and Improving Recall in Efficient Language ModelsabstractAttention-free language models that combine gating and convolutions are growing in popularity due to their efficiency and increasingly competitive performance. To better understand these architectures, we pretrain a suite of 17 attention and gated-convolution language models, finding that SoTA gated-convolution architectures still underperform attention by up to 2.1 perplexity points on the Pile. In fine-grained analysis, we find 82% of the gap is explained by each model's ability to recall information that is previously mentioned in-context, e.g. "Hakuna Matata means no worries Hakuna Matata it means no" -> ??. On this task, termed "associative recall", we find that attention outperforms gated-convolutions by a large margin: a 70M parameter attention model outperforms a 1.4 billion parameter gated-convolution model on associative recall. This is surprising because prior work shows gated convolutions can perfectly solve synthetic tests for AR capability. To close the gap between synthetics and real language, we develop a new formalization of the task called multi-query associative recall (MQAR) that better reflects actual language. We perform an empirical and theoretical study of MQAR that elucidates differences in the parameter-efficiency of attention and gated-convolution recall. Informed by our analysis, we evaluate simple convolution-attention hybrids and show that hybrids with input-dependent sparse attention patterns can close 97.4% of the gap to attention, while maintaining sub-quadratic scaling. Code is at: https://github.com/HazyResearch/zoology. Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou 0001, Atri Rudra, Christopher Ré |
ICLR | 2 |
| 2024 | Simple linear attention language models balance the recall-throughput tradeoffabstractRecent work has shown that attention-based language models excel at "recall", the ability to ground generations in tokens previously seen in context. However, the efficiency of attention-based models is bottle-necked during inference by the KV-cache's aggressive memory consumption. In this work, we explore whether we can improve language model efficiency (e.g. by reducing memory consumption) without compromising on recall. By applying experiments and theory to a broad set of architectures, we identify a key tradeoff between a model's recurrent state size and recall ability. We show that efficient alternatives to attention (e.g. H3, Mamba, RWKV) maintain a fixed-size recurrent state, but struggle at recall. We propose BASED a simple architecture combining linear and sliding window attention. By varying BASED window size and linear attention feature dimension, we can dial the state size and traverse the Pareto frontier of the recall-memory tradeoff curve, recovering the full quality of attention on one end and the small state size of attention-alternatives on the other. We train language models up to $1.3$b parameters and show that BASED matches the strongest sub-quadratic models (e.g. Mamba) in perplexity and outperforms them on real-world recall-intensive tasks by 10.36 accuracy points. We further develop IO-aware algorithms that enable BASED to provide 24× higher throughput on language generation than FlashAttention-2, when generating 1024 tokens using 1.3b parameter models. Overall, BASED expands the Pareto frontier of the throughput-recall tradeoff space beyond prior architectures. Simran Arora, Sabri Eyuboglu, Aman Timalsina, Silas Alberti, James Zou 0001, Atri Rudra, Christopher Ré |
ICML | 2 |
| 2023 | HAPI Explorer: Comprehension, Discovery, and Explanation on History of ML APIsabstractMachine learning prediction APIs offered by Google, Microsoft, Amazon, and many other providers have been continuously adopted in a plethora of applications, such as visual object detection, natural language comprehension, and speech recognition. Despite the importance of a systematic study and comparison of different APIs over time, this topic is currently under-explored because of the lack of data and user-friendly exploration tools. To address this issue, we present HAPI Explorer (History of API Explorer), an interactive system that offers easy access to millions of instances of commercial API applications collected in three years, prioritize attention on user-defined instance regimes, and explain interesting patterns across different APIs, subpopulations, and time periods via visual and natural languages. HAPI Explorer can facilitate further comprehension and exploitation of ML prediction APIs. Lingjiao Chen, Zhihua Jin, Sabri Eyuboglu, Huamin Qu, Christopher Ré, Matei Zaharia, James Zou 0001 |
AAAI | 3 |
| 2023 | Monarch Mixer: A Simple Sub-Quadratic GEMM-Based ArchitectureabstractMachine learning models are increasingly being scaled in both sequence length and model dimension to reach longer contexts and better performance. However, existing architectures such as Transformers scale quadratically along both these axes. We ask: are there performant architectures that can scale sub-quadratically along sequence length and model dimension? We introduce Monarch Mixer (M2), a new architecture that uses the same sub-quadratic primitive along both sequence length and model dimension: Monarch matrices, a simple class of expressive structured matrices that captures many linear transforms, achieves high hardware efficiency on GPUs, and scales sub-quadratically. As a proof of concept, we explore the performance of M2 in three domains: non-causal BERT-style language modeling, ViT-style image classification, and causal GPT-style language modeling. For non-causal BERT-style modeling, M2 matches BERT-base and BERT-large in downstream GLUE quality with up to 27% fewer parameters, and achieves up to 9.1$\times$ higher throughput at sequence length 4K. On ImageNet, M2 outperforms ViT-b by 1% in accuracy, with only half the parameters. Causal GPT-style models introduce a technical challenge: enforcing causality via masking introduces a quadratic bottleneck. To alleviate this bottleneck, we develop a novel theoretical view of Monarch matrices based on multivariate polynomial evaluation and interpolation, which lets us parameterize M2 to be causal while remaining sub-quadratic. Using this parameterization, M2 matches GPT-style Transformers at 360M parameters in pretraining perplexity on The PILE—showing for the first time that it may be possible to match Transformer quality without attention or MLPs. Daniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson, Sabri Eyuboglu, Armin W. Thomas, Benjamin Spector, Michael Poli, Atri Rudra, Christopher Ré |
NeurIPS | 5 |
| 2023 | DataPerf: Benchmarks for Data-Centric AI DevelopmentabstractMachine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry. Mark Mazumder, Colby R. Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Frederick Diamos, Gregory Frederick Diamos, Lynn He, Alicia Parrish, Hannah Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett D. Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller 0001, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung-Levy, Newsha Ardalani, Praveen K. Paritosh, Ce Zhang 0001, James Zou 0001, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi |
NeurIPS | 24 |
| 2023 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesabstractA long standing goal in the data management community is developing systems that input documents and output queryable tables without user effort. Given the sheer variety of potential documents, state-of-the art systems make simplifying assumptions and use domain specific training. In this work, we ask whether we can maintain generality by using the in-context learning abilities of large language models (LLMs). We propose and evaluate Evaporate, a prototype system powered by LLMs. We identify two strategies for implementing this system: prompt the LLM to directly extract values from documents or prompt the LLM to synthesize code that performs the extraction. Our evaluations show a cost-quality tradeoff between these two approaches. Code synthesis is cheap, but far less accurate than directly processing each document with the LLM. To improve quality while maintaining low cost, we propose an extended implementation, Evaporate-Code+, which achieves better quality than direct extraction. Our insight is to generate many candidate functions and ensemble their extractions using weak supervision. Evaporate-Code+ outperforms the state-of-the art systems using a sublinear pass over the documents with the LLM. This equates to a 110X reduction in the number of documents the LLM needs to process across our 16 real-world evaluation settings. Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, Christopher Ré |
Proc. VLDB Endow. | 3 |
| 2022 | Domino: Discovering Systematic Errors with Cross-Modal Embeddings
Sabri Eyuboglu, Maya Varma, Khaled Saab 0002, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou 0001, Christopher Ré |
ICLR | 1 |
| 2022 | HAPI: A Large-scale Longitudinal Dataset of Commercial ML API PredictionsabstractCommercial ML APIs offered by providers such as Google, Amazon and Microsoft have dramatically simplified ML adoptions in many applications. Numerous companies and academics pay to use ML APIs for tasks such as object detection, OCR and sentiment analysis. Different ML APIs tackling the same task can have very heterogeneous performances. Moreover, the ML models underlying the APIs also evolve over time. As ML APIs rapidly become a valuable marketplace and an integral part of analytics, it is critical to systematically study and compare different APIs with each other and to characterize how individual APIs change over time. However, this practically important topic is currently underexplored due to the lack of data. In this paper, we present HAPI (History of APIs), a longitudinal dataset of 1,761,417 instances of commercial ML API applications (involving APIs from Amazon, Google, IBM, Microsoft and other providers) across diverse tasks including image tagging, speech recognition, and text mining from 2020 to 2022. Each instance consists of a query input for an API (e.g., an image or text) along with the API’s output prediction/annotation and confidence scores. HAPI is the first large-scale dataset of ML API usages and is a unique resource for studying ML as-a-service (MLaaS). As examples of the types of analyses that HAPI enables, we show that ML APIs’ performance changes substantially over time—several APIs’ accuracies dropped on specific benchmark datasets. Even when the API’s aggregate performance stays steady, its error modes can shift across different subtypes of data between 2020 and 2022. Such changes can substantially impact the entire analytics pipelines that use some ML API as a component. We further use HAPI to study commercial APIs’ performance disparities across demographic subgroups over time. HAPI can stimulate more research in the growing field of MLaaS. Lingjiao Chen, Zhihua Jin, Sabri Eyuboglu, Christopher Ré, Matei Zaharia, James Zou 0001 |
NeurIPS | 3 |