EDBT 2026 Demo / reviewers in the wild / expert
Clara Mohri
dblp:331/5522
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Learning theory · 34% Deep learning architectures and training · 19% Language models and text generation · 19% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | Mixture of Parrots: Experts improve memorization more than reasoning · ICLR 2025 |
Computer vision › Vision and language › language-guided learning › language-guided vision
language-guided visual recognition |
0.7 | 1 | 2023 | Using Language to Extend to Unseen Domains · ICLR 2023 |
Machine learning › Learning theory
matrix completion |
0.7 | 1 | 2023 | Partial Matrix Completion · NeurIPS 2023 |
Machine learning › Learning theory
online learning |
0.7 | 1 | 2023 | Partial Matrix Completion · NeurIPS 2023 |
Machine learning › Learning theory › neural network theory › neural network analysis
memorization versus reasoning |
0.3 | 1 | 2025 | Mixture of Parrots: Experts improve memorization more than reasoning · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
theoretical analysis · 0.9mixture of experts · 0.9vision-language model · 0.7language model · 0.7iterative gradient updates · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mixture of Parrots: Experts improve memorization more than reasoningabstractThe Mixture-of-Experts (MoE) architecture enables a significant increase in the total number of model parameters with minimal computational overhead.
However, it is not clear what performance tradeoffs, if any, exist between MoEs and standard dense transformers.
In this paper,
we show that as we increase the number of experts (while fixing the number of active parameters), the memorization performance consistently increases while the reasoning capabilities saturate.
We begin by analyzing the theoretical limitations of MoEs at reasoning. We prove that there exist graph problems that cannot be solved by any number of experts of a certain width; however, the same task can be easily solved by a dense model with a slightly larger width.
On the other hand, we find that on memory-intensive tasks, MoEs can effectively leverage a small number of active parameters with a large number of experts to memorize the data.
We empirically validate these findings on synthetic graph problems and memory-intensive closed book retrieval tasks.
Lastly, we pre-train a series of MoEs and dense transformers and evaluate them on commonly used benchmarks in math and natural language.
We find that increasing the number of experts helps solve knowledge-intensive tasks, but fails to yield the same benefits for reasoning tasks. Samy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu, Nikhil Vyas 0001, Nikhil Anand, David Alvarez-Melis, Yuanzhi Li, Sham M. Kakade, Eran Malach |
ICLR | 2 |
| 2023 | Using Language to Extend to Unseen Domains
Lisa Dunlap, Clara Mohri, Devin Guillory, Trevor Darrell, Joseph Gonzalez 0001, Aditi Raghunathan, Anna Rohrbach |
ICLR | 2 |
| 2023 | Partial Matrix CompletionabstractThe matrix completion problem involves reconstructing a low-rank matrix by using a given set of revealed (and potentially noisy) entries. Although existing methods address the completion of the entire matrix, the accuracy of the completed entries can vary significantly across the matrix, due to differences in the sampling distribution. For instance, users may rate movies primarily from their country or favorite genres, leading to inaccurate predictions for the majority of completed entries.
We propose a novel formulation of the problem as Partial Matrix Completion, where the objective is to complete a substantial subset of the entries with high confidence. Our algorithm efficiently handles the unknown and arbitrarily complex nature of the sampling distribution, ensuring high accuracy for all completed entries and sufficient coverage across the matrix. Additionally, we introduce an online version of the problem and present a low-regret efficient algorithm based on iterative gradient updates. Finally, we conduct a preliminary empirical evaluation of our methods. Elad Hazan, Adam Tauman Kalai, Varun Kanade, Clara Mohri, Y. Jennifer Sun |
NeurIPS | 4 |