Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Clara Mohri

dblp:331/5522 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Learning theory · 34% Deep learning architectures and training · 19% Language models and text generation · 19%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Mixture of Parrots: Experts improve memorization more than reasoning · ICLR 2025
Computer vision › Vision and language › language-guided learning › language-guided vision
language-guided visual recognition
0.712023
Using Language to Extend to Unseen Domains · ICLR 2023
Machine learning › Learning theory
matrix completion
0.712023
Partial Matrix Completion · NeurIPS 2023
Machine learning › Learning theory
online learning
0.712023
Partial Matrix Completion · NeurIPS 2023
Machine learning › Learning theory › neural network theory › neural network analysis
memorization versus reasoning
0.312025
Mixture of Parrots: Experts improve memorization more than reasoning · ICLR 2025

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 0.9mixture of experts · 0.9vision-language model · 0.7language model · 0.7iterative gradient updates · 0.7
YearPublicationVenuePosition
2025 Mixture of Parrots: Experts improve memorization more than reasoning
abstract
The Mixture-of-Experts (MoE) architecture enables a significant increase in the total number of model parameters with minimal computational overhead. However, it is not clear what performance tradeoffs, if any, exist between MoEs and standard dense transformers. In this paper, we show that as we increase the number of experts (while fixing the number of active parameters), the memorization performance consistently increases while the reasoning capabilities saturate. We begin by analyzing the theoretical limitations of MoEs at reasoning. We prove that there exist graph problems that cannot be solved by any number of experts of a certain width; however, the same task can be easily solved by a dense model with a slightly larger width. On the other hand, we find that on memory-intensive tasks, MoEs can effectively leverage a small number of active parameters with a large number of experts to memorize the data. We empirically validate these findings on synthetic graph problems and memory-intensive closed book retrieval tasks. Lastly, we pre-train a series of MoEs and dense transformers and evaluate them on commonly used benchmarks in math and natural language. We find that increasing the number of experts helps solve knowledge-intensive tasks, but fails to yield the same benefits for reasoning tasks.
Samy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu, Nikhil Vyas 0001, Nikhil Anand, David Alvarez-Melis, Yuanzhi Li, Sham M. Kakade, Eran Malach
ICLR2
2023 Using Language to Extend to Unseen Domains
Lisa Dunlap, Clara Mohri, Devin Guillory, Trevor Darrell, Joseph Gonzalez 0001, Aditi Raghunathan, Anna Rohrbach
ICLR2
2023 Partial Matrix Completion
abstract
The matrix completion problem involves reconstructing a low-rank matrix by using a given set of revealed (and potentially noisy) entries. Although existing methods address the completion of the entire matrix, the accuracy of the completed entries can vary significantly across the matrix, due to differences in the sampling distribution. For instance, users may rate movies primarily from their country or favorite genres, leading to inaccurate predictions for the majority of completed entries. We propose a novel formulation of the problem as Partial Matrix Completion, where the objective is to complete a substantial subset of the entries with high confidence. Our algorithm efficiently handles the unknown and arbitrarily complex nature of the sampling distribution, ensuring high accuracy for all completed entries and sufficient coverage across the matrix. Additionally, we introduce an online version of the problem and present a low-regret efficient algorithm based on iterative gradient updates. Finally, we conduct a preliminary empirical evaluation of our methods.
Elad Hazan, Adam Tauman Kalai, Varun Kanade, Clara Mohri, Y. Jennifer Sun
NeurIPS4