EDBT 2026 Demo / reviewers in the wild / expert
Parsa Mirtaheri
dblp:400/5482
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 67% 3D vision · 17% Reinforcement learning · 17% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Direct Alignment with Heterogeneous Preferences · NeurIPS 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025 |
Computer vision › 3D vision
direct alignment |
0.9 | 1 | 2025 | Direct Alignment with Heterogeneous Preferences · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model inference
inference-time computation |
0.9 | 1 | 2025 | Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025 |
Machine learning › Reinforcement learning
sample efficiency |
0.9 | 1 | 2025 | Direct Alignment with Heterogeneous Preferences · NeurIPS 2025 |
Natural language and speech › Language models and text generation
test-time scaling |
0.9 | 1 | 2025 | Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025 |
Graph algorithms and graph theory
graph algorithms |
0.3 | 1 | 2025 | Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025 |
Graph algorithms and graph theory
graph connectivity |
0.3 | 1 | 2025 | Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
majority voting · 1.7chain-of-thought · 1.7direct policy alignment · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short OnesabstractInference-time computation has emerged as a promising scaling axis for improving large language model reasoning. However, despite yielding impressive performance, the optimal allocation of inference-time computation remains poorly understood. A central question is whether to prioritize sequential scaling (e.g., longer chains of thought) or parallel scaling (e.g., majority voting across multiple short chains of thought). In this work, we seek to illuminate the landscape of test-time scaling by demonstrating the existence of reasoning settings where sequential scaling offers an exponential advantage over parallel scaling. These settings are based on graph connectivity problems in challenging distributions of graphs. We validate our theoretical findings with comprehensive experiments across a range of language models, including models trained from scratch for graph connectivity with different chain of thought strategies as well as large reasoning models. Parsa Mirtaheri, Ezra Edelman, Samy Jelassi, Eran Malach, Enric Boix-Adserà |
NeurIPS | 1 |
| 2025 | Direct Alignment with Heterogeneous PreferencesabstractAlignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity by introducing user types and examine the limits of the homogeneity assumption.
We show that aligning to heterogeneous preferences with a single policy is best achieved using the average reward across user types. However, this requires additional information about annotators. We examine improvements under different information settings, focusing on direct alignment methods. We find that minimal information can yield first-order improvements, while full feedback from each user type leads to consistent learning of the optimal policy. Surprisingly, however, no sample-efficient consistent direct loss exists in this latter setting. These results reveal a fundamental tension between consistency and sample efficiency in direct policy alignment. Ali Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri, Rediet Abebe, Ariel D. Procaccia |
NeurIPS | 4 |