Parsa Mirtaheri

dblp:400/5482 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 67% 3D vision · 17% Reinforcement learning · 17%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
Direct Alignment with Heterogeneous Preferences · NeurIPS 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Computer vision › 3D vision
direct alignment
0.912025
Direct Alignment with Heterogeneous Preferences · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model inference
inference-time computation
0.912025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Machine learning › Reinforcement learning
sample efficiency
0.912025
Direct Alignment with Heterogeneous Preferences · NeurIPS 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Graph algorithms and graph theory
graph algorithms
0.312025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025
Graph algorithms and graph theory
graph connectivity
0.312025
Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

majority voting · 1.7chain-of-thought · 1.7direct policy alignment · 0.9
YearPublicationVenuePosition
2025 Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short Ones
abstract
Inference-time computation has emerged as a promising scaling axis for improving large language model reasoning. However, despite yielding impressive performance, the optimal allocation of inference-time computation remains poorly understood. A central question is whether to prioritize sequential scaling (e.g., longer chains of thought) or parallel scaling (e.g., majority voting across multiple short chains of thought). In this work, we seek to illuminate the landscape of test-time scaling by demonstrating the existence of reasoning settings where sequential scaling offers an exponential advantage over parallel scaling. These settings are based on graph connectivity problems in challenging distributions of graphs. We validate our theoretical findings with comprehensive experiments across a range of language models, including models trained from scratch for graph connectivity with different chain of thought strategies as well as large reasoning models.
Parsa Mirtaheri, Ezra Edelman, Samy Jelassi, Eran Malach, Enric Boix-Adserà
NeurIPS1
2025 Direct Alignment with Heterogeneous Preferences
abstract
Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity by introducing user types and examine the limits of the homogeneity assumption. We show that aligning to heterogeneous preferences with a single policy is best achieved using the average reward across user types. However, this requires additional information about annotators. We examine improvements under different information settings, focusing on direct alignment methods. We find that minimal information can yield first-order improvements, while full feedback from each user type leads to consistent learning of the optimal policy. Surprisingly, however, no sample-efficient consistent direct loss exists in this latter setting. These results reveal a fundamental tension between consistency and sample efficiency in direct policy alignment.
Ali Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri, Rediet Abebe, Ariel D. Procaccia
NeurIPS4