Amirreza Neshaei Moghaddam

dblp:375/1665 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Motion planning and robot control · 25% Learning theory · 25% Reinforcement learning · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
gradient estimation
0.912025
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens · J. Mach. Learn. Res. 2025
Robotics › Motion planning and robot control › robot control › optimal control
linear quadratic regulator
0.912025
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens · J. Mach. Learn. Res. 2025
Machine learning › Learning theory
sample complexity
0.912025
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens · J. Mach. Learn. Res. 2025

Methods — techniques the papers use, named apart from their topics

policy gradient · 0.9function evaluations · 0.9
YearPublicationVenuePosition
2025 Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
abstract
We provide the first known algorithm that provably achieves $\varepsilon$-optimality within $\widetilde{O}(1/\varepsilon)$ function evaluations for the discounted discrete-time linear quadratic regulator problem with unknown parameters, without relying on two-point gradient estimates. These estimates are known to be unrealistic in many settings, as they depend on using the exact same initialization, which is to be selected randomly, for two different policies. Our results substantially improve upon the existing literature outside the realm of two-point gradient estimates, which either leads to $\widetilde{O}(1/\varepsilon^2)$ rates or heavily relies on stability assumptions.
Amirreza Neshaei Moghaddam, Alexander Olshevsky, Bahman Gharesifard
J. Mach. Learn. Res.1