Yonatan Glassner

dblp:144/7804 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Optimization for machine learning · 50% Reinforcement learning · 50%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
conditional value-at-risk
0.212015
Optimizing the CVaR via Sampling · AAAI 2015
Machine learning › Optimization for machine learning › stochastic optimization
risk-sensitive optimization
0.212015
Optimizing the CVaR via Sampling · AAAI 2015
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning
0.212015
Optimizing the CVaR via Sampling · AAAI 2015
Machine learning › Optimization for machine learning
stochastic gradient descent
0.212015
Optimizing the CVaR via Sampling · AAAI 2015
Mathematical optimization › risk measures
conditional value at risk
0.112015
Optimizing the CVaR via Sampling · AAAI 2015
Mathematical optimization
gradient estimation
0.112015
Optimizing the CVaR via Sampling · AAAI 2015

Methods — techniques the papers use, named apart from their topics

sampling-based gradient estimator · 0.4likelihood-ratio method · 0.2likelihood ratio method · 0.2
YearPublicationVenuePosition
2015 Optimizing the CVaR via Sampling
abstract
Conditional Value at Risk (CVaR) is a prominent risk measure that is being used extensively in various domains. We develop a new formula for the gradient of the CVaR in the form of a conditional expectation. Based on this formula, we propose a novel sampling-based estimator for the gradient of the CVaR, in the spirit of the likelihood-ratio method. We analyze the bias of the estimator, and prove the convergence of a corresponding stochastic gradient descent algorithm to a local CVaR optimum. Our method allows to consider CVaR optimization in new domains. As an example, we consider a reinforcement learning application, and learn a risk-sensitive controller for the game of Tetris.
Aviv Tamar, Yonatan Glassner, Shie Mannor
AAAI2