Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Boris Lesner

dblp:54/8009 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 100%
Human-computer interaction and pervasive computing
1 paper
Games and playful interaction · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
dynamic programming
0.422015
Approximate modified policy iteration and its application to the game of Tetris · J. Mach. Learn. Res. 2015
Non-Stationary Approximate Modified Policy Iteration · ICML 2015
Machine learning › Reinforcement learning › non-stationary reinforcement learning › continual reinforcement learning
non-stationary policy
0.422015
Non-Stationary Approximate Modified Policy Iteration · ICML 2015
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes · NIPS 2012
Machine learning › Reinforcement learning › dynamic programming
approximate dynamic programming
0.212015
Approximate modified policy iteration and its application to the game of Tetris · J. Mach. Learn. Res. 2015
Machine learning › Reinforcement learning › markov decision process › infinite-horizon markov decision process
infinite-horizon discounted MDP
0.112012
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes · NIPS 2012
Machine learning › Reinforcement learning
markov decision process
0.112012
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes · NIPS 2012
Machine learning › Reinforcement learning › dynamic programming
policy iteration
0.112012
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes · NIPS 2012
Machine learning › Reinforcement learning › dynamic programming
value iteration
0.112012
On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes · NIPS 2012

Methods — techniques the papers use, named apart from their topics

approximate modified policy iteration · 0.4value iteration · 0.4policy iteration · 0.4approximate dynamic programming · 0.2
YearPublicationVenuePosition
2015 Non-Stationary Approximate Modified Policy Iteration
abstract
We consider the infinite-horizon γ-discounted optimal control problem formalized by Markov Decision Processes. Running any instance of Modified Policy Iteration—a family of algorithms that can interpolate between Value and Policy Iteration—with an error εat each iteration is known to lead to stationary policies that are at least \frac2γε(1-γ)^2-optimal. Variations of Value and Policy Iteration, that build \ell-periodic non-stationary policies, have recently been shown to display a better \frac2γε(1-γ)(1-γ^\ell)-optimality guarantee. Our first contribution is to describe a new algorithmic scheme, Non-Stationary Modified Policy Iteration, a family of algorithms parameterized by two integers m \ge 0 and \ell \ge 1 that generalizes all the above mentionned algorithms. While m allows to interpolate between Value-Iteration-style and Policy-Iteration-style updates, \ell specifies the period of the non-stationary policy that is output. We show that this new family of algorithms also enjoys the improved \frac2γε(1-γ)(1-γ^\ell)-optimality guarantee. Perhaps more importantly, we show, by exhibiting an original problem instance, that this guarantee is tight for all m and \ell; this tightness was to our knowledge only proved two specific cases, Value Iteration (m=0,\ell=1) and Policy Iteration (m=∞,\ell=1).
Boris Lesner, Bruno Scherrer
ICML1
2015 Approximate modified policy iteration and its application to the game of Tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, Matthieu Geist
J. Mach. Learn. Res.4
2012 On the Use of Non-Stationary Policies for Stationary Infinite-Horizon Markov Decision Processes
abstract
We consider infinite-horizon stationary $\gamma$-discounted Markov Decision Processes, for which it is known that there exists a stationary optimal policy. Using Value and Policy Iteration with some error $\epsilon$ at each iteration, it is well-known that one can compute stationary policies that are $\frac{2\gamma{(1-\gamma)^2}\epsilon$-optimal. After arguing that this guarantee is tight, we develop variations of Value and Policy Iteration for computing non-stationary policies that can be up to $\frac{2\gamma}{1-\gamma}\epsilon$-optimal, which constitutes a significant improvement in the usual situation when $\gamma$ is close to $1$. Surprisingly, this shows that the problem of ``computing near-optimal non-stationary policies'' is much simpler than that of ``computing near-optimal stationary policies''.
Bruno Scherrer, Boris Lesner
NIPS2
2010 Language-Independent Clone Detection Applied to Plagiarism Detection
abstract
Clone detection is usually applied in the context of detecting small-to medium scale fragments of duplicated code in large software systems. In this paper, we address the problem of clone detection applied to plagiarism detection in the context of source code assignments done by computer science students. Plagiarism detection comes with a distinct set of constraints to usual clone detection approaches, which influenced the design of the approach we present in this paper. For instance, the source code can be heavily changed at a superficial level (in an attempt to look genuine), yet be functionally very similar. Since assignments turned in by computer science students can be in a variety of languages, we work at the syntactic level and do not consider the source-code semantics. Consequently, the approach we propose is endogenous and makes no assumption about the programming language being analysed. It is based on an alignment method using the parallel principle at local resolution (character level) to compute similarities between documents. We tested our framework on hundreds of real source files, involving a wide array of programming languages (Java, C, Python, PHP, Haskell, bash). Our approach allowed us to discover previously undetected frauds, and to empirically evaluate its accuracy and robustness.
Romain Brixtel, Mathieu Fontaine 0001, Boris Lesner, Cyril Bazin, Romain Robbes
SCAM3