A. I. Lvovsky 0001

dblp:52/10152 · also Alexander I. Lvovsky · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0003-3165-6654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Transfer learning and domain adaptation · 46% Reinforcement learning · 23% Robot manipulation · 23%
Theoretical computer science
1 paper
Mathematical optimization · 87% Graph algorithms and graph theory · 13%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization
0.412020
Interferobot: aligning an optical interferometer by a reinforcement learning agent · NeurIPS 2020
Machine learning › Reinforcement learning
reinforcement learning for combinatorial optimization
0.412020
Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.412020
Interferobot: aligning an optical interferometer by a reinforcement learning agent · NeurIPS 2020
Mathematical optimization
combinatorial optimization
0.412020
Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020
Mathematical optimization › combinatorial optimization
graph combinatorial optimization
0.412020
Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020
Graph algorithms and graph theory › graph cut
max-cut
0.112020
Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020

Methods — techniques the papers use, named apart from their topics

random search · 0.9exploratory search · 0.9deep q-network · 0.9domain randomization · 0.4deep reinforcement learning · 0.4
YearPublicationVenuePosition
2020 Exploratory Combinatorial Optimization with Reinforcement Learning
abstract
Many real-world problems can be reduced to combinatorial optimization on a graph, where the subset or ordering of vertices that maximize some objective function must be found. With such tasks often NP-hard and analytically intractable, reinforcement learning (RL) has shown promise as a framework with which efficient heuristic methods to tackle these problems can be learned. Previous works construct the solution subset incrementally, adding one element at a time, however, the irreversible nature of this approach prevents the agent from revising its earlier decisions, which may be necessary given the complexity of the optimization task. We instead propose that the agent should seek to continuously improve the solution by learning to explore at test time. Our approach of exploratory combinatorial optimization (ECO-DQN) is, in principle, applicable to any combinatorial problem that can be defined on a graph. Experimentally, we show our method to produce state-of-the-art RL performance on the Maximum Cut problem. Moreover, because ECO-DQN can start from any arbitrary configuration, it can be combined with other search methods to further improve performance, which we demonstrate using a simple random search.
Thomas D. Barrett, William R. Clements, Jakob N. Foerster, A. I. Lvovsky 0001
AAAI4
2020 Interferobot: aligning an optical interferometer by a reinforcement learning agent
abstract
Limitations in acquiring training data restrict potential applications of deep reinforcement learning (RL) methods to the training of real-world robots. Here we train an RL agent to align a Mach-Zehnder interferometer, which is an essential part of many optical experiments, based on images of interference fringes acquired by a monocular camera. The agent is trained in a simulated environment, without any hand-coded features or a priori information about the physics, and subsequently transferred to a physical interferometer. Thanks to a set of domain randomizations simulating uncertainties in physical measurements, the agent successfully aligns this interferometer without any fine-tuning, achieving a performance level of a human expert.
Dmitry Igorevich Sorokin, Alexander E. Ulanov, Ekaterina A. Sazhina, A. I. Lvovsky 0001
NeurIPS4