Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Victoria Krakovna

dblp:191/5983 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 94% Multi-agent systems · 6%
Databases, data mining, and information retrieval
1 paper
Data mining · 50% Information retrieval · 50%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reward design
0.412020
Avoiding Side Effects By Considering Future Tasks · NeurIPS 2020
Machine learning › Reinforcement learning › safe reinforcement learning
side effect avoidance
0.412020
Avoiding Side Effects By Considering Future Tasks · NeurIPS 2020
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.312017
Reinforcement Learning with a Corrupted Reward Channel · IJCAI 2017
Knowledge, reasoning and agents › Multi-agent systems
grid environments
0.112020
Avoiding Side Effects By Considering Future Tasks · NeurIPS 2020
Information retrieval
compact coding
0.112010
A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010
Data mining › pattern mining › formal concept analysis
concept lattice
0.112010
A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010
Logic in computer science › knowledge representation and reasoning
formal concept analysis
0.012010
A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010

Methods — techniques the papers use, named apart from their topics

reward shaping · 0.4baseline policy · 0.4reinforcement learning · 0.3randomisation · 0.3zero-preserving encoding · 0.2
YearPublicationVenuePosition
2020 Avoiding Side Effects By Considering Future Tasks
abstract
Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided while completing the task). To alleviate the burden on the reward designer, we propose an algorithm to automatically generate an auxiliary reward function that penalizes side effects. This auxiliary objective rewards the ability to complete possible future tasks, which decreases if the agent causes side effects during the current task. The future task reward can also give the agent an incentive to interfere with events in the environment that make future tasks less achievable, such as irreversible actions by other agents. To avoid this interference incentive, we introduce a baseline policy that represents a default course of action (such as doing nothing), and use it to filter out future tasks that are not achievable by default. We formally define interference incentives and show that the future task approach with a baseline policy avoids these incentives in the deterministic case. Using gridworld environments that test for side effects and interference, we show that our method avoids interference and is more effective for avoiding side effects than the common approach of penalizing irreversible actions.
Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic, Shane Legg
NeurIPS1
2017 Reinforcement Learning with a Corrupted Reward Channel
abstract
No real-world reward function is perfect. Sensory errors and software bugs may result in agents getting higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formalise this problem as a generalised Markov Decision Problem called Corrupt Reward MDP. Traditional RL methods fare poorly in CRMDPs, even under strong simplifying assumptions and when trying to compensate for the possibly corrupt rewards. Two ways around the problem are investigated. First, by giving the agent richer data, such as in inverse reinforcement learning and semi-supervised reinforcement learning, reward corruption stemming from systematic sensory errors may sometimes be completely managed. Second, by using randomisation to blunt the agent's optimisation, reward corruption can be partially managed under some assumptions.
Tom Everitt, Victoria Krakovna, Laurent Orseau, Shane Legg
IJCAI2
2016 Memory-Bounded Left-Corner Unsupervised Grammar Induction on Child-Directed Input
abstract
This paper presents a new memory-bounded left-corner parsing model for unsupervised raw-text syntax induction, using unsupervised hierarchical hidden Markov models (UHHMM). We deploy this algorithm to shed light on the extent to which human language learners can discover hierarchical syntax through distributional statistics alone, by modeling two widely-accepted features of human language acquisition and sentence processing that have not been simultaneously modeled by any existing grammar induction algorithm: (1) a left-corner parsing strategy and (2) limited working memory capacity. To model realistic input to human language learners, we evaluate our system on a corpus of child-directed speech rather than typical newswire corpora. Results beat or closely match those of three competing systems.
Cory Shain, William Bryce, Lifeng Jin, Victoria Krakovna, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz
COLING4
2010 A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices
Matthew Skala, Victoria Krakovna, János Kramár, Gerald Penn
ACL2