Sergiu Goschin

dblp:64/5591 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 50% Deep learning architectures and training · 30% Motion planning and robot control · 20%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
cross-entropy optimization
0.212013
The Cross-Entropy Method Optimizes for Quantiles · ICML (3) 2013
Machine learning › Reinforcement learning
policy search
0.212013
The Cross-Entropy Method Optimizes for Quantiles · ICML (3) 2013
Mathematical optimization › stochastic optimization
cross-entropy method
0.212013
The Cross-Entropy Method Optimizes for Quantiles · ICML (3) 2013
Mathematical optimization
stochastic optimization
0.212013
The Cross-Entropy Method Optimizes for Quantiles · ICML (3) 2013
Machine learning › Reinforcement learning
model-based reinforcement learning
0.112010
Integrating Sample-Based Planning and Model-Based Reinforcement Learning · AAAI 2010
Robotics › Motion planning and robot control › motion planning
sampling-based motion planning
0.112010
Integrating Sample-Based Planning and Model-Based Reinforcement Learning · AAAI 2010

Methods — techniques the papers use, named apart from their topics

quantile optimization · 0.3cross-entropy method · 0.3sample-based planning · 0.1MDP · 0.1
YearPublicationVenuePosition
2013 The Cross-Entropy Method Optimizes for Quantiles
abstract
Cross-entropy optimization (CE) has proven to be a powerful tool for search in control environments. In the basic scheme, a distribution over proposed solutions is repeatedly adapted by evaluating a sample of solutions and refocusing the distribution on a percentage of those with the highest scores. We show that, in the kind of noisy evaluation environments that are common in decision-making domains, this percentage-based refocusing does not optimize the expected utility of solutions, but instead a quantile metric. We provide a variant of CE (Proportional CE) that effectively optimizes the expected value. We show using variants of established noisy environments that Proportional CE can be used in place of CE and can improve solution quality.
Sergiu Goschin, Ari Weinstein, Michael L. Littman
ICML (3)1
2012 Dynamic Teaching in Sequential Decision Making Environments
Thomas J. Walsh 0001, Sergiu Goschin
UAI2
2011 The effects of selection on noisy fitness optimization
abstract
This paper examines how the choice of the selection mechanism in an evolutionary algorithm impacts the objective function it optimizes, specifically when the fitness function is noisy. We provide formal results showing that, in an abstract infinite-population model, proportional selection optimizes expected fitness, truncation selection optimizes order statistics, and tournament selection can oscillate. The "winner" in a population depends on the choice of selection rule, especially when fitness distributions differ between individuals resulting in variable risk. These findings are further developed through empirical results on a novel stochastic optimization problem called "Die4", which, while simple, extends existing benchmark problems by admitting a variety of interpretations of optimality.
Sergiu Goschin, Michael L. Littman, David H. Ackley
GECCO1
2010 Integrating Sample-Based Planning and Model-Based Reinforcement Learning
abstract
Recent advancements in model-based reinforcement learning have shown that the dynamics of many structured domains (e.g. DBNs) can be learned with tractable sample complexity, despite their exponentially large state spaces. Unfortunately, these algorithms all require access to a planner that computes a near optimal policy, and while many traditional MDP algorithms make this guarantee, their computation time grows with the number of states. We show how to replace these over-matched planners with a class of sample-based planners — whose computation time is independent of the number of states — without sacrificing the sample-efficiency guarantees of the overall learning algorithms. To do so, we define sufficient criteria for a sample-based planner to be used in such a learning system and analyze two popular sample-based approaches from the literature. We also introduce our own sample-based planner, which combines the strategies from these algorithms and still meets the criteria for integration into our learning system. In doing so, we define the first complete RL solution for compactly represented (exponentially sized) state spaces with efficiently learnable dynamics that is both sample efficient and whose computation time does not grow rapidly with the number of states.
Thomas J. Walsh 0001, Sergiu Goschin, Michael L. Littman
AAAI2
2007 Combine and compare evolutionary robotics and reinforcement Learning as methods of designing autonomous robots
abstract
The purpose of this paper is to present a comparison between two methods of building adaptive controllers for robots. In spite of the wide range of techniques which are used for defining autonomous robot architectures, few attempts have been made in comparing their performance under similar circumstances. This comparison is particularly important in establishing benchmarks and in determining the best approach methods. The robotic tasks in our research concern mainly the convergence of behaviors like obstacle avoidance, hitting targets and shortest path finding using various methods of synthesizing control architectures' parameters. The first approach that has been used combines Neural Networks and Genetic Algorithms in a simple yet robust controller using an Evolutionary Robotics technique. The second one introduces a manner of using Reinforcement Learning with a Neural Network based architecture. The experiments take place in a simulated 3D environment, which was designed to allow the development, testing and comparison of various controllers in terms of advantages and disadvantages in order to establish a benchmark for autonomous robots.
Sergiu Goschin, Eduard Franti, Monica Dascalu, Sanda Maiduc
IEEE Congress on Evolutionary Computation1