Simone Parisi

dblp:147/4917 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 10 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 52% Representation and self-supervised learning · 30% Robot manipulation · 13%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
1.322024
Beyond Optimism: Exploration With Partially Observable Rewards · NeurIPS 2024
Interesting Object, Curious Agent: Learning Task-Agnostic Exploration · NeurIPS 2021
Machine learning › Representation and self-supervised learning › pre-training
pre-trained representations
0.612022
The Unsurprising Effectiveness of Pre-Trained Vision Models for Control · ICML 2022
Machine learning › Representation and self-supervised learning › pre-training
visual pre-training
0.612022
The Unsurprising Effectiveness of Pre-Trained Vision Models for Control · ICML 2022
Machine learning › Representation and self-supervised learning
visual representation
0.612022
The Unsurprising Effectiveness of Pre-Trained Vision Models for Control · ICML 2022
Robotics › Robot manipulation
visuomotor control
0.612022
The Unsurprising Effectiveness of Pre-Trained Vision Models for Control · ICML 2022
Machine learning › Reinforcement learning › exploration › exploration in markov decision processes
task-agnostic exploration
0.512021
Interesting Object, Curious Agent: Learning Task-Agnostic Exploration · NeurIPS 2021
Machine learning › Reinforcement learning › policy search
contextual policy search
0.312017
Policy Search with High-Dimensional Context Variables · AAAI 2017
Machine learning › Reinforcement learning
policy search
0.312017
Policy Search with High-Dimensional Context Variables · AAAI 2017
Algorithmic game theory and mechanism design
multi-armed bandit
0.212024
Beyond Optimism: Exploration With Partially Observable Rewards · NeurIPS 2024
Algorithmic game theory and mechanism design › regret minimization
partial monitoring
0.212024
Beyond Optimism: Exploration With Partially Observable Rewards · NeurIPS 2024
Machine learning › Reinforcement learning
multi-objective reinforcement learning
0.212015
Multi-Objective Reinforcement Learning with Continuous Pareto Frontier Approximation · AAAI 2015
Machine learning › Optimization for machine learning › multi-objective optimization
pareto front approximation
0.212015
Multi-Objective Reinforcement Learning with Continuous Pareto Frontier Approximation · AAAI 2015
Machine learning › Deep learning architectures and training › regularization › spectral regularization
nuclear norm regularization
0.112017
Policy Search with High-Dimensional Context Variables · AAAI 2017

Methods — techniques the papers use, named apart from their topics

optimism-based exploration · 1.5monitored markov decision process · 1.5self-supervised learning · 0.6representation learning · 0.6curiosity-driven exploration · 0.5agent-centric and environment-centric exploration · 0.5relative entropy stochastic search · 0.3principal component analysis · 0.3nuclear norm regularization · 0.3gradient ascent · 0.2
YearPublicationVenuePosition
2024 Beyond Optimism: Exploration With Partially Observable Rewards
abstract
Exploration in reinforcement learning (RL) remains an open challenge. RL algorithms rely on observing rewards to train the agent, and if informative rewards are sparse the agent learns slowly or may not learn at all. To improve exploration and reward discovery, popular algorithms rely on optimism. But what if sometimes rewards are unobservable, e.g., situations of partial monitoring in bandits and the recent formalism of monitored Markov decision process? In this case, optimism can lead to suboptimal behavior that does not explore further to collapse uncertainty. With this paper, we present a novel exploration strategy that overcomes the limitations of existing methods and guarantees convergence to an optimal policy even when rewards are not always observable. We further propose a collection of tabular environments for benchmarking exploration in RL (with and without unobservable rewards) and show that our method outperforms existing ones.
Simone Parisi, Alireza Kazemipour, Michael H. Bowling
NeurIPS1
2023 Binary Classification of Agricultural Crops Using Sentinel Satellite Data and Machine Learning Techniques
abstract
The automated process of determining the crop type carried on plots of land, leveraging data provided by earth observation satellites, represents a highly valuable ability that can serve as a foundation for subsequent analyses or as input for calibrating models, such as Decision Support Systems.This paper presents a study on the task of crop classification starting from indices derived from imagery data provided by ESA Satellites Sentinel 1 and 2. We create a valuable tool to verify farmers' claims, especially in relation to state subsidies for specific crops of interest.To this purpose, we focus on perfecting a binary classification for each of five crops of interest (Tomatoes, Soy, Sugar Beet, Rice, and Wheat), aimed to accurately discern the target crop against any other possible crop.The paper investigates various preprocessing techniques to create a dataset suitable for traditional machine learning methods, which presumes that each land plot to classify is represented by a fixed set of features.To deal with inevitable missing observations caused by clouds or other environmental factors, we investigate different imputation strategies (linear interpolation and constant value filling).Complementary, we study the impact of imbalanced classification labels and evaluate the effectiveness of standard balancing techniques.The findings offer practical implications for monitoring and optimizing agricultural practices in the context of precision farming and sustainable agriculture.
Paolo Bertellini, Gianluca D'Addese, Giorgia Franchini, Simone Parisi, Carmelo Scribano, Daniele Zanirato, Marko Bertogna
FedCSIS4
2022 The Unsurprising Effectiveness of Pre-Trained Vision Models for Control
abstract
Recent years have seen the emergence of pre-trained representations as a powerful abstraction for AI applications in computer vision, natural language, and speech. However, policy learning for control is still dominated by a tabula-rasa learning paradigm, with visuo-motor policies often trained from scratch using data from deployment environments. In this context, we revisit and study the role of pre-trained visual representations for control, and in particular representations trained on large-scale computer vision datasets. Through extensive empirical evaluation in diverse control domains (Habitat, DeepMind Control, Adroit, Franka Kitchen), we isolate and study the importance of different representation training methods, data augmentations, and feature hierarchies. Overall, we find that pre-trained visual representations can be competitive or even better than ground-truth state representations to train control policies. This is in spite of using only out-of-domain data from standard vision datasets, without any in-domain data from the deployment environments.
Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, Abhinav Gupta 0001
ICML1
2021 Interesting Object, Curious Agent: Learning Task-Agnostic Exploration
abstract
Common approaches for task-agnostic exploration learn tabula-rasa --the agent assumes isolated environments and no prior knowledge or experience. However, in the real world, agents learn in many environments and always come with prior experiences as they explore new ones. Exploration is a lifelong process. In this paper, we propose a paradigm change in the formulation and evaluation of task-agnostic exploration. In this setup, the agent first learns to explore across many environments without any extrinsic goal in a task-agnostic manner.Later on, the agent effectively transfers the learned exploration policy to better explore new environments when solving tasks. In this context, we evaluate several baseline exploration strategies and present a simple yet effective approach to learning task-agnostic exploration policies. Our key idea is that there are two components of exploration: (1) an agent-centric component encouraging exploration of unseen parts of the environment based on an agent’s belief; (2) an environment-centric component encouraging exploration of inherently interesting objects. We show that our formulation is effective and provides the most consistent exploration across several training-testing environment pairs. We also introduce benchmarks and metrics for evaluating task-agnostic exploration strategies. The source code is available at https://github.com/sparisi/cbet/.
Simone Parisi, Victoria Dean, Deepak Pathak, Abhinav Gupta 0001
NeurIPS1
2019 TD-regularized actor-critic methods
abstract
Actor-critic methods can achieve incredible performance on difficult reinforcement learning problems, but they are also prone to instability. This is partly due to the interaction between the actor and critic during learning, e.g., an inaccurate step taken by one of them might adversely affect the other and destabilize the learning. To avoid such issues, we propose to regularize the learning objective of the actor by penalizing the temporal difference (TD) error of the critic. This improves stability by avoiding large steps in the actor update whenever the critic is highly inaccurate. The resulting method, which we call the TD-regularized actor-critic method, is a simple plug-and-play approach to improve stability and overall performance of the actor-critic methods. Evaluations on standard benchmarks confirm this. Source code can be found at https://github.com/sparisi/td-reg .
Simone Parisi, Voot Tangkaratt, Jan Peters 0001, Mohammad Emtiyaz Khan
Mach. Learn.1
2017 Policy Search with High-Dimensional Context Variables
abstract
Direct contextual policy search methods learn to improve policy parameters and simultaneously generalize these parameters to different context or task variables. However, learning from high-dimensional context variables, such as camera images, is still a prominent problem in many real-world tasks. A naive application of unsupervised dimensionality reduction methods to the context variables, such as principal component analysis, is insufficient as task-relevant input may be ignored. In this paper, we propose a contextual policy search method in the model-based relative entropy stochastic search framework with integrated dimensionality reduction. We learn a model of the reward that is locally quadratic in both the policy parameters and the context variables. Furthermore, we perform supervised linear dimensionality reduction on the context variables by nuclear norm regularization. The experimental results show that the proposed method outperforms naive dimensionality reduction via principal component analysis and a state-of-the-art contextual policy search method.
Voot Tangkaratt, Herke van Hoof, Simone Parisi, Gerhard Neumann, Jan Peters 0001, Masashi Sugiyama
AAAI3
2017 Goal-driven dimensionality reduction for reinforcement learning
abstract
Defining a state representation on which optimal control can perform well is a tedious but crucial process. It typically requires expert knowledge, does not generalize straightforwardly over different tasks and strongly influences the quality of the learned controller. In this paper, we present an autonomous feature construction method for learning low-dimensional manifolds of goal-relevant features jointly with an optimal controller using reinforcement learning. Our method combines information-theoretic algorithms with principal component analysis to performs a return-weighted reduction of the state representation. The method does not require any preprocessing of the data, does not assume strong restrictions on the state representation, and substantially improves the performance of learning by reducing the number of samples required. We show that our method can learn high quality controller in redundant spaces, even from pixels, and outperforms both classical and state-of-the-art deep learning approaches.
Simone Parisi, Simon Ramstedt, Jan Peters 0001
IROS1
2017 Manifold-based multi-objective policy search with sample reuse
Simone Parisi, Matteo Pirotta, Jan Peters 0001
Neurocomputing1
2016 Multi-objective Reinforcement Learning through Continuous Pareto Manifold Approximation
abstract
Many real-world control applications, from economics to robotics, are characterized by the presence of multiple conflicting objectives. In these problems, the standard concept of optimality is replaced by Pareto-optimality and the goal is to find the Pareto frontier, a set of solutions representing different compromises among the objectives. Despite recent advances in multi-objective optimization, achieving an accurate representation of the Pareto frontier is still an important challenge. In this paper, we propose a reinforcement learning policy gradient approach to learn a continuous approximation of the Pareto frontier in multi-objective Markov Decision Problems (MOMDPs). Differently from previous policy gradient algorithms, where n optimization routines are executed to have n solutions, our approach performs a single gradient ascent run, generating at each step an improved continuous approximation of the Pareto frontier. The idea is to optimize the parameters of a function defining a manifold in the policy parameters space, so that the corresponding image in the objectives space gets as close as possible to the true Pareto frontier. Besides deriving how to compute and estimate such gradient, we will also discuss the non-trivial issue of defining a metric to assess the quality of the candidate Pareto frontiers. Finally, the properties of the proposed approach are empirically evaluated on two problems, a linear-quadratic Gaussian regulator and a water reservoir control task.
Simone Parisi, Matteo Pirotta, Marcello Restelli
J. Artif. Intell. Res.1
2015 Multi-Objective Reinforcement Learning with Continuous Pareto Frontier Approximation
abstract
This paper is about learning a continuous approximation of the Pareto frontier in Multi-Objective Markov Decision Problems (MOMDPs).We propose a policy-based approach that exploits gradient information to generate solutions close to the Pareto ones.Differently from previous policy-gradient multi-objective algorithms, where n optimization routines are used to have n solutions, our approach performs a single gradient-ascent run that at each step generates an improved continuous approximation of the Pareto frontier.The idea is to exploit a gradient-based approach to optimize the parameters of a function that defines a manifold in the policy parameter space so that the corresponding image in the objective space gets as close as possible to the Pareto frontier.Besides deriving how to compute and estimate such gradient, we will also discuss the non-trivial issue of defining a metric to assess the quality of the candidate Pareto frontiers.Finally, the properties of the proposed approach are empirically evaluated on two interesting MOMDPs.
Matteo Pirotta, Simone Parisi, Marcello Restelli
AAAI2
2015 Reinforcement learning vs human programming in tetherball robot games
abstract
Reinforcement learning of motor skills is an important challenge in order to endow robots with the ability to learn a wide range of skills and solve complex tasks. However, comparing reinforcement learning against human programming is not straightforward. In this paper, we create a motor learning framework consisting of state-of-the-art components in motor skill learning and compare it to a manually designed program on the task of robot tetherball. We use dynamical motor primitives for representing the robot's trajectories and relative entropy policy search to train the motor framework and improve its behavior by trial and error. These algorithmic components allow for high-quality skill learning while the experimental setup enables an accurate evaluation of our framework as robot players can compete against each other. In the complex game of robot tetherball, we show that our learning approach outperforms and wins a match against a high quality hand-crafted system.
Simone Parisi, Hany Abdulsamad, Alexandros Paraschos, Christian Daniel, Jan Peters 0001
IROS1
2014 Policy gradient approaches for multi-objective sequential decision making: A comparison
abstract
This paper investigates the use of policy gradient techniques to approximate the Pareto frontier in Multi-Objective Markov Decision Processes (MOMDPs). Despite the popularity of policy-gradient algorithms and the fact that gradient-ascent algorithms have been already proposed to numerically solve multi-objective optimization problems, especially in combination with multi-objective evolutionary algorithms, so far little attention has been paid to the use of gradient information to face multi-objective sequential decision problems. Three different Multi-Objective Reinforcement-Learning (MORL) approaches are here presented. The first two, called radial and Pareto following, start from an initial policy and perform gradient-based policy-search procedures aimed at finding a set of non-dominated policies. Differently, the third approach performs a single gradient-ascent run that, at each step, generates an improved continuous approximation of the Pareto frontier. The parameters of a function that defines a manifold in the policy parameter space are updated following the gradient of some performance criterion so that the sequence of candidate solutions gets as close as possible to the Pareto front. Besides reviewing the three different approaches and discussing their main properties, we empirically compare them with other MORL algorithms on two interesting MOMDPs.
Simone Parisi, Matteo Pirotta, Nicola Smacchia, Luca Bascetta, Marcello Restelli
ADPRL1
2014 Policy gradient approaches for multi-objective sequential decision making
abstract
This paper investigates the use of policy gradient techniques to approximate the Pareto frontier in Multi-Objective Markov Decision Processes (MOMDPs). Despite the popularity of policy gradient algorithms and the fact that gradient ascent algorithms have been already proposed to numerically solve multi-objective optimization problems, especially in combination with multi-objective evolutionary algorithms, so far little attention has been paid to the use of gradient information to face multi-objective sequential decision problems. Two different Multi-Objective Reinforcement-Learning (MORL) approaches, called radial and Pareto following, that, starting from an initial policy, perform gradient-based policy-search procedures aimed at finding a set of non-dominated policies are here presented. Both algorithms are empirically evaluated and compared to state-of-the-art MORL algorithms on three MORL benchmark problems.
Simone Parisi, Matteo Pirotta, Nicola Smacchia, Luca Bascetta, Marcello Restelli
IJCNN1