Alberto Silvio Chiappa

dblp:269/4002 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0001-2764-6552ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 52% Motion planning and robot control · 26% Legged, aerial and field robots · 22%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
0.912025
Auto-Bidding in Real-Time Auctions via Oracle Imitation Learning · KDD (2) 2025
Algorithmic game theory and mechanism design › mechanism design
auction design
0.912025
Auto-Bidding in Real-Time Auctions via Oracle Imitation Learning · KDD (2) 2025
Algorithmic game theory and mechanism design › online advertising
real-time bidding
0.912025
Auto-Bidding in Real-Time Auctions via Oracle Imitation Learning · KDD (2) 2025
Machine learning › Reinforcement learning
exploration
0.712023
Latent exploration for Reinforcement Learning · NeurIPS 2023
Robotics › Motion planning and robot control › robot control › actuator control
motor control
0.712023
Latent exploration for Reinforcement Learning · NeurIPS 2023
Robotics › Motion planning and robot control
musculoskeletal control
0.712023
Latent exploration for Reinforcement Learning · NeurIPS 2023
Robotics › Legged, aerial and field robots › locomotion
adaptive locomotion
0.612022
DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing body · NeurIPS 2022
Machine learning › Reinforcement learning › deep reinforcement learning
attention-based policy
0.612022
DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing body · NeurIPS 2022
Robotics › Legged, aerial and field robots
locomotion
0.612022
DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing body · NeurIPS 2022
Machine learning › Reinforcement learning › policy learning › policy parameterization
policy architecture
0.612022
DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing body · NeurIPS 2022
Computational finance and economics › online advertising
auto-bidding
0.312025
Auto-Bidding in Real-Time Auctions via Oracle Imitation Learning · KDD (2) 2025
Computational finance and economics
online advertising
0.312025
Auto-Bidding in Real-Time Auctions via Oracle Imitation Learning · KDD (2) 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.6imitation learning · 2.6multivariate gaussian noise · 0.7latent time-correlated exploration · 0.7SAC · 0.7PPO · 0.7proprioceptive processing · 0.6distributed policy · 0.6attention mechanism · 0.6
YearPublicationVenuePosition
2025 Adaptive Budget Optimization for Multichannel Advertising Using Combinatorial Bandits
Briti Gangopadhyay, Zhao Wang 0009, Alberto Silvio Chiappa, Shingo Takamatsu
AAMAS3
2025 Auto-Bidding in Real-Time Auctions via Oracle Imitation Learning
Alberto Silvio Chiappa, Briti Gangopadhyay, Zhao Wang 0009, Shingo Takamatsu
KDD (2)1
2023 Latent exploration for Reinforcement Learning
abstract
In Reinforcement Learning, agents learn policies by exploring and interacting with the environment. Due to the curse of dimensionality, learning policies that map high-dimensional sensory input to motor output is particularly challenging. During training, state of the art methods (SAC, PPO, etc.) explore the environment by perturbing the actuation with independent Gaussian noise. While this unstructured exploration has proven successful in numerous tasks, it can be suboptimal for overactuated systems. When multiple actuators, such as motors or muscles, drive behavior, uncorrelated perturbations risk diminishing each other's effect, or modifying the behavior in a task-irrelevant way. While solutions to introduce time correlation across action perturbations exist, introducing correlation across actuators has been largely ignored. Here, we propose LATent TIme-Correlated Exploration (Lattice), a method to inject temporally-correlated noise into the latent state of the policy network, which can be seamlessly integrated with on- and off-policy algorithms. We demonstrate that the noisy actions generated by perturbing the network's activations can be modeled as a multivariate Gaussian distribution with a full covariance matrix. In the PyBullet locomotion tasks, Lattice-SAC achieves state of the art results, and reaches 18\% higher reward than unstructured exploration in the Humanoid environment. In the musculoskeletal control environments of MyoSuite, Lattice-PPO achieves higher reward in most reaching and object manipulation tasks, while also finding more energy-efficient policies with reductions of 20-60\%. Overall, we demonstrate the effectiveness of structured action noise in time and actuator space for complex motor control tasks. The code is available at: https://github.com/amathislab/lattice.
Alberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander Mathis
NeurIPS1
2022 DMAP: a Distributed Morphological Attention Policy for learning to locomote with a changing body
abstract
Biological and artificial agents need to deal with constant changes in the real world. We study this problem in four classical continuous control environments, augmented with morphological perturbations. Learning to locomote when the length and the thickness of different body parts vary is challenging, as the control policy is required to adapt to the morphology to successfully balance and advance the agent. We show that a control policy based on the proprioceptive state performs poorly with highly variable body configurations, while an (oracle) agent with access to a learned encoding of the perturbation performs significantly better. We introduce DMAP, a biologically-inspired, attention-based policy network architecture. DMAP combines independent proprioceptive processing, a distributed policy with individual controllers for each joint, and an attention mechanism, to dynamically gate sensory information from different body parts to different controllers. Despite not having access to the (hidden) morphology information, DMAP can be trained end-to-end in all the considered environments, overall matching or surpassing the performance of an oracle agent. Thus DMAP, implementing principles from biological motor control, provides a strong inductive bias for learning challenging sensorimotor tasks. Overall, our work corroborates the power of these principles in challenging locomotion tasks. The code is available at the following link: https://github.com/amathislab/dmap
Alberto Silvio Chiappa, Alessandro Marin Vargas, Alexander Mathis
NeurIPS1