Daniele Meli

dblp:147/0071 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-3162-388XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Symbolic Knowledge Transfer for Sample-Efficient Deep Reinforcement Learning
abstract
Reinforcement Learning (RL) provides a principled framework for sequential decision-making in complex environments. However, state-of-the-art Deep Reinforcement Learning (DRL) algorithms typically require large amounts of training data and often fail to generalize beyond small-scale training scenarios, even on standard benchmarks. We propose a neuro-symbolic DRL approach that incorporates background symbolic knowledge to improve both sample efficiency and generalization to more challenging, unseen tasks. Specifically, partial policies learned in simple domain instances, where high performance can be achieved reliably, are transferred as structured priors to accelerate learning in more complex environments, eliminating the need to tune DRL parameters from scratch. Our method represents partial policies as logical rules in the Answer Set Programming (ASP) formalism and performs online reasoning to guide training through two complementary mechanisms: (i) biasing the action distribution during exploration, and (ii) rescaling Q-values during exploitation. This integration of ASP reasoning with DRL enhances interpretability and trustworthiness while accelerating convergence, particularly in sparse-reward settings and tasks with long planning horizons, without introducing significant computational overhead. We empirically evaluate our approach on challenging variants of gridworld environments under both fully and partially observable settings. Results demonstrate consistent performance improvements over a state-of-the-art reward machine baseline.
Celeste Veronese, Alessandro Farinelli, Daniele Meli
KR3
2025 Learning Logic Specifications for Policy Guidance in POMDPs: an Inductive Logic Programming Approach
abstract
Partially Observable Markov Decision Processes (POMDPs) are a powerful framework for planning under uncertainty. They allow to model state uncertainty as a belief probability distribution. Approximate solvers based on Monte Carlo sampling show great success to relax the computational demand and perform online planning. However, scaling to complex realistic domains with many actions and long planning horizons is still a major challenge, and a key point to achieve good performance is guiding the action-selection process with domain-dependent policy heuristics which are tailored for the specific application domain. We propose to learn high-quality heuristics from POMDP traces of executions generated by any solver. We convert the belief-action pairs to a logical semantics, and exploit data- and time-efficient Inductive Logic Programming (ILP) to generate interpretable belief-based policy specifications, which are then used as online heuristics. We evaluate thoroughly our methodology on two notoriously challenging POMDP problems, involving large action spaces and long planning horizons, namely, rocksample and pocman. Considering different state-of-the-art online POMDP solvers, including POMCP, DESPOT and AdaOPS, we show that learned heuristics expressed in Answer Set Programming (ASP) yield performance superior to neural networks and similar to optimal handcrafted task-specific heuristics within lower computational time. Moreover, they well generalize to more challenging scenarios not experienced in the training phase (e.g., increasing rocks and grid size in rocksample, incrementing the size of the map and the aggressivity of ghosts in pocman).
Daniele Meli, Alberto Castellini, Alessandro Farinelli
AAAI1
2025 Monte Carlo Tree Search with Velocity Obstacles for Safe and Efficient Motion Planning in Dynamic Environments
Lorenzo Bonanni, Daniele Meli, Alberto Castellini, Alessandro Farinelli
AAMAS2
2025 Learning Symbolic Persistent Macro-Actions for POMDP Solving Over Time
abstract
This paper proposes an integration of temporal logical reasoning and Partially Observable Markov Decision Processes (POMDPs) to achieve interpretable decision-making under uncertainty with macro-actions. Our method leverages a fragment of Linear Temporal Logic (LTL) based on Event Calculus (EC) to generate persistent (i.e., constant) macro-actions, which guide Monte Carlo Tree Search (MCTS)-based POMDP solvers over a time horizon, significantly reducing inference time while ensuring robust performance. Such macro-actions are learnt via Inductive Logic Programming (ILP) from a few traces of execution (belief-action pairs), thus eliminating the need for manually designed heuristics and requiring only the specification of the POMDP transition model. In the Pocman and Rocksample benchmark scenarios, our learned macro-actions demonstrate increased expressiveness and generality when compared to time-independent heuristics, indeed offering substantial computational efficiency improvements.
Celeste Veronese, Daniele Meli, Alessandro Farinelli
NeSy2
2025 Inductive learning of robot task knowledge from raw data and online expert feedback
Daniele Meli, Paolo Fiorini
Mach. Learn.1
2024 Learning Logic Specifications for Policy Guidance in POMDPs: an Inductive Logic Programming Approach
abstract
Partially Observable Markov Decision Processes (POMDPs) are a powerful framework for planning under uncertainty. They allow to model state uncertainty as a belief probability distribution. Approximate solvers based on Monte Carlo sampling show great success to relax the computational demand and perform online planning. However, scaling to complex realistic domains with many actions and long planning horizons is still a major challenge, and a key point to achieve good performance is guiding the action-selection process with domain-dependent policy heuristics which are tailored for the specific application domain. We propose to learn high-quality heuristics from POMDP traces of executions generated by any solver. We convert the belief-action pairs to a logical semantics, and exploit data- and time-efficient Inductive Logic Programming (ILP) to generate interpretable belief-based policy specifications, which are then used as online heuristics. We evaluate thoroughly our methodology on two notoriously challenging POMDP problems, involving large action spaces and long planning horizons, namely, rocksample and pocman. Considering different state-of-the-art online POMDP solvers, including POMCP, DESPOT and AdaOPS, we show that learned heuristics expressed in Answer Set Programming (ASP) yield performance superior to neural networks and similar to optimal handcrafted task-specific heuristics within lower computational time. Moreover, they well generalize to more challenging scenarios not experienced in the training phase (e.g., increasing rocks and grid size in rocksample, incrementing the size of the map and the aggressivity of ghosts in pocman).
Daniele Meli, Alberto Castellini, Alessandro Farinelli
J. Artif. Intell. Res.1
2023 Mapping natural language procedures descriptions to linear temporal logic templates: an application in the surgical robotic domain
abstract
Abstract Natural language annotations and manuals can provide useful procedural information and relations for the highly specialized scenario of autonomous robotic task planning. In this paper, we propose and publicly release AUTOMATE, a pipeline for automatic task knowledge extraction from expert-written domain texts. AUTOMATE integrates semantic sentence classification, semantic role labeling, and identification of procedural connectors, in order to extract templates of Linear Temporal Logic (LTL) relations that can be directly implemented in any sufficiently expressive logic programming formalism for autonomous reasoning, assuming some low-level commonsense and domain-independent knowledge is available. This is the first work that bridges natural language descriptions of complex LTL relations and the automation of full robotic tasks. Unlike most recent similar works that assume strict language constraints in substantially simplified domains, we test our pipeline on texts that reflect the expressiveness of natural language used in available textbooks and manuals. In fact, we test AUTOMATE in the surgical robotic scenario, defining realistic language constraints based on a publicly available dataset. In the context of two benchmark training tasks with texts constrained as above, we show that automatically extracted LTL templates, after translation to a suitable logic programming paradigm, achieve comparable planning success in reduced time, with respect to logic programs written by expert programmers.
Marco Bombieri, Daniele Meli, Diego Dall'Alba, Marco Rospocher, Paolo Fiorini
Appl. Intell.2
2022 Deliberation in autonomous robotic surgery: a framework for handling anatomical uncertainty
abstract
Autonomous robotic surgery requires deliberation, i.e. the ability to plan and execute a task adapting to uncer-tain and dynamic environments. Uncertainty in the surgical domain is mainly related to the partial pre-operative knowledge about patient-specific anatomical properties. In this paper, we introduce a logic-based framework for surgical tasks with deliberative functions of monitoring and learning. The DE-liberative Framework for Robot-Assisted Surgery (DEFRAS) estimates a pre-operative patient-specific plan, and executes it while continuously measuring the applied force obtained from a biomechanical pre-operative model. Monitoring module compares this model with the actual situation reconstructed from sensors. In case of significant mismatch, the learning module is invoked to update the model, thus improving the estimate of the exerted force. DEFRAS is validated both in simulated and real environment with da Vinci Research Kit executing soft tissue retraction. Compared with state-of-the-art related works, the success rate of the task is improved while minimizing the interaction with the tissue to prevent unintentional damage.
Eleonora Tagliabue, Daniele Meli, Diego Dall'Alba, Paolo Fiorini
ICRA2
2021 Inductive learning of answer set programs for autonomous surgical task planning
abstract
Abstract The quality of robot-assisted surgery can be improved and the use of hospital resources can be optimized by enhancing autonomy and reliability in the robot’s operation. Logic programming is a good choice for task planning in robot-assisted surgery because it supports reliable reasoning with domain knowledge and increases transparency in the decision making. However, prior knowledge of the task and the domain is typically incomplete, and it often needs to be refined from executions of the surgical task(s) under consideration to avoid sub-optimal performance. In this paper, we investigate the applicability of inductive logic programming for learning previously unknown axioms governing domain dynamics. We do so under answer set semantics for a benchmark surgical training task, the ring transfer. We extend our previous work on learning the immediate preconditions of actions and constraints, to also learn axioms encoding arbitrary temporal delays between atoms that are effects of actions under the event calculus formalism. We propose a systematic approach for learning the specifications of a generic robotic task under the answer set semantics, allowing easy knowledge refinement with iterative learning. In the context of 1000 simulated scenarios, we demonstrate the significant improvement in performance obtained with the learned axioms compared with the hand-written ones; specifically, the learned axioms address some critical issues related to the plan computation time, which is promising for reliable real-time performance during surgery.
Daniele Meli, Mohan Sridharan, Paolo Fiorini
Mach. Learn.1
2020 Autonomous task planning and situation awareness in robotic surgery
abstract
The use of robots in minimally invasive surgery has improved the quality of standard surgical procedures. So far, only the automation of simple surgical actions has been investigated by researchers, while the execution of structured tasks requiring reasoning on the environment and the choice among multiple actions is still managed by human surgeons. In this paper, we propose a framework to implement surgical task automation. The framework consists of a task-level reasoning module based on answer set programming, a low-level motion planning module based on dynamic movement primitives, and a situation awareness module. The logic-based reasoning module generates explainable plans and is able to recover from failure conditions, which are identified and explained by the situation awareness module interfacing to a human supervisor, for enhanced safety. Dynamic Movement Primitives allow to replicate the dexterity of surgeons and to adapt to obstacles and changes in the environment. The framework is validated on different versions of the standard surgical training peg-and-ring task.
Michele Ginesi, Daniele Meli, Andrea Roberti, Nicola Sansonetto, Paolo Fiorini
IROS2
2020 Towards inductive learning of surgical task knowledge: a preliminary case study of the peg transfer task
abstract
Autonomy in robotic surgery will significantly improve the quality of interventions in terms of safety and recovery time for the patient, and reduce fatigue of surgeons and hospital costs. A key requirement for such autonomy is the ability of the surgical system to encode and reason with commonsense task knowledge, and to adapt to variations introduced by the surgical scenarios and the individual patients. However, it is difficult to encode all the variability in surgical scenarios and in the anatomy of individual patients a priori, and new knowledge often needs to be acquired and merged with the existing knowledge. At the same time, it is not possible to provide a large number of labeled training examples in the robotic surgery. This paper presents a framework based on inductive logic programming and answer set semantics for incrementally learning domain knowledge from a limited number of executions of basic surgical tasks. As an illustrative example, we focus on the peg transfer task, and learn state constraints and the preconditions of actions starting from different levels of prior knowledge. We do so using a small dataset comprising human and robotic executions with the da Vinci surgical robot in a challenging simulated scenario.
Daniele Meli, Paolo Fiorini, Mohan Sridharan
KES1
1999 Buffer control technique for transmission frequency recovery of CBR connections over ATM networks
abstract
The transmission of audio-video coded signals in real time applications over ATM networks requires sophisticated techniques for synchronization and buffer control. The presence of ATM cell delay variation (CDV) represents the major jitter source affecting the reconstruction of the time reference signal associated to the actual real-time service. In this paper is presented a buffer control technique and the related implementation aspects, based on both measure and utilization of CDV statistics or alternatively making use of buffer occupation statistics. The system allows setting target jitter attenuation in a way to have pre-established buffer underflow and overflow probabilities and its optimal utilization. The presented technique can be further extended to any asynchronous network context and is particularly suitable for high demanding professional audio-video applications.
Gianmarco Panza, Silvio Cucchi, Daniele Meli
ICASSP3