Alberto Sardinha

dblp:91/515 · also José Alberto R. P. Sardinha, José Alberto Rodrigues Pereira Sardinha · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-5782-3142ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 13 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Centralized Training with Hybrid Execution in Multi-Agent Reinforcement Learning via Predictive Observation Imputation (Abstract Reprint)
abstract
We study hybrid execution in multi-agent reinforcement learning (MARL), a paradigm where agents aim to complete cooperative tasks with arbitrary communication levels at execution time by taking advantage of information-sharing among the agents. Under hybrid execution, the communication level can range from a setting in which no communication is allowed between agents (fully decentralized), to a setting featuring full communication (fully centralized), but the agents do not know beforehand which communication level they will encounter at execution time. We contribute MARO, an approach that makes use of an auto-regressive predictive model, trained in a centralized manner, to estimate missing agents' observations at execution time. We evaluate MARO on standard scenarios and extensions of previous benchmarks tailored to emphasize the impact of partial observability in MARL. Experimental results show that our method consistently outperforms relevant baselines, allowing agents to act with faulty communication while successfully exploiting shared information.
Pedro P. Santos, Diogo S. Carvalho, Miguel Vasco, Alberto Sardinha, Pedro Santos 0001, Ana Paiva 0001, Francisco S. Melo
AAAI4
2025 The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes
abstract
The general-utility Markov decision processes (GUMDPs) framework generalizes the MDPs framework by considering objective functions that depend on the frequency of visitation of state-action pairs induced by a given policy. In this work, we contribute with the first analysis on the impact of the number of trials, i.e., the number of randomly sampled trajectories, in infinite-horizon GUMDPs. We show that, as opposed to standard MDPs, the number of trials plays a key-role in infinite-horizon GUMDPs and the expected performance of a given policy depends, in general, on the number of trials. We consider both discounted and average GUMDPs, where the objective function depends, respectively, on discounted and average frequencies of visitation of state-action pairs. First, we study policy evaluation under discounted GUMDPs, proving lower and upper bounds on the mismatch between the finite and infinite trials formulations for GUMDPs. Second, we address average GUMDPs, studying how different classes of GUMDPs impact the mismatch between the finite and infinite trials formulations. Third, we provide a set of empirical results to support our claims, highlighting how the number of trajectories and the structure of the underlying GUMDP influence policy evaluation.
Pedro P. Santos, Alberto Sardinha, Francisco S. Melo
ICML2
2025 Networked Agents in the Dark: Team Value Learning under Partial Observability
Guilherme S. Varela, Alberto Sardinha, Francisco S. Melo
AAMAS2
2025 Distributed Value Decomposition Networks with Networked Agents
Guilherme S. Varela, Alberto Sardinha, Francisco S. Melo
AAMAS2
2025 Implicit Repair with Reinforcement Learning in Emergent Communication
Fábio Vital, Alberto Sardinha, Francisco S. Melo
AAMAS2
2025 Centralized training with hybrid execution in multi-agent reinforcement learning via predictive observation imputation
abstract
We study hybrid execution in multi-agent reinforcement learning (MARL), a paradigm where agents aim to complete cooperative tasks with arbitrary communication levels at execution time by taking advantage of information-sharing among the agents. Under hybrid execution, the communication level can range from a setting in which no communication is allowed between agents (fully decentralized), to a setting featuring full communication (fully centralized), but the agents do not know beforehand which communication level they will encounter at execution time. We contribute MARO, an approach that makes use of an auto-regressive predictive model, trained in a centralized manner, to estimate missing agents' observations at execution time. We evaluate MARO on standard scenarios and extensions of previous benchmarks tailored to emphasize the impact of partial observability in MARL. Experimental results show that our method consistently outperforms relevant baselines, allowing agents to act with faulty communication while successfully exploiting shared information.
Pedro P. Santos, Diogo S. Carvalho, Miguel Vasco, Alberto Sardinha, Pedro Santos 0001, Ana Paiva 0001, Francisco S. Melo
Artif. Intell.4
2024 TEAMSTER: Model-Based Reinforcement Learning for Ad Hoc Teamwork (Abstract Reprint)
abstract
This paper investigates the use of model-based reinforcement learning in the context of ad hoc teamwork. We introduce a novel approach, named TEAMSTER, where we propose learning both the environment's model and the model of the teammates' behavior separately. Compared to the state-of-the-art PLASTIC algorithms, our results in four different domains from the multi-agent systems literature show that TEAMSTER is more flexible than the PLASTIC-Model, by learning the environment's model instead of assuming a perfect hand-coded model, and more robust/efficient than PLASTIC-Policy, by being able to continuously adapt to newly encountered teams, without implicitly learning a new environment model from scratch.
João G. Ribeiro, Gonçalo Rodrigues, Alberto Sardinha, Francisco S. Melo
AAAI3
2024 The impact of data distribution on Q-learning with function approximation
abstract
Abstract We study the interplay between the data distribution and Q-learning-based algorithms with function approximation. We provide a unified theoretical and empirical analysis as to how different properties of the data distribution influence the performance of Q-learning-based algorithms. We connect different lines of research, as well as validate and extend previous results, being primarily focused on offline settings. First, we analyze the impact of the data distribution by using optimization as a tool to better understand which data distributions yield low concentrability coefficients. We motivate high-entropy distributions from a game-theoretical point of view and propose an algorithm to find the optimal data distribution from the point of view of concentrability. Second, from an empirical perspective, we introduce a novel four-state MDP specifically tailored to highlight the impact of the data distribution in the performance of Q-learning-based algorithms with function approximation. Finally, we experimentally assess the impact of the data distribution properties on the performance of two offline Q-learning-based algorithms under different environments. Our results attest to the importance of different properties of the data distribution such as entropy, coverage, and data quality (closeness to optimal policy).
Pedro P. Santos, Diogo S. Carvalho, Alberto Sardinha, Francisco S. Melo
Mach. Learn.3
2023 Making Friends in the Dark: Ad Hoc Teamwork Under Partial Observability
abstract
This paper introduces a formal definition of the setting of ad hoc teamwork under partial observability and proposes a first-principled model-based approach which relies only on prior knowledge and partial observations of the environment in order to perform ad hoc teamwork. We make three distinct assumptions that set it apart previous works, namely: i) the state of the environment is always partially observable, ii) the actions of the teammates are always unavailable to the ad hoc agent and iii) the ad hoc agent has no access to a reward signal which could be used to learn the task from scratch. Our results in 70 POMDPs from 11 domains show that our approach is not only effective in assisting unknown teammates in solving unknown tasks but is also robust in scaling to more challenging problems. Supplementary material is available at https://github.com/jmribeiro/adhoc-teamwork-under-partial-observability.
João G. Ribeiro, Cassandro Martinho, Alberto Sardinha, Francisco S. Melo
ECAI3
2023 TEAMSTER: Model-based reinforcement learning for ad hoc teamwork
João G. Ribeiro, Gonçalo Rodrigues, Alberto Sardinha, Francisco S. Melo
Artif. Intell.3
2023 Onception: Active Learning with Expert Advice for Real World Machine Translation
abstract
Active learning can play an important role in low-resource settings (i.e., where annotated data is scarce), by selecting which instances may be more worthy to annotate. Most active learning approaches for Machine Translation assume the existence of a pool of sentences in a source language, and rely on human annotators to provide translations or post-edits, which can still be costly. In this article, we apply active learning to a real-world human-in-the-loop scenario in which we assume that: (1) the source sentences may not be readily available, but instead arrive in a stream; (2) the automatic translations receive feedback in the form of a rating, instead of a correct/edited translation, since the human-in-the-loop might be a user looking for a translation, but not be able to provide one. To tackle the challenge of deciding whether each incoming pair source–translations is worthy to query for human feedback, we resort to a number of stream-based active learning query strategies. Moreover, because we do not know in advance which query strategy will be the most adequate for a certain language pair and set of Machine Translation models, we propose to dynamically combine multiple strategies using prediction with expert advice. Our experiments on different language pairs and feedback settings show that using active learning allows us to converge on the best Machine Translation systems with fewer human interactions. Furthermore, combining multiple strategies using prediction with expert advice outperforms several individual active learning strategies with even fewer interactions, particularly in partial feedback settings.
Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha
Comput. Linguistics4
2022 Perceive, Represent, Generate: Translating Multimodal Information to Robotic Motion Trajectories
abstract
We present Perceive-Represent-Generate (PRG), a novel three-stage framework that maps perceptual information of different modalities (e.g., visual or sound), corresponding to a series of instructions, to a sequence of movements to be executed by a robot. In the first stage, we perceive and preprocess the given inputs, isolating individual commands from the complete instruction provided by a human user. In the second stage we encode the individual commands into a multimodal latent space, employing a deep generative model. Finally, in the third stage we convert the latent samples into individual trajectories and combine them into a single dynamic movement primitive, allowing its execution by a robotic manipulator. We evaluate our pipeline in the context of a novel robotic handwriting task, where the robot receives as input a word through different perceptual modalities (e.g., image, sound), and generates the corresponding motion trajectory to write it, creating coherent and high-quality handwritten words.
Fábio Vital, Miguel Vasco, Alberto Sardinha, Francisco S. Melo
IROS3
2021 Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort
abstract
Vânia Mendonça, Ricardo Rei, Luisa Coheur, Alberto Sardinha, Ana Lúcia Santos. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha, Ana Lúcia Santos
ACL/IJCNLP (1)4
2021 MARE: an Active Learning Approach for Requirements Classification
abstract
Several studies indicate that poor requirements practices, that result in incomplete or inaccurate requirements, poorly managed requirement changes, and missed requirements, are the most common factors in project failure. Possible solutions for better requirements definition include better requirements documentation, and requirements reuse. In this paper, we present a novel application of machine learning and active learning to classify the requirements of a given dataset. This approach can accelerate project development. By organizing the requirements into categories, developers can easily see what requirements were already implemented, and where they need to focus on the next step of development.
Cláudia Magalhães, João Araújo 0001, Alberto Sardinha
RE3
2020 Query Strategies, Assemble! Active Learning with Expert Advice for Low-resource Natural Language Processing
abstract
Active learning plays an important role in low-resource scenarios, i.e., when only a small amount of annotated instances is available. However, one does not know what is the best active learning strategy before actually testing a handful of strategies on a labeled set, which might not be viable in a real world low-resource scenario. Instead, it would be desirable to dynamically obtain the results from the best strategy on a given scenario, while using as little annotated resources as possible.In this paper, we present a novel application of prediction with expert advice to combine different query strategies as experts, giving a greater weight to those which select the most useful instances. We evaluated our approach in two Natural Language Processing (NLP) tasks: Part-of-Speech tagging (for English) and Named Entity Recognition (for Portuguese). Results show that our solution keeps up with the results of the best strategy in each scenario, nearly reaching fully supervised performance with only half of the annotated data.
Vânia Mendonça, Alberto Sardinha, Luísa Coheur, Ana Lúcia Santos
FUZZ-IEEE2
2019 Project INSIDE: towards autonomous semi-unstructured human-robot social interaction in autism therapy
Francisco S. Melo, Alberto Sardinha, David Belo, Marta Couto, Miguel Faria 0001, Anabela Farias, Hugo Gamboa, Cátia Jesus, Mithun Kinarullathil, Pedro U. Lima, Luís Luz, André Mateus 0001, Isabel Melo, Plinio Moreno, Daniel Faustino de Noronha Osório, Ana Paiva 0001, Jhielson M. Pimentel, Rodrigo M. M. Ventura
Artif. Intell. Medicine2
2017 NAPP: Connecting Mentors and Students at Técnico Lisboa
Pedro Veiga, Alberto Sardinha, Ana Moura Santos, Carla Boura
EC-TEL2
2016 Ad hoc teamwork by learning teammates' task
Francisco S. Melo, Alberto Sardinha
Auton. Agents Multi Agent Syst.2
2015 VITHEA-Kids: a Platform for Improving Language Skills of Children with Autism Spectrum Disorder
abstract
In this work, we present a platform designed for children with Autism Spectrum Disorder to develop language and generalization skills, in response to the lack of applications tailored for the unique abilities, symptoms, and challenges of the autistic children. This platform allows caregivers to build customized multiple choice exercises while taking into account specific needs/characteristics of each child. We also propose a module for the automatic generation of exercises, aiming to ease the task of exercise creation for caregivers.
Vânia Mendonça, Luísa Coheur, Alberto Sardinha
ASSETS3
2013 EA-Analyzer: automating conflict detection in a large set of textual aspect-oriented requirements
Alberto Sardinha, Ruzanna Chitchyan, Nathan Weston, Phil Greenwood, Awais Rashid
Autom. Softw. Eng.1
2009 EA-Analyzer: Automating Conflict Detection in Aspect-Oriented Requirements
abstract
One of the aims of aspect-oriented requirements engineering is to address the composability and subsequent analysis of crosscutting and non-crosscutting concerns during requirements engineering. Composing concerns may help to reveal conflicting dependencies that need to be identified and resolved. However, detecting conflicts in a large set of textual aspect-oriented requirements is an error-prone and time-consuming task. This paper presents EA-analyzer, the first automated tool for identifying conflicts in aspect-oriented requirements specified in natural-language text. The tool is based on a novel application of a Bayesian learning method that has been effective at classifying text. We present an empirical evaluation of the tool with three industrial-strength requirements documents from different real-life domains. We show that the tool achieves up to 92.97% accuracy when one of the case study documents is used as a training set and the other two as a validation set.
Alberto Sardinha, Ruzanna Chitchyan, Nathan Weston, Phil Greenwood, Awais Rashid
ASE1
2009 A meta-control architecture for orchestrating policy enforcement across heterogeneous information sources
Jinghai Rao, Alberto Sardinha, Norman M. Sadeh
J. Web Semant.2
2006 CMieux: adaptive strategies for competitive supply chain trading
abstract
Supply chains are a central element of today's global economy. Existing management practices consist primarily of static interactions between established partners. Global competition, shorter product life cycles and the emergence of Internet-mediated business solutions create an incentive for exploring more dynamic supply chain practices. The Supply Chain Trading Agent Competition (TAC SCM) was designed to explore approaches to dynamic supply chain trading. TAC SCM pits against one another trading agents developed by teams from around the world. Each agent is responsible for running the procurement, planning and bidding operations of a PC assembly company, while competing with others for both customer orders and supplies under varying market conditions. This paper presents Carnegie Mellon University's 2005 TAC SCM entry, the CMieux supply chain trading agent. CMieux implements a novel approach to coordinating supply chain bidding, procurement and planning, with an emphasis on the ability to rapidly adapt to changing market conditions. We present empirical results based on 200 games involving agents entered by 25 different teams during what can be seen as the most competitive phase of the 2005 tournament. Not only did CMieux perform among the top five agents, it significantly outperformed these agents in procurement while matching their bidding performance.
Michael Benisch, Alberto Sardinha, James Andrews, Norman M. Sadeh
ICEC2
2006 A combined specification language and development framework for agent-based application engineering
Alberto Sardinha, Ricardo Choren, Viviane Torres da Silva, Ruy Milidiú, Carlos José Pereira de Lucena
J. Syst. Softw.1
2003 Software Engineering for Large-Scale Multi-Agent Systems - SELMAS'2003
abstract
Objects and agents are abstractions that exhibit Points of similarity, but the development of multi-agent systems (MASs) poses other challenges to Software Engineering since software agents are inherently more complex entities. In addition, a large MAS needs to satisfy multiple stringent requirements such as reliability, trustability, security, interoperability, scalability, reusability, and maintainability. This workshop brings together researchers and practitioners to discuss the current state of the art and the future research directions in software engineering for large-scale MASs. A particular interest is to understand those issues in the agent technology that make it difficult and/or improve the production of complex distributed systems.
Carlos José Pereira de Lucena, Alberto Sardinha, Alessandro F. Garcia 0001, Alexander B. Romanovsky, Jaelson Brelaz de Castro, Paulo S. C. Alencar, Donald D. Cowan
ICSE2