VLDB 2026 Research / reviewers in the wild / expert
Anna Helena Reali Costa
dblp:r/AnnaHelenaRealiCosta · also Anna H. R. Costa
· DBLP profile ↗
42ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-7309-4528ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When to Prune? The Importance of Timing in Data Efficiency Training
Vinicius Yuiti Fukase, Heitor Gama, Bárbara Fernandes Dias Bueno, Lucas Libanio, Anna Helena Reali Costa, Artur Jordão |
ICPR (7) | 5 |
| 2026 | Layer-Wise LoRA Fine-Tuning: A Similarity Metric Approach
Keith Ogawa, Bruno Yamamoto, Lucas Lauton de Alcantara, Lucas F. A. O. Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordão |
ICPR (8) | 7 |
| 2025 | Efficient LLMs with AMP: Attention Heads and MLP PruningabstractDeep learning drives a new wave in computing systems and triggers the automation of increasingly complex problems. In particular, Large Language Models (LLMs) have significantly advanced cognitive tasks, often matching or even surpassing human-level performance. However, their extensive parameters result in high computational costs and slow inference, posing challenges for deployment in resource-limited settings. Among the strategies to overcome the aforementioned challenges, pruning emerges as a successful mechanism since it reduces model size while maintaining predictive ability. In this paper, we introduce AMP: Attention Heads and MLP Pruning, a novel structured pruning method that efficiently compresses LLMs by removing less critical structures within Multi-Head Attention (MHA) and Multilayer Perceptron (MLP). By projecting the input data onto weights, AMP assesses structural importance and overcomes the limitations of existing techniques, which often fall short in flexibility or efficiency. In particular, AMP surpasses the current state-of-the-art on commonsense reasoning tasks by up to 1.49 percentage points, achieving a 30% pruning ratio with minimal impact on zero-shot task performance. Moreover, AMP also improves inference speeds, making it well-suited for deployment in resource-constrained environments. We confirm the flexibility of AMP on different families of LLMs, including LLaMA and Phi. Leandro Giusti Mugnaini, Bruno Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias, Edson Bollis, Lucas F. A. O. Pellicer, Anna Helena Reali Costa, Artur Jordão |
IJCNN | 7 |
| 2025 | Pruning Everything, Everywhere, All at OnceabstractDeep learning stands as the modern paradigm for solving cognitive tasks. However, as the problem complexity increases, models grow deeper and computationally prohibitive, hindering advancements in real-world and resource-constrained applications. Extensive studies reveal that pruning structures in these models efficiently reduces model complexity and improves computational efficiency. Successful strategies in this sphere include removing neurons (i.e., filters, heads) or layers, but not both together. Therefore, simultaneously pruning different structures remains an open problem. To fill this gap and leverage the benefits of eliminating neurons and layers at once, we propose a new method capable of pruning different structures within a model as follows. Given two candidate subnetworks (pruned models), one from layer pruning and the other from neuron pruning, our method decides which to choose by selecting the one with the highest representation similarity to its parent (the network that generates the subnetworks) using the Centered Kernel Alignment (CKA) metric. Iteratively repeating this process provides highly sparse models that preserve the original predictive ability. Throughout extensive experiments on standard architectures and benchmarks, we confirm the effectiveness of our approach and show that it outperforms state-of-the-art layer and filter pruning techniques. At high levels of Floating Point Operations (FLOPs) reduction, most state-of-the-art methods degrade accuracy, whereas our approach either improves it or experiences only a minimal drop. Notably, on the popular ResNet56 and ResNet110, we achieve a milestone of 86.37% and 95.82% FLOPs reduction. Besides, our pruned models obtain robustness to adversarial and out-of-distribution samples and take an important step towards GreenAI, reducing carbon emissions by up to 83.31%. Overall, we believe our work opens a new chapter in pruning. Code is available at: https://github.com/NascimentoG/PruningEverything. Gustavo H. do Nascimento, Ian Pons, Anna Helena Reali Costa, Artur Jordão |
IJCNN | 3 |
| 2025 | Fuzzy-Based Ensemble Method for Robust Concept Drift Detection in Multivariate Time SeriesabstractConcept drift detection (CDD) is the general problem of identifying significant changes in streaming data distribution over time. Effective drift detection is important in industrial processes such as oil and gas exploration to mitigate financial losses, ensure personnel safety, and reduce environmental risks. However, current CDD methods face challenges in large-scale, multivariate datasets, where single drift detectors (DD) often fail to capture variable interdependencies. While ensemble drift detectors (EDD) are usually adopted to mitigate the adoption of a single DD, EDD may suffer when detections do not converge. This misalignment can cause voting mechanisms to neglect critical intervals with high detection rates. To address this issue, we propose a fuzzy ensemble drift detector (FEDD) that integrates unsupervised threshold voting with fuzzy logic to provide time tolerance and reconcile minor temporal misalignments in drift detection. FEDD is evaluated using the 3W dataset, a realistic public benchmark with rare undesirable real events in oil wells. The results demonstrate that FEDD outperforms existing approaches by improving detection robustness and coverage, ensuring more reliable drift detection in high-dimensional, noisy environments. Lucas Giusti Tavares, Janio Lima, Matheus Melo, Chao Chen 0007, Jonathan M. Garibaldi, Gabriel dos Santos Scatena, Anna Helena Reali Costa, Edson S. Gomi, Rebecca Salles, Esther Pacitti, Ismael H. F. dos Santos, Isabela Guimarães Siqueira, Diego Carvalho 0001, Rafaelli de C. Coutinho, Fábio Porto 0001, Eduardo S. Ogasawara |
IJCNN | 7 |
| 2024 | Early Detection of Extreme Storm Tide Events Using Multimodal Data ProcessingabstractSea-level rise is a well-known consequence of climate change. Several studies have estimated the social and economic impact of the increase in extreme flooding. An efficient way to mitigate its consequences is the development of a flood alert and prediction system, based on high-resolution numerical models and robust sensing networks. However, current models use various simplifying assumptions that compromise accuracy to ensure solvability within a reasonable timeframe, hindering more regular and cost-effective forecasts for various locations along the shoreline. To address these issues, this work proposes a hybrid model for multimodal data processing that combines physics-based numerical simulations, data obtained from a network of sensors, and satellite images to provide refined wave and sea-surface height forecasts, with real results obtained in a critical location within the Port of Santos (the largest port in Latin America). Our approach exhibits faster convergence than data-driven models while achieving more accurate predictions. Moreover, the model handles irregularly sampled time series and missing data without the need for complex preprocessing mechanisms or data imputation while keeping low computational costs through a combination of time encoding, recurrent and graph neural networks. Enabling raw sensor data to be easily combined with existing physics-based models opens up new possibilities for accurate extreme storm tide events forecast systems that enhance community safety and aid policymakers in their decision-making processes. Marcel R. de Barros, Andressa Pinto, Andres Monroy, Felipe M. Moreno, Jefferson F. Coelho, Aldomar Pietro Silva, Caio F. D. Netto, José Roberto Leite, Marlon S. Mathias, Eduardo Aoun Tannuri, Artur Jordão, Edson S. Gomi, Fábio G. Cozman, Marcelo Dottori, Anna Helena Reali Costa |
AAAI | 15 |
| 2024 | Effective Layer Pruning Through Similarity Metric Perspective
Ian Pons, Bruno Yamamoto, Anna Helena Reali Costa, Artur Jordão |
ICPR (5) | 3 |
| 2022 | Outperforming algorithmic trading reinforcement learning systems: A supervised approach to the cryptocurrency marketabstractThe interdisciplinary relationship between machine learning and financial markets has long been a theme of great interest among both research communities. Recently, reinforcement learning and deep learning methods gained prominence in the active asset trading task, aiming to achieve outstanding performances compared with classical benchmarks, such as the Buy and Hold strategy. This paper explores both the supervised learning and reinforcement learning approaches applied to active asset trading, drawing attention to the benefits of both approaches. This work extends the comparison between the supervised approach and reinforcement learning by using state-of-the-art strategies with both techniques. We propose adopting the ResNet architecture, one of the best deep learning approaches for time series classification, into the ResNet-LSTM actor (RSLSTM-A). We compare RSLSTM-A against classical and recent reinforcement learning techniques, such as recurrent reinforcement learning, deep Q-network, and advantage actor–critic. We simulated a currency exchange market environment with the price time series of the Bitcoin, Litecoin, Ethereum, Monero, Nxt, and Dash cryptocurrencies to run our tests. We show that our approach achieves better overall performance, confirming that supervised learning can outperform reinforcement learning for trading. We also present a graphic representation of the features extracted from the ResNet neural network to identify which type of characteristics each residual block generates. Leonardo Kanashiro Felizardo, Francisco Caio Lima Paiva, Catharine de Vita Graves, Elia Yathie Matsumoto, Anna Helena Reali Costa, Emilio Del-Moral-Hernandez, Paolo Brandimarte |
Expert Syst. Appl. | 5 |
| 2021 | Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the OceanabstractCurrent research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pirá dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pirá is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pirá dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pirá, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation. André F. A. Paschoal, Paulo Pirozelli, Valdinei Freire, Karina Valdivia Delgado, Sarajane Marques Peres, Marcos M. José, Flávio Nakasato Cação, André Seidel Oliveira, Anarosa A. F. Brandão, Anna Helena Reali Costa, Fábio G. Cozman |
CIKM | 10 |
| 2020 | Agents teaching agents: a survey on inter-agent transfer learning
Felipe Leno da Silva, Garrett Warnell, Anna Helena Reali Costa, Peter Stone 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2020 | Qualitative case-based reasoning and learningabstractThe development of autonomous agents that perform tasks with the same dexterity as performed by humans is one of the challenges of artificial intelligence and robotics. This motivates the research on intelligent agents, since the agent must choose the best action in a dynamic environment in order to maximise the final score. In this context, the present paper introduces a novel algorithm for Qualitative Case-Based Reasoning and Learning (QCBRL), which is a case-based reasoning system that uses qualitative spatial representations to retrieve and reuse cases by means of relations between objects in the environment. Combined with reinforcement learning, QCBRL allows the agent to learn new qualitative cases at runtime, without assuming a pre-processing step. In order to avoid cases that do not lead to the maximum performance, QCBRL executes case-base maintenance, excluding these cases and obtaining new (more suitable) ones. Experimental evaluation of QCBRL was conducted in a simulated robot-soccer environment, in a real humanoid-robot environment and on simple tasks in two distinct gridworld domains. Results show that QCBRL outperforms traditional RL methods. As a result of running QCBRL in autonomous soccer matches, the robots performed a higher average number of goals than those obtained when using pure numerical models. In the gridworlds considered, the agent was able to learn optimal and safety policies. Thiago Pedro Donadon Homem, Paulo E. Santos, Anna Helena Reali Costa, Reinaldo Augusto da Costa Bianchi, Ramón López de Mántaras |
Artif. Intell. | 3 |
| 2020 | A framework to shift basins of attraction of gene regulatory networks through batch reinforcement learningabstractA major challenge in gene regulatory networks (GRN) of biological systems is to discover when and what interventions should be applied to shift them to healthy phenotypes. A set of gene activity profiles, called basin of attraction (BOA), takes this network to a specific phenotype; therefore, a healthy BOA leads the GRN to a healthy phenotype. However, without the complete observability of the genes, it is not possible to identify whether the current BOA is healthy. In this article we investigate external interventions in GRN with partial observability aiming to bring it to healthy BOAs. We propose a new batch reinforcement learning method (BRL), called mSFQI, to define intervention strategies based on the probabilities of the gene activity profiles being in healthy BOAs, which are calculated from a set of previous observed experiences. BRL uses approximation functions and repeated applications of previous experiences to accelerate learning. Results demonstrate that our proposal can quickly shift a partially observable GRN to healthy BOAs, while reducing the number of interventions. In addition, when observability is poor, mSFQI produces better results when the probabilities for a greater amount of previous observations are available. Cyntia Eico Hayama Nishida, Reinaldo Augusto da Costa Bianchi, Anna Helena Reali Costa |
Artif. Intell. Medicine | 3 |
| 2020 | Combining novelty and popularity on personalised recommendations via user profile learningabstractRecommender systems have been widely used by large companies in the e-commerce segment as aid tools in the search for relevant contents according to the user’s particular preferences. A wide variety of algorithms have been proposed in the literature aiming at improving the process of generating recom- mendations; in particular, a collaborative, diffusion-based hybrid algorithm has been proposed in the lit- erature to solve the problem of sparse data, which affects the quality of recommendations. This algorithm was the basis for several others that effectively solved the sparse data problem. However, this family of algorithms does not differentiate users according to their profiles. In this paper, a new algorithm is pro- posed for learning the user profile and, consequently, generating personalised recommendations through diffusion, combining novelty with the popularity of items. Experiments performed in well-known datasets show that the results of the proposed algorithm outperform those from both diffusion-based hybrid al- gorithm and traditional collaborative filtering algorithm, in the same settings. Ricardo Mitollo Bertani, Reinaldo Augusto da Costa Bianchi, Anna Helena Reali Costa |
Expert Syst. Appl. | 3 |
| 2020 | DECAF: Deep Case-based Policy Inference for knowledge transfer in Reinforcement LearningabstractHaving the ability to solve increasingly complex problems using Reinforcement Learning (RL) has prompted researchers to start developing a greater interest in systematic approaches to retain and reuse knowledge over a variety of tasks. With Case-based Reasoning (CBR) there exists a general methodology that provides a framework for knowledge transfer which has been underrepresented in the RL literature so far. We formulate a terminology for the CBR framework targeted towards RL researchers with the goal of facilitating communication between the respective research communities. Based on this framework, we propose the Deep Case-based Policy Inference (DECAF) algorithm to accelerate learning by building a library of cases and reusing them if they are similar to a new task when training a new policy. DECAF guides the training by dynamically selecting and blending policies according to their usefulness for the current target task, reusing previously learned policies for a more effective exploration but still enabling the adaptation to particularities of the new task. We show an empirical evaluation in the Atari game playing domain depicting the benefits of our algorithm with regards to sample efficiency, robustness against negative transfer, and performance increase when compared to state-of-the-art methods. Ruben Glatt, Felipe Leno da Silva, Reinaldo Augusto da Costa Bianchi, Anna Helena Reali Costa |
Expert Syst. Appl. | 4 |
| 2019 | A Survey on Transfer Learning for Multiagent Reinforcement Learning SystemsabstractMultiagent Reinforcement Learning (RL) solves complex tasks that require coordination with other agents through autonomous exploration of the environment. However, learning a complex task from scratch is impractical due to the huge sample complexity of RL algorithms. For this reason, reusing knowledge that can come from previous experience or other agents is indispensable to scale up multiagent RL algorithms. This survey provides a unifying view of the literature on knowledge reuse in multiagent RL. We define a taxonomy of solutions for the general knowledge reuse problem, providing a comprehensive discussion of recent progress on knowledge reuse in Multiagent Systems (MAS) and of techniques for knowledge reuse across agents (that may be actuating in a shared environment or not). We aim at encouraging the community to work towards reusing all the knowledge sources available in a MAS. For that, we provide an in-depth discussion of current lines of research and open questions. Felipe Leno da Silva, Anna Helena Reali Costa |
J. Artif. Intell. Res. | 2 |
| 2019 | MOO-MDP: An Object-Oriented Representation for Cooperative Multiagent Reinforcement LearningabstractReinforcement learning (RL) is a widely known technique to enable autonomous learning. Even though RL methods achieved successes in increasingly large and complex problems, scaling solutions remains a challenge. One way to simplify (and consequently accelerate) learning is to exploit regularities in a domain, which allows generalization and reduction of the learning space. While object-oriented Markov decision processes (OO-MDPs) provide such generalization opportunities, we argue that the learning process may be further simplified by dividing the workload of tasks amongst multiple agents, solving problems as multiagent systems (MAS). In this paper, we propose a novel combination of OO-MDP and MAS, called multiagent OO-MDP (MOO-MDP). Our proposal accrues the benefits of both OO-MDP and MAS, better addressing scalability issues. We formalize the general model MOO-MDP and present an algorithm to solve deterministic cooperative MOO-MDPs. We show that our algorithm learns optimal policies while reducing the learning space by exploiting state abstractions. We experimentally compare our results with earlier approaches in three domains and evaluate the advantages of our approach in sample efficiency and memory requirements. Felipe Leno da Silva, Ruben Glatt, Anna Helena Reali Costa |
IEEE Trans. Cybern. | 3 |
| 2018 | Autonomously Reusing Knowledge in Multiagent Reinforcement LearningabstractAutonomous agents are increasingly required to solve complex tasks; hard-coding behaviors has become infeasible. Hence, agents must learn how to solve tasks via interactions with the environment. In many cases, knowledge reuse will be a core technology to keep training times reasonable, and for that, agents must be able to autonomously and consistently reuse knowledge from multiple sources, including both their own previous internal knowledge and from other agents. In this paper, we provide a literature review of methods for knowledge reuse in Multiagent Reinforcement Learning. We define an important challenge problem for the AI community, survey the existent methods, and discuss how they can all contribute to this challenging problem. Moreover, we highlight gaps in the current literature, motivating "low-hanging fruit'' for those interested in the area. Our ambition is that this paper will encourage the community to work on this difficult and relevant research challenge. Felipe Leno da Silva, Matthew E. Taylor, Anna Helena Reali Costa |
IJCAI | 3 |
| 2018 | Pairwise registration in indoor environments using adaptive combination of 2D and 3D cues
Juan Carlos Perafan Villota, Felipe Leno da Silva, Ricardo de Souza Jacomini, Anna Helena Reali Costa |
Image Vis. Comput. | 4 |
| 2017 | Learning Options in Multiobjective Reinforcement LearningabstractReinforcement Learning (RL) is a successful technique to train autonomous agents. However, the classical RL methods take a long time to learn how to solve tasks. Option-based solutions can be used to accelerate learning and transfer learned behaviors across tasks by encapsulating a partial policy into an action. However, the literature report only single-agent and single-objective option-based methods, but many RL tasks, especially real-world problems, are better described through multiple objectives. We here propose a method to learn options in Multiobjective Reinforcement Learning domains in order to accelerate learning and reuse knowledge across tasks. Our initial experiments in the Goldmine Domain show that our proposal learn useful options that accelerate learning in multiobjective domains. Our next steps are to use the learned options to transfer knowledge across tasks and evaluate this method with stochastic policies. Rodrigo Cesar Bonini, Felipe Leno da Silva, Anna Helena Reali Costa |
AAAI | 3 |
| 2017 | Policy Reuse in Deep Reinforcement LearningabstractDriven by recent developments in Artificial Intelligence research, a promising new technology for building intelligent agents has evolved. The approach is termed Deep Reinforcement Learning and combines the classic field of Reinforcement Learning (RL) with the representational power of modern Deep Learning approaches. It is very well suited for single task learning but needs a long time to learn any new task. To speed up this process, we propose to extend the concept to multi-task learning by adapting Policy Reuse, a Transfer Learning approach from classic RL, to use with Deep Q-Networks. Ruben Glatt, Anna Helena Reali Costa |
AAAI | 2 |
| 2017 | Improving Deep Reinforcement Learning with Knowledge TransferabstractRecent successes in applying Deep Learning techniques on Reinforcement Learning algorithms have led to a wave of breakthrough developments in agent theory and established the field of Deep Reinforcement Learning (DRL). While DRL has shown great results for single task learning, the multi-task case is still underrepresented in the available literature. This D.Sc. research proposal aims at extending DRL to the multi- task case by leveraging the power of Transfer Learning algorithms to improve the training time and results for multi-task learning. Our focus lies on defining a novel framework for scalable DRL agents that detects similarities between tasks and balances various TL techniques, like parameter initialization, policy or skill transfer. Ruben Glatt, Anna Helena Reali Costa |
AAAI | 2 |
| 2017 | Accelerating Multiagent Reinforcement Learning through Transfer LearningabstractReinforcement Learning (RL) is a widely used solution for sequential decision-making problems and has been used in many complex domains. However, RL algorithms suffer from scalability issues, especially when multiple agents are acting in a shared environment. This research intends to accelerate learning in multiagent sequential decision-making tasks by reusing previous knowledge, both from past solutions and advising between agents. We intend to contribute a Transfer Learning framework focused on Multiagent RL, requiring as few domain-specific hand-coded parameters as possible. Felipe Leno da Silva, Anna Helena Reali Costa |
AAAI | 2 |
| 2017 | An Advising Framework for Multiagent Reinforcement Learning SystemsabstractReinforcement Learning has long been employed to solve sequential decision-making problems with minimal input data. However, the classical approach requires a long time to learn a suitable policy, especially in Multiagent Systems. The teacher-student framework proposes to mitigate this problem by integrating an advising procedure in the learning process, in which an experienced agent (human or not) can advise a student to guide her exploration. However, the teacher is assumed to be an expert in the learning task. We here propose an advising framework where multiple agents advise each other while learning in a shared environment, and the advisor is not expected to necessarily act optimally. Our experiments in a simulated Robot Soccer environment show that the learning process is improved by incorporating this kind of advice. Felipe Leno da Silva, Ruben Glatt, Anna Helena Reali Costa |
AAAI | 3 |
| 2016 | Transfer Learning for Multiagent Reinforcement Learning Systems
Felipe Leno da Silva, Anna Helena Reali Costa |
IJCAI | 2 |
| 2015 | Batch Reinforcement Learning for Smart Home Energy Management
Heider Berlink, Anna Helena Reali Costa |
IJCAI | 2 |
| 2015 | Stochastic Abstract Policies: Generalizing Knowledge to Improve Reinforcement LearningabstractReinforcement learning (RL) enables an agent to learn behavior by acquiring experience through trial-and-error interactions with a dynamic environment. However, knowledge is usually built from scratch and learning to behave may take a long time. Here, we improve the learning performance by leveraging prior knowledge; that is, the learner shows proper behavior from the beginning of a target task, using the knowledge from a set of known, previously solved, source tasks. In this paper, we argue that building stochastic abstract policies that generalize over past experiences is an effective way to provide such improvement and this generalization outperforms the current practice of using a library of policies. We achieve that contributing with a new algorithm, AbsProb-PI-multiple and a framework for transferring knowledge represented as a stochastic abstract policy in new RL tasks. Stochastic abstract policies offer an effective way to encode knowledge because the abstraction they provide not only generalizes solutions but also facilitates extracting the similarities among tasks. We perform experiments in a robotic navigation environment and analyze the agent's behavior throughout the learning process and also assess the transfer ratio for different amounts of source tasks. We compare our method with the transfer of a library of policies, and experiments show that the use of a generalized policy produces better results by more effectively guiding the agent when learning a target task. Marcelo Li Koga, Valdinei Freire, Anna Helena Reali Costa |
IEEE Trans. Cybern. | 3 |
| 2014 | Heuristically-Accelerated Multiagent Reinforcement LearningabstractThis paper presents a novel class of algorithms, called Heuristically-Accelerated Multiagent Reinforcement Learning (HAMRL), which allows the use of heuristics to speed up well-known multiagent reinforcement learning (RL) algorithms such as the Minimax-Q. Such HAMRL algorithms are characterized by a heuristic function, which suggests the selection of particular actions over others. This function represents an initial action selection policy, which can be handcrafted, extracted from previous experience in distinct domains, or learnt from observation. To validate the proposal, a thorough theoretical analysis proving the convergence of four algorithms from the HAMRL class (HAMMQ, HAMQ(λ), HAMQS, and HAMS) is presented. In addition, a comprehensive systematical evaluation was conducted in two distinct adversarial domains. The results show that even the most straightforward heuristics can produce virtually optimal action selection policies in much fewer episodes, significantly improving the performance of the HAMRL over vanilla RL algorithms. Reinaldo Augusto da Costa Bianchi, Murilo Fernandes Martins, Carlos H. C. Ribeiro, Anna Helena Reali Costa |
IEEE Trans. Cybern. | 4 |
| 2013 | Reusing Risk-Aware Stochastic Abstract Policies in Robotic Navigation Learning
Valdinei Freire, Marcelo Li Koga, Fábio G. Cozman, Anna Helena Reali Costa |
RoboCup | 4 |
| 2013 | Corisco: Robust edgel-based orientation estimation for generic camera models
Nicolau Leal Werneck, Anna Helena Reali Costa |
Image Vis. Comput. | 2 |
| 2012 | Finding Memoryless Probabilistic Relational Policies for Inter-task Reuse
Valdinei Freire, Fernando A. Pereira, Anna Helena Reali Costa |
IPMU (2) | 3 |
| 2011 | A Geometric Approach to Find Nondominated Policies to Imprecise Reward MDPs
Valdinei Freire, Anna Helena Reali Costa |
ECML/PKDD (1) | 2 |
| 2008 | Hybrid and Incremental Fuzzy Learning for Human Skin DetectionabstractIn this paper, a framework for detection of human skin in digital images is proposed. This framework is composed of a training phase and a detection phase. A skin class model is learned during the training phase by processing several training images in a hybrid and incremental fuzzy learning scheme. This scheme combines unsupervised- and supervised-learning: unsupervised, by fuzzy clustering, to obtain clusters of color groups from training images; and supervised to select groups that represent skin color. At the end of the training phase, aggregation operators are used to provide combinations of selected groups into a skin model. In the detection phase, the learned skin model is used to detect human skin in an efficient way. Experimental results show robust and accurate human skin detection performed by the proposed framework. Waldemar Bonventi Jr., Anna Helena Reali Costa |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2007 | Heuristic Selection of Actions in Multiagent Reinforcement Learning
Reinaldo Augusto da Costa Bianchi, Carlos H. C. Ribeiro, Anna Helena Reali Costa |
IJCAI | 3 |
| 2007 | Fast loopy belief propagation for topological SamabstractSLAM has been one of the main focuses of attention in robotics research. In the last years, some new graphical solutions for this problem have been proposed, which are concerned about jointly determining the environment map and the robot localization history, a more specific problem known as Smoothing and Mapping (SAM). When applied to topological maps, Loopy Belief Propagation (LBP) provides an incremental and distributed solution to this problem, but eventually may incur time-consuming convergence. This work introduces the concept of starting points of belief propagation, a technique that can be used to reduce the convergence time of the LBP algorithm. We then propose an approach for determining starting points using information about the specific SAM graph structure in order to limit the number of iterations needed by LBP to provide an approximated global Maximum a Posteriori (MAP) estimate of the map and the robot trajectory. The experiments presented, performed with real-world data, confirm the adequacy of the proposed approach and encourage further investigation on it. Antonio Henrique Pinto Selvatici, Anna Helena Reali Costa |
IROS | 2 |
| 2007 | Eliciting preferences over observed behaviours based on relative evaluationsabstractReinforcement learning addresses the question of programming an autonomous agent to execute tasks that are described as reinforcement functions. Then, the agent is responsible for discovering the best actions to fulfil such task. Most of the work on reinforcement learning considers that reinforcements are given by the environment, not addressing the problem of how to describe tasks as reinforcement functions. Preference elicitation addresses the problem of describing a human preference through utility functions, from which reinforcement functions are special cases. This paper proposes an approach where preference elicitation and reinforcement learning are handled in an integrated manner, providing an autonomous method of programming an agent. The agent is programmed through pairwise evaluations over observed behaviours of the agent, where the evaluations are summarised in the reinforcement function. In this paper we present an approach to solve such a problem based on evaluations over observed behaviours. We propose a new algorithm, PEOB-RS, that can be shown to converge towards an optimal policy, providing the number of trials for each behaviour tends to infinity. Experimental results from learning in a grid stochastic environment are used to obtain a reinforcement function, illustrating the effectiveness of PEOB-RS, even if requiring too many evaluations. Such reinforcement function is then transferred to a more real-like environment simulating a pioneer robot, showing the abstraction property of utility functions. Valdinei Freire, Pedro U. Lima, Anna Helena Reali Costa |
IROS | 3 |
| 2007 | Heuristic Reinforcement Learning Applied to RoboCup Simulation Agents
Luiz A. Celiberto, Carlos H. C. Ribeiro, Anna Helena Reali Costa, Reinaldo Augusto da Costa Bianchi |
RoboCup | 3 |
| 2006 | Inverse Reinforcement Learning with EvaluationabstractReinforcement learning (RL) is a method that helps programming an autonomous agent through human-like objectives as reinforcements, where the agent is responsible for discovering the best actions to fulfil the objectives. Nevertheless, it is not easy to disentangle human objectives in reinforcement like objectives. Inverse reinforcement learning (IRL) determines the reinforcements that a given agent behaviour is fulfilling from the observation of the desired behaviour. In this paper we present a variant of IRL, which is called IRL with evaluation (IRLE) where instead of observing the desired agent behaviour, the relative evaluation between different behaviours is known by the access to an evaluator. We present also a solution for this problem under the assumption that a relative linear function that preserves the order assumed by the evaluator exists and that the evaluator evaluates policies instead of behaviours. This is posed as a linear feasibility problem, whose solution is well known. Results of simulations of a set of heterogeneous robots in a search and rescue scenario are presented to illustrate the method and the possibility to transfer the learned reinforcement function among robots Valdinei Freire, Anna Helena Reali Costa, Pedro U. Lima |
ICRA | 2 |
| 2005 | A Hybrid Adaptive Architecture for Mobile Robots Based on Reactive BehaviorsabstractIt is desirable that mobile robots applied to real world applications perform their tasks in previously unknown environments. Thus, a mobile robot architecture capable of adaptation is very suitable. This work presents a hybrid adaptive architecture for mobile robots called AAREACT that has the ability of learning how to coordinate primitive behaviors codified by the potential fields method by using reinforcement learning. The proposed architecture is evaluated in terms of its performance curve when the robot is moved from a scenario to another. Experiments were performed on a Pioneer robot simulator, from ActivMedia Robotics/spl reg/. Results suggest that AAREACT has good adaptation skills for specific environment and task. Antonio Henrique Pinto Selvatici, Anna Helena Reali Costa |
HIS | 2 |
| 2005 | Cooperative multi-robot localization: using communication to reduce localization error
Valguima Odakura, Anna Helena Reali Costa |
ICINCO | 2 |
| 2002 | Optimal control of ship unloaders using reinforcement learning
Leonardo Azevedo Scardua, José Jaime Da Cruz, Anna Helena Reali Costa |
Adv. Eng. Informatics | 3 |
| 2001 | Implementing Computer Vision Algorithms in Hardware: An FPGA/VHDL-Based Vision System for a Mobile Robot
Reinaldo Augusto da Costa Bianchi, Anna Helena Reali Costa |
RoboCup | 2 |
| 1999 | Learning to Behave by Environment Reinforcement
Leonardo Azevedo Scardua, Anna Helena Reali Costa, José Jaime Da Cruz |
RoboCup | 2 |