VLDB 2026 Research / reviewers in the wild / expert
Fernando Fernández 0001
dblp:147/3696 · also Fernando Fernández Rebollo
· DBLP profile ↗
40ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-3801-6801ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing market-making strategies: A multi-objective reinforcement learning approach with pareto frontsabstractFinancial markets are complex ecosystems with multiple participants where market makers (MM) play a crucial role in providing liquidity. They improve market efficiency and stability while reducing price volatility. In this operation, MMs strike a balance between maximizing profitability and minimizing inventory levels. This paper first formulates the market-making problem as a pure multi-objective reinforcement learning task and then introduces M 3 ORL (Market-Maker based on Multi-Objective RL), an approach that handles competing objectives independently rather than combining them into a single reward function. Traditionally, multi-objective problems like this one have been approached with reinforcement learning (RL) as a single-objective framework by combining both goals into a parameterized function. M 3 O R L is a pure multi-objective MM agent based on deep RL to simultaneously optimize both objectives independently. In the proposed solution, the MM agent vectorizes both sub-goals instead of aggregating them, utilizing two different pairs of neural networks: one pair for inventory management and another pair for profitability. The utilization of a multi-objective MM instead of a single objective allows for visualizing the Pareto front and analyzing the trade-offs between objectives. This approach offers advantages over traditional reward engineering-based methods, which are discussed in detail. We compare our approach to classic alternatives based on reward engineering and demonstrate its effectiveness in balancing profitability and inventory control. Additionally, we evaluate our method using well-known multi-objective metrics such as hypervolume, the sparsity of solutions, and the number of undominated solutions, showcasing its superior performance when compared to the rest of reward engineering techniques. Moreover, we demonstrate that this approach is promising in addressing diverse financial challenges beyond market-making. The ability to simultaneously optimize multiple objectives using RL and Pareto fronts opens up new possibilities for tackling complex problems in the financial domain. Oscar Fernández Vicente, Javier García 0001, Fernando Fernández 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Dataset Reduction for Offline Reinforcement Learning using Genetic Algorithms with Image-Based HeuristicsabstractIn offline Reinforcement Learning (RL), the size and quality of the training dataset play a crucial role in determining policy performance. Large datasets can lead to excessive training times, while low-quality data can result in sub-optimal policies, particularly for deep learning-based RL frameworks. To address these challenges, we propose a novel approach that leverages genetic algorithms for efficient dataset reduction, paired with image-based learning using Convolutional Neural Networks (CNNs) to reduce the evaluation time of the fitness function. Specifically, our method predicts the performance of policies (fitness) learned from offline RL datasets (phenotype) and identifies optimized subsets that preserve or enhance policy quality. We evaluate our approach across three well-established RL domains, demonstrating that it effectively reduces dataset size while maintaining or improving policy performance. Furthermore, we show the transferability of the learned models to similar tasks, enabling efficient dataset optimization via transfer learning. Enrique Mateos-Melero, Miguel Iglesias Alcázar, Raquel Fuentetaja 0001, Fernando Fernández 0001 |
GECCO | 4 |
| 2025 | On the Gains from Using Action Observations in Domain RepairabstractDesigning a PDDL planning domain is an error-prone task, which can result in unsolvable planning tasks or unexpected plans. Existing domain repair methods either rely on a complete plan to identify unsatisfied preconditions or operate without any input plan by compiling the flawed planning task into a new planning task with self-repair actions. In contrast, learning approaches often benefit from a range of input observations to infer domain models. In this paper, we extend the self-repair compilation to also accept as input a variable number of action observations. Experimental results show improved domain repair quality and generally strong performance compared to previous domain repair and learning methods. Alba Gragera, Raquel Fuentetaja 0001, Angel García Olaya, Fernando Fernández 0001 |
ICAPS | 4 |
| 2025 | Where is the Nearest EV Charging Station? Evolutionary Optimization of the Gas/charging Stations Topology
Enrique Mateos-Melero, Javier Moralejo-Piñas, Ángela Durán-Pinto, Francisco Martinez-Gil, María Soriano, Fernando Fernández 0001 |
AAMAS | 6 |
| 2025 | Towards a No Code Deployment of Social Robotics Use CasesabstractABSTRACT Social Autonomous Robotics aims to deploy robots in scenarios that involve intensive and continuous interaction with humans. To control the behaviour of robotic platforms in such environments, the use of automated planning (AP) within a control architecture has been proposed as an effective mechanism. However, the design of AP models is time‐consuming and typically carried out by domain experts and engineers. A significant amount of knowledge must be acquired in order to properly define the use case description by specifying the different tasks performed by the robot. In this paper, we present DeVPlan , a framework for graphically designing robotic use cases and configuring the platform for the desired execution. DeVPlan provides an interface that allows domain experts, in collaboration with knowledge engineers, to use state transition diagrams to specify the tasks a robot can perform and define recovery strategies for exogenous events that disrupt normal execution. This graphical design is automatically translated into the standard Planning Domain Definition Language (PDDL). Additionally, to facilitate the integration of the AP model with the robot's control architecture, DeVPlan includes a module for generating the configuration files required to set up the control system. The proposed framework has been successfully used to design and deploy two different use cases in a real environment in a retirement home. Alba Gragera, Carmen Díaz-de-mera, Juan Pedro Bandera Rubio, Angel García Olaya, Fernando Fernández 0001 |
Expert Syst. J. Knowl. Eng. | 5 |
| 2025 | On the Combination of Classical Knowledge Engineering Tools and LLMs to Build Automated Planning ModelsabstractAutomated Planning (AP) is a problem-solving technique applicable to a wide range of scenarios and goals. It typically requires a complete and accurate description of the planning task expressed in a formal language to generate a solution plan that achieves the goals. However, creating these descriptions can be time-consuming and error-prone, often resulting in unsolvable planning tasks. Planning systems often lack the ability to explain why a task is deemed unsolvable. In this work, we present an integrated knowledge engineering system that allows users to graphically depict AP use cases using transition diagrams, which are automatically converted into a formal language. To facilitate the debugging process, we propose connecting the system with large language models (LLMs) to explore their capabilities in assisting with flawed planning tasks, fixing the model and making the tasks solvable. Alba Gragera, Angel García Olaya, Fernando Fernández 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2024 | Clustering-based attack detection for adversarial reinforcement learningabstractAbstract Detecting malicious attacks presents a major challenge in the field of reinforcement learning (RL), as such attacks can force the victim to perform abnormal actions, with potentially severe consequences. To mitigate these risks, current research focuses on the enhancement of RL algorithms with efficient detection mechanisms, especially for real-world applications. Adversarial attacks have the potential to alter the environmental dynamics of a Markov Decision Process (MDP) perceived by an RL agent. Leveraging these changes in dynamics, we propose a novel approach to detect attacks. Our contribution can be summarized in two main aspects. Firstly, we propose a novel formalization of the attack detection problem that entails analyzing modifications made by attacks to the transition and reward dynamics within the environment. This problem can be framed as a context change detection problem, where the goal is to identify the transition from a “free-of-attack” situation to an “under-attack” scenario. To solve this problem, we propose a groundbreaking “model-free” clustering-based countermeasure. This approach consists of two essential steps: first, partitioning the transition space into clusters, and then using this partitioning to identify changes in environmental dynamics caused by adversarial attacks. To assess the efficiency of our detection method, we performed experiments on four established RL domains (grid-world, mountain car, carpole, and acrobot) and subjected them to four advanced attack types. Uniform, Strategically-timed, Q-value, and Multi-objective. Our study proves that our technique has a high potential for perturbation detection, even in scenarios where attackers employ more sophisticated strategies. Rubén Majadas, Javier García 0001, Fernando Fernández 0001 |
Appl. Intell. | 3 |
| 2023 | Automated market maker inventory management with deep reinforcement learningabstractAbstract Stock markets are the result of the interaction of multiple participants, and market makers are one of them. Their main goal is to provide liquidity and market depth to the stock market by streaming bids and offers at both sides of the order book, at different price levels. This activity allows the rest of the participants to have more available prices to buy or sell stocks. In the last years, reinforcement learning market maker agents have been able to be profitable. But profit is not the only measure to evaluate the quality of a market maker. Inventory management arises as a risk source that must be under control. In this paper, we focus on inventory risk management designing an adaptive reward function able to control inventory depending on designer preferences. To achieve this, we introduce two control coefficients, AIIF (Alpha Inventory Impact Factor) and DITF (Dynamic Inventory Threshold Factor), which modulate dynamically the behavior of the market maker agent according to its evolving liquidity with good results. In addition, we analyze the impact of these factors in the trading operative, detailing the underlying strategies performed by these intelligent agents in terms of operative, profitability and inventory management. Last, we present a comparison with other existing reward functions to illustrate the robustness of our approach. Graphic Abstract Oscar Fernández Vicente, Fernando Fernández 0001, Javier García 0001 |
Appl. Intell. | 2 |
| 2022 | A taxonomy for similarity metrics between Markov decision processesabstractAbstract Although the notion of task similarity is potentially interesting in a wide range of areas such as curriculum learning or automated planning, it has mostly been tied to transfer learning. Transfer is based on the idea of reusing the knowledge acquired in the learning of a set of source tasks to a new learning process in a target task, assuming that the target and source tasks areclose enough. In recent years, transfer learning has succeeded in making reinforcement learning (RL) algorithms more efficient (e.g., by reducing the number of samples needed to achieve (near-)optimal performance). Transfer in RL is based on the core concept ofsimilarity: whenever the tasks aresimilar, the transferred knowledge can be reused to solve the target task and significantly improve the learning performance. Therefore, the selection of good metrics to measure these similarities is a critical aspect when building transfer RL algorithms, especially when this knowledge is transferred from simulation to the real world. In the literature, there are many metrics to measure the similarity between MDPs, hence, many definitions ofsimilarityor its complementdistancehave been considered. In this paper, we propose a categorization of these metrics and analyze the definitions ofsimilarityproposed so far, taking into account such categorization. We also follow this taxonomy to survey the existing literature, as well as suggesting future directions for the construction of new metrics. Javier García 0001, Álvaro Visús, Fernando Fernández 0001 |
Mach. Learn. | 3 |
| 2021 | Probabilistic Multi-knowledge Transfer in Reinforcement LearningabstractTransfer in Reinforcement Learning (RL) aims to remedy the problem of learning complex RL tasks from scratch, which is impractical in most of the cases due to the huge sample requirements. To overcome this problem, transferring the knowledge acquired from a set of source tasks to a new target task is a core idea. This knowledge can be the policy, the model (state transition and/or reward function), or the value function learned in the source tasks. However, algorithms in transfer learning focus on transferring a single type of knowledge at a time, although intuitively it might be interesting to reuse several types of this knowledge. For this reason, in this paper we propose a multi-knowledge transfer RL algorithm which we call Probabilistic Transfer of Policies and Models (PTPM). PTPM, unlike single-knowledge transfer approaches, combines the transfer of two types of knowledge: policies and models. We show through different experiments on two well-known domains (Grid World and Mountain Car) how this novel multi-knowledge transfer algorithm improves the results of the two methods in which it is inspired separately. As an additional result, we show that sequential learning of multiple tasks is generally better than learning from a library of previously learned tasks from scratch. Fernando Fernández 0001, Javier García 0001 |
ICMLA | 2 |
| 2021 | Extending the Evaluation of Social Assistive Robots With Accessibility Indicators: The AUSUS Evaluation FrameworkabstractThe introduction of robots in the real world requires previous evaluation of the envisaged performance. Factors like usability, user experience, social acceptance, or societal impact, among others, have been taken into account in evaluation frameworks defined during the last years. However, one of the most important factors that need to be evaluated in any kind of interaction is whether all the users are able to work with the system with the same opportunities and easiness, and none of the current human–robot interaction (HRI) evaluation frameworks include this factor yet. This article proposes an extension of a popular HRI evaluation framework, including accessibility as a new evaluation factor. The proposed approach, named AUSUS, considers Accessibility, Usability, Social acceptance, User experience, and Societal impact. This article presents a use case of the framework, which is evaluated through a socially assistive robotic platform created to perform comprehensive geriatric assessment: the CLARC system. The details of the evaluation process in a hospital and a retirement home are reported, and the main difficulties and recommendations of using AUSUS are discussed. Ana Iglesias 0001, Javier García 0001, Angel García Olaya, Raquel Fuentetaja 0001, Fernando Fernández 0001, Adrián Romero-Garcés, Rebeca Marfil, Antonio Bandera, Karine Lan Hing Ting, Dimitri Voilmy, Alvaro Dueñas-Ruiz, Cristina Suarez-Mejias |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2020 | Learning adversarial attack policies through multi-objective reinforcement learning
Javier García 0001, Rubén Majadas, Fernando Fernández 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2019 | Reinforcement learning for pricing strategy optimization in the insurance industry
Elena Krasheninnikova, Javier García 0001, Roberto Maestre, Fernando Fernández 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2019 | Probabilistic Policy Reuse for Safe Reinforcement LearningabstractThis work introduces Policy Reuse for Safe Reinforcement Learning , an algorithm that combines Probabilistic Policy Reuse and teacher advice for safe exploration in dangerous and continuous state and action reinforcement learning problems in which the dynamic behavior is reasonably smooth and the space is Euclidean. The algorithm uses a continuously increasing monotonic risk function that allows for the identification of the probability to end up in failure from a given state. Such a risk function is defined in terms of how far such a state is from the state space known by the learning agent. Probabilistic Policy Reuse is used to safely balance the exploitation of actual learned knowledge, the exploration of new actions, and the request of teacher advice in parts of the state space considered dangerous. Specifically, the π-reuse exploration strategy is used. Using experiments in the helicopter hover task and a business management problem, we show that the π-reuse exploration strategy can be used to completely avoid the visit to undesirable situations while maintaining the performance (in terms of the classical long-term accumulated reward) of the final policy achieved. Javier García 0001, Fernando Fernández 0001 |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2018 | Towards a robust robotic assistant for Comprehensive Geriatric Assessment procedures: updating the CLARC systemabstractSocially assistive robots appear as a powerful tool in the upcoming silver society. They are among the technologies for Assisted Living, offering a natural interface with smart environments, while helping people through social interaction. The CLARC project aims to develop a socially assistive robot to help clinicians perform Comprehensive Geriatric Assessment (CGA) procedures. This robot autonomously drives some tests and processes, saving time for the clinician to perform more added-value activities, like designing care plans. The project has recently finished its first two phases, and now it faces its final one. This paper details the current prototype of the CLARC system and the main results collected so far during its evaluation. Then, it describes the updates and modifications planned for the next year, in which long term extensive evaluations will be conducted to validate its acceptability and utility. Jesús Martínez, Adrián Romero-Garcés, Cristina Suarez-Mejias, Rebeca Marfil, Karine Lan Hing Ting, Ana Iglesias 0001, Javier García 0001, Fernando Fernández 0001, Alvaro Dueñas-Ruiz, Luis Calderita, Antonio Bandera, Juan Pedro Bandera Rubio |
RO-MAN | 8 |
| 2016 | The IBaCoP Planning System: Instance-Based Configured PortfoliosabstractSequential planning portfolios are very powerful in exploiting the complementary strength of different automated planners. The main challenge of a portfolio planner is to define which base planners to run, to assign the running time for each planner and to decide in what order they should be carried out to optimize a planning metric. Portfolio configurations are usually derived empirically from training benchmarks and remain fixed for an evaluation phase. In this work, we create a per-instance configurable portfolio, which is able to adapt itself to every planning task. The proposed system pre-selects a group of candidate planners using a Pareto-dominance filtering approach and then it decides which planners to include and the time assigned according to predictive models. These models estimate whether a base planner will be able to solve the given problem and, if so, how long it will take. We define different portfolio strategies to combine the knowledge generated by the models. The experimental evaluation shows that the resulting portfolios provide an improvement when compared with non-informed strategies. One of the proposed portfolios was the winner of the Sequential Satisficing Track of the International Planning Competition held in 2014. Isabel Cenamor, Tomás de la Rosa, Fernando Fernández 0001 |
J. Artif. Intell. Res. | 3 |
| 2015 | Strategies for simulating pedestrian navigation with multiple reinforcement learning agents
Francisco Martinez-Gil, Miguel Lozano 0001, Fernando Fernández 0001 |
Auton. Agents Multi Agent Syst. | 3 |
| 2015 | A comprehensive survey on safe reinforcement learning
Javier García 0001, Fernando Fernández 0001 |
J. Mach. Learn. Res. | 2 |
| 2014 | Emergent Collective Behaviors in a Multi-agent Reinforcement Learning Pedestrian Simulation: A Case Study
Francisco Martinez-Gil, Miguel Lozano 0001, Fernando Fernández 0001 |
MABS | 3 |
| 2013 | Integrating Planning, Execution, and Learning to Improve Plan ExecutionabstractAlgorithms for planning under uncertainty require accurate action models that explicitly capture the uncertainty of the environment. Unfortunately, obtaining these models is usually complex. In environments with uncertainty, actions may produce countless outcomes and hence, specifying them and their probability is a hard task. As a consequence, when implementing agents with planning capabilities, practitioners frequently opt for architectures that interleave classical planning and execution monitoring following a replanning when failure paradigm. Though this approach is more practical, it may produce fragile plans that need continuous replanning episodes or even worse, that result in execution dead‐ends. In this paper, we propose a new architecture to relieve these shortcomings. The architecture is based on the integration of a relational learning component and the traditional planning and execution monitoring components. The new component allows the architecture to learn probabilistic rules of the success of actions from the execution of plans and to automatically upgrade the planning model with these rules. The upgraded models can be used by any classical planner that handles metric functions or, alternatively, by any probabilistic planner. This architecture proposal is designed to integrate off‐the‐shelf interchangeable planning and learning components so it can profit from the last advances in both fields without modifying the architecture. Sergio Jiménez Celorrio, Fernando Fernández 0001, Daniel Borrajo |
Comput. Intell. | 2 |
| 2012 | Calibrating a Motion Model Based on Reinforcement Learning for Pedestrian Simulation
Francisco Martinez-Gil, Miguel Lozano 0001, Fernando Fernández 0001 |
MIG | 3 |
| 2012 | A Meta-Tool to Support the Development of Knowledge Engineering Methodologies and ProjectsabstractKnowledge-based systems (KBSs) or expert systems (ESs) are able to solve problems generally through the application of knowledge representing a domain and a set of inference rules. In knowledge engineering (KE), the use of KBSs in the real world, three principal disadvantages have been encountered. First, the knowledge acquisition process has a very high cost in terms of money and time. Second, processing information provided by experts is often difficult and tedious. Third, the establishment of mark times associated with each project phase is difficult due to the complexity described in the previous two points. In response to these obstacles, many methodologies have been developed, most of which include a tool to support the application of the given methodology. Nevertheless, there are advantages and disadvantages inherent in KE methodologies, as well. For instance, particular phases or components of certain methodologies seem to be better equipped than others to respond to a given problem. However, since KE tools currently available support just one methodology the joint use of these phases or components from different methodologies for the solution of a particular problem is hindered. This paper presents KEManager, a generic meta-tool that facilitates the definition and combined application of phases or components from different methodologies. Although other methodologies could be defined and combined in the KEManager, this paper focuses on the combination of two well-known KE methodologies, CommonKADS and IDEAL, together with the most commonly-applied knowledge acquisition methods. The result is an example of the ad hoc creation of a new methodology from pre-existing methodologies, allowing for the adaptation of the KE process to an organization or domain-specific characteristics. The tool was evaluated by students at Carlos III University of Madrid (Spain). José Eloy Flórez, Javier Ignacio Carbó Rubiera, Fernando Fernández 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2012 | Safe Exploration of State and Action Spaces in Reinforcement LearningabstractIn this paper, we consider the important problem of safe exploration in reinforcement learning. While reinforcement learning is well-suited to domains with complex transition dynamics and high-dimensional state-action spaces, an additional challenge is posed by the need for safe and efficient exploration. Traditional exploration techniques are not particularly useful for solving dangerous tasks, where the trial and error process may lead to the selection of actions whose execution in some states may result in damage to the learning system (or any other system). Consequently, when an agent begins an interaction with a dangerous and high-dimensional state-action space, an important question arises; namely, that of how to avoid (or at least minimize) damage caused by the exploration of the state-action space. We introduce the PI-SRL algorithm which safely improves suboptimal albeit robust behaviors for continuous state and action control tasks and which efficiently learns from the experience gained from the environment. We evaluate the proposed method in four complex tasks: automatic car parking, pole-balancing, helicopter hovering, and business management. Javier García 0001, Fernando Fernández 0001 |
J. Artif. Intell. Res. | 2 |
| 2012 | A prototype-based method for classification with time constraints: a case study on automated planning
Rocío García-Durán, Fernando Fernández 0001, Daniel Borrajo |
Pattern Anal. Appl. | 2 |
| 2011 | Safe reinforcement learning in high-risk tasks through policy improvementabstractReinforcement Learning (RL) methods are widely used for dynamic control tasks. In many cases, these are high risk tasks where the trial and error process may select actions which execution from unsafe states can be catastrophic. In addition, many of these tasks have continuous state and action spaces, making the learning problem harder and unapproachable with conventional RL algorithms. So, when the agent begins to interact with a risky and large state-action space environment, an important question arises: how can we avoid that the exploration of the state-action space causes damages in the learning (or other) systems. In this paper, we define the concept of risk and address the problem of safe exploration in the context of RL. Our notion of safety is concerned with states that can lead to damage. Moreover, we introduce an algorithm that safely improves suboptimal but robust behaviors for continuous state and action control tasks, and that learns efficiently from the experience gathered from the environment. We report experimental results using the helicopter hovering task from the RL Competition. Javier García 0001, Fernando Fernández 0001 |
ADPRL | 2 |
| 2011 | A Similarity Function with Local Feature Weighting for Structured Data
Rubén Suárez, Rocío García-Durán, Fernando Fernández 0001 |
ESANN | 3 |
| 2010 | A Reinforcement Learning Approach for Multiagent Navigation
Francisco Martinez-Gil, Fernando Barber, Miguel Lozano 0001, Francisco Grimaldo 0001, Fernando Fernández 0001 |
ICAART (1) | 5 |
| 2010 | SIMBA: A simulator for business education and research
Fernando Borrajo, Yolanda Bueno, Isidro de Pablo, Begoña Santos, Fernando Fernández 0001, Javier García 0001, Ismael Sagredo-Olivenza |
Decis. Support Syst. | 5 |
| 2009 | Learning teaching strategies in an Adaptive and Intelligent Educational System through Reinforcement Learning
Ana Iglesias 0001, Paloma Martínez, Ricardo Aler, Fernando Fernández 0001 |
Appl. Intell. | 4 |
| 2009 | Reinforcement learning of pedagogical policies in adaptive and intelligent educational systems
Ana Iglesias 0001, Paloma Martínez, Ricardo Aler, Fernando Fernández 0001 |
Knowl. Based Syst. | 4 |
| 2008 | The PELA Architecture: Integrating Planning and Learning to Improve Execution
Sergio Jiménez Celorrio, Fernando Fernández 0001, Daniel Borrajo |
AAAI | 2 |
| 2008 | Two steps reinforcement learningabstractWhen applying reinforcement learning in domains with very large or continuous state spaces, the experience obtained by the learning agent in the interaction with the environment must be generalized. The generalization methods are usually based on the approximation of the value functions used to compute the action policy and tackled in two different ways. On the one hand by using an approximation of the value functions based on a supervized learning method. On the other hand, by discretizing the environment to use a tabular representation of the value functions. In this work, we propose an algorithm that uses both approaches to use the benefits of both mechanisms, allowing a higher performance. The approach is based on two learning phases. In the first one, a learner is used as a supervized function approximator, but using a machine learning technique which also outputs a state space discretization of the environment, such as nearest prototype classifiers or decision trees do. In the second learning phase, the space discretization computed in the first phase is used to obtain a tabular representation of the value function computed in the previous phase, allowing a tuning of such value function approximation. Experiments in different domains show that executing both learning phases improves the results obtained executing only the first one. The results take into account the resources used and the performance of the learned behavior. © 2008 Wiley Periodicals, Inc. Fernando Fernández 0001, Daniel Borrajo |
Int. J. Intell. Syst. | 1 |
| 2008 | Local Feature Weighting in Nearest Prototype ClassificationabstractThe distance metric is the corner stone of nearest neighbor (NN)-based methods, and therefore, of nearest prototype (NP) algorithms. That is because they classify depending on the similarity of the data. When the data is characterized by a set of features which may contribute to the classification task in different levels, feature weighting or selection is required, sometimes in a local sense. However, local weighting is typically restricted to NN approaches. In this paper, we introduce local feature weighting (LFW) in NP classification. LFW provides each prototype its own weight vector, opposite to typical global weighting methods found in the NP literature, where all the prototypes share the same one. Providing each prototype its own weight vector has a novel effect in the borders of the Voronoi regions generated: They become nonlinear. We have integrated LFW with a previously developed evolutionary nearest prototype classifier (ENPC). The experiments performed both in artificial and real data sets demonstrate that the resulting algorithm that we call LFW in nearest prototype classification (LFW-NPC) avoids overfitting on training data in domains where the features may have different contribution to the classification task in different areas of the feature space. This generalization capability is also reflected in automatically obtaining an accurate and reduced set of prototypes. Fernando Fernández 0001, Pedro Isasi Viñuela |
IEEE Trans. Neural Networks | 1 |
| 2006 | Combining Macro-operators with Control Knowledge
Rocío García-Durán, Fernando Fernández 0001, Daniel Borrajo |
ILP | 2 |
| 2006 | Roboskeleton: An architecture for coordinating robot soccer agents
David Camacho, Fernando Fernández 0001, Miguel A. Rodelgo |
Eng. Appl. Artif. Intell. | 2 |
| 2005 | Machine Learning of Plan Robustness Knowledge About Instances
Sergio Jiménez Celorrio, Fernando Fernández 0001, Daniel Borrajo |
ECML | 2 |
| 2004 | Learning Content Sequencing in an Educational Environment According to Student Needs
Ana Iglesias 0001, Paloma Martínez, Ricardo Aler, Fernando Fernández 0001 |
ALT | 4 |
| 2002 | On Determinism Handling While Learning Reduced State Space Representations
Fernando Fernández 0001, Daniel Borrajo |
ECAI | 1 |
| 2001 | Designing nearest neighbour classifiers by the evolution of a population of prototypes
Fernando Fernández 0001, Pedro Isasi Viñuela |
ESANN | 1 |
| 1999 | VQQL. Applying Vector Quantization to Reinforcement Learning
Fernando Fernández 0001, Daniel Borrajo |
RoboCup | 1 |