Fernando Martínez-Plumed

dblp:76/8771 · DBLP profile ↗
← Back
37ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0003-2902-6477ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Semi-supervised Soft Clustering with Flexible Cardinality
Diego Vallejo-Huanga, Mateo Montenegro, Brenda Simbaña, Cèsar Ferri, Fernando Martínez-Plumed
ICPR (10)5
2026 Predictable artificial intelligence
abstract
Many areas of artificial intelligence, and machine learning in particular, aim at being probably correct, i.e., valid on average, rather than pursuing the idealistic goal of being provably valid for all inputs. However, AI systems could still be predictably valid, such as an imperfect robot deliverer for which we can reliably and precisely predict the task instances for which it is correct and safe, its valid operating range. “Predictable AI” is a nascent research area that explores ways of anticipating key validity indicators (e.g., performance, safety) of present and future AI ecosystems. We argue that achieving predictability is crucial for fostering trust, liability, control, alignment and safety of AI, and thus should be prioritised over performance. We formally characterise predictability, explore its most relevant components, illustrate what can be predicted, describe alternative candidates for predictors, as well as the trade-offs between maximising validity and predictability. To illustrate these concepts, we bring an array of illustrative examples covering diverse ecosystem configurations. “Predictable AI” is related to other areas of technical and non-technical AI research, but have distinctive questions, hypotheses, techniques and challenges. This paper aims to elucidate them, calls for identifying paths towards a landscape of predictably valid AI systems and outlines the potential impact of this emergent field.
Lexin Zhou, P. A. M. Casares, Fernando Martínez-Plumed, John Burden, Ryan Burnell, Lucy Cheke, Cèsar Ferri, Alexandru Marcoci, Behzad Mehrbakhsh, Yael Moros-Daval, Seán Ó hÉigeartaigh, Danaja Rutar, Wout Schellaert, Konstantinos Voudouris, José Hernández-Orallo
Artif. Intell.3
2026 What Should an AI Assessor Optimise for?
abstract
Abstract An AI assessor is an external, ideally independent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can leverage information from the test results of many other AI systems and have the flexibility of being trained on any loss function or scoring rule: from squared error to toxicity metrics. Here we address the question: is it always optimal to train the assessor for the target metric? Or could it be better to train for a different metric and then map predictions back to the target metric? Using twenty regression and classification problems with tabular data, we experimentally explore this question for, respectively, regression losses and classification scores with monotonic and nonmonotonic mappings and find that, contrary to intuition, optimising for more informative metrics (i.e., yielding a better-conditioned supervision signal) is not universally preferred. Surprisingly, some monotonic transformations are promising. For example, logistic loss is useful for minimising absolute or quadratic errors in regression, and logarithmic score helps maximise quadratic or spherical scores in classification.
Daniel Romero-Alvarado, Fernando Martínez-Plumed, José Hernández-Orallo
Mach. Learn.2
2026 Human or Machine: A Novel Deep Learning Framework for Autonomous Driver Identification Based on Vehicle Trajectories
abstract
Monitoring traffic streams through vehicle trajectories offers valuable insights into traffic flow characteristics. In recent years, there has been a surge in the availability of vehicle trajectory datasets. At the same time, the number of autonomously-driven vehicles on the road is increasing, largely due to the adoption of systems like adaptive cruise control. However, distinguishing system-controlled vehicles from human-driven vehicles remains challenging, despite its potential to enable valuable applications and informed policy-making. The differences between the driving behavior of human-driven (HDs) and automated (ADs) vehicles in the longitudinal direction are highlighted in the literature and hold promise for novel methodologies that exploit them to identify the type of driver. Here, we propose a novel online-offline framework with three key contributions. First, a feature design component performs feature disentanglement to increase the performance of downstream deep learning models. Second, a bidirectional LSTM (bLSTM) architecture demonstrates excellent accuracy in differentiating between HD and AD vehicles. Third, a data drift detection component identifies changes in data distributions, enabling the framework to generalize effectively to unseen datasets with minimal new labeled observations.
Andres L. Marin, Fernando Martínez-Plumed, María José Ramírez-Quintana, Konstantinos Mattas, Georgios Fontaras, Anastasios Kouvelas, Michael Makridis
IEEE Trans. Intell. Transp. Syst.2
2025 ClustSize: An Algorithmic Framework for Size-Constrained Clustering
Diego Vallejo-Huanga, Cèsar Ferri, Fernando Martínez-Plumed
DATA3
2025 Contamination Budget: Trade-offs Between Breadth, Depth and Difficulty
abstract
Contamination in large language models (LLMs), and machine learning more broadly, refers to the inclusion of equal --or very similar-- examples in both training and test sets. This phenomenon usually translates into better test performance. Here we explore when this contamination is performed intentionally, for purposes that can be malicious (e.g., get better scores in evaluations) or benevolent (e.g., fix some mistakes). These interventions, usually in the form of fine-tuning memorisations, come with a budget in the size of the fine-tuning dataset. Several trade-offs appear between the breadth of the intervention (how many examples to be memorised), its depth (how many repetitions of each example) and the difficulty of the examples. By studying several LLMs and datasets, we observe some monotonic behaviour (more difficult items require more depth to be `fixed') but also some non-monotonic phenomena (very high depth levels have negative effects on non-contaminated examples). This suggests that trade-offs should be found not only in terms of the budget but also according to model specifics, the task and the item difficulty at hand.
Behzad Mehrbakhsh, Fernando Martínez-Plumed, José Hernández-Orallo
IJCAI2
2025 Refining Community Detection in Social Networks: Agglomerative and Divisive Methods with Size Constraints
Diego Vallejo-Huanga, Erlend Eindride Fasmer, Cèsar Ferri, Fernando Martínez-Plumed
MDAI4
2025 Cracking black-box models: Revealing hidden machine learning techniques behind their predictions
abstract
The quest for transparency in black-box models has gained significant momentum in recent years. In particular, discovering the underlying machine learning technique type (or model family) from the performance of a black-box model is a real important problem both for better understanding its behaviour and for developing strategies to attack it by exploiting the weaknesses intrinsic to the learning technique. In this paper, we tackle the challenging task of identifying which kind of machine learning model is behind the predictions when we interact with a black-box model. Our innovative method involves systematically querying a black-box model (oracle) to label an artificially generated dataset, which is then used to train different surrogate models using machine learning techniques from different families (each one trying to partially approximate the oracle’s behaviour). We present two approaches based on similarity measures, one selecting the most similar family and the other using a conveniently constructed meta-model. In both cases, we use both crisp and soft classifiers and their corresponding similarity metrics. By experimentally comparing all these methods, we gain valuable insights into the explanatory and predictive capabilities of our model family concept. This provides a deeper understanding of the black-box models and increases their transparency and interpretability, paving the way for more effective decision making.
Raül Fabra-Boluda, Cèsar Ferri, José Hernández-Orallo, M. José Ramrez-Quintana, Fernando Martínez-Plumed
Intell. Data Anal.5
2025 Analysing the Predictability of Language Model Performance
abstract
Can a language model predict for which questions another language model will answer successfully? We investigate the extent to which performance prediction is possible and dissect various factors that influence it. Our experimental setting fine-tunes DeBERTa models, which we call assessors , on the evaluation results of generative language models with up to 128 billion parameters, which we refer to as subject systems . Our analysis spans more than 100 tasks from BIG-bench. We find that the assessors can match and even exceed the subjects’ confidence in both refinement and calibration, anticipating failures at near perfect levels for some tasks. We also find that for performance prediction it can be beneficial to learn from the scores on multiple tasks or to learn from the scores of multiple subjects, but both depend on the task at hand. Lastly, we find that large and small subject systems are equally predictable, showing promise for the scalability of the predictability problem.
Wout Schellaert, Fernando Martínez-Plumed, José Hernández-Orallo
ACM Trans. Intell. Syst. Technol.2
2024 Your Prompt Is My Command: On Assessing the Human-Centred Generality of Multimodal Models (Abstract Reprint)
abstract
Even with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully autonomous one. HCG implies that —for a specific user— a system is only as general as it is effective for the user’s relevant tasks and their prevalent ways of prompting. A human-centred evaluation of general-purpose AI systems therefore needs to reflect the personal nature of interaction, tasks and cognition. We argue that the best way to understand these systems is as highly-coupled cognitive extenders, and to analyse the bidirectional cognitive adaptations between them and humans. In this paper, we give a formulation of HCG, as well as a high-level overview of the elements and trade-offs involved in the prompting process. We end the paper by outlining some essential research questions and suggestions for improving evaluation practices, which we envision as characteristic for the evaluation of general artificial intelligence in the future.
Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, John Burden, P. A. M. Casares, Bao Sheng Loe, Roi Reichart, Seán Ó hÉigeartaigh, Anna Korhonen, José Hernández-Orallo
AAAI2
2024 Distilling the Effects of Language Model Contamination
abstract
The proportion of AI-generated content permeating the well of knowledge is increasing significantly. Large language models (LLMs) contribute to that contamination but they also suffer from it. However, it is yet to be clarified the effect of different sources of error, be it human-generated or LLM-generated. Controlling for the percentage of error, we explore the impact on LLM fine-tuning when errors come from humans, from other language models or are generated randomly using an aleatoric or epistemic source. In this paper, we compare these different types of error for in-distribution and out-of-distribution experimental settings. By analysing the levels of errors and their distribution, we find a nuanced view: while in-distribution human-generated noise seems more benign than the LLM-generated counterpart, in the out-of-distribution case the model-generated noise may not be necessarily worse.
Behzad Mehrbakhsh, Fernando Martínez-Plumed, José Hernández-Orallo
ECAI2
2024 Language Task Difficulty Prediction Through LLM-Annotated Meta-Features
abstract
Assessing the capabilities of large language models (LLMs) is increasingly challenging due to their generality and uneven task performance. Often, we do not know how much of the success or failure on a particular task is due to the ‘loading’ of the language elements in the task, such as narrative understanding, or some other intrinsic (non-linguistic) components, such as domain-specific common sense or reasoning capabilities. Understanding what tasks are most loaded on language and determine the predictability of LLMs on these tasks is crucial for improving benchmarks, designing better LLMs, and ensuring their safe deployment. We present an innovative methodology that uses LLMs to annotate linguistic meta-features, allowing us to predict task difficulty and understand linguistic loadings more accurately than traditional readability scores. Using GPT-4 for automated annotation, we show strong predictability for a variety of tasks and language models (e.g., MMLU with R2 from 0.68 to 0.83), but observe limited predictability for other tasks (e.g., LSAT with R2 of -0.07).
Yael Moros-Daval, Fernando Martínez-Plumed, José Hernández-Orallo
ECAI2
2024 Automatic PDF Document Classification with Machine Learning
Sócrates Llácer Luna, Darío Garigliotti, Fernando Martínez-Plumed, Cèsar Ferri
IDEAL (1)3
2024 Noise Tolerance and Robustness Ranking in Machine Learning Models
Cristina Padró-Ferragut, María José Ramírez-Quintana, Fernando Martínez-Plumed
IDEAL (2)3
2024 How Resilient are Language Models to Text Perturbations?
Daniel Romero-Alvarado, José Hernández-Orallo, Fernando Martínez-Plumed
IDEAL (1)3
2023 Adversarial Benchmark Evaluation Rectified by Controlling for Difficulty
abstract
Adversarial benchmark construction, where harder instances challenge new generations of AI systems, is becoming the norm. While this approach may lead to better machine learning models —on average and for the new benchmark—, it is unclear how these models behave on the original distribution. Two opposing effects are intertwined here. On the one hand, the adversarial benchmark has a higher proportion of difficult instances, with lower expected performance. On the other hand, models trained on the adversarial benchmark may improve on these difficult instances (but may also neglect some easy ones). To disentangle these two effects we can control for difficulty, showing that we can recover the performance on the original distribution, provided the harder instances were obtained from this distribution in the first place. We show this difficulty-aware rectification works in practice, through a series of experiments with several benchmark construction schemas and the use of a populational difficulty metric. As a take-away message, instead of distributional averages we recommend using difficulty-conditioned characteristic curves when evaluating models built with adversarial benchmarks.
Behzad Mehrbakhsh, Fernando Martínez-Plumed, José Hernández-Orallo
ECAI2
2023 Your Prompt is My Command: On Assessing the Human-Centred Generality of Multimodal Models
abstract
Even with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully autonomous one. HCG implies that —for a specific user— a system is only as general as it is effective for the user’s relevant tasks and their prevalent ways of prompting. A human-centred evaluation of general-purpose AI systems therefore needs to reflect the personal nature of interaction, tasks and cognition. We argue that the best way to understand these systems is as highly-coupled cognitive extenders, and to analyse the bidirectional cognitive adaptations between them and humans. In this paper, we give a formulation of HCG, as well as a high-level overview of the elements and trade-offs involved in the prompting process. We end the paper by outlining some essential research questions and suggestions for improving evaluation practices, which we envision as characteristic for the evaluation of general artificial intelligence in the future. This paper appears in the AI & Society track.
Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, John Burden, P. A. M. Casares, Bao Sheng Loe, Roi Reichart, Seán Ó hÉigeartaigh, Anna Korhonen, José Hernández-Orallo
J. Artif. Intell. Res.2
2023 Can language models automate data wrangling?
abstract
Abstract The automation of data science and other data manipulation processes depend on the integration and formatting of ‘messy’ data. Data wrangling is an umbrella term for these tedious and time-consuming tasks. Tasks such as transforming dates, units or names expressed in different formats have been challenging for machine learning because (1) users expect to solve them with short cues or few examples, and (2) the problems depend heavily on domain knowledge. Interestingly, large language models today (1) can infer from very few examples or even a short clue in natural language, and (2) can integrate vast amounts of domain knowledge. It is then an important research question to analyse whether language models are a promising approach for data wrangling, especially as their capabilities continue growing. In this paper we apply different variants of the language model Generative Pre-trained Transformer (GPT) to five batteries covering a wide range of data wrangling problems. We compare the effect of prompts and few-shot regimes on their results and how they compare with specialised data wrangling systems and other tools. Our major finding is that they appear as a powerful tool for a wide range of data wrangling tasks. We provide some guidelines about how they can be integrated into data processing pipelines, provided the users can take advantage of their flexibility and the diversity of tasks to be addressed. However, reliability is still an important issue to overcome.
Gonzalo Jaimovitch-López, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, María José Ramírez-Quintana
Mach. Learn.4
2022 Training on the Test Set: Mapping the System-Problem Space in AI
abstract
Many present and future problems associated with artificial intelligence are not due to its limitations, but to our poor assessment of its behaviour. Our evaluation procedures produce aggregated performance metrics that lack detail and quantified uncertainty about the following question: how will an AI system, with a particular profile \pi, behave for a new problem, characterised by a particular situation \mu? Instead of just aggregating test results, we can use machine learning methods to fully capitalise on this evaluation information. In this paper, we introduce the concept of an assessor model, \hat{R}(r|\pi,\mu), a conditional probability estimator trained on test data. We discuss how these assessors can be built by using information of the full system-problem space and illustrate a broad range of applications that derive from varied inferences and aggregations from \hat{R}. Building good assessor models will change the predictive and explanatory power of AI evaluation and will lead to new research directions for building and using them. We propose accompanying every deployed AI system with its own assessor.
José Hernández-Orallo, Wout Schellaert, Fernando Martínez-Plumed
AAAI3
2022 When AI Difficulty Is Easy: The Explanatory Power of Predicting IRT Difficulty
abstract
One of challenges of artificial intelligence as a whole is robustness. Many issues such as adversarial examples, out of distribution performance, Clever Hans phenomena, and the wider areas of AI evaluation and explainable AI, have to do with the following question: Did the system fail because it is a hard instance or because something else? In this paper we address this question with a generic method for estimating IRT-based instance difficulty for a wide range of AI domains covering several areas, from supervised feature-based classification to automated reasoning. We show how to estimate difficulty systematically using off-the-shelf machine learning regression models. We illustrate the usefulness of this estimation for a range of applications.
Fernando Martínez-Plumed, David Castellano Falcón, Carlos Monserrat Aranda, José Hernández-Orallo
AAAI1
2022 Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks (Extended Abstract)
abstract
We present a framework for analysing the impact of AI on occupations. This framework maps 59 generic tasks from different occupational datasets to 14 cognitive abilities and these to a comprehensive list of 328 AI benchmarks used to evaluate research intensity in AI. The use of cognitive abilities as an intermediate layer allows for an identification of potential AI exposure for tasks for which AI applications have not been explicitly programmed. We provide insights into the abilities through which AI is most likely to affect jobs, and we show how some of the abilities where AI research is currently very intense are linked to tasks with comparatively limited labour input in the labour markets of advanced economies.
Songül Tolan, Annarosa Pesole, Fernando Martínez-Plumed, Enrique Fernández-Macías, José Hernández-Orallo, Emilia Gómez
IJCAI3
2021 Missing the missing values: The ugly duckling of fairness in machine learning
abstract
Nowadays, there is an increasing concern in machine learning about the causes underlying unfair decision making, that is, algorithmic decisions discriminating some groups over others, especially with groups that are defined over protected attributes, such as gender, race and nationality. Missing values are one frequent manifestation of all these latent causes: protected groups are more reluctant to give information that could be used against them, sensitive information for some groups can be erased by human operators, or data acquisition may simply be less complete and systematic for minority groups. However, most recent techniques, libraries and experimental results dealing with fairness in machine learning have simply ignored missing data. In this paper, we present the first comprehensive analysis of the relation between missing values and algorithmic fairness for machine learning: (1) we analyse the sources of missing data and bias, mapping the common causes, (2) we find that rows containing missing values are usually fairer than the rest, which should discourage the consideration of missing values as the uncomfortable ugly data that different techniques and libraries for handling algorithmic bias get rid of at the first occasion, (3) we study the trade-off between performance and fairness when the rows with missing values are used (either because the technique deals with them directly or by imputation methods), and (4) we show that the sensitivity of six different machine-learning techniques to missing values is usually low, which reinforces the view that the rows with missing data contribute more to fairness through the other, nonmissing, attributes. We end the paper with a series of recommended procedures about what to do with missing data when aiming for fair decision making.
Fernando Martínez-Plumed, Cèsar Ferri, David Nieves, José Hernández-Orallo
Int. J. Intell. Syst.1
2021 Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks
abstract
In this paper we develop a framework for analysing the impact of Artificial Intelligence (AI) on occupations. This framework maps 59 generic tasks from worker surveys and an occupational database to 14 cognitive abilities (that we extract from the cognitive science literature) and these to a comprehensive list of 328 AI benchmarks used to evaluate research intensity across a broad range of different AI areas. The use of cognitive abilities as an intermediate layer, instead of mapping work tasks to AI benchmarks directly, allows for an identification of potential AI exposure for tasks for which AI applications have not been explicitly created. An application of our framework to occupational databases gives insights into the abilities through which AI is most likely to affect jobs and allows for a ranking of occupations with respect to AI exposure. Moreover, we show that some jobs that were not known to be affected by previous waves of automation may now be subject to higher AI exposure. Finally, we find that some of the abilities where AI research is currently very intense are linked to tasks with comparatively limited labour input in the labour markets of advanced economies (e.g., visual and auditory processing using deep learning, and sensorimotor interaction through (deep) reinforcement learning). This article appears in the special track on AI and Society.
Songül Tolan, Annarosa Pesole, Fernando Martínez-Plumed, Enrique Fernández-Macías, José Hernández-Orallo, Emilia Gómez
J. Artif. Intell. Res.3
2021 CRISP-DM Twenty Years Later: From Data Mining Processes to Data Science Trajectories
abstract
CRISP-DM(CRoss-Industry Standard Process for Data Mining) has its origins in the second half of the nineties and is thus about two decades old. According to many surveys and user polls it is still the de facto standard for developing data mining and knowledge discovery projects. However, undoubtedly the field has moved on considerably in twenty years, with data science now the leading term being favoured over data mining. In this paper we investigate whether, and in what contexts, CRISP-DM is still fit for purpose for data science projects. We argue that if the project is goal-directed and process-driven the process model view still largely holds. On the other hand, when data science projects become more exploratory the paths that the project can take become more varied, and a more flexible model is called for. We suggest what the outlines of such a trajectory-based model might look like and how it can be used to categorise data science projects (goal-directed, exploratory or data management). We examine seven real-life exemplars where exploratory activities play an important role and compare them against 51 use cases extracted from the NIST Big Data Public Working Group. We anticipate this categorisation can help project planning in terms of time and cost characteristics.
Fernando Martínez-Plumed, Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Meelis Kull, Nicolas Lachiche, María José Ramírez-Quintana, Peter A. Flach
IEEE Trans. Knowl. Data Eng.1
2020 Does AI Qualify for the Job?: A Bidirectional Model Mapping Labour and AI Intensities
abstract
In this paper we present a setting for examining the relation be-tween the distribution of research intensity in AI research and the relevance for a range of work tasks (and occupations) in current and simulated scenarios. We perform a mapping between labourand AI using a set of cognitive abilities as an intermediate layer. This setting favours a two-way interpretation to analyse (1) what impact current or simulated AI research activity has or would have on labour-related tasks and occupations, and (2) what areas of AI research activity would be responsible for a desired or undesired effect on specific labour tasks and occupations. Concretely, in our analysis we map 59 generic labour-related tasks from several worker surveys and databases to 14 cognitive abilities from the cognitive science literature, and these to a comprehensive list of 328 AI benchmarks used to evaluate progress in AI techniques. We provide this model and its implementation as a tool for simulations. We also show the effectiveness of our setting with some illustrative examples.
Fernando Martínez-Plumed, Songül Tolan, Annarosa Pesole, José Hernández-Orallo, Enrique Fernández-Macías, Emilia Gómez
AIES1
2020 Family and Prejudice: A Behavioural Taxonomy of Machine Learning Techniques
abstract
One classical way of characterising the rich range of machine learning techniques is by defining 'families', according to their formulation and learning strategy (e.g., neural networks, Bayesian methods, etc.).However, this taxonomy of learning techniques does not consider the extent to which models built with techniques from the same or different family agree on their outputs, especially when their predictions have to extrapolate in sparse zones where insufficient training data was available.In this paper we present a new taxonomy of machine learning techniques for classification, where families are clustered according to their degree of (dis)agreement in behaviour considering both dense and sparse zones, using Cohen's kappa statistic.To this end, we use a representative collection of datasets and learning techniques.We finally validate the taxonomy by performing a number of experiments for technique selection.We show that ranking techniques by only following prejudice -the reputation they have for other problems-is worse than selecting techniques based on family diversity.
Raül Fabra-Boluda, Cèsar Ferri, Fernando Martínez-Plumed, José Hernández-Orallo, María José Ramírez-Quintana
ECAI3
2020 AI Paradigms and AI Safety: Mapping Artefacts and Techniques to Safety Issues
abstract
AI safety often analyses a risk or safety issue, such as interruptibility, under a particular AI paradigm, such as reinforcement learning.But what is an AI paradigm and how does it affect the understanding and implications of the safety issue?Is AI safety research covering the most representative paradigms and the right combinations of paradigms with safety issues?Will current research directions in AI safety be able to anticipate more capable and powerful systems yet to come?In this paper we analyse these questions, introducing a distinction between two types of paradigms in AI: artefacts and techniques.We then use experimental data of research and media documents from AI Topics, an official publication of the AAAI, to examine how safety research is distributed across artefacts and techniques.We observe that AI safety research is not sufficiently anticipatory, and is heavily weighted towards certain research paradigms.We identify a need for AI safety to be more explicit about the artefacts and techniques for which a particular issue may be applicable, in order to identify gaps and cover a broader range of issues.
José Hernández-Orallo, Fernando Martínez-Plumed, Shahar Avin, Jess Whittlestone, Seán Ó hÉigeartaigh
ECAI2
2020 Tracking AI: The Capability Is (Not) Near
abstract
AI is an area of strategic importance with potential to be a key driver of economic development and with a wide range of potential social\nimplications. In order to assess present and future impact, there is a need to analyse what AI can (and will) achieve. But, what is AI capable of? This question is as crucial as elusive, as AI is progressing in ways that are open-ended about the techniques and resources AI can operate with. The truth is that whenever a task is solved, researchers find increasingly challenging to extrapolate whether this task can be reproduced, even when only a few things change: the data, the domain knowledge, the level of uncertainty, the (hyper)parameters, the techniques, the team, the compute, etc. In the end, we would like to infer whether a good result (or a breakthrough) in task A transfers to a similar good result in task B. This extrapolation is precisely what the notion of capability, borrowed from psychology, tries to answer. However, we lack the tools, and the data, to do similarly in AI. Benchmarks, competitions and challenges are behind much of the recent progress in AI, especially in machine learning (ML) [10], but the dynamics of rushing breakthroughs at the expense of massive data, compute, specialisation, etc., has led to a more complex AI landscape, in terms of what can be achieved and how. As a result, policy makers and other stakeholders have no way of assessing what AI systems can do today and in the future. This does not mean that we must disregard or understate the valuable information that is provided by a plethora of benchmarks. On the contrary, the analysis of the progress of AI must be based on data-grounded evidence, relying on finding and testing hypotheses through the computational analysis of big amounts of shared data [6], using open data science tools [11]. But this analysis must be abstracted from tasks to capabilities, for the purposes of integration3 and evaluation [8]. In this paper, we identify a series of problems to track and understand what AI is capable of, surveying some previous initiatives. We present the AIcollaboratory, a data-driven framework to collect\nand explore data about AI results, progress and ultimately capabilities, being developed in the context of AI WATCH, the European\nCommission (EC) knowledge service to monitor the development, uptake and impact of AI in Europe4. We close the paper with some\nchallenges for the community emerging around the collaboratory.
Fernando Martínez-Plumed, José Hernández-Orallo, Emilia Gómez
ECAI1
2020 Dual Indicators to Analyze AI Benchmarks: Difficulty, Discrimination, Ability, and Generality
abstract
With the purpose of better analyzing the result of artificial intelligence (AI) benchmarks, we present two indicators on the side of the AI problems, difficulty and discrimination, and two indicators on the side of the AI systems, ability and generality. The first three are adapted from psychometric models in item response theory (IRT), whereas generality is defined as a new metric that evaluates whether an agent is consistently good at easy problems and bad at difficult ones. We illustrate how these key indicators give us more insight on the results of two popular benchmarks in AI, the Arcade Learning Environment (Atari 2600 games) and the General Video Game AI competition, and we include some guidelines to estimate and interpret these indicators for other AI benchmarks and competitions.
Fernando Martínez-Plumed, José Hernández-Orallo
IEEE Trans. Games1
2019 Automated Data Transformation with Inductive Programming and Dynamic Background Knowledge
Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, María José Ramírez-Quintana, Susumu Katayama
ECML/PKDD (3)4
2019 BK-ADAPT: Dynamic Background Knowledge for Automating Data Transformation
Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, María José Ramírez-Quintana, Susumu Katayama
ECML/PKDD (3)4
2019 Item response theory in AI: Analysing machine learning classifiers at the instance level
Fernando Martínez-Plumed, Ricardo B. C. Prudêncio, Adolfo Martínez Usó, José Hernández-Orallo
Artif. Intell.1
2018 The Facets of Artificial Intelligence: A Framework to Track the Evolution of AI
abstract
We present nine facets for the analysis of the past and future evolution of AI. Each facet has also a set of edges that can summarise different trends and contours in AI. With them, we first conduct a quantitative analysis using the information from two decades of AAAI/IJCAI conferences and around 50 years of documents from AI topics, an official database from the AAAI, illustrated by several plots. We then perform a qualitative analysis using the facets and edges, locating AI systems in the intelligence landscape and the discipline as a whole. This analytical framework provides a more structured and systematic way of looking at the shape and boundaries of AI.
Fernando Martínez-Plumed, Bao Sheng Loe, Peter A. Flach, Seán Ó hÉigeartaigh, Karina Vold, José Hernández-Orallo
IJCAI1
2017 Computer Models Solving Intelligence Test Problems: Progress and Implications (Extended Abstract)
abstract
While some computational models of intelligence test problems were proposed throughout the second half of the XXth century, in the first years of the XXIst century we have seen an increasing number of computer systems being able to score well on particular intelligence test tasks. However, despitethis increasing trend there has been no general account of all these works in terms of how theyrelate to each other and what their real achievements are. In this paper, we provide some insighton these issues by giving a comprehensive account of about thirty computer models, from the 1960sto nowadays, and their relationships, focussing on the range of intelligence test tasks they address, thepurpose of the models, how general or specialised these models are, the AI techniques they use in eachcase, their comparison with human performance, and their evaluation of item difficulty.
José Hernández-Orallo, Fernando Martínez-Plumed, Ute Schmid, Michael Siebers, David L. Dowe
IJCAI2
2016 Making Sense of Item Response Theory in Machine Learning
abstract
Item response theory (IRT) is widely used to measure latent abilities of subjects (specially for educational testing) based on their responses to items with different levels of difficulty. The adaptation of IRT has been recently suggested as a novel perspective for a better understanding of the results of machine learning experiments and, by extension, other artificial intelligence experiments. For instance, IRT suits classification tasks perfectly, where instances correspond to items and classifiers correspond to subjects. By adopting IRT, item (i.e., instance) characteristic curves can be estimated using logistic models, for which several parameters characterise each dataset instance: difficulty, discrimination and guessing. IRT looks promising for the analysis of instance hardness, noise, classifier dominances, etc. However, some caveats have been found when trying to interpret the IRT parameters in a machine learning setting, especially when we include some artificial classifiers in the pool of classifiers to be evaluated: the optimal and pessimal classifiers, a random classifier and the majority and minority classifiers. In this paper we perform a series of experiments with a range of datasets and classification methods to fully understand how IRT works and what their parameters really mean in the context of machine learning. This better understanding will hopefully pave the way to a myriad of potential applications in machine learning and artificial intelligence.
Fernando Martínez-Plumed, Ricardo B. C. Prudêncio, Adolfo Martínez Usó, José Hernández-Orallo
ECAI1
2016 Computer models solving intelligence test problems: Progress and implications
abstract
While some computational models of intelligence test problems were proposed throughout the second half of the XXth century, in the first years of the XXIst century we have seen an increasing number of computer systems being able to score well on particular intelligence test tasks. However, despite this increasing trend there has been no general account of all these works in terms of how they relate to each other and what their real achievements are. Also, there is poor understanding about what intelligence tests measure in machines, whether they are useful to evaluate AI systems, whether they are really challenging problems, and whether they are useful to understand (human) intelligence. In this paper, we provide some insight on these issues, in the form of nine specific questions, by giving a comprehensive account of about thirty computer models, from the 1960s to nowadays, and their relationships, focussing on the range of intelligence test tasks they address, the purpose of the models, how general or specialised these models are, the AI techniques they use in each case, their comparison with human performance, and their evaluation of item difficulty. As a conclusion, these tests and the computer models attempting them show that AI is still lacking general techniques to deal with a variety of problems at the same time. Nonetheless, a renewed attention on these problems and a more careful understanding of what intelligence tests offer for AI may help build new bridges between psychometrics, cognitive science, and AI; and may motivate new kinds of problem repositories.
José Hernández-Orallo, Fernando Martínez-Plumed, Ute Schmid, Michael Siebers, David L. Dowe
Artif. Intell.2
2014 A Knowledge Growth and Consolidation Framework for Lifelong Machine Learning Systems
abstract
A more effective vision of machine learning systems entails tools that are able to improve task after task and to reuse the patterns and knowledge that are acquired previously for future tasks. This incremental, long-life view of machine learning goes beyond most of state-of-the-art machine learning techniques that learn throw-away models. In this paper we present a long-life knowledge acquisition, evaluation and consolidation framework that is designed to work with any rule-based machine learning or inductive inference engine and integrate it into a long-life learner. In order to do that we work over the graph of working memory rules and introduce several topological metrics over it from which we derive an oblivion criterion to drop useless rules from working memory and a consolidation process to promote the rules to the knowledge base. We evaluate the framework on a series of tasks in a chess rule learning domain.
Fernando Martínez-Plumed, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana
ICMLA1