VLDB 2026 Research / reviewers in the wild / expert
José Hernández-Orallo
dblp:h/JoseHernandezOrallo
· DBLP profile ↗
92ranked-venue papers
21as first author
38since 2021 · last 2026
0000-0001-9746-7632ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 17 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 20 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Theory of computation · 3Software engineering, systems software and programming languages · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRACE: A Corpus of Team Creative DiscussionsabstractUnderstanding how discussion dynamics shape team creativity has been limited by the difficulty of measuring process at scale.We introduce TRACE, a corpus of 309 group discussions from 103 teams (421 participants) across six creative problem-solving tasks.The dataset follows an input-process-output framework, integrating team composition (demographics, personalities), full discussion transcripts, and creativity outcomes.Using sentence embeddings and factor analysis, we identify four interpretable discussion dimensions:Coherence, Exploration, Convergence, and Participation.Analysis reveals a depth-breadth trade-off: coherent idea development inversely relates to semantic exploration.Larger teams explore more broadly but converge less effectively while team diversity shapes participation patterns more than discussion content.Novelty and usefulness in the creativity outcomes follow distinct pathways: Exploration and Convergence predict novelty, whereas Coherence predicts usefulness.These findings ground our understanding of how teams talk their way to creative solutions and provide guidance for designing multiagent systems. Yixuan Jiang, Tiancheng Hu, José Hernández-Orallo, David Stillwell, Luning Sun 0001 |
ACL (1) | 3 |
| 2026 | Predictable artificial intelligenceabstractMany areas of artificial intelligence, and machine learning in particular, aim at being probably correct, i.e., valid on average, rather than pursuing the idealistic goal of being provably valid for all inputs. However, AI systems could still be predictably valid, such as an imperfect robot deliverer for which we can reliably and precisely predict the task instances for which it is correct and safe, its valid operating range. “Predictable AI” is a nascent research area that explores ways of anticipating key validity indicators (e.g., performance, safety) of present and future AI ecosystems. We argue that achieving predictability is crucial for fostering trust, liability, control, alignment and safety of AI, and thus should be prioritised over performance. We formally characterise predictability, explore its most relevant components, illustrate what can be predicted, describe alternative candidates for predictors, as well as the trade-offs between maximising validity and predictability. To illustrate these concepts, we bring an array of illustrative examples covering diverse ecosystem configurations. “Predictable AI” is related to other areas of technical and non-technical AI research, but have distinctive questions, hypotheses, techniques and challenges. This paper aims to elucidate them, calls for identifying paths towards a landscape of predictably valid AI systems and outlines the potential impact of this emergent field. Lexin Zhou, P. A. M. Casares, Fernando Martínez-Plumed, John Burden, Ryan Burnell, Lucy Cheke, Cèsar Ferri, Alexandru Marcoci, Behzad Mehrbakhsh, Yael Moros-Daval, Seán Ó hÉigeartaigh, Danaja Rutar, Wout Schellaert, Konstantinos Voudouris, José Hernández-Orallo |
Artif. Intell. | 15 |
| 2026 | When Redundancy Matters: Machine Teaching of RepresentationsabstractAbstract In traditional machine teaching, a teacher needs to teach a concept to a learner by means of a finite set of examples, the witness set. But concepts can have many equivalent representations. This redundancy strongly affects the search space, to the extent that teacher and learner may not be able to easily determine the equivalence class of each representation. In this common situation, instead of teaching concepts, we explore the idea of teaching representations. We work with several teaching schemas that exploit representation and witness size (Eager, Greedy and Optimal) and analyze the gains in teaching effectiveness, both theoretically, and also experimentally for languages where redundancy can vary (DNF expressions and Turing-complete P3 programs). Our theoretical and experimental results indicate that there are various types of redundancy, related, e.g,. to the spread of the redundant representations, handled better by the new Greedy schema introduced here than by the Eager schema. For P3 programs witness sets found by Greedy are usually smaller than the programs they identify, corroborating previous results that conveying information efficiently is a leitmotif of machine teaching. Cèsar Ferri, Darío Garigliotti, José Hernández-Orallo, Brigt Håvardstun, Jan Arne Telle |
Mach. Learn. | 3 |
| 2026 | What Should an AI Assessor Optimise for?abstractAbstract An AI assessor is an external, ideally independent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can leverage information from the test results of many other AI systems and have the flexibility of being trained on any loss function or scoring rule: from squared error to toxicity metrics. Here we address the question: is it always optimal to train the assessor for the target metric? Or could it be better to train for a different metric and then map predictions back to the target metric? Using twenty regression and classification problems with tabular data, we experimentally explore this question for, respectively, regression losses and classification scores with monotonic and nonmonotonic mappings and find that, contrary to intuition, optimising for more informative metrics (i.e., yielding a better-conditioned supervision signal) is not universally preferred. Surprisingly, some monotonic transformations are promising. For example, logistic loss is useful for minimising absolute or quadratic errors in regression, and logarithmic score helps maximise quadratic or spherical scores in classification. Daniel Romero-Alvarado, Fernando Martínez-Plumed, José Hernández-Orallo |
Mach. Learn. | 3 |
| 2025 | Relative Drawing Identification Complexity Is Invariant to Modality in Vision-Language ModelsabstractLarge language models have become multimodal, and many of them are said to integrate their modalities using common representations. If this were true, a drawing of a car as an image, for instance, should map to a similar area in the latent space as a textual description of the strokes that form the drawing. To explore this in a black-box access regime to these models, we propose the use of machine teaching, a theory that studies the minimal set of examples a teacher needs to choose so that the learner captures the concept. In this paper, we evaluate the complexity of teaching vision-language models a subset of objects in the Quick, Draw! dataset using two presentations: raw images as bitmaps and trace coordinates in TikZ format. The results indicate that image-based representations generally require fewer segments and achieve higher accuracy than coordinate-based representations. But, surprisingly, the teaching size usually ranks concepts similarly across both modalities, even when controlling for (a human proxy of) concept priors, suggesting that the simplicity of concepts may be an inherent property that transcends modality representations. Diogo Freitas, Brigt Håvardstun, Darío Garigliotti, Jan Arne Telle, Cèsar Ferri, José Hernández-Orallo |
ECAI | 6 |
| 2025 | Paradigms of AI Evaluation: Mapping Goals, Methodologies and CultureabstractResearch in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other's contributions. This fragmentation has led to insular research trajectories and communication barriers both among different paradigms and with the general public, contributing to unmet expectations for deployed AI systems. To help bridge this insularity, in this paper we survey recent work in the AI evaluation landscape and identify six main paradigms. We characterise major recent contributions within each paradigm across key dimensions related to their goals, methodologies and research cultures. By clarifying the unique combination of questions and approaches associated with each paradigm, we aim to increase awareness of the breadth of current evaluation approaches and foster cross-pollination between different paradigms. We also identify potential gaps in the field to inspire future research directions. John Burden, Marko Tesic, Lorenzo Pacchiardi, José Hernández-Orallo |
IJCAI | 4 |
| 2025 | Contamination Budget: Trade-offs Between Breadth, Depth and DifficultyabstractContamination in large language models (LLMs), and machine learning more broadly, refers to the inclusion of equal --or very similar-- examples in both training and test sets. This phenomenon usually translates into better test performance. Here we explore when this contamination is performed intentionally, for purposes that can be malicious (e.g., get better scores in evaluations) or benevolent (e.g., fix some mistakes). These interventions, usually in the form of fine-tuning memorisations, come with a budget in the size of the fine-tuning dataset. Several trade-offs appear between the breadth of the intervention (how many examples to be memorised), its depth (how many repetitions of each example) and the difficulty of the examples. By studying several LLMs and datasets, we observe some monotonic behaviour (more difficult items require more depth to be `fixed') but also some non-monotonic phenomena (very high depth levels have negative effects on non-contaminated examples). This suggests that trade-offs should be found not only in terms of the budget but also according to model specifics, the task and the item difficulty at hand. Behzad Mehrbakhsh, Fernando Martínez-Plumed, José Hernández-Orallo |
IJCAI | 3 |
| 2025 | Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using ConcordiaabstractLarge Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In this paper, we introduce a method for evaluating the ability of LLM-based agents to cooperate in zero-shot, mixed-motive environments using Concordia, a natural language multi-agent simulation environment. Our method measures general cooperative intelligence by testing an agent's ability to identify and exploit opportunities for mutual gain across diverse partners and contexts. We present empirical results from the NeurIPS 2024 Concordia Contest, where agents were evaluated on their ability to achieve mutual gains across a suite of diverse scenarios ranging from negotiation to collective action problems. Our findings reveal significant gaps between current agent capabilities and the robust generalization required for reliable cooperation, particularly in scenarios demanding persuasion and norm enforcement. Chandler Smith, Marwa Abdulhai, Manfred Diaz, Marko Tesic, Rakshit S. Trivedi, Alexander Vezhnevets, Lewis Hammond, Jesse Clifton, Minsuk Chang, Edgar A. Duéñez-Guzmán, John P. Agapiou, Jayd Matyas, Danny Karmon, Beining Zhang, Jim Dilkes, Akash Kundu, Emanuel Tewolde, Jebish Purbey, Ram Mohan Rao Kadiyala, Siddhant Gupta, Aliaksei Korshuk, Buyantuev Alexander, Ilya Makarov, Rolando Fernandez, Zhihan Wang, Caroline Wang, Jiaxun Cui, Lingyun Xiao, Yoonchang Sung, Muhammad Arrasy Rahman, Peter Stone 0001, Yipeng Kang, Hyeonggeun Yun, Ananya, Taehun Cha, Elizaveta Tennant, Olivia Macmillan-Scott, Marta Segura, Diana Riazi, Fuyang Cui, Sriram Ganapathi, Toryn Q. Klassen, Nico Schiavone, Mogtaba Alim, Sheila A. McIlraith, Manuel Ríos, Oswaldo Peña, Manuela Chacon-Chamorro, Rubén Manrique, Luis Felipe Giraldo, Nicanor Quijano, Fangwei Zhong, Wenming Tu, Zhaowei Zhang 0001, Zixia Jia, Zilong Zheng, Chichen Lin, Weijian Fan, Chenao Liu, Sneheel Sarangi, Shuqing Shi, Yali Du 0001, Avinaash Anand Kulandaivel, Yang Liu 0266, Ruiyang Wu 0007, Chetan Talele, Sunjia Lu, Gema Parreno, Shamika Dhuri, Bain McHale, Tim Baarslag, Dylan Hadfield-Menell, Natasha Jaques, José Hernández-Orallo, Joel Z. Leibo |
NeurIPS | 85 |
| 2025 | Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent ApproachabstractLarge language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely.
Existing safety evaluations primarily rely on context-independent metrics—such as factuality, bias, or toxicity—overlooking the fact that the same response may carry divergent risks depending on the user's background or condition.
We introduce ``personalized safety'' to fill this gap and present PENGUIN—a benchmark comprising 14,000 scenarios across seven sensitive domains with both context-rich and context-free variants. Evaluating six leading LLMs, we demonstrate that personalized user information significantly improves safety scores by 43.2%, confirming the effectiveness of personalization in safety alignment. However, not all context attributes contribute equally to safety enhancement. To address this, we develop RAISE—a training-free, two-stage agent framework that strategically acquires user-specific background. RAISE improves safety scores by up to 31.6% over six vanilla LLMs, while maintaining a low interaction cost of just 2.7 user queries on average. Our findings highlight the importance of selective information gathering in safety-critical domains and offer a practical solution for personalizing LLM responses without model retraining. This work establishes a foundation for safety research that adapts to individual user contexts rather than assuming a universal harm standard. Edward Sun, Kaijie Zhu, Jianxun Lian, José Hernández-Orallo, Aylin Caliskan, Jindong Wang 0001 |
NeurIPS | 5 |
| 2025 | Cracking black-box models: Revealing hidden machine learning techniques behind their predictionsabstractThe quest for transparency in black-box models has gained significant momentum in recent years. In particular, discovering the underlying machine learning technique type (or model family) from the performance of a black-box model is a real important problem both for better understanding its behaviour and for developing strategies to attack it by exploiting the weaknesses intrinsic to the learning technique. In this paper, we tackle the challenging task of identifying which kind of machine learning model is behind the predictions when we interact with a black-box model. Our innovative method involves systematically querying a black-box model (oracle) to label an artificially generated dataset, which is then used to train different surrogate models using machine learning techniques from different families (each one trying to partially approximate the oracle’s behaviour). We present two approaches based on similarity measures, one selecting the most similar family and the other using a conveniently constructed meta-model. In both cases, we use both crisp and soft classifiers and their corresponding similarity metrics. By experimentally comparing all these methods, we gain valuable insights into the explanatory and predictive capabilities of our model family concept. This provides a deeper understanding of the black-box models and increases their transparency and interpretability, paving the way for more effective decision making. Raül Fabra-Boluda, Cèsar Ferri, José Hernández-Orallo, M. José Ramrez-Quintana, Fernando Martínez-Plumed |
Intell. Data Anal. | 3 |
| 2025 | The impact of sociality regimes on heterogeneous cooperative-competitive multi-agent reinforcement learning: a study with the predator-prey gameabstractThe performance in multi-agent reinforcement learning (MARL) scenarios has usually been analysed in homogeneous teams with a few choices for the sociality regime (selfish, egalitarian, or altruistic). In this paper we analyse both homogeneous and heterogeneous teams in a variation of sociality regimes in the predator-prey game, using a novel normalisation of the weights so that the sum of all rewards is independent of the sociality regime. We find that the selfish regime is advantageous for both predator and prey teams, and for both homogeneous and heterogeneous teams. In particular, rewards are about 100% higher for the predator team when switching from the egalitarian to selfish regime and more than 400% higher from the altruistic regime. For the prey, the increase is around 40% and 100% respectively. The results are similar for homogeneous and heterogeneous situations. The takeaway message is that any study of homogeneous and heterogeneous cooperative-competitive multi-agent reinforcement learning teams should also take into account the sociality regimes before making conclusions on the preference of any algorithm. Yue Zhao 0023, José Hernández-Orallo |
J. Exp. Theor. Artif. Intell. | 2 |
| 2025 | Analysing the Predictability of Language Model PerformanceabstractCan a language model predict for which questions another language model will answer successfully? We investigate the extent to which performance prediction is possible and dissect various factors that influence it. Our experimental setting fine-tunes DeBERTa models, which we call assessors , on the evaluation results of generative language models with up to 128 billion parameters, which we refer to as subject systems . Our analysis spans more than 100 tasks from BIG-bench. We find that the assessors can match and even exceed the subjects’ confidence in both refinement and calibration, anticipating failures at near perfect levels for some tasks. We also find that for performance prediction it can be beneficial to learn from the scores on multiple tasks or to learn from the scores of multiple subjects, but both depend on the task at hand. Lastly, we find that large and small subject systems are equally predictable, showing promise for the scalability of the predictability problem. Wout Schellaert, Fernando Martínez-Plumed, José Hernández-Orallo |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | Your Prompt Is My Command: On Assessing the Human-Centred Generality of Multimodal Models (Abstract Reprint)abstractEven with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully autonomous one. HCG implies that —for a specific user— a system is only as general as it is effective for the user’s relevant tasks and their prevalent ways of prompting. A human-centred evaluation of general-purpose AI systems therefore needs to reflect the personal nature of interaction, tasks and cognition. We argue that the best way to understand these systems is as highly-coupled cognitive extenders, and to analyse the bidirectional cognitive adaptations between them and humans. In this paper, we give a formulation of HCG, as well as a high-level overview of the elements and trade-offs involved in the prompting process. We end the paper by outlining some essential research questions and suggestions for improving evaluation practices, which we envision as characteristic for the evaluation of general artificial intelligence in the future. Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, John Burden, P. A. M. Casares, Bao Sheng Loe, Roi Reichart, Seán Ó hÉigeartaigh, Anna Korhonen, José Hernández-Orallo |
AAAI | 10 |
| 2024 | Caveats and Solutions for Characterising General-Purpose AIabstractThe concept of General-Purpose AI (GPAI) has recently been permeating research papers, policy reports and legal regulations, as a way of referring to current and future models with high levels of capability and generality. Yet precisely characterising GPAI models remains elusive. Current definitions often describe GPAI models as those that ‘competently perform a wide range of distinct tasks’. To properly characterise GPAI we need well-grounded definitions of capability and generality. In this paper, I will briefly introduce –or revisit– the concept of capability, going well beyond aggregate performance on benchmarks, and discuss practical procedures to evaluate the capability profile of AI systems, and derive generality metrics from them. José Hernández-Orallo |
ECAI | 1 |
| 2024 | Distilling the Effects of Language Model ContaminationabstractThe proportion of AI-generated content permeating the well of knowledge is increasing significantly. Large language models (LLMs) contribute to that contamination but they also suffer from it. However, it is yet to be clarified the effect of different sources of error, be it human-generated or LLM-generated. Controlling for the percentage of error, we explore the impact on LLM fine-tuning when errors come from humans, from other language models or are generated randomly using an aleatoric or epistemic source. In this paper, we compare these different types of error for in-distribution and out-of-distribution experimental settings. By analysing the levels of errors and their distribution, we find a nuanced view: while in-distribution human-generated noise seems more benign than the LLM-generated counterpart, in the out-of-distribution case the model-generated noise may not be necessarily worse. Behzad Mehrbakhsh, Fernando Martínez-Plumed, José Hernández-Orallo |
ECAI | 3 |
| 2024 | Language Task Difficulty Prediction Through LLM-Annotated Meta-FeaturesabstractAssessing the capabilities of large language models (LLMs) is increasingly challenging due to their generality and uneven task performance. Often, we do not know how much of the success or failure on a particular task is due to the ‘loading’ of the language elements in the task, such as narrative understanding, or some other intrinsic (non-linguistic) components, such as domain-specific common sense or reasoning capabilities. Understanding what tasks are most loaded on language and determine the predictability of LLMs on these tasks is crucial for improving benchmarks, designing better LLMs, and ensuring their safe deployment. We present an innovative methodology that uses LLMs to annotate linguistic meta-features, allowing us to predict task difficulty and understand linguistic loadings more accurately than traditional readability scores. Using GPT-4 for automated annotation, we show strong predictability for a variety of tasks and language models (e.g., MMLU with R2 from 0.68 to 0.83), but observe limited predictability for other tasks (e.g., LSAT with R2 of -0.07). Yael Moros-Daval, Fernando Martínez-Plumed, José Hernández-Orallo |
ECAI | 3 |
| 2024 | How Resilient are Language Models to Text Perturbations?
Daniel Romero-Alvarado, José Hernández-Orallo, Fernando Martínez-Plumed |
IDEAL (1) | 2 |
| 2024 | Melting Pot Contest: Charting the Future of Generalized Cooperative IntelligenceabstractMulti-agent AI research promises a path to develop human-like and human-compatible intelligent technologies that complement the solipsistic view of other approaches, which mostly do not consider interactions between agents. Aiming to make progress in this direction, the Melting Pot contest 2023 focused on the problem of cooperation among interacting agents and challenged researchers to push the boundaries of multi-agent reinforcement learning (MARL) for mixed-motive games. The contest leveraged the Melting Pot environment suite to rigorously evaluate how well agents can adapt their cooperative skills to interact with novel partners in unforeseen situations. Unlike other reinforcement learning challenges, this challenge focused on social rather than environmental generalization. In particular, a population of agents performs well in Melting Pot when its component individuals are adept at finding ways to cooperate both with others in their population and with strangers. Thus Melting Pot measures cooperative intelligence.The contest attracted over 600 participants across 100+ teams globally and was a success on multiple fronts: (i) it contributed to our goal of pushing the frontiers of MARL towards building more cooperatively intelligent agents, evidenced by several submissions that outperformed established baselines; (ii) it attracted a diverse range of participants, from independent researchers to industry affiliates and academic labs, both with strong background and new interest in the area alike, broadening the field’s demographic and intellectual diversity; and (iii) analyzing the submitted agents provided important insights, highlighting areas for improvement in evaluating agents' cooperative intelligence. This paper summarizes the design aspects and results of the contest and explores the potential of Melting Pot as a benchmark for studying Cooperative AI. We further analyze the top solutions and conclude with a discussion on promising directions for future research. Rakshit S. Trivedi, Akbir Khan, Jesse Clifton, Lewis Hammond, Edgar A. Duéñez-Guzmán, Dipam Chakraborty, John P. Agapiou, Jayd Matyas, Alexander Vezhnevets, Barna Pásztor, Yunke Ao, Omar G. Younis, Benjamin Swain, Haoyuan Qin, Mian Deng, Ziwei Deng, Utku Erdoganaras, Yue Zhao 0023, Marko Tesic, Natasha Jaques, Jakob N. Foerster, Vincent Conitzer, José Hernández-Orallo, Dylan Hadfield-Menell, Joel Z. Leibo |
NeurIPS | 24 |
| 2023 | Adversarial Benchmark Evaluation Rectified by Controlling for DifficultyabstractAdversarial benchmark construction, where harder instances challenge new generations of AI systems, is becoming the norm. While this approach may lead to better machine learning models —on average and for the new benchmark—, it is unclear how these models behave on the original distribution. Two opposing effects are intertwined here. On the one hand, the adversarial benchmark has a higher proportion of difficult instances, with lower expected performance. On the other hand, models trained on the adversarial benchmark may improve on these difficult instances (but may also neglect some easy ones). To disentangle these two effects we can control for difficulty, showing that we can recover the performance on the original distribution, provided the harder instances were obtained from this distribution in the first place. We show this difficulty-aware rectification works in practice, through a series of experiments with several benchmark construction schemas and the use of a populational difficulty metric. As a take-away message, instead of distributional averages we recommend using difficulty-conditioned characteristic curves when evaluating models built with adversarial benchmarks. Behzad Mehrbakhsh, Fernando Martínez-Plumed, José Hernández-Orallo |
ECAI | 3 |
| 2023 | XAI with Machine Teaching When Humans Are (Not) Informed About the Irrelevant Features
Brigt Håvardstun, Cèsar Ferri, José Hernández-Orallo, Pekka Parviainen, Jan Arne Telle |
ECML/PKDD (3) | 3 |
| 2023 | Your Prompt is My Command: On Assessing the Human-Centred Generality of Multimodal ModelsabstractEven with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully autonomous one. HCG implies that —for a specific user— a system is only as general as it is effective for the user’s relevant tasks and their prevalent ways of prompting. A human-centred evaluation of general-purpose AI systems therefore needs to reflect the personal nature of interaction, tasks and cognition. We argue that the best way to understand these systems is as highly-coupled cognitive extenders, and to analyse the bidirectional cognitive adaptations between them and humans. In this paper, we give a formulation of HCG, as well as a high-level overview of the elements and trade-offs involved in the prompting process. We end the paper by outlining some essential research questions and suggestions for improving evaluation practices, which we envision as characteristic for the evaluation of general artificial intelligence in the future. This paper appears in the AI & Society track. Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, John Burden, P. A. M. Casares, Bao Sheng Loe, Roi Reichart, Seán Ó hÉigeartaigh, Anna Korhonen, José Hernández-Orallo |
J. Artif. Intell. Res. | 10 |
| 2023 | Heuristic search of optimal machine teaching curriculaabstractAbstract In curriculum learning the order of concepts is determined by the teacher but not the examples for each concept, while in machine teaching it is the examples that are chosen by the teacher to minimise the learning effort, though the concepts are taught in isolation. Curriculum teaching is the natural combination of both, where both concept order and the set of examples can be chosen to minimise the size of the whole teaching session. Yet, this simultaneous minimisation of teaching sets and concept order is computationally challenging, facing issues such as the “interposition” phenomenon: previous knowledge may be counter-productive. We build on a machine-teaching framework based on simplicity priors that can achieve short teaching sizes for large classes of languages. Given a set of concepts, we identify an inequality relating the sizes of example sets and concept descriptions. This leverages the definition of admissible heuristics for A* search to spot the optimal curricula by avoiding interposition, being able to find the shortest teaching sessions in a more efficient way than an exhaustive search and with the guarantees we do not have with a greedy algorithm. We illustrate these theoretical findings through case studies in a drawing domain, polygonal strokes on a grid described by a simple language implementing compositionality and recursion. Manuel Garcia-Piqueras, José Hernández-Orallo |
Mach. Learn. | 2 |
| 2023 | Can language models automate data wrangling?abstractAbstract The automation of data science and other data manipulation processes depend on the integration and formatting of ‘messy’ data. Data wrangling is an umbrella term for these tedious and time-consuming tasks. Tasks such as transforming dates, units or names expressed in different formats have been challenging for machine learning because (1) users expect to solve them with short cues or few examples, and (2) the problems depend heavily on domain knowledge. Interestingly, large language models today (1) can infer from very few examples or even a short clue in natural language, and (2) can integrate vast amounts of domain knowledge. It is then an important research question to analyse whether language models are a promising approach for data wrangling, especially as their capabilities continue growing. In this paper we apply different variants of the language model Generative Pre-trained Transformer (GPT) to five batteries covering a wide range of data wrangling problems. We compare the effect of prompts and few-shot regimes on their results and how they compare with specialised data wrangling systems and other tools. Our major finding is that they appear as a powerful tool for a wide range of data wrangling tasks. We provide some guidelines about how they can be integrated into data processing pipelines, provided the users can take advantage of their flexibility and the diversity of tasks to be addressed. However, reliability is still an important issue to overcome. Gonzalo Jaimovitch-López, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, María José Ramírez-Quintana |
Mach. Learn. | 3 |
| 2022 | How General-Purpose Is a Language Model? Usefulness and Safety with Human Prompters in the WildabstractThe new generation of language models is reported to solve some extraordinary tasks the models were never trained for specifically, in few-shot or zero-shot settings. However, these reports usually cherry-pick the tasks, use the best prompts, and unwrap or extract the solutions leniently even if they are followed by nonsensical text. In sum, they are specialised results for one domain, a particular way of using the models and interpreting the results. In this paper, we present a novel theoretical evaluation framework and a distinctive experimental study assessing language models as general-purpose systems when used directly by human prompters --- in the wild. For a useful and safe interaction in these increasingly more common conditions, we need to understand when the model fails because of a lack of capability or a misunderstanding of the user's intents. Our results indicate that language models such as GPT-3 have limited understanding of the human command; far from becoming general-purpose systems in the wild. P. A. M. Casares, Bao Sheng Loe, John Burden, Seán Ó hÉigeartaigh, José Hernández-Orallo |
AAAI | 5 |
| 2022 | Training on the Test Set: Mapping the System-Problem Space in AIabstractMany present and future problems associated with artificial intelligence are not due to its limitations, but to our poor assessment of its behaviour. Our evaluation procedures produce aggregated performance metrics that lack detail and quantified uncertainty about the following question: how will an AI system, with a particular profile \pi, behave for a new problem, characterised by a particular situation \mu? Instead of just aggregating test results, we can use machine learning methods to fully capitalise on this evaluation information. In this paper, we introduce the concept of an assessor model, \hat{R}(r|\pi,\mu), a conditional probability estimator trained on test data. We discuss how these assessors can be built by using information of the full system-problem space and illustrate a broad range of applications that derive from varied inferences and aggregations from \hat{R}. Building good assessor models will change the predictive and explanatory power of AI evaluation and will lead to new research directions for building and using them. We propose accompanying every deployed AI system with its own assessor. José Hernández-Orallo, Wout Schellaert, Fernando Martínez-Plumed |
AAAI | 1 |
| 2022 | When AI Difficulty Is Easy: The Explanatory Power of Predicting IRT DifficultyabstractOne of challenges of artificial intelligence as a whole is robustness. Many issues such as adversarial examples, out of distribution performance, Clever Hans phenomena, and the wider areas of AI evaluation and explainable AI, have to do with the following question: Did the system fail because it is a hard instance or because something else? In this paper we address this question with a generic method for estimating IRT-based instance difficulty for a wide range of AI domains covering several areas, from supervised feature-based classification to automated reasoning. We show how to estimate difficulty systematically using off-the-shelf machine learning regression models. We illustrate the usefulness of this estimation for a range of applications. Fernando Martínez-Plumed, David Castellano Falcón, Carlos Monserrat Aranda, José Hernández-Orallo |
AAAI | 4 |
| 2022 | Not a Number: Identifying Instance Features for Capability-Oriented EvaluationabstractIn AI evaluation, performance is often calculated by averaging across various instances. But to fully understand the capabilities of an AI system, we need to understand the factors that cause its pattern of success and failure. In this paper, we present a new methodology to identify and build informative instance features that can provide explanatory and predictive power to analyse the behaviour of AI systems more robustly. The methodology builds on these relevant features that should relate monotonically with success, and represents patterns of performance in a new type of plots known as ‘agent characteristic grids’. We illustrate this methodology with the Animal-AI competition as a representative example of how we can revisit existing competitions and benchmarks in AI—even when evaluation data is sparse. Agents with the same average performance can show very different patterns of performance at the instance level. With this methodology, these patterns can be visualised, explained and predicted, progressing towards a capability-oriented evaluation rather than relying on a less informative average performance score. Ryan Burnell, John Burden, Danaja Rutar, Konstantinos Voudouris, Lucy Cheke, José Hernández-Orallo |
IJCAI | 6 |
| 2022 | Non-Cheating Teaching Revisited: A New Probabilistic Machine Teaching ModelabstractOver the past decades in the field of machine teaching, several restrictions have been introduced to avoid ‘cheating’, such as collusion-free or non-clashing teaching. However, these restrictions forbid several teaching situations that we intuitively consider natural and fair, especially those ‘changes of mind’ of the learner as more evidence is given, affecting the likelihood of concepts and ultimately their posteriors. Under a new generalised probabilistic teaching, not only do these non-cheating constraints look too narrow but we also show that the most relevant machine teaching models are particular cases of this framework: the consistency graph between concepts and elements simply becomes a joint probability distribution. We show a simple procedure that builds the witness joint distribution from the ground joint distribution. We prove a chain of relations, also with a theoretical lower bound, on the teaching dimension of the old and new models. Overall, this new setting is more general than the traditional machine teaching models, yet at the same time more intuitively capturing a less abrupt notion of non-cheating teaching. Cèsar Ferri, José Hernández-Orallo, Jan Arne Telle |
IJCAI | 2 |
| 2022 | Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI Benchmarks (Extended Abstract)abstractWe present a framework for analysing the impact of AI on occupations. This framework maps 59 generic tasks from different occupational datasets to 14 cognitive abilities and these to a comprehensive list of 328 AI benchmarks used to evaluate research intensity in AI. The use of cognitive abilities as an intermediate layer allows for an identification of potential AI exposure for tasks for which AI applications have not been explicitly programmed. We provide insights into the abilities through which AI is most likely to affect jobs, and we show how some of the abilities where AI research is currently very intense are linked to tasks with comparatively limited labour input in the labour markets of advanced economies. Songül Tolan, Annarosa Pesole, Fernando Martínez-Plumed, Enrique Fernández-Macías, José Hernández-Orallo, Emilia Gómez |
IJCAI | 5 |
| 2022 | Heterogeneity Breaks the Game: Evaluating Cooperation-Competition with Multisets of Agents
Yue Zhao 0023, José Hernández-Orallo |
ECML/PKDD (4) | 2 |
| 2021 | Muppets: Multipurpose Table Segmentation
Gust Verbruggen, Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Luc De Raedt |
IDA | 4 |
| 2021 | Think Big, Teach Small: Do Language Models Distil Occam's Razor?abstractLarge language models have recently shown a remarkable ability for few-shot learning, including patterns of algorithmic nature. However, it is still an open question to determine what kind of patterns these models can capture and how many examples they need in their prompts. We frame this question as a teaching problem with strong priors, and study whether language models can identify simple algorithmic concepts from small witness sets. In particular, we explore how several GPT architectures, program induction systems and humans perform in terms of the complexity of the concept and the number of additional examples, and how much their behaviour differs. This first joint analysis of language models and machine teaching can address key questions for artificial intelligence and machine learning, such as whether some strong priors, and Occam’s razor in particular, can be distilled from data, making learning from a few examples possible. Gonzalo Jaimovitch-López, David Castellano Falcón, Cèsar Ferri, José Hernández-Orallo |
NeurIPS | 4 |
| 2021 | Optimal Teaching Curricula with Compositional Simplicity Priors
Manuel Garcia-Piqueras, José Hernández-Orallo |
ECML/PKDD (1) | 2 |
| 2021 | Making sense of sensory inputabstractThis paper attempts to answer a central question in unsupervised learning: what does it mean to “make sense” of a sensory sequence? In our formalization, making sense involves constructing a symbolic causal theory that both explains the sensory sequence and also satisfies a set of unity conditions. The unity conditions insist that the constituents of the causal theory – objects, properties, and laws – must be integrated into a coherent whole. On our account, making sense of sensory input is a type of program synthesis, but it is unsupervised program synthesis. Our second contribution is a computer implementation, the Apperception Engine, that was designed to satisfy the above requirements. Our system is able to produce interpretable human-readable causal theories from very small amounts of data, because of the strong inductive bias provided by the unity conditions. A causal theory produced by our system is able to predict future sensor readings, as well as retrodict earlier readings, and impute (fill in the blanks of) missing sensory readings, in any combination. In fact, it is able to do all three tasks simultaneously. We tested the engine in a diverse variety of domains, including cellular automata, rhythms and simple nursery tunes, multi-modal binding problems, occlusion tasks, and sequence induction intelligence tests. In each domain, we test our engine's ability to predict future sensor values, retrodict earlier sensor values, and impute missing sensory data. The Apperception Engine performs well in all these domains, significantly out-performing neural net baselines. We note in particular that in the sequence induction intelligence tests, our system achieved human-level performance. This is notable because our system is not a bespoke system designed specifically to solve intelligence tests, but a general-purpose system that was designed to make sense of any sensory sequence. Richard Evans 0001, José Hernández-Orallo, Johannes Welbl, Pushmeet Kohli, Marek J. Sergot |
Artif. Intell. | 2 |
| 2021 | Missing the missing values: The ugly duckling of fairness in machine learningabstractNowadays, there is an increasing concern in machine learning about the causes underlying unfair decision making, that is, algorithmic decisions discriminating some groups over others, especially with groups that are defined over protected attributes, such as gender, race and nationality. Missing values are one frequent manifestation of all these latent causes: protected groups are more reluctant to give information that could be used against them, sensitive information for some groups can be erased by human operators, or data acquisition may simply be less complete and systematic for minority groups. However, most recent techniques, libraries and experimental results dealing with fairness in machine learning have simply ignored missing data. In this paper, we present the first comprehensive analysis of the relation between missing values and algorithmic fairness for machine learning: (1) we analyse the sources of missing data and bias, mapping the common causes, (2) we find that rows containing missing values are usually fairer than the rest, which should discourage the consideration of missing values as the uncomfortable ugly data that different techniques and libraries for handling algorithmic bias get rid of at the first occasion, (3) we study the trade-off between performance and fairness when the rows with missing values are used (either because the technique deals with them directly or by imputation methods), and (4) we show that the sensitivity of six different machine-learning techniques to missing values is usually low, which reinforces the view that the rows with missing data contribute more to fairness through the other, nonmissing, attributes. We end the paper with a series of recommended procedures about what to do with missing data when aiming for fair decision making. Fernando Martínez-Plumed, Cèsar Ferri, David Nieves, José Hernández-Orallo |
Int. J. Intell. Syst. | 4 |
| 2021 | Measuring the Occupational Impact of AI: Tasks, Cognitive Abilities and AI BenchmarksabstractIn this paper we develop a framework for analysing the impact of Artificial Intelligence (AI) on occupations. This framework maps 59 generic tasks from worker surveys and an occupational database to 14 cognitive abilities (that we extract from the cognitive science literature) and these to a comprehensive list of 328 AI benchmarks used to evaluate research intensity across a broad range of different AI areas. The use of cognitive abilities as an intermediate layer, instead of mapping work tasks to AI benchmarks directly, allows for an identification of potential AI exposure for tasks for which AI applications have not been explicitly created. An application of our framework to occupational databases gives insights into the abilities through which AI is most likely to affect jobs and allows for a ranking of occupations with respect to AI exposure. Moreover, we show that some jobs that were not known to be affected by previous waves of automation may now be subject to higher AI exposure. Finally, we find that some of the abilities where AI research is currently very intense are linked to tasks with comparatively limited labour input in the labour markets of advanced economies (e.g., visual and auditory processing using deep learning, and sensorimotor interaction through (deep) reinforcement learning). This article appears in the special track on AI and Society. Songül Tolan, Annarosa Pesole, Fernando Martínez-Plumed, Enrique Fernández-Macías, José Hernández-Orallo, Emilia Gómez |
J. Artif. Intell. Res. | 5 |
| 2021 | AUTOMAT[R]IX: learning simple matrix pipelinesabstractAbstract Matrices are a very common way of representing and working with data in data science and artificial intelligence. Writing a small snippet of code to make a simple matrix transformation is frequently frustrating, especially for those people without an extensive programming expertise. We present AUTOMATIX, a system that is able to induce R program snippets from a single (and possibly partial) matrix transformation example provided by the user. Our learning algorithm is able to induce the correct matrix pipeline snippet by composing primitives from a library. Because of the intractable search space—exponential on the size of the library and the number of primitives to be combined in the snippet, we speed up the process with (1) a typed system that excludes all combinations of primitives with inconsistent mapping between input and output matrix dimensions, and (2) a probabilistic model to estimate the probability of each sequence of primitives from their frequency of use and a text hint provided by the user. We validate AUTOMATIX with a set of real programming queries involving matrices from Stack Overflow, showing that we can learn the transformations efficiently, from just one partial example. Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo |
Mach. Learn. | 3 |
| 2021 | CRISP-DM Twenty Years Later: From Data Mining Processes to Data Science TrajectoriesabstractCRISP-DM(CRoss-Industry Standard Process for Data Mining) has its origins in the second half of the nineties and is thus about two decades old. According to many surveys and user polls it is still the de facto standard for developing data mining and knowledge discovery projects. However, undoubtedly the field has moved on considerably in twenty years, with data science now the leading term being favoured over data mining. In this paper we investigate whether, and in what contexts, CRISP-DM is still fit for purpose for data science projects. We argue that if the project is goal-directed and process-driven the process model view still largely holds. On the other hand, when data science projects become more exploratory the paths that the project can take become more varied, and a more flexible model is called for. We suggest what the outlines of such a trajectory-based model might look like and how it can be used to categorise data science projects (goal-directed, exploratory or data management). We examine seven real-life exemplars where exploratory activities play an important role and compare them against 51 use cases extracted from the NIST Big Data Public Working Group. We anticipate this categorisation can help project planning in terms of time and cost characteristics. Fernando Martínez-Plumed, Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Meelis Kull, Nicolas Lachiche, María José Ramírez-Quintana, Peter A. Flach |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Does AI Qualify for the Job?: A Bidirectional Model Mapping Labour and AI IntensitiesabstractIn this paper we present a setting for examining the relation be-tween the distribution of research intensity in AI research and the relevance for a range of work tasks (and occupations) in current and simulated scenarios. We perform a mapping between labourand AI using a set of cognitive abilities as an intermediate layer. This setting favours a two-way interpretation to analyse (1) what impact current or simulated AI research activity has or would have on labour-related tasks and occupations, and (2) what areas of AI research activity would be responsible for a desired or undesired effect on specific labour tasks and occupations. Concretely, in our analysis we map 59 generic labour-related tasks from several worker surveys and databases to 14 cognitive abilities from the cognitive science literature, and these to a comprehensive list of 328 AI benchmarks used to evaluate progress in AI techniques. We provide this model and its implementation as a tool for simulations. We also show the effectiveness of our setting with some illustrative examples. Fernando Martínez-Plumed, Songül Tolan, Annarosa Pesole, José Hernández-Orallo, Enrique Fernández-Macías, Emilia Gómez |
AIES | 4 |
| 2020 | Family and Prejudice: A Behavioural Taxonomy of Machine Learning TechniquesabstractOne classical way of characterising the rich range of machine learning techniques is by defining 'families', according to their formulation and learning strategy (e.g., neural networks, Bayesian methods, etc.).However, this taxonomy of learning techniques does not consider the extent to which models built with techniques from the same or different family agree on their outputs, especially when their predictions have to extrapolate in sparse zones where insufficient training data was available.In this paper we present a new taxonomy of machine learning techniques for classification, where families are clustered according to their degree of (dis)agreement in behaviour considering both dense and sparse zones, using Cohen's kappa statistic.To this end, we use a representative collection of datasets and learning techniques.We finally validate the taxonomy by performing a number of experiments for technique selection.We show that ranking techniques by only following prejudice -the reputation they have for other problems-is worse than selecting techniques based on family diversity. Raül Fabra-Boluda, Cèsar Ferri, Fernando Martínez-Plumed, José Hernández-Orallo, María José Ramírez-Quintana |
ECAI | 4 |
| 2020 | Finite and Confident Teaching in Expectation: Sampling from Infinite Concept ClassesabstractWe investigate the teaching of infinite concept classes through the effect of the learning prior (which is used by the learner to derive posteriors giving preference of some concepts over others and by the teacher to devise the teaching examples) and the sampling prior (which determines how the concepts are sampled from the class). We analyse two important classes: Turing machines and finite-state machines. We derive bounds for the teaching dimension when the learning prior is derived from a complexity measure (Kolmogorov complexity and minimal number of states respectively) and analyse the sampling distributions that lead to finite expected teaching dimensions. The learning prior goes beyond a complexity or preference choice when we use it to increase the confidence of identification, expressed as a posterior, which increases as more examples are given. We highlight the existing trade-off between three elements: the bound on teaching dimension, the representativeness of the sample and the certainty of the identification. This has implications for the understanding of what teaching from rich concept classes to machines (and humans) entails. José Hernández-Orallo, Jan Arne Telle |
ECAI | 1 |
| 2020 | AI Paradigms and AI Safety: Mapping Artefacts and Techniques to Safety IssuesabstractAI safety often analyses a risk or safety issue, such as interruptibility, under a particular AI paradigm, such as reinforcement learning.But what is an AI paradigm and how does it affect the understanding and implications of the safety issue?Is AI safety research covering the most representative paradigms and the right combinations of paradigms with safety issues?Will current research directions in AI safety be able to anticipate more capable and powerful systems yet to come?In this paper we analyse these questions, introducing a distinction between two types of paradigms in AI: artefacts and techniques.We then use experimental data of research and media documents from AI Topics, an official publication of the AAAI, to examine how safety research is distributed across artefacts and techniques.We observe that AI safety research is not sufficiently anticipatory, and is heavily weighted towards certain research paradigms.We identify a need for AI safety to be more explicit about the artefacts and techniques for which a particular issue may be applicable, in order to identify gaps and cover a broader range of issues. José Hernández-Orallo, Fernando Martínez-Plumed, Shahar Avin, Jess Whittlestone, Seán Ó hÉigeartaigh |
ECAI | 1 |
| 2020 | Tracking AI: The Capability Is (Not) NearabstractAI is an area of strategic importance with potential to be a key driver of economic development and with a wide range of potential social\nimplications. In order to assess present and future impact, there is a need to analyse what AI can (and will) achieve. But, what is AI capable of? This question is as crucial as elusive, as AI is progressing in ways that are open-ended about the techniques and resources AI can operate with. The truth is that whenever a task is solved, researchers find increasingly challenging to extrapolate whether this task can be reproduced, even when only a few things change: the data, the domain knowledge, the level of uncertainty, the (hyper)parameters, the techniques, the team, the compute, etc. In the end, we would like to infer whether a good result (or a breakthrough) in task A transfers to a similar good result in task B. This extrapolation is precisely what the notion of capability, borrowed from psychology, tries to answer. However, we lack the tools, and the data, to do similarly in AI. Benchmarks, competitions and challenges are behind much of the recent progress in AI, especially in machine learning (ML) [10], but the dynamics of rushing breakthroughs at the expense of massive data, compute, specialisation, etc., has led to a more complex AI landscape, in terms of what can be achieved and how. As a result, policy makers and other stakeholders have no way of assessing what AI systems can do today and in the future. This does not mean that we must disregard or understate the valuable information that is provided by a plethora of benchmarks. On the contrary, the analysis of the progress of AI must be based on data-grounded evidence, relying on finding and testing hypotheses through the computational analysis of big amounts of shared data [6], using open data science tools [11]. But this analysis must be abstracted from tasks to capabilities, for the purposes of integration3 and evaluation [8]. In this paper, we identify a series of problems to track and understand what AI is capable of, surveying some previous initiatives. We present the AIcollaboratory, a data-driven framework to collect\nand explore data about AI results, progress and ultimately capabilities, being developed in the context of AI WATCH, the European\nCommission (EC) knowledge service to monitor the development, uptake and impact of AI in Europe4. We close the paper with some\nchallenges for the community emerging around the collaboratory. Fernando Martínez-Plumed, José Hernández-Orallo, Emilia Gómez |
ECAI | 2 |
| 2020 | Learning alternative ways of performing a task
David Nieves, María José Ramírez-Quintana, Carlos Monserrat Aranda, Cèsar Ferri, José Hernández-Orallo |
Expert Syst. Appl. | 5 |
| 2020 | Dual Indicators to Analyze AI Benchmarks: Difficulty, Discrimination, Ability, and GeneralityabstractWith the purpose of better analyzing the result of artificial intelligence (AI) benchmarks, we present two indicators on the side of the AI problems, difficulty and discrimination, and two indicators on the side of the AI systems, ability and generality. The first three are adapted from psychometric models in item response theory (IRT), whereas generality is defined as a new metric that evaluates whether an agent is consistently good at easy problems and bad at difficult ones. We illustrate how these key indicators give us more insight on the results of two popular benchmarks in AI, the Arcade Learning Environment (Atari 2600 games) and the General Video Game AI competition, and we include some guidelines to estimate and interpret these indicators for other AI benchmarks and competitions. Fernando Martínez-Plumed, José Hernández-Orallo |
IEEE Trans. Games | 2 |
| 2019 | AI Extenders: The Ethical and Societal Implications of Humans Cognitively Extended by AIabstractHumans and AI systems are usually portrayed as separate systems that we need to align in values and goals. However, there is a great deal of AI technology found in non-autonomous systems that are used as cognitive tools by humans. Under the extended mind thesis, the functional contributions of these tools become as essential to our cognition as our brains. But AI can take cognitive extension towards totally new capabilities, posing new philosophical, ethical and technical challenges. To analyse these challenges better, we define and place AI extenders in a continuum between fully-externalized systems, loosely coupled with humans, and fully internalized processes, with operations ultimately performed by the brain, making the tool redundant. We dissect the landscape of cognitive capabilities that can foreseeably be extended by AI and examine their ethical implications.We suggest that cognitive extenders using AI be treated as distinct from other cognitive enhancers by all relevant stakeholders, including developers, policy makers, and human users. José Hernández-Orallo, Karina Vold |
AIES | 1 |
| 2019 | Automated Data Transformation with Inductive Programming and Dynamic Background Knowledge
Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, María José Ramírez-Quintana, Susumu Katayama |
ECML/PKDD (3) | 3 |
| 2019 | BK-ADAPT: Dynamic Background Knowledge for Automating Data Transformation
Lidia Contreras Ochando, Cèsar Ferri, José Hernández-Orallo, Fernando Martínez-Plumed, María José Ramírez-Quintana, Susumu Katayama |
ECML/PKDD (3) | 3 |
| 2019 | Item response theory in AI: Analysing machine learning classifiers at the instance level
Fernando Martínez-Plumed, Ricardo B. C. Prudêncio, Adolfo Martínez Usó, José Hernández-Orallo |
Artif. Intell. | 4 |
| 2019 | Setting decision thresholds when operating conditions are uncertainabstractThe quality of the decisions made by a machine learning model depends on the data and the operating conditions during deployment. Often, operating conditions such as class distribution and misclassification costs have changed during the time since the model was trained and evaluated. When deploying a binary classifier that outputs scores, once we know the new class distribution and the new cost ratio between false positives and false negatives, there are several methods in the literature to help us choose an appropriate threshold for the classifier’s scores. However, on many occasions, the information that we have about this operating condition is uncertain . Previous work has considered ranges or distributions of operating conditions during deployment, with expected costs being calculated for ranges or intervals, but still the decision for each point is made as if the operating condition were certain. The implications of this assumption have received limited attention: a threshold choice that is best suited without uncertainty may be suboptimal under uncertainty. In this paper we analyse the effect of operating condition uncertainty on the expected loss for different threshold choice methods, both theoretically and experimentally. We model uncertainty as a second conditional distribution over the actual operation condition and study it theoretically in such a way that minimum and maximum uncertainty are both seen as special cases of this general formulation. This is complemented by a thorough experimental analysis investigating how different learning algorithms behave for a range of datasets according to the threshold choice method and the uncertainty level. Cèsar Ferri, José Hernández-Orallo, Peter A. Flach |
Data Min. Knowl. Discov. | 2 |
| 2019 | AI Generality and Spearman's Law of Diminishing ReturnsabstractMany areas of AI today use benchmarks and competitions with larger and wider sets of tasks. This tries to deter AI systems (and research effort) from specialising to a single task, and encourage them to be prepared to solve previously unseen tasks. It is unclear, however, whether the methods with best performance are actually those that are most general and, in perspective, whether the trend moves towards more general AI systems. This question has a striking similarity with the analysis of the so-called positive manifold and general factors in the area of human intelligence. In this paper, we first show how the existence of a manifold (positive average pairwise task correlation) can also be analysed in AI, and how this relates to the notion of agent generality, from the individual and the populational points of view. From the populational perspective, we analyse the following question: is this manifold correlation higher for the most or for the least able group of agents? We contrast this analysis with one of the most controversial issues in human intelligence research, the so-called Spearman's Law of Diminishing Returns (SLODR), which basically states that the relevance of a general factor diminishes for most able human groups. We perform two empirical studies on these issues in AI. We analyse the results of the 2015 general video game AI (GVGAI) competition, with games as tasks and "controllers" as agents, and the results of a synthetic setting, with modified elementary cellular automata (ECA) rules as tasks and simple interactive programs as agents. In both cases, we see that SLODR doesnot appear. The data, and the use of just two scenarios, does not clearly support the reverse either, a Universal Law of Augmenting Returns (ULOAR), but calls for more experiments on this question. José Hernández-Orallo |
J. Artif. Intell. Res. | 1 |
| 2019 | The teaching size: computable teachers and learners for universal languages
Jan Arne Telle, José Hernández-Orallo, Cèsar Ferri |
Mach. Learn. | 2 |
| 2018 | The Facets of Artificial Intelligence: A Framework to Track the Evolution of AIabstractWe present nine facets for the analysis of the past and future evolution of AI. Each facet has also a set of edges that can summarise different trends and contours in AI. With them, we first conduct a quantitative analysis using the information from two decades of AAAI/IJCAI conferences and around 50 years of documents from AI topics, an official database from the AAAI, illustrated by several plots. We then perform a qualitative analysis using the facets and edges, locating AI systems in the intelligence landscape and the discipline as a whole. This analytical framework provides a more structured and systematic way of looking at the shape and boundaries of AI. Fernando Martínez-Plumed, Bao Sheng Loe, Peter A. Flach, Seán Ó hÉigeartaigh, Karina Vold, José Hernández-Orallo |
IJCAI | 6 |
| 2017 | Computer Models Solving Intelligence Test Problems: Progress and Implications (Extended Abstract)abstractWhile some computational models of intelligence test problems were proposed throughout the second half of the XXth century, in the first years of the XXIst century we have seen an increasing number of computer systems being able to score well on particular intelligence test tasks. However, despitethis increasing trend there has been no general account of all these works in terms of how theyrelate to each other and what their real achievements are. In this paper, we provide some insighton these issues by giving a comprehensive account of about thirty computer models, from the 1960sto nowadays, and their relationships, focussing on the range of intelligence test tasks they address, thepurpose of the models, how general or specialised these models are, the AI techniques they use in eachcase, their comparison with human performance, and their evaluation of item difficulty. José Hernández-Orallo, Fernando Martínez-Plumed, Ute Schmid, Michael Siebers, David L. Dowe |
IJCAI | 1 |
| 2016 | Is Spearman's Law of Diminishing Returns (SLODR) Meaningful for Artificial Agents?abstractThe progress of artificial intelligence is reaching a point that some research questions that were only relevant for human and other animal agents are becoming relevant for artificial agents as well. One of those questions comes from human intelligence research and is known as Spearman's Law of Diminishing Returns (SLODR). Charles Spearman, the father of factor analysis and the g factor (a dominant factor explaining most of the variance in cognitive tests for human populations), observed that when the analysis was restricted to the subpopulation of most able subjects, the relevance of this dominant factor diminished, as if the power of general intelligence were saturated or not fully used by the most able individuals. In about a century, there have been numerous theoretical explanations and experiments to confirm or reject Spearman's hypothesis. However, all of them have been based on human or animal populations. In this paper, we analyse for the first time whether the SLODR makes sense for artificial agents and what its role should be in the analysis of general-purpose AI. We use a synthetic scenario based on modified elementary cellular automata (ECA) where the ECA rules work as tasks and the population of agents is generated with an agent policy language. Different slices of the population by ability and of the tasks by difficulty are analysed, showing that SLODR does not really appear. Indeed, even if very slightly, we find the reverse, i.e., that more correlation takes place for more able subpopulations, what we conjecture as the Universal Law of Augmenting Returns (ULOAR). José Hernández-Orallo |
ECAI | 1 |
| 2016 | Making Sense of Item Response Theory in Machine LearningabstractItem response theory (IRT) is widely used to measure latent abilities of subjects (specially for educational testing) based on their responses to items with different levels of difficulty. The adaptation of IRT has been recently suggested as a novel perspective for a better understanding of the results of machine learning experiments and, by extension, other artificial intelligence experiments. For instance, IRT suits classification tasks perfectly, where instances correspond to items and classifiers correspond to subjects. By adopting IRT, item (i.e., instance) characteristic curves can be estimated using logistic models, for which several parameters characterise each dataset instance: difficulty, discrimination and guessing. IRT looks promising for the analysis of instance hardness, noise, classifier dominances, etc. However, some caveats have been found when trying to interpret the IRT parameters in a machine learning setting, especially when we include some artificial classifiers in the pool of classifiers to be evaluated: the optimal and pessimal classifiers, a random classifier and the majority and minority classifiers. In this paper we perform a series of experiments with a range of datasets and classification methods to fully understand how IRT works and what their parameters really mean in the context of machine learning. This better understanding will hopefully pave the way to a myriad of potential applications in machine learning and artificial intelligence. Fernando Martínez-Plumed, Ricardo B. C. Prudêncio, Adolfo Martínez Usó, José Hernández-Orallo |
ECAI | 4 |
| 2016 | Computer models solving intelligence test problems: Progress and implicationsabstractWhile some computational models of intelligence test problems were proposed throughout the second half of the XXth century, in the first years of the XXIst century we have seen an increasing number of computer systems being able to score well on particular intelligence test tasks. However, despite this increasing trend there has been no general account of all these works in terms of how they relate to each other and what their real achievements are. Also, there is poor understanding about what intelligence tests measure in machines, whether they are useful to evaluate AI systems, whether they are really challenging problems, and whether they are useful to understand (human) intelligence. In this paper, we provide some insight on these issues, in the form of nine specific questions, by giving a comprehensive account of about thirty computer models, from the 1960s to nowadays, and their relationships, focussing on the range of intelligence test tasks they address, the purpose of the models, how general or specialised these models are, the AI techniques they use in each case, their comparison with human performance, and their evaluation of item difficulty. As a conclusion, these tests and the computer models attempting them show that AI is still lacking general techniques to deal with a variety of problems at the same time. Nonetheless, a renewed attention on these problems and a more careful understanding of what intelligence tests offer for AI may help build new bridges between psychometrics, cognitive science, and AI; and may motivate new kinds of problem repositories. José Hernández-Orallo, Fernando Martínez-Plumed, Ute Schmid, Michael Siebers, David L. Dowe |
Artif. Intell. | 1 |
| 2016 | Binarised regression tasks: methods and evaluation metrics
José Hernández-Orallo, Cèsar Ferri, Nicolas Lachiche, Adolfo Martínez Usó, María José Ramírez-Quintana |
Data Min. Knowl. Discov. | 1 |
| 2015 | Multidimensional Prediction Models When the Resolution Context Changes
Adolfo Martínez Usó, José Hernández-Orallo |
ECML/PKDD (2) | 2 |
| 2015 | On environment difficulty and discriminating power
José Hernández-Orallo |
Auton. Agents Multi Agent Syst. | 1 |
| 2014 | A Knowledge Growth and Consolidation Framework for Lifelong Machine Learning SystemsabstractA more effective vision of machine learning systems entails tools that are able to improve task after task and to reuse the patterns and knowledge that are acquired previously for future tasks. This incremental, long-life view of machine learning goes beyond most of state-of-the-art machine learning techniques that learn throw-away models. In this paper we present a long-life knowledge acquisition, evaluation and consolidation framework that is designed to work with any rule-based machine learning or inductive inference engine and integrate it into a long-life learner. In order to do that we work over the graph of working memory rules and introduce several topological metrics over it from which we derive an oblivion criterion to drop useless rules from working memory and a consolidation process to promote the rules to the knowledge base. We evaluate the framework on a series of tasks in a chess rule learning domain. Fernando Martínez-Plumed, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
ICMLA | 3 |
| 2014 | Bridging the Gap between Distance and GeneralizationabstractDistance‐based and generalization‐based methods are two families of artificial intelligence techniques that have been successfully used over a wide range of real‐world problems. In the first case, general algorithms can be applied to any data representation by just changing the distance. The metric space sets the search and learning space, which is generally instance‐oriented. In the second case, models can be obtained for a given pattern language, which can be comprehensible. The generality‐ordered space sets the search and learning space, which is generally model‐oriented. However, the concepts of distance and generalization clash in many different ways, especially when knowledge representation is complex (e.g., structured data). This work establishes a framework where these two fields can be integrated in a consistent way. We introduce the concept of distance‐based generalization, which connects all the generalized examples in such a way that all of them are reachable inside the generalization by using straight paths in the metric space. This makes the metric space and the generality‐ordered space coherent (or even dual). Additionally, we also introduce a definition of minimal distance‐based generalization that can be seen as the first formulation of the Minimum Description Length (MDL)/Minimum Message Length (MML) principle in terms of a distance function. We instantiate and develop the framework for the most common data representations and distances, where we show that consistent instances can be found for numerical data, nominal data, sets, lists, tuples, graphs, first‐order atoms, and clauses. As a result, general learning methods that integrate the best from distance‐based and generalization‐based methods can be defined and adapted to any specific problem by appropriately choosing the distance, the pattern language and the generalization operator. Vicent Estruch, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
Comput. Intell. | 3 |
| 2014 | Aggregative quantification for regression
Antonio Bella, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
Data Min. Knowl. Discov. | 3 |
| 2014 | Probabilistic Reframing for Cost-Sensitive RegressionabstractCommon-day applications of predictive models usually involve the full use of the available contextual information. When the operating context changes, one may fine-tune the by-default (incontextual) prediction or may even abstain from predicting a value (a reject). Global reframing solutions, where the same function is applied to adapt the estimated outputs to a new cost context, are possible solutions here. An alternative approach, which has not been studied in a comprehensive way for regression in the knowledge discovery and data mining literature, is the use of a local (e.g., probabilistic) reframing approach, where decisions are made according to the estimated output and a reliability, confidence, or probability estimation. In this article, we advocate for a simple two-parameter (mean and variance) approach, working with a normal conditional probability density. Given the conditional mean produced by any regression technique, we develop lightweight “enrichment” methods that produce good estimates of the conditional variance, which are used by the probabilistic (local) reframing methods. We apply these methods to some very common families of cost-sensitive problems, such as optimal predictions in (auction) bids, asymmetric loss scenarios, and rejection rules. José Hernández-Orallo |
ACM Trans. Knowl. Discov. Data | 1 |
| 2013 | On the effect of calibration in classifier combination
Antonio Bella, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
Appl. Intell. | 3 |
| 2013 | ROC curves in cost space
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri |
Mach. Learn. | 1 |
| 2013 | ROC curves for regression
José Hernández-Orallo |
Pattern Recognit. | 1 |
| 2012 | A unified view of performance metrics: translating threshold choice into expected classification loss
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri |
J. Mach. Learn. Res. | 1 |
| 2011 | A Coherent Interpretation of AUC as a Measure of Aggregated Classification Performance
Peter A. Flach, José Hernández-Orallo, Cèsar Ferri |
ICML | 2 |
| 2011 | Brier Curves: a New Cost-Based Visualisation of Classifier Performance
José Hernández-Orallo, Peter A. Flach, Cèsar Ferri |
ICML | 1 |
| 2010 | Quantification via Probability EstimatorsabstractQuantification is the name given to a novel machine learning task which deals with correctly estimating the number of elements of one class in a set of examples. The output of a quantifier is a real value, since training instances are the same as a classification problem, a natural approach is to train a classifier and to derive a quantifier from it. Some previous works have shown that just classifying the instances and counting the examples belonging to the class of interest classify count typically yields bad quantifiers, especially when the class distribution may vary between training and test. Hence, adjusted versions of classify count have been developed by using modified thresholds. However, previous works have explicitly discarded (without a deep analysis) any possible approach based on the probability estimations of the classifier. In this paper, we present a method based on averaging the probability estimations of a classifier with a very simple scaling that does perform reasonably well, showing that probability estimators for quantification capture a richer view of the problem than methods based on a threshold. Antonio Bella, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
ICDM | 3 |
| 2010 | Data Mining Strategies for CRM Negotiation Prescription Problems
Antonio Bella, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
IEA/AIE (1) | 3 |
| 2010 | Measuring universal intelligence: Towards an anytime intelligence test
José Hernández-Orallo, David L. Dowe |
Artif. Intell. | 1 |
| 2009 | Similarity-Binning Averaging: A Generalisation of Binning Calibration
Antonio Bella, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
IDEAL | 3 |
| 2009 | An Instantiation of Hierarchical Distance-Based Conceptual Clustering for Propositional Learning
Ana Funes, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
PAKDD | 3 |
| 2009 | An experimental comparison of performance measures for classification
Cèsar Ferri, José Hernández-Orallo, R. Modroiu |
Pattern Recognit. Lett. | 2 |
| 2008 | Hierarchical Distance-Based Conceptual Clustering
Ana Maria Funes, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
ECML/PKDD (1) | 3 |
| 2007 | Joint Cutoff Probabilistic Estimation Using Simulation: A Mailing Campaign Application
Antonio Bella, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
IDEAL | 3 |
| 2006 | Minimal Distance-Based Generalisation Operators for First-Order Objects
Vicent Estruch, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
ILP | 3 |
| 2005 | Knowledge acquisition through machine learning: minimising expert's effortabstractMachine learning can be applied to solve the knowledge acquisition bottleneck in many areas where an expert makes predictions to single cases, such as diagnosis, estimation, etc. The idea is to query the expert with as many cases as possible and get their answers. With this data we train a machine learning model which mimics the expert's behaviour. This is just a simple application of a modelling technique known as "mimetism", which has many other applications. This "soft" approach to knowledge acquisition has many advantages: any machine learning technique can be used, the expert must only answer simple questions (cases) and we can combine the decisions of several experts easily. However, one problem of this approach is that we do not know in advance how many cases we will need to ask in order to get a good model which is accurate wrt. the expert's knowledge. Obviously, as more data is labelled by the expert better results are obtained. However, asking thousands of cases to the expert is usually impractical. In this paper, we analyse the behaviour of knowledge acquisition through mimetic learning according to two factors: accuracy and comprehensibility of the resulting model and we devise a method to compute the minimum number of cases that we need to ask the expert to attain a certain quality level. Ricardo Blanco-Vega, José Hernández-Orallo, María José Ramírez-Quintana |
ICMLA | 2 |
| 2005 | Distance Based Generalisation
Vicent Estruch, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
ILP | 3 |
| 2004 | Analysing the Trade-Off Between Comprehensibility and Accuracy in Mimetic Models
Ricardo Blanco-Vega, José Hernández-Orallo, María José Ramírez-Quintana |
Discovery Science | 2 |
| 2004 | Delegating classifiersabstractA sensible use of classifiers must be based on the estimated reliability of their predictions. A cautious classifier would delegate the difficult or uncertain predictions to other, possibly more specialised, classifiers. In this paper we analyse and develop this idea of delegating classifiers in a systematic way. First, we design a two-step scenario where a first classifier chooses which examples to classify and delegates the difficult examples to train a second classifier. Secondly, we present an iterated scenario involving an arbitrary number of chained classifiers. We compare these scenarios to classical ensemble methods, such as bagging and boosting. We show experimentally that our approach is not far behind these methods in terms of accuracy, but with several advantages: (i) improved efficiency, since each classifier learns from fewer examples than the previous one; (ii) improved comprehensibility, since each classification derives from a single classifier; and (iii) the possibility to simplify the overall multi-classifier by removing the parts that lead to delegation. Cèsar Ferri, Peter A. Flach, José Hernández-Orallo |
ICML | 3 |
| 2003 | Improving the AUC of Probabilistic Estimation Trees
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo |
ECML | 3 |
| 2003 | Volume under the ROC Surface for Multi-class Problems
Cèsar Ferri, José Hernández-Orallo, Miguel A. Salido |
ECML | 2 |
| 2002 | From Ensemble Methods to Comprehensible Models
Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
Discovery Science | 2 |
| 2002 | Learning Decision Trees Using the Area Under the ROC Curve
Cèsar Ferri, Peter A. Flach, José Hernández-Orallo |
ICML | 3 |
| 2002 | SMILES: A Multi-purpose Learning System
Vicent Estruch, Cèsar Ferri, José Hernández-Orallo, María José Ramírez-Quintana |
JELIA | 3 |
| 2001 | Predictive Software
José Hernández-Orallo, María José Ramírez-Quintana |
Autom. Softw. Eng. | 1 |
| 2000 | Software as Learning: Quality Factors and Life-Cycle Revised
José Hernández-Orallo, María José Ramírez-Quintana |
FASE | 1 |
| 2000 | Truth from Trash. How Learning Makes Sense by Chris Thornton
José Hernández-Orallo |
Artif. Intell. | 1 |
| 2000 | Constructive reinforcement learningabstractThis paper presents an operative measure of reinforcement for constructive learning methods, i.e., eager learning methods using highly expressible (or universal) representation languages. These evaluation tools allow a further insight in the study of the growth of knowledge, theory revision, and abduction. The final approach is based on an apportionment of credit wrt the “course” that the evidence makes through the learned theory. Our measure of reinforcement is shown to be justified by cross-validation and by the connection with other successful evaluation criteria, like the minimum description length principle. Finally, the relation with the classical view of reinforcement is studied, where the actions of an intelligent system can be rewarded or penalized, and we discuss whether this should affect our distribution of reinforcement. The most important result of this paper is that the way we distribute reinforcement into knowledge results in a rated ontology, instead of a single prior distribution. Therefore, this detailed information can be exploited for guiding the space search of inductive learning algorithms. Likewise, knowledge revision may be done to the part of the theory which is not justified by the evidence. ©2000 John Wiley & Sons, Inc. José Hernández-Orallo |
Int. J. Intell. Syst. | 1 |