EDBT 2026 Demo / reviewers in the wild / expert
Joaquin Vanschoren
dblp:85/5045
· DBLP profile ↗
49ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0001-7044-9805ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 3 first-author · 22 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Machine Learning for Unsupervised Tabular TasksabstractAbstract In this work, we present Learning to Learn with Optimal Transport for Unsupervised Scenarios (LOTUS), a simple yet effective method to perform model selection for multiple unsupervised machine learning (ML) tasks such as outlier detection and clustering. Our intuition behind this work is that a machine learning pipeline will perform well in a new dataset if it previously worked well on datasets with a similar underlying data distribution. We use Optimal Transport distances to find this similarity between unlabeled tabular datasets and recommend machine learning pipelines with one unified single method on two downstream unsupervised tasks: outlier detection and clustering. We present the effectiveness of our approach with experiments against strong baselines and show that LOTUS is a very promising first step toward model selection for multiple unsupervised ML tasks. Prabhant Singh, Pieter Gijsbers, Elif Ceren Gok Yildirim, Murat Onur Yildirim, Joaquin Vanschoren |
Mach. Learn. | 5 |
| 2026 | Coresets are more than replay: a data-centric view of continual learningabstractAbstract Continual Learning (CL) addresses the challenge of enabling models to adapt to evolving data and tasks while retaining previously acquired knowledge. The main challenge in this paradigm is catastrophic forgetting, where models lose prior knowledge upon learning new tasks. While much of the CL literature has focused on model-centric innovations, we argue for the substantial potential of a data-centric approach, specifically by revisiting the ‘learn-it-all’ assumption prevalent in current CL paradigms. This paper presents the first empirical study systematically evaluating the impact of different coreset methods for training samples in combination with CL methods. Unlike conventional uses of coreset selection, which restrict its role to populating a small rehearsal buffer, we present the first empirical study showing that coreset methods can be applied directly to the full training set, substantially reducing the amount of data needed for learning. Our results reveal that training on carefully selected coreset substantially enhances incremental accuracy while reducing computational overhead. We demonstrate that this performance improvement is primarily driven by an improved stability-plasticity trade-off, largely attributable to the enhanced retention of prior knowledge. This study not only highlights the significant benefits of data-centric strategies in CL but also advocates for a shift in research focus towards these approaches to stimulate and guide future advancements in the field. Code is available at https://github.com/ElifCerenGokYildirim/Coreset-CL . Elif Ceren Gok Yildirim, Murat Onur Yildirim, Joaquin Vanschoren |
Neural Comput. Appl. | 3 |
| 2025 | Score Matching on Large Geometric Graphs for Cosmology Generation
Diana-Alexandra Onutu, Yue Zhao 0016, Joaquin Vanschoren, Vlado Menkovski |
DS | 3 |
| 2025 | Unsupervised Meta-Learning via In-Context LearningabstractUnsupervised meta-learning aims to learn feature representations from unsupervised datasets that can transfer to downstream tasks with limited labeled data.
In this paper, we propose a novel approach to unsupervised meta-learning that leverages the generalization abilities of in-context learning observed in transformer architectures. Our method reframes meta-learning as a sequence modeling problem, enabling the transformer encoder to learn task context from support images and utilize it to predict query images.
At the core of our approach lies the creation of diverse tasks generated using a combination of data augmentations and a mixing strategy that challenges the model during training while fostering generalization to unseen tasks at test time.
Experimental results on benchmark datasets showcase the superiority of our approach over existing unsupervised meta-learning baselines, establishing it as the new state-of-the-art. Remarkably, our method achieves competitive results with supervised and self-supervised approaches, underscoring its efficacy in leveraging generalization over memorization. Anna Vettoruzzo, Lorenzo Braccaioli, Joaquin Vanschoren, Marlena Nowaczyk |
ICLR | 3 |
| 2025 | CrypticBio: A Large Multimodal Dataset for Visually Confusing SpeciesabstractWe present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models in the context of biodiversity applications. Visually confusing or cryptic species are groups of two or more taxa that are nearly indistinguishable based on visual characteristics alone. While much existing work addresses taxonomic identification in a broad sense, datasets that directly address the morphological confusion of cryptic species are small, manually curated, and target only a single taxon. Thus, the challenge of identifying such subtle differences in a wide range of taxa remains unaddressed. Curated from real-world trends in species misidentification among community annotators of iNaturalist, CrypticBio contains 52K unique cryptic groups spanning 67K species represented in 166 million images. Records in the dataset include research-grade image annotations—scientific, multicultural, and multilingual species terminology, hierarchical taxonomy, spatiotemporal context, and associated cryptic groups. To facilitate easy subset curation from CrypticBio, we provide an open-source pipeline, CrypticBio-Curate. The multimodal design of the dataset provides complementary cues such as spatiotemporal context that support the identification of cryptic species. To highlight the importance of the dataset, we benchmark a suite of state-of-the-art foundation models across CrypticBio subsets of common, unseen, endangered, and invasive species, and demonstrate the substantial impact of spatiotemporal context on vision-language zero-shot learning for cryptic species. By introducing CrypticBio, we aim to catalyze progress toward real-world-ready fine-grained species classification models for biodiversity monitoring capable of handling the nuanced challenges of species ambiguity. The data and the code are publicly available in the project website https://georgianagmanolache.github.io/crypticbio. Georgiana Manolache, Gerard Schouten, Joaquin Vanschoren |
NeurIPS | 3 |
| 2024 | HyTAS: A Hyperspectral Image Transformer Architecture Search Benchmark and Analysis
Fangqin Zhou, Mert Kilickaya, Joaquin Vanschoren, Ran Piao |
ECCV (32) | 3 |
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 41 |
| 2024 | MALIBO: Meta-learning for Likelihood-free Bayesian OptimizationabstractBayesian optimization (BO) is a popular method to optimize costly black-box functions, and meta-learning has emerged as a way to leverage knowledge from related tasks to optimize new tasks faster. However, existing meta-learning methods for BO rely on surrogate models that are not scalable or are sensitive to varying input scales and noise types across tasks. Moreover, they often overlook the uncertainty associated with task similarity, leading to unreliable task adaptation when a new task differs significantly or has not been sufficiently explored yet. We propose a novel meta-learning BO approach that bypasses the surrogate model and directly learns the utility of queries across tasks. It explicitly models task uncertainty and includes an auxiliary model to enable robust adaptation to new tasks. Extensive experiments show that our method achieves strong performance and outperforms multiple meta-learning BO methods across various benchmarks. Jiarong Pan, Stefan Falkner, Felix Berkenkamp, Joaquin Vanschoren |
ICML | 4 |
| 2024 | Croissant: A Metadata Format for ML-Ready DatasetsabstractData is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise. Mubashara Akhtar, Omar Benjelloun, Costanza Conforti, Luca Foschini 0002, Joan Giner-Miguelez, Pieter Gijsbers, Sujata S. Goswami, Nitisha Jain, Michalis Karamousadakis, Michael Kuchnik, Satyapriya Krishna, Sylvain Lesage, Quentin Lhoest, Pierre Marcenac, Manil Maskey, Peter Mattson, Luis Oala, Hamidah Oderinwale, Pierre Ruyssen, Tim Santos, Rajat Shinde, Elena Simperl, Arjun Suresh, Goeffry Thomas, Slava Tykhonov, Joaquin Vanschoren, Susheel Varma, Jos van der Velde, Steffen Vogler, Carole-Jean Wu |
NeurIPS | 26 |
| 2024 | Better trees: an empirical study on hyperparameter tuning of classification decision tree induction algorithms
Rafael Gomes Mantovani, Tomás Horváth, André Luis Debiaso Rossi, Ricardo Cerri, Sylvio Barbon Junior, Joaquin Vanschoren, André C. P. L. F. de Carvalho |
Data Min. Knowl. Discov. | 6 |
| 2024 | Can Fairness be Automated? Guidelines and Opportunities for Fairness-aware AutoMLabstractThe field of automated machine learning (AutoML) introduces techniques that automate parts of the development of machine learning (ML) systems, accelerating the process and reducing barriers for novices. However, decisions derived from ML models can reproduce, amplify, or even introduce unfairness in our societies, causing harm to (groups of) individuals. In response, researchers have started to propose AutoML systems that jointly optimize fairness and predictive performance to mitigate fairness-related harm. However, fairness is a complex and inherently interdisciplinary subject, and solely posing it as an optimization problem can have adverse side effects. With this work, we aim to raise awareness among developers of AutoML systems about such limitations of fairness-aware AutoML, while also calling attention to the potential of AutoML as a tool for fairness research. We present a comprehensive overview of different ways in which fairness-related harm can arise and the ensuing implications for the design of fairness-aware AutoML. We conclude that while fairness cannot be automated, fairness-aware AutoML can play an important role in the toolbox of ML practitioners. We highlight several open technical challenges for future work in this direction. Additionally, we advocate for the creation of more user-centered assistive systems designed to tackle challenges encountered in fairness work. This article appears in the AI & Society track. Hilde J. P. Weerts, Florian Pfisterer, Matthias Feurer 0001, Katharina Eggensperger, Edward Bergman, Noor H. Awad, Joaquin Vanschoren, Mykola Pechenizkiy, Bernd Bischl, Frank Hutter |
J. Artif. Intell. Res. | 7 |
| 2024 | AMLB: an AutoML BenchmarkabstractComparing different AutoML frameworks is notoriously challenging and often done incorrectly. We introduce an open and extensible benchmark that follows best practices and avoids common mistakes when comparing AutoML frameworks. We conduct a thorough comparison of 9 well-known AutoML frameworks across 71 classification and 33 regression tasks. The differences between the AutoML frameworks are explored with a multi-faceted analysis, evaluating model accuracy, its trade-offs with inference time, and framework failures. We also use Bradley-Terry trees to discover subsets of tasks where the relative AutoML framework rankings differ. The benchmark comes with an open-source tool that integrates with many AutoML frameworks and automates the empirical evaluation process end-to-end: from framework installation and resource allocation to in-depth evaluation. The benchmark uses public data sets, can be easily extended with other AutoML frameworks and tasks, and has a website with up-to-date results. Pieter Gijsbers, Marcos L. P. Bueno, Stefan Coors, Erin LeDell, Sébastien Poirier, Janek Thomas, Bernd Bischl, Joaquin Vanschoren |
J. Mach. Learn. Res. | 8 |
| 2024 | Towards efficient AutoML: a pipeline synthesis approach leveraging pre-trained transformers for multimodal dataabstractAbstract This paper introduces an Automated Machine Learning (AutoML) framework specifically designed to efficiently synthesize end-to-end multimodal machine learning pipelines. Traditional reliance on the computationally demanding Neural Architecture Search is minimized through the strategic integration of pre-trained transformer models. This innovative approach enables the effective unification of diverse data modalities into high-dimensional embeddings, streamlining the pipeline development process. We leverage an advanced Bayesian Optimization strategy, informed by meta-learning, to facilitate the warm-starting of the pipeline synthesis, thereby enhancing computational efficiency. Our methodology demonstrates its potential to create advanced and custom multimodal pipelines within limited computational resources. Extensive testing across 23 varied multimodal datasets indicates the promise and utility of our framework in diverse scenarios. The results contribute to the ongoing efforts in the AutoML field, suggesting new possibilities for efficiently handling complex multimodal data. This research represents a step towards developing more efficient and versatile tools in multimodal machine learning pipeline development, acknowledging the collaborative and ever-evolving nature of this field. Ambarish Moharil, Joaquin Vanschoren, Prabhant Singh, Damian A. Tamburri |
Mach. Learn. | 2 |
| 2024 | Advances and Challenges in Meta-Learning: A Technical ReviewabstractMeta-learning empowers learning systems with the ability to acquire knowledge from multiple tasks, enabling faster adaptation and generalization to new tasks. This review provides a comprehensive technical overview of meta-learning, emphasizing its importance in real-world applications where data may be scarce or expensive to obtain. The article covers the state-of-the-art meta-learning approaches and explores the relationship between meta-learning and multi-task learning, transfer learning, domain adaptation and generalization, self-supervised learning, personalized federated learning, and continual learning. By highlighting the synergies between these topics and the field of meta-learning, the article demonstrates how advancements in one area can benefit the field as a whole, while avoiding unnecessary duplication of efforts. Additionally, the article delves into advanced meta-learning topics such as learning from complex multi-modal task distributions, unsupervised meta-learning, learning to efficiently adapt to data distribution shifts, and continual meta-learning. Lastly, the article highlights open problems and challenges for future research in the field. By synthesizing the latest research developments, this article provides a thorough understanding of meta-learning and its potential impact on various machine learning applications. We believe that this technical overview will contribute to the advancement of meta-learning and its practical implications in addressing real-world problems. Anna Vettoruzzo, Mohamed-Rafik Bouguelia, Joaquin Vanschoren, Thorsteinn S. Rögnvaldsson, KC Santosh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Locality-Aware Hyperspectral Classification
Fangqin Zhou, Mert Kilickaya, Joaquin Vanschoren |
BMVC | 3 |
| 2023 | AutoML for Outlier Detection with Optimal Transport DistancesabstractAutomated machine learning (AutoML) has been widely researched and adopted for supervised problems, but progress in unsupervised settings has been limited. We propose `"LOTUS", a novel framework to automate outlier detection based on meta-learning. Our premise is that the selection of the optimal outlier detection technique depends on the inherent properties of the data distribution. We leverage optimal transport to find the dataset with the most similar underlying distribution, and then apply the outlier detection techniques that proved to work best for that data distribution. We evaluate the robustness of our framework and find that it outperforms all state-of-the-art automated outlier detection tools. This approach can also be easily generalized to automate other unsupervised settings. Prabhant Singh, Joaquin Vanschoren |
IJCAI | 2 |
| 2023 | Efficient-DASH: Automated Radar Neural Network Design Across Tasks and DatasetsabstractThis work shows the benefit of Neural Architecture Search (NAS) for robust and efficient radar neural network design. Radar-based neural networks have received increasing attention in the development of Advanced Driver-Assistance Systems. However, many state-of-the-art networks are designed for specific tasks and datasets, often disregarding network efficiency, leading to a lack of robust methods for efficient radar network design. This paper leverages a differentiable NAS method called DASH to automate the design of radar neural networks. We evaluate our method across multiple tasks and datasets to assess its robustness. Additionally, we introduce Efficient-DASH, which leverages compound scaling to discover more efficient networks. Our results show that DASH outperforms a state-of-the-art radar network across tasks, while maintaining strong performance when generalizing to other datasets. Efficient-DASH finds architectures that are more efficient than the baseline network, while still reporting performance gains over 4.5%. This demonstrates the potential of NAS for robust and efficient radar neural network design. Thomas Boot, Nicolas Cazin, Willem P. Sanberg, Joaquin Vanschoren |
IV | 4 |
| 2023 | DataPerf: Benchmarks for Data-Centric AI DevelopmentabstractMachine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry. Mark Mazumder, Colby R. Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Frederick Diamos, Gregory Frederick Diamos, Lynn He, Alicia Parrish, Hannah Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett D. Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller 0001, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung-Levy, Newsha Ardalani, Praveen K. Paritosh, Ce Zhang 0001, James Zou 0001, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi |
NeurIPS | 35 |
| 2023 | An Analysis of Evolutionary Migration Models for Multi-Objective, Multi-Fidelity AutomlabstractMethods have been proposed to maximize the efficiency of Automated Machine Learning (AutoML) systems while simplifying the search for solutions. Multi-fidelity approaches have been shown as an alternative to achieve this since they are straightforward to implement and reduce the computational cost when modeling large datasets. However, they are not suited to shifting data distributions, and they remove configurations too fast. Therefore, they must be combined with other techniques. Combining multi-fidelity methods with evolutionary algorithms and island models helps to adapt better and can contribute to maintaining diversity. This paper presents a comparative analysis of 10 network topologies distributed over a generalized island model for AutoML. A dynamic migration model with multi-objective and multi-fidelity evaluation is proposed to reduce the complexity of the tasks. This proposal is compared against state-of-the-art AutoML frameworks. It was found that Hypercube, Grid 2-dim, and Grid 3-dim topologies have the best performance as they maintain a balance in the number of connections. Furthermore, these topologies were shown to be competitive against other frameworks in the state of the art. Israel Campero-Jurado, Joaquin Vanschoren |
SMC | 2 |
| 2023 | Online AutoML: an adaptive AutoML framework for online learning
Bilge Celik, Prabhant Singh, Joaquin Vanschoren |
Mach. Learn. | 3 |
| 2022 | Meta-Album: Multi-domain Meta-Dataset for Few-Shot Image ClassificationabstractWe introduce Meta-Album, an image classification meta-dataset designed to facilitate few-shot learning, transfer learning, meta-learning, among other tasks. It includes 40 open datasets, each having at least 20 classes with 40 examples per class, with verified licences. They stem from diverse domains, such as ecology (fauna and flora), manufacturing (textures, vehicles), human actions, and optical character recognition, featuring various image scales (microscopic, human scales, remote sensing). All datasets are preprocessed, annotated, and formatted uniformly, and come in 3 versions (Micro $\subset$ Mini $\subset$ Extended) to match users’ computational resources. We showcase the utility of the first 30 datasets on few-shot learning problems. The other 10 will be released shortly after. Meta-Album is already more diverse and larger (in number of datasets) than similar efforts, and we are committed to keep enlarging it via a series of competitions. As competitions terminate, their test data are released, thus creating a rolling benchmark, available through OpenML.org. Our website https://meta-album.github.io/ contains the source code of challenge winning methods, baseline methods, data loaders, and instructions for contributing either new datasets or algorithms to our expandable meta-dataset. Dustin Carrión-Ojeda, Sergio Escalera, Isabelle Guyon, Mike Huisman, Felix Mohr, Jan N. van Rijn, Haozhe Sun, Joaquin Vanschoren, Phan Anh Vu |
NeurIPS | 9 |
| 2022 | Meta-features for meta-learningabstracta b s t r a c tMeta-learning is increasingly used to support the recommendation of machine learning algorithms and their configurations.These recommendations are made based on meta-data, consisting of performance evaluations of algorithms and characterizations on prior datasets.These characterizations, also called meta-features, describe properties of the data which are predictive for the performance of machine learning algorithms trained on them.Unfortunately, despite being used in many studies, meta-features are not uniformly described, organized and computed, making many empirical studies irreproducible and hard to compare.This paper aims to deal with this by systematizing and standardizing data characterization measures for classification datasets used in meta-learning.Moreover, it presents an extensive list of meta-features and characterization tools, which can be used as a guide for new practitioners.By identifying particularities and subtle issues related to the characterization measures, this survey points out possible future directions that the development of meta-features for meta-learning can assume. Adriano Rivolli, Luís Paulo F. Garcia, Carlos Soares, Joaquin Vanschoren, André C. P. L. F. de Carvalho |
Knowl. Based Syst. | 4 |
| 2022 | Theory-based habit modeling for enhancing behavior prediction in behavior change support systemsabstractAbstract Psychological theories of habit posit that when a strong habit is formed through behavioral repetition, it can trigger behavior automatically in the same environment. Given the reciprocal relationship between habit and behavior, changing lifestyle behaviors is largely a task of breaking old habits and creating new and healthy ones. Thus, representing users’ habit strengths can be very useful for behavior change support systems, for example, to predict behavior or to decide when an intervention reaches its intended effect. However, habit strength is not directly observable and existing self-report measures are taxing for users. In this paper, building on recent computational models of habit formation, we propose a method to enable intelligent systems to compute habit strength based on observable behavior. The hypothesized advantage of using computed habit strength for behavior prediction was tested using data from two intervention studies on dental behavior change ( $$N = 36$$ N=36 and $$N = 75$$ N=75 ), where we instructed participants to brush their teeth twice a day for three weeks and monitored their behaviors using accelerometers. The results showed that for the task of predicting future brushing behavior, the theory-based model that computed habit strength achieved an accuracy of 68.6% (Study 1) and 76.1% (Study 2), which outperformed the model that relied on self-reported behavioral determinants but showed no advantage over models that relied on past behavior. We discuss the implications of our results for research on behavior change support systems and habit formation. Chao Zhang 0071, Joaquin Vanschoren, Arlette van Wissen, Daniël Lakens, Boris E. R. de Ruyter, Wijnand A. IJsselsteijn |
User Model. User Adapt. Interact. | 2 |
| 2021 | OpenML-Python: an extensible Python API for OpenMLabstractOpenML is an online platform for open science collaboration in machine learning, used to share datasets and results of machine learning experiments. In this paper, we introduce OpenML-Python, a client API for Python, which opens up the OpenML platform for a wide range of Python-based machine learning tools. It provides easy access to all datasets, tasks and experiments on OpenML from within Python. It also provides functionality to conduct machine learning experiments, upload the results to OpenML, and reproduce results which are stored on OpenML. Furthermore, it comes with a scikit-learn extension and an extension mechanism to easily integrate other machine learning libraries written in Python into the OpenML ecosystem. Source code and documentation are available at https://github.com/openml/openml-python/. Matthias Feurer 0001, Jan N. van Rijn, Arlind Kadra, Pieter Gijsbers, Neeratyoy Mallik, Sahithya Ravi, Andreas C. Müller 0001, Joaquin Vanschoren, Frank Hutter |
J. Mach. Learn. Res. | 8 |
| 2021 | Adaptation Strategies for Automated Machine Learning on Evolving DataabstractAutomated Machine Learning (AutoML) systems have been shown to efficiently build good models for new datasets. However, it is often not clear how well they can adapt when the data evolves over time. The main goal of this study is to understand the effect of concept drift on the performance of AutoML methods, and which adaptation strategies can be employed to make them more robust to changes in the underlying data. To that end, we propose 6 concept drift adaptation strategies and evaluate their effectiveness on a variety of AutoML approaches for building machine learning pipelines, including Bayesian optimization, genetic programming, and random search with automated stacking. These are evaluated empirically on real-world and synthetic data streams with different types of concept drift. Based on this analysis, we propose ways to develop more sophisticated and robust AutoML techniques. Bilge Celik, Joaquin Vanschoren |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Towards Scalable Online Machine Learning Collaborations with OpenMLabstractIs massively collaborative machine learning possible? Can we share and organize our collective knowledge of machine learning to solve ever more challenging problems? In a way, yes: as a community, we are already very successful at developing high-quality open-source machine learning libraries, thanks to frictionless collaboration platforms for software development. However, code is only one aspect. The answer is much less clear when we also consider the data that goes into these algorithms and the exact models that are produced. A tremendous amount of work and experience goes into the collection, cleaning, and preprocessing of data and the design, evaluation, and finetuning of models, yet very little of this is shared and organized in a way so that others can easily build on it. Suppose one had a global platform for sharing machine learning datasets, models, and reproducible experiments in a frictionless way so that anybody could chip in at any time to share a good model, add or improve data, or suggest an idea. OpenML is an open-source initiative to create such a platform. It allows anyone to share datasets, machine learning pipelines, and full experiments, organizes all of it online with rich metadata, and enables anyone to reuse and build on them in novel and unexpected ways. All data is open and accessible through APIs, and it is readily integrated into popular machine learning tools to allow easy sharing of models and experiments. This openness also allows a budding ecosystem of automated processes to scale up machine learning further, such as discovering similar datasets, creating systematic benchmarks, or learning from all collected results how to build the best machine learning models and even automatically doing so for any new dataset. We welcome all of you to become a part of it. Joaquin Vanschoren |
Proc. VLDB Endow. | 1 |
| 2020 | Guest editors' introduction to the special issue on Discovery Science
Larisa N. Soldatova, Joaquin Vanschoren |
Mach. Learn. | 2 |
| 2019 | Beyond Bag-of-Concepts: Vectors of Locally Aggregated Concepts
Maarten Grootendorst, Joaquin Vanschoren |
ECML/PKDD (2) | 2 |
| 2019 | A meta-learning recommender system for hyperparameter tuning: Predicting when tuning improves SVM classifiers
Rafael Gomes Mantovani, André Luis Debiaso Rossi, Edesio Alcobaça, Joaquin Vanschoren, André C. P. L. F. de Carvalho |
Inf. Sci. | 4 |
| 2018 | Data Augmentation using Conditional Generative Adversarial Networks for Leaf Counting in Arabidopsis Plants
Yezi Zhu, Marc Aoun, Marcel Krijn, Joaquin Vanschoren |
BMVC | 4 |
| 2018 | Speeding up algorithm selection using average ranking and active testing by introducing runtime
Salisu Mamman Abdulrahman, Pavel Brazdil, Jan N. van Rijn, Joaquin Vanschoren |
Mach. Learn. | 4 |
| 2018 | Meta-QSAR: a large-scale application of meta-learning to drug design and discoveryabstractWe investigate the learning of quantitative structure activity relationships (QSARs) as a case-study of meta-learning. This application area is of the highest societal importance, as it is a key step in the development of new medicines. The standard QSAR learning problem is: given a target (usually a protein) and a set of chemical compounds (small molecules) with associated bioactivities (e.g. inhibition of the target), learn a predictive mapping from molecular representation to activity. Although almost every type of machine learning method has been applied to QSAR learning there is no agreed single best way of learning QSARs, and therefore the problem area is well-suited to meta-learning. We first carried out the most comprehensive ever comparison of machine learning methods for QSAR learning: 18 regression methods, 3 molecular representations, applied to more than 2700 QSAR problems. (These results have been made publicly available on OpenML and represent a valuable resource for testing novel meta-learning methods.) We then investigated the utility of algorithm selection for QSAR problems. We found that this meta-learning approach outperformed the best individual QSAR learning method (random forests using a molecular fingerprint representation) by up to 13%, on average. We conclude that meta-learning outperforms base-learning methods for QSAR learning, and as this investigation is one of the most extensive ever comparisons of base and meta-learning methods ever made, it provides evidence for the general effectiveness of meta-learning over base-learning. Iván Olier, Noureddin Sadawi, G. Richard J. Bickerton, Joaquin Vanschoren, Crina Grosan, Larisa N. Soldatova, Ross D. King |
Mach. Learn. | 4 |
| 2018 | The online performance estimation framework: heterogeneous ensemble learning for data streamsabstractEnsembles of classifiers are among the best performing classifiers available in many data mining applications, including the mining of data streams. Rather than training one classifier, multiple classifiers are trained, and their predictions are combined according to a given voting schedule. An important prerequisite for ensembles to be successful is that the individual models are diverse. One way to vastly increase the diversity among the models is to build an heterogeneous ensemble, comprised of fundamentally different model types. However, most ensembles developed specifically for the dynamic data stream setting rely on only one type of base-level classifier, most often Hoeffding Trees . We study the use of heterogeneous ensembles for data streams. We introduce the Online Performance Estimation framework, which dynamically weights the votes of individual classifiers in an ensemble. Using an internal evaluation on recent training data, it measures how well ensemble members performed on this and dynamically updates their weights. Experiments over a wide range of data streams show performance that is competitive with state of the art ensemble techniques, including Online Bagging and Leveraging Bagging , while being significantly faster. All experimental results from this work are easily reproducible and publicly available online. Jan N. van Rijn, Geoff Holmes 0001, Bernhard Pfahringer, Joaquin Vanschoren |
Mach. Learn. | 4 |
| 2016 | ASlib: A benchmark library for algorithm selection
Bernd Bischl, Pascal Kerschke, Lars Kotthoff, Marius Lindauer, Yuri Malitsky, Alexandre Fréchette, Holger H. Hoos, Frank Hutter, Kevin Leyton-Brown, Kevin Tierney, Joaquin Vanschoren |
Artif. Intell. | 11 |
| 2015 | Who is More Positive in Private? Analyzing Sentiment Differences across Privacy Levels and Demographic Factors in Facebook Chats and PostsabstractUnderstanding users' sentiments in social media is important in many domains, such as marketing and online applications. Is one demographic group inherently different from another? Does a group express the same sentiment both in private and public? How can we compare the sentiments of different groups composed of multiple attributes? In this paper, we take an interdisciplinary approach towards mining the patterns of textual sentiments and metadata. First, we look into several existing hypotheses in social science on the interplay between user characteristics and sentiments, as well as the related evidence in the field of social network data analysis. Second, we present a dataset with unique features (Facebook users' chats and posts in multiple languages) and a procedure to process the data. Third, we test our hypotheses on this dataset and interpret the results. Fourth, under the subgroup-discovery paradigm, we present an approach with two algorithms that generalizes single-attribute testing. This approach provides more detailed insight into the relationships among attributes, and reveals interesting attribute-value combinations with distinct sentiments. Furthermore, it offers novel hypotheses for examination in future studies. Bo Gao 0002, Bettina Berendt, Joaquin Vanschoren |
ASONAM | 3 |
| 2015 | Having a Blast: Meta-Learning and Heterogeneous Ensembles for Data StreamsabstractEnsembles of classifiers are among the best performing classifiers available in many data mining applications. However, most ensembles developed specifically for the dynamic data stream setting rely on only one type of base-level classifier, most often Hoeffding Trees. In this paper, we study the use of heterogeneous ensembles, comprised of fundamentally different model types. Heterogeneous ensembles have proven successful in the classical batch data setting, however they do not easily transfer to the data stream setting. We therefore introduce the Online Performance Estimation framework, which can be used in data stream ensembles to weight the votes of (heterogeneous) ensemble members differently across the stream. Experiments over a wide range of data streams show performance that is competitive with state of the art ensemble techniques, including Online Bagging and Leveraging Bagging. All experimental results from this work are easily reproducible and publicly available on OpenML for further analysis. Jan N. van Rijn, Geoff Holmes 0001, Bernhard Pfahringer, Joaquin Vanschoren |
ICDM | 4 |
| 2015 | Fast Algorithm Selection Using Learning Curves
Jan N. van Rijn, Salisu Mamman Abdulrahman, Pavel Brazdil, Joaquin Vanschoren |
IDA | 4 |
| 2015 | To tune or not to tune: Recommending when to adjust SVM hyper-parameters via meta-learningabstractMany classification algorithms, such as Neural Networks and Support Vector Machines, have a range of hyper-parameters that may strongly affect the predictive performance of the models induced by them. Hence, it is recommended to define the values of these hyper-parameters using optimization techniques. While these techniques usually converge to a good set of values, they typically have a high computational cost, because many candidate sets of values are evaluated during the optimization process. It is often not clear whether this will result in parameter settings that are significantly better than the default settings. When training time is limited, it may help to know when these parameters should definitely be tuned. In this study, we use meta-learning to predict when optimization techniques are expected to lead to models whose predictive performance is better than those obtained by using default parameter settings. Hence, we can choose to employ optimization techniques only when they are expected to improve performance, thus reducing the overall computational cost. We evaluate these meta-learning techniques on more than one hundred data sets. The experimental results show that it is possible to accurately predict when optimization techniques should be used instead of default values suggested by some machine learning libraries. Rafael Gomes Mantovani, André Luis Debiaso Rossi, Joaquin Vanschoren, Bernd Bischl, André C. P. L. F. de Carvalho |
IJCNN | 3 |
| 2015 | Effectiveness of Random Search in SVM hyper-parameter tuningabstractClassification is one of the most common machine learning tasks. SVMs have been frequently applied to this task. In general, the values chosen for the hyper-parameters of SVMs affect the performance of their induced predictive models. Several studies use optimization techniques to find a set of hyper-parameter values that induces classifiers with good predictive performance. This paper investigates the hypothesis that a simple Random Search method is sufficient to adjust the hyper-parameters of SVMs. A set of experiments compared the performance of five tuning techniques: three meta-heuristics commonly used, Random Search and Grid Search. The experimental results show that the predictive performance of models using Random Search is equivalent to those obtained using meta-heuristics and Grid Search, but with a lower computational cost. Rafael Gomes Mantovani, André Luis Debiaso Rossi, Joaquin Vanschoren, Bernd Bischl, André C. P. L. F. de Carvalho |
IJCNN | 3 |
| 2014 | Algorithm Selection on Data Streams
Jan N. van Rijn, Geoff Holmes 0001, Bernhard Pfahringer, Joaquin Vanschoren |
Discovery Science | 4 |
| 2013 | OpenML: A Collaborative Science Platform
Jan N. van Rijn, Bernd Bischl, Luís Torgo, Bo Gao 0002, Venkatesh Umaashankar, Simon Fischer 0001, Patrick Winter, Bernd Wiswedel, Michael R. Berthold, Joaquin Vanschoren |
ECML/PKDD (3) | 10 |
| 2012 | Scientific Workflow Management with ADAMS
Peter Reutemann, Joaquin Vanschoren |
ECML/PKDD (2) | 2 |
| 2012 | MDL-Based Analysis of Time Series at Multiple Time-Scales
Ugo Vespier, Arno J. Knobbe, Siegfried Nijssen, Joaquin Vanschoren |
ECML/PKDD (2) | 4 |
| 2012 | Experiment databases - A new way to share, organize and learn from experimentsabstractThousands of machine learning research papers contain extensive experimental comparisons. However, the details of those experiments are often lost after publication, making it impossible to reuse these experiments in further research, or reproduce them to verify the claims made. In this paper, we present a collaboration framework designed to easily share machine learning experiments with the community, and automatically organize them in public databases. This enables immediate reuse of experiments for subsequent, possibly much broader investigation and offers faster and more thorough analysis based on a large set of varied results. We describe how we designed such an experiment database, currently holding over 650,000 classification experiments, and demonstrate its use by answering a wide range of interesting research questions and by verifying a number of recent studies. Joaquin Vanschoren, Hendrik Blockeel, Bernhard Pfahringer, Geoff Holmes 0001 |
Mach. Learn. | 1 |
| 2011 | Traffic Events Modeling for Structural Health Monitoring
Ugo Vespier, Arno J. Knobbe, Joaquin Vanschoren, Shengfa Miao, Arne Koopman, Bas Obladen, Carlos Bosma |
IDA | 3 |
| 2009 | A Community-Based Platform for Machine Learning Experimentation
Joaquin Vanschoren, Hendrik Blockeel |
ECML/PKDD (2) | 1 |
| 2008 | Organizing the World's Machine Learning Information
Joaquin Vanschoren, Hendrik Blockeel, Bernhard Pfahringer, Geoff Holmes 0001 |
ISoLA | 1 |
| 2008 | Learning from the Past with Experiment Databases
Joaquin Vanschoren, Bernhard Pfahringer, Geoff Holmes 0001 |
PRICAI | 1 |
| 2007 | Experiment Databases: Towards an Improved Experimental Methodology in Machine Learning
Hendrik Blockeel, Joaquin Vanschoren |
PKDD | 2 |