VLDB 2026 Research / reviewers in the wild / expert
Mario A. Muñoz
dblp:29/5811 · also Mario Andrés Muñoz, Mario Andrés Muñoz Acosta
· DBLP profile ↗
30ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0002-7254-2808ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 11 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalisation of Automated Algorithm Selection in Black-Box Optimisation: The Role of Algorithm Portfolio and Learning Model
Behzad Moradi, Mario A. Muñoz, Michael Kirley |
EvoApplications (1) | 2 |
| 2026 | Constructing Streams of Optimization Instances for Benchmarking Algorithm Selection and Configuration Approaches in Streaming Scenarios
Margherita Battistotti, Manuel López-Ibáñez 0001, Kate Smith-Miles, Julia Handl, Mario A. Muñoz |
GECCO | 5 |
| 2026 | An Efficient Hybrid Racing Method for Portfolio ConfigurationabstractPortfolio configurators tune an algorithm's parameters to produce a portfolio of parameterisations, with complementary strengths on different problem instances. However, as the portfolio size increases, the marginal contribution of each additional configuration decreases, making portfolio configuration computationally expensive. We propose a hybrid racing method that combines statistical testing (to early eliminate poorly performing configurations) and successive halving (which progressively allocates more resources to promising configurations by discarding fractions of the configuration candidate set in rounds). We connect greedy portfolio configuration approaches to the theory of submodular function optimisation, which implies an approximation ratio of 1 − 1/e for the problem in idealised form. We experimentally compare the hybrid method with pure statistical racing and pure successive halving, on algorithm configuration benchmarks from the literature. We show that the proposed hybrid racing method achieves portfolios with performance matching those configured on the full evaluation data, while usually using roughly 5% of the runs of a full evaluation, and that the method is more efficient than pure statistical racing or successive halving for portfolio configuration. Our experiments demonstrate that adaptive elimination strategies can significantly improve the efficiency of portfolio configuration methods. Anthony Rasulo, Manuel López-Ibáñez 0001, Julia Handl, Mario A. Muñoz, Kate Smith-Miles |
GECCO | 4 |
| 2026 | Instance space of clustering validation measuresabstractAbstract In clustering, selecting the most appropriate partitioning of a dataset is often guided by clustering validity indexes. However, with numerous competing indexes each with its own strengths and weaknesses, choosing the right one can be challenging and may significantly affect clustering outcomes. Despite their widespread use, limited research has explored how index performance varies across problem types, with traditional benchmarks focusing on ground-truth properties that cannot be known prior to clustering. Instance Space Analysis (ISA) is a visual meta-learning methodology that provides tools to examine the relationship between problem features and algorithmic performance. This study presents the first application of ISA to clustering validity indexes, analysing the behaviour of nine indexes across a diverse set of 18,351 synthetic benchmark datasets and eight clustering algorithms. The results uncover distinct performance patterns and offer data-driven guidance for selecting appropriate indexes based on measurable problem characteristics, providing insights into the relative strengths and weaknesses of commonly used indexes. Connor Simpson, Mario A. Muñoz, Ricardo J. G. B. Campello |
Data Min. Knowl. Discov. | 2 |
| 2025 | Tracing the Interactions of Modular CMA-ES Configurations Across Problem LandscapesabstractThis paper leverages the recently introduced concept of algorithm footprints to investigate the interplay between algorithm configurations and problem characteristics. Performance footprints are calculated for six modular variants of the CMA-ES algorithm (modCMA), evaluated on 24 benchmark problems from the BBOB suite, across two-dimensional settings: 5-dimensional and 30-dimensional. These footprints provide insights into why different configurations of the same algorithm exhibit varying performance and identify the problem features influencing these outcomes. Our analysis uncovers shared behavioral patterns across configurations due to common interactions with problem properties, as well as distinct behaviors on the same problem driven by differing problem features. The results demonstrate the effectiveness of algorithm footprints in enhancing interpretability and guiding configuration choices. Ana Nikolikj, Mario A. Muñoz, Eva Tuba, Tome Eftimov |
CEC | 2 |
| 2025 | Why We Should be Benchmarking Evolutionary Algorithms on Neural Network Training Tasks
Katherine M. Malan, Mario A. Muñoz |
GECCO | 2 |
| 2025 | Instance Space Analysis of the Capacitated Vehicle Routing Problem with Mixture Discriminant AnalysisabstractIn this paper, we attempt a deeper understanding of the relative performance of two state-of-the-art metaheuristic solvers for the capacitated vehicle routing problem (CVRP). To this end, we employ a novel CVRP instance generator to expand the set of CVRP instances used to assess heuristics. This generator modifies existing problem instances using the outliers of node clusters to produce relevant new CVRP instances. We consider a large number of features to characterise each problem instance, and propose to use mixture discriminant analysis (MDA) to obtain both a low dimensional representation of the instance space and a classifier of algorithm performance. MDA has not been previously used in instance space analysis, and as a supervised dimension reduction method has the advantage that the tasks of dimension reduction and classification are handled in a unified framework (rather than two separate steps). The resulting predictive models perform as well as more complex classifiers that involve more tuning parameters and are computationally more expensive (like support vector machines). Our analysis highlights that the performance comparison between the two CVRP metaheuristics is nuanced and the best algorithm depends on the time budget, as well as certain key characteristics of the problem instance. Danielle Notice, Hamed Soleimani, Nicos G. Pavlidis, Ahmed Kheiri, Mario A. Muñoz |
GECCO | 5 |
| 2025 | ISA3: a 3-dimensional expansion of instance space analysis
Connor Simpson, Mario A. Muñoz, Sevvandi Kandanaarachchi, Ricardo J. G. B. Campello |
Mach. Learn. | 2 |
| 2024 | Characterising harmful data sources when constructing multi-fidelity surrogate models
Nicolau Andrés-Thió, Mario A. Muñoz, Kate Smith-Miles |
Artif. Intell. | 2 |
| 2024 | Optimal selection of benchmarking datasets for unbiased machine learning algorithm evaluation
João Luiz Junho Pereira, Kate Smith-Miles, Mario A. Muñoz, Ana Carolina Lorena |
Data Min. Knowl. Discov. | 3 |
| 2023 | Algorithm Instance Footprint: Separating Easily Solvable and Challenging Problem InstancesabstractIn black-box optimization, it is essential to understand why an algorithm instance works on a set of problem instances while failing on others and provide explanations of its behavior. We propose a methodology for formulating an algorithm instance footprint that consists of a set of problem instances that are easy to be solved and a set of problem instances that are difficult to be solved, for an algorithm instance. This behavior of the algorithm instance is further linked to the landscape properties of the problem instances to provide explanations of which properties make some problem instances easy or challenging. The proposed methodology uses meta-representations that embed the landscape properties of the problem instances and the performance of the algorithm into the same vector space. These meta-representations are obtained by training a supervised machine learning regression model for algorithm performance prediction and applying model explainability techniques to assess the importance of the landscape features to the performance predictions. Next, deterministic clustering of the meta-representations demonstrates that using them captures algorithm performance across the space and detects regions of poor and good algorithm performance, together with an explanation of which landscape properties are leading to it. Ana Nikolikj, Saso Dzeroski, Mario A. Muñoz, Carola Doerr, Peter Korosec, Tome Eftimov |
GECCO | 3 |
| 2023 | An Instance Space Analysis of Constrained Multiobjective Optimization ProblemsabstractConstrained multiobjective optimization problems (CMOPs) are generally more challenging than unconstrained problems. This in part can be attributed to the infeasible region generated by the constraint functions, the interaction between constraints and objectives, or both. In this article, we explore the relationship between the performance of constrained multiobjective evolutionary algorithms (CMOEAs) and the instance characteristics of CMOP using instance space analysis (ISA). To do this, we extend recent work on Landscape Analysis features for characterizing CMOPs. Specifically, we introduce new features to describe the multiobjective-violation landscape, formed by the interaction between constraint violation and multiobjective fitness. The detailed evaluation of the algorithm footprints, spanning eight CMOP benchmark suites and 15 CMOEAs, demonstrates that ISA effectively captures the strength and weakness of the CMOEAs. We conclude that two characteristics, the isolation of nondominate set and the correlation between constraints and objectives evolvability, have the greatest impact on algorithm performance. However, the current benchmarks problems lack of diversity to represent the real-world problems and to fully reveal the efficacy of CMOEAs evaluated. Hanan Alsouly, Michael Kirley, Mario A. Muñoz |
IEEE Trans. Evol. Comput. | 3 |
| 2023 | Instance Space Analysis of Search-Based Software TestingabstractSearch-based software testing (SBST) is now a mature area, with numerous techniques developed to tackle the challenging task of software testing. SBST techniques have shown promising results and have been successfully applied in the industry to automatically generate test cases for large and complex software systems. Their effectiveness, however, has been shown to be problem dependent. In this paper, we revisit the problem of objective performance evaluation of SBST techniques in light of recent methodological advances – in the form of Instance Space Analysis (ISA) – enabling the strengths and weaknesses of SBST techniques to be visualised and assessed across the broadest possible space of problem instances (software classes) from common benchmark datasets. We identify features of SBST problems that explain why a particular instance is hard for an SBST technique, reveal areas of hard and easy problems in the instance space of existing benchmark datasets, and identify the strengths and weaknesses of state-of-the-art SBST techniques. In addition, we examine the diversity and quality of common benchmark datasets used in experimental evaluations. Neelofar, Kate Smith-Miles, Mario A. Muñoz, Aldeida Aleti |
IEEE Trans. Software Eng. | 3 |
| 2022 | Bifidelity Surrogate Modelling: Showcasing the Need for New Test InstancesabstractIn recent years, multifidelity expensive black-box (Mf-EBB) methods have received increasing attention due to their strong applicability to industrial design problems. The challenge, however, is that knowledge of the relationship between decisions and objective values is limited to a small set of sample observations of variable quality. In the field of Mf-EBB, a problem instance consists of an expensive yet accurate source of information, and one or more cheap yet less accurate sources of information. The field aims to provide techniques either to accurately explain how decisions affect design outcome, or to find the best decisions to optimise design outcomes. Many techniques that use surrogate models have been developed to provide solutions to both aims. Only in recent years, however, have researchers begun to explore the conditions under which these new techniques are reliable, often focusing on problems with a single low-fidelity function, known as bifidelity expensive black-box (Bf-EBB) problems. This study extends the existing Bf-EBB test instances found in the literature, as well as the features used to determine when the low-fidelity information source should be used. A literature test suite is constructed and augmented with new instances to demonstrate the potentially misleading results that could be reached using only the instances currently found in the literature, and to expose the criticality of a more heterogeneous test suite for algorithm assessment. Addressing the shortcomings of the existing literature, a new set of features is presented, as well as a new instance creation procedure, and a study of their impact on algorithm assessment is conducted. The low-fidelity information source is shown to be valuable if it is often locally accurate, even when its overall accuracy is relatively low. This contradicts the existing literature guidelines, which indicate the low-fidelity information is only useful if it has a high overall accuracy. History: Accepted by Antonio Frangioni, Area Editor for Design & Analysis of Algorithms – Continuous. Funding: This work was supported by Australian Research Council [Grant IC200100009] for the ARC Training Centre in Optimisation Technologies, Integrated Methodologies and Applications (OPTIMA), and the University of Melbourne Research Computing Services and Petascale Campus Initiative. N. Andrés-Thió is also supported by a Research Training Program scholarship from the University of Melbourne. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplementary Information [ https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.1217 ] or is available from the IJOC GitHub software repository ( https://github.com/INFORMSJoC ) at [ http://dx.doi.org/10.5281/zenodo.6578060 ]. Nicolau Andrés-Thió, Mario A. Muñoz, Kate Smith-Miles |
INFORMS J. Comput. | 2 |
| 2022 | Analyzing randomness effects on the reliability of exploratory landscape analysis
Mario A. Muñoz, Michael Kirley, Kate Smith-Miles |
Nat. Comput. | 1 |
| 2022 | Informing Multiobjective Optimization Benchmark Construction Through Instance Space AnalysisabstractThe role of carefully constructed benchmark suites in algorithm design and testing is critical. Within the continuous multiobjective optimization domain, existing suites include the general purpose ZDT, DTLZ, and WFG suites, and more recent ones specifically designed to explore the impacts of a particular problem characteristic. However, the relationship between existing suites is not clear, and the field would benefit from a “stock-take” assessment. This article investigates the coverage of current continuous multiobjective suites using the instance space analysis (ISA) methodology. Exploratory landscape analysis is used to measure critical features of each problem suite. Thereafter, we generate a 2-D visualization of the existing problem instances by locating them in the instance space, assessing their diversity, and identifying whether there are sparse areas of value to fill with new problem instances. Our findings show that the current suites are restricted in diversity when representing the entire problem instance space. We propose and evaluate three problem construction methods: 1) problem tuning; 2) toolkit hybridization; and 3) new function injection. Problem tuning is shown to generate problems surrounding existing instances, while hybridization creates problems falling between existing suites. Furthermore, utilizing the insights afforded by ISA, we show how problem features can be identified to inform the creation of new functions which fill gaps toward the boundaries of the instance space. Estefania Yap, Mario A. Muñoz, Kate Smith-Miles |
IEEE Trans. Evol. Comput. | 2 |
| 2021 | An Instance Space Analysis of Regression ProblemsabstractThe quest for greater insights into algorithm strengths and weaknesses, as revealed when studying algorithm performance on large collections of test problems, is supported by interactive visual analytics tools. A recent advance is Instance Space Analysis, which presents a visualization of the space occupied by the test datasets, and the performance of algorithms across the instance space. The strengths and weaknesses of algorithms can be visually assessed, and the adequacy of the test datasets can be scrutinized through visual analytics. This article presents the first Instance Space Analysis of regression problems in Machine Learning, considering the performance of 14 popular algorithms on 4,855 test datasets from a variety of sources. The two-dimensional instance space is defined by measurable characteristics of regression problems, selected from over 26 candidate features. It enables the similarities and differences between test instances to be visualized, along with the predictive performance of regression algorithms across the entire instance space. The purpose of creating this framework for visual analysis of an instance space is twofold: one may assess the capability and suitability of various regression techniques; meanwhile the bias, diversity, and level of difficulty of the regression problems popularly used by the community can be visually revealed. This article shows the applicability of the created regression instance space to provide insights into the strengths and weaknesses of regression algorithms, and the opportunities to diversify the benchmark test instances to support greater insights. Mario A. Muñoz, Matheus R. Leal, Kate Smith-Miles, Ana Carolina Lorena, Gisele L. Pappa, Rômulo Madureira Rodrigues |
ACM Trans. Knowl. Discov. Data | 1 |
| 2020 | Instance Space Analysis of Combinatorial Multi-objective Optimization ProblemsabstractIn recent years, there has been a continuous stream of development in evolutionary multi-objective optimization (EMO) algorithms. The large quantity of existing algorithms introduces difficulty in selecting suitable algorithms for a given problem instance. In this paper, we perform instance space analysis on discrete multi-objective optimization problems (MOPs) for the first time under three different conditions. We create visualizations of the relationship between problem instances and algorithm performance for instance features previously identified using decision trees, as well an independent feature selection. The suitability of these features in discriminating between algorithm performance and understanding strengths and weaknesses is investigated. Furthermore, we explore the impact of various definitions of “good” performance. The visualization of the instance space provides an alternative method of algorithm discrimination by showing clusters of instances where algorithms perform well across the instance space. We validate the suitability of existing features and identify opportunities for future development. Estefania Yap, Mario A. Muñoz, Kate Smith-Miles, Arnaud Liefooghe |
CEC | 2 |
| 2020 | On normalization and algorithm selection for unsupervised outlier detection
Sevvandi Kandanaarachchi, Mario A. Muñoz, Rob J. Hyndman, Kate Smith-Miles |
Data Min. Knowl. Discov. | 2 |
| 2020 | Generating New Space-Filling Test Instances for Continuous Black-Box OptimizationabstractThis article presents a method to generate diverse and challenging new test instances for continuous black-box optimization. Each instance is represented as a feature vector of exploratory landscape analysis measures. By projecting the features into a two-dimensional instance space, the location of existing test instances can be visualized, and their similarities and differences revealed. New instances are generated through genetic programming which evolves functions with controllable characteristics. Convergence to selected target points in the instance space is used to drive the evolutionary process, such that the new instances span the entire space more comprehensively. We demonstrate the method by generating two-dimensional functions to visualize its success, and ten-dimensional functions to test its scalability. We show that the method can recreate existing test functions when target points are co-located with existing functions, and can generate new functions with entirely different characteristics when target points are located in empty regions of the instance space. Moreover, we test the effectiveness of three state-of-the-art algorithms on the new set of instances. The results demonstrate that the new set is not only more diverse than a well-known benchmark set, but also more challenging for the tested algorithms. Hence, the method opens up a new avenue for developing test instances with controllable characteristics, necessary to expose the strengths and weaknesses of algorithms, and drive algorithm development. Mario A. Muñoz, Kate Smith-Miles |
Evol. Comput. | 1 |
| 2018 | Instance spaces for machine learning classification
Mario A. Muñoz, Laura Villanova, Davaatseren Baatar, Kate Smith-Miles |
Mach. Learn. | 1 |
| 2017 | Performance Analysis of Continuous Black-Box Optimization Algorithms via Footprints in Instance SpaceabstractThis article presents a method for the objective assessment of an algorithm's strengths and weaknesses. Instead of examining the performance of only one or more algorithms on a benchmark set, or generating custom problems that maximize the performance difference between two algorithms, our method quantifies both the nature of the test instances and the algorithm performance. Our aim is to gather information about possible phase transitions in performance, that is, the points in which a small change in problem structure produces algorithm failure. The method is based on the accurate estimation and characterization of the algorithm footprints, that is, the regions of instance space in which good or exceptional performance is expected from an algorithm. A footprint can be estimated for each algorithm and for the overall portfolio. Therefore, we select a set of features to generate a common instance space, which we validate by constructing a sufficiently accurate prediction model. We characterize the footprints by their area and density. Our method identifies complementary performance between algorithms, quantifies the common features of hard problems, and locates regions where a phase transition may lie. Mario A. Muñoz, Kate Smith-Miles |
Evol. Comput. | 1 |
| 2016 | ICARUS: Identification of complementary algorithms by uncovered setsabstractSince there is no single best performing algorithm for all problems, an algorithm portfolio would leverage the strengths of complementary algorithms to achieve the best performance. In this paper, we present and evaluate a new technique for designing algorithm portfolios for continuous black-box optimization problems, based on social choice and voting theory concepts. Our technique, which we call ICARUS, models the portfolio design task as an election, in which each problem `votes' for a subset of preferred algorithms guided by a performance metric such as the number of fitness evaluations. The resulting `uncovered set' of algorithms forms the portfolio. We demonstrate the efficacy of ICARUS using a suite of state-of-the-art evolutionary algorithms and benchmark continuous optimization problems. Our analysis confirms that ICARUS creates an algorithm portfolio where the expected performance is superior to a manually constructed portfolio. Mario A. Muñoz, Michael Kirley |
CEC | 1 |
| 2015 | Effects of function translation and dimensionality reduction on landscape analysisabstractExploratory Landscape Analysis (ELA) measures have been shown to predict algorithm performance; hence, they are being applied on critical tasks such as automatic algorithm selection and problem generation. This paper provides a cautionary examination on their use in black-box continuous optimization. We explore the effect that translations have on the measures, when the cost function is defined within a bound-constrained region. Furthermore, we examine the robustness of the neighborhood structure after dimensionality reduction. The results demonstrate that a measure may transition abruptly due a translation. Therefore, we should not generalize the measures of an instance nor report average values of a measure as belonging to the generating function. Moreover, dimensionality reduction could alter the neighborhood structure, such that the regions corresponding to significantly different functions overlap. Mario A. Muñoz, Kate Smith-Miles |
CEC | 1 |
| 2015 | Algorithm selection for black-box continuous optimization problems: A survey on methods and challenges
Mario A. Muñoz, Yuan Sun 0003, Michael Kirley, Saman K. Halgamuge |
Inf. Sci. | 1 |
| 2015 | Exploratory Landscape Analysis of Continuous Space Optimization Problems Using Information ContentabstractData-driven analysis methods, such as the information content of a fitness sequence, characterize a discrete fitness landscape by quantifying its smoothness, ruggedness, or neutrality. However, enhancements to the information content method are required when dealing with continuous fitness landscapes. One typically employed adaptation is to sample the fitness landscape using random walks with variable step size. However, this adaptation has significant limitations: random walks may produce biased samples, and uncertainty is added because the distance between observations is not accounted for. In this paper, we introduce a robust information content-based method for continuous fitness landscapes, which addresses these limitations. Our method generates four measures related to the landscape features. Numerical simulations are used to evaluate the efficacy of the proposed method. We calculate the Pearson correlation coefficient between the new measures and other well-known exploratory landscape analysis measures. Significant differences on the measures between benchmark functions are subsequently identified. We then demonstrate the practical relevance of the new measures using them as class predictors on a machine learning model, which classifies the benchmark functions into five groups. Classification accuracy greater than 90% was obtained, with computational costs bounded between 1% and 10% of the maximum function evaluation budget. The results demonstrate that our method provides relevant information, at a low cost in terms of function evaluations. Mario A. Muñoz, Michael Kirley, Saman K. Halgamuge |
IEEE Trans. Evol. Comput. | 1 |
| 2012 | Landscape characterization of numerical optimization problems using biased scattered dataabstractThe characterization of optimization problems over continuous parameter spaces plays an important role in optimization. A form of “fitness landscape” analysis is often carried out to describe the problem space in terms of modality, smoothness and variable separability. The outcomes of this analysis can then be used as a measure of problem difficulty and to predict the behaviour of a given algorithm. However, the metric value estimates of the landscape characterization are dependent upon the representation scheme adopted and the sampling method used. Consequently, the development of a complete classification of problem structure and complexity has proven to be challenging. In this paper, we continue this line of research. We present a methodology for the characterization of two dimensional numerical optimization problems. In our approach, data extracted during the search process is analyzed and the dependency of the results to the nominated sampling method are corrected. We show via computational simulations that the calculated metric values using our approach are consistent with the results from random experiments. As such, this study provides a first step towards the on-line calculation of fitness landscape characterization metrics and the development of empirical performance models of search algorithms. Advances in these areas would provide answers to the algorithm selection and portfolio configuration problems. Mario A. Muñoz, Michael Kirley, Saman K. Halgamuge |
IEEE Congress on Evolutionary Computation | 1 |
| 2012 | A Meta-learning Prediction Model of Algorithm Performance for Continuous Optimization Problems
Mario A. Muñoz, Michael Kirley, Saman K. Halgamuge |
PPSN (1) | 1 |
| 2010 | Simplifying the Bacteria Foraging Optimization AlgorithmabstractThe Bacterial Foraging Optimization Algorithm is a swarm intelligence technique which models the individual and group foraging policies of the E. Coli bacteria as a distributed optimization process. The algorithm is structurally complex due to its nested loop architecture and includes several parameters whose selection deeply influences the result. This paper presents some modifications to the original algorithm that simplifies the algorithm structure, and the inclusion of best member information into the search strategy, which improves the performance. The results on several benchmarks show reasonable performance in most tests and a considerable improvement in some complex functions. Also, with the use of the T-Test we were able to confirm that the performance enhancement is statistically significant. Mario A. Muñoz, Saman K. Halgamuge, Wilfredo Alfonso, Eduardo Caicedo Bravo |
IEEE Congress on Evolutionary Computation | 1 |
| 2009 | An artificial beehive algorithm for continuous optimizationabstractThis paper presents an artificial beehive algorithm for optimization in continuous search spaces based on a model aimed at individual bee behavior. The algorithm defines a set of behavioral rules for each agent to determine what kind of actions must be carried out. Also, the algorithm proposed includes some adaptations not considered in the biological model to increase the performance in the search for better solutions. To compare the performance of the algorithm with other swarm-based Techniques, we conducted statistical analyses by using the so-called t test. This comparison is done with several common benchmark functions. © 2009 Wiley Periodicals, Inc. Mario A. Muñoz, Jesús A. López, Eduardo Caicedo Bravo |
Int. J. Intell. Syst. | 1 |