David Martens

dblp:20/3975 · DBLP profile ↗
← Back
38ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 On the definition and detection of cherry-picking in counterfactual explanations
James Hinns, Sofie Goethals, Stephan Van der Veeken, Theodoros Evgeniou, David Martens
Data Min. Knowl. Discov.5
2026 Would a large language model pay extra for a view? Inferring willingness to pay from subjective choices
Manon Reusens, Sofie Goethals, Toon Calders, David Martens
Expert Syst. Appl.4
2025 Exposing Shortcuts in Image Classification by Aggregating Counterfactuals
James Hinns, David Martens
IJCCI (3)2
2025 Tell me a story! Narrative-driven XAI with Large Language Models
David Martens, James Hinns, Camille Dams, Mark Vergouwen, Theodoros Evgeniou
Decis. Support Syst.1
2024 NICE: an algorithm for nearest instance counterfactual explanations
Dieter Brughmans, Pieter Leyman, David Martens
Data Min. Knowl. Discov.3
2024 PreCoF: counterfactual explanations for fairness
Sofie Goethals, David Martens, Toon Calders
Mach. Learn.2
2024 Can metafeatures help improve explanations of prediction models when using behavioral and textual data?
Yanou Ramon, David Martens, Theodoros Evgeniou, Stiene Praet
Mach. Learn.2
2023 The Privacy Issue of Counterfactual Explanations: Explanation Linkage Attacks
abstract
Black-box machine learning models are used in an increasing number of high-stakes domains, and this creates a growing need for Explainable AI (XAI). However, the use of XAI in machine learning introduces privacy risks, which currently remain largely unnoticed. Therefore, we explore the possibility of an explanation linkage attack , which can occur when deploying instance-based strategies to find counterfactual explanations. To counter such an attack, we propose k -anonymous counterfactual explanations and introduce pureness as a metric to evaluate the validity of these k -anonymous counterfactual explanations. Our results show that making the explanations, rather than the whole dataset, k -anonymous, is beneficial for the quality of the explanations.
Sofie Goethals, Kenneth Sörensen, David Martens
ACM Trans. Intell. Syst. Technol.3
2022 Explainable image classification with evidence counterfactual
abstract
Abstract The complexity of state-of-the-art modeling techniques for image classification impedes the ability to explain model predictions in an interpretable way. A counterfactual explanation highlights the parts of an image which, when removed, would change the predicted class. Both legal scholars and data scientists are increasingly turning to counterfactual explanations as these provide a high degree of human interpretability, reveal what minimal information needs to be changed in order to come to a different prediction and do not require the prediction model to be disclosed. Our literature review shows that existing counterfactual methods for image classification have strong requirements regarding access to the training data and the model internals, which often are unrealistic. Therefore, SEDC is introduced as a model-agnostic instance-level explanation method for image classification that does not need access to the training data. As image classification tasks are typically multiclass problems, an additional contribution is the introduction of the SEDC-T method that allows specifying a target counterfactual class. These methods are experimentally tested on ImageNet data, and with concrete examples, we illustrate how the resulting explanations can give insights in model decisions. Moreover, SEDC is benchmarked against existing model-agnostic explanation methods, demonstrating stability of results, computational efficiency and the counterfactual nature of the explanations.
Tom Vermeire, Dieter Brughmans, Sofie Goethals, Raphael Mazzine, David Martens
Pattern Anal. Appl.5
2021 Node classification over bipartite graphs through projection
Marija Stankova, Stiene Praet, David Martens, Foster J. Provost
Mach. Learn.3
2020 Customs fraud detection
Jellis Vanhoeyveld, David Martens, Bruno Peeters
Pattern Anal. Appl.2
2018 Imbalanced classification in sparse and large behaviour datasets
Jellis Vanhoeyveld, David Martens
Data Min. Knowl. Discov.2
2018 Wallenius Bayes
Enric Junqué de Fortuny, David Martens, Foster J. Provost
Mach. Learn.2
2017 Bankruptcy prediction for SMEs using relational data
Ellen Tobback, Tony Bellotti, Julie Moeyersoms, Marija Stankova, David Martens
Decis. Support Syst.5
2015 Iteratively refining SVMs using priors
abstract
Research on scalable machine learning algorithms has gained a considerable amount of traction since the exponential growth in data assets during the past decades. Many Big Data applications resort to somewhat "simple" data modelling techniques due to the computational constraints associated with more complex models. Simple models, while being very efficient to estimate, often fail to capture some of the finer details of more complex datasets. In this manuscript, we explore the idea that complex large scale classification can be tractable using a process of iterative refining. In such a process, we focus on non-linearities of the data only after having first found an approximate linear model. This knowledge is then incorporated into the nonlinear model implicitly, allowing the non-linear model to focus on important parts of the data after a rough first estimation. This in turn reduces overall training time and allows for a richer model representation, eventually leading to more predictive power.
Enric Junqué de Fortuny, Theodoros Evgeniou, David Martens, Foster J. Provost
IEEE BigData3
2015 To tune or not to tune: rule evaluation for metaheuristic-based sequential covering algorithms
Bart Minnaert, David Martens, Manu De Backer, Bart Baesens
Data Min. Knowl. Discov.2
2015 Loyal to your city? A data mining analysis of a public service loyalty program
Sofie De Cnudde, David Martens
Decis. Support Syst.2
2015 Including high-cardinality attributes in predictive models: A case study in churn prediction in the energy sector
Julie Moeyersoms, David Martens
Decis. Support Syst.2
2015 Comprehensible software fault and effort prediction: A data mining approach
Julie Moeyersoms, Enric Junqué de Fortuny, Karel Dejaeger, Bart Baesens, David Martens
J. Syst. Softw.5
2015 Active Learning-Based Pedagogical Rule Extraction
abstract
Many of the state-of-the-art data mining techniques introduce nonlinearities in their models to cope with complex data relationships effectively. Although such techniques are consistently included among the top classification techniques in terms of predictive power, their lack of transparency renders them useless in any domain where comprehensibility is of importance. Rule-extraction algorithms remedy this by distilling comprehensible rule sets from complex models that explain how the classifications are made. This paper considers a new rule extraction technique, based on active learning. The technique generates artificial data points around training data with low confidence in the output score, after which these are labeled by the black-box model. The main novelty of the proposed method is that it uses a pedagogical approach without making any architectural assumptions of the underlying model. It can therefore be applied to any black-box technique. Furthermore, it can generate any rule format, depending on the chosen underlying rule induction technique. In a large-scale empirical study, we demonstrate the validity of our technique to extract trees and rules from artificial neural networks, support vector machines, and random forests, on 25 data sets of varying size and dimensionality. Our results show that not only do the generated rules explain the black-box models well (thereby facilitating the acceptance of such models), the proposed algorithm also performs significantly better than traditional rule induction techniques in terms of accuracy as well as fidelity.
Enric Junqué de Fortuny, David Martens
IEEE Trans. Neural Networks Learn. Syst.2
2014 Corporate residence fraud detection
abstract
With the globalisation of the world's economies and ever-evolving financial structures, fraud has become one of the main dissipaters of government wealth and perhaps even a major contributor in the slowing down of economies in general. Although corporate residence fraud is known to be a major factor, data availability and high sensitivity have caused this domain to be largely untouched by academia. The current Belgian government has pledged to tackle this issue at large by using a variety of in-house approaches and cooperations with institutions such as academia, the ultimate goal being a fair and efficient taxation system. This is the first data mining application specifically aimed at finding corporate residence fraud, where we show the predictive value of using both structured and fine-grained invoicing data. We further describe the problems involved in building such a fraud detection system, which are mainly data-related (e.g. data asymmetry, quality, volume, variety and velocity) and deployment-related (e.g. the need for explanations of the predictions made).
Enric Junqué de Fortuny, Marija Stankova, Julie Moeyersoms, Bart Minnaert, Foster J. Provost, David Martens
KDD6
2014 Evaluating and understanding text-based stock price prediction models
Enric Junqué de Fortuny, Tom De Smedt, David Martens, Walter Daelemans
Inf. Process. Manag.3
2014 A Comment on "Correlation as a Heuristic for Accurate and Comprehensible Ant Colony Optimization-Based Classifiers"
abstract
The paper provides a comment from the authors regarding the correlation as a heuristic for accurate and comprehensible ant colony optimization-based classifiers. Baig et al. proposed a new classification rule mining algorithm in IEEE TEC. The technique introduces a correlation-based function for the ant colony optimization (ACO) component. The authors compare variants of the technique with several existing ACO-based algorithms and the state-of-the art RIPPER algorithm. The AntMiner+ algorithm, proposed by Martens et al. (2007) in IEEE TEC, was also included in the comparison. However, the reported performance is far below what is expected and is more akin to making random predictions.
Bart Minnaert, David Martens
IEEE Trans. Evol. Comput.2
2012 Media coverage in times of political crisis: A text mining approach
Enric Junqué de Fortuny, Tom De Smedt, David Martens, Walter Daelemans
Expert Syst. Appl.3
2012 Data Mining Techniques for Software Effort Estimation: A Comparative Study
abstract
A predictive model is required to be accurate and comprehensible in order to inspire confidence in a business setting. Both aspects have been assessed in a software effort estimation setting by previous studies. However, no univocal conclusion as to which technique is the most suited has been reached. This study addresses this issue by reporting on the results of a large scale benchmarking study. Different types of techniques are under consideration, including techniques inducing tree/rule-based models like M5 and CART, linear models such as various types of linear regression, nonlinear models (MARS, multilayered perceptron neural networks, radial basis function networks, and least squares support vector machines), and estimation techniques that do not explicitly induce a model (e.g., a case-based reasoning approach). Furthermore, the aspect of feature subset selection by using a generic backward input selection wrapper is investigated. The results are subjected to rigorous statistical testing and indicate that ordinary least squares regression in combination with a logarithmic transformation performs best. Another key finding is that by selecting a subset of highly predictive attributes such as project size, development, and environment related attributes, typically a significant increase in estimation accuracy can be obtained.
Karel Dejaeger, Wouter Verbeke, David Martens, Bart Baesens
IEEE Trans. Software Eng.3
2011 Performance of classification models from a user perspective
David Martens, Jan Vanthienen, Wouter Verbeke, Bart Baesens
Decis. Support Syst.1
2011 Identifying financially successful start-up profiles with data mining
David Martens, Christine Vanhoutte, Sophie De Winne, Bart Baesens, Luc Sels, Christophe Mues
Expert Syst. Appl.1
2011 Building comprehensible customer churn prediction models with advanced rule induction techniques
Wouter Verbeke, David Martens, Christophe Mues, Bart Baesens
Expert Syst. Appl.2
2011 Editorial survey: swarm intelligence for data mining
David Martens, Bart Baesens, Tom Fawcett
Mach. Learn.1
2011 Guest Editorial White Box Nonlinear Prediction Models
abstract
The five papers in this special section focus on white-box nonlinear prediction models.
Bart Baesens, David Martens, Rudy Setiono, Jacek M. Zurada
IEEE Trans. Neural Networks2
2010 Software Effort Prediction Using Regression Rule Extraction from Neural Networks
abstract
Neural networks are often selected as tool for software effort prediction because of their capability to approximate any continuous function with arbitrary accuracy. A major drawback of neural networks is the complex mapping between inputs and output, which is not easily understood by a user. This paper describes a rule extraction technique that derives a set of comprehensible IF-THEN rules from a trained neural network applied to the domain of software effort prediction. The suitability of this technique is tested on the ISBSG R11 data set by a comparison with linear regression, radial basis function networks, and CART. It is found that the most accurate results are obtained by CART, though the large number of rules limits comprehensibility. Considering comprehensible models only, the concise set of extracted rules outperform the pruned CART tree, making neural network rule extraction the most suitable technique for software effort prediction when comprehensibility is important.
Rudy Setiono, Karel Dejaeger, Wouter Verbeke, David Martens, Bart Baesens
ICTAI (2)4
2010 From linear to non-linear kernel based classifiers for bankruptcy prediction
Tony Van Gestel, Bart Baesens, David Martens
Neurocomputing3
2009 Inferring comprehensible business/ICT alignment rules
Bjorn Cumps, David Martens, Manu De Backer, Raf Haesen, Stijn Viaene, Guido Dedene, Bart Baesens, Monique Snoeck
Inf. Manag.2
2009 Robust Process Discovery with Artificial Negative Events
Stijn Goedertier, David Martens, Jan Vanthienen, Bart Baesens
J. Mach. Learn. Res.2
2009 Decompositional Rule Extraction from Support Vector Machines by Active Learning
abstract
Support vector machines (SVMs) are currently state-of-the-art for the classification task and, generally speaking, exhibit good predictive performance due to their ability to model nonlinearities. However, their strength is also their main weakness, as the generated nonlinear models are typically regarded as incomprehensible black-box models. In this paper, we propose a new Active Learning-Based Approach (ALBA) to extract comprehensible rules from opaque SVM models. Through rule extraction, some insight is provided into the logics of the SVM model. ALBA extracts rules from the trained SVM model by explicitly making use of key concepts of the SVM: the support vectors, and the observation that these are typically close to the decision boundary. Active learning implies the focus on apparent problem areas, which for rule induction techniques are the regions close to the SVM decision boundary where most of the noise is found. By generating extra data close to these support vectors that are provided with a class label by the trained SVM model, rule induction techniques are better able to discover suitable discrimination rules. This performance increase, both in terms of predictive accuracy as comprehensibility, is confirmed in our experiments where we apply ALBA on several publicly available data sets.
David Martens, Bart Baesens, Tony Van Gestel
IEEE Trans. Knowl. Data Eng.1
2008 Predicting going concern opinion with data mining
David Martens, Liesbeth Bruynseels, Bart Baesens, Marleen Willekens, Jan Vanthienen
Decis. Support Syst.1
2008 Mining software repositories for comprehensible software fault prediction models
Olivier Vandecruys, David Martens, Bart Baesens, Christophe Mues, Manu De Backer, Raf Haesen
J. Syst. Softw.2
2007 Classification With Ant Colony Optimization
abstract
Ant colony optimization (ACO) can be applied to the data mining field to extract rule-based classifiers. The aim of this paper is twofold. On the one hand, we provide an overview of previous ant-based approaches to the classification task and compare them with state-of-the-art classification techniques, such as C4.5, RIPPER, and support vector machines in a benchmark study. On the other hand, a new ant-based classification technique is proposed, named AntMiner+. The key differences between the proposed AntMiner+ and previous AntMiner versions are the usage of the better performing MAX-MIN ant system, a clearly defined and augmented environment for the ants to walk through, with the inclusion of the class variable to handle multiclass problems, and the ability to include interval rules in the rule list. Furthermore, the commonly encountered problem in ACO of setting system parameters is dealt with in an automated, dynamic manner. Our benchmarking experiments show an AntMiner+ accuracy that is superior to that obtained by the other AntMiner versions, and competitive or better than the results achieved by the compared classification techniques.
David Martens, Manu De Backer, Raf Haesen, Jan Vanthienen, Monique Snoeck, Bart Baesens
IEEE Trans. Evol. Comput.1