A. J. Feelders

dblp:23/3319 · also Ad Feelders · DBLP profile ↗
← Back
42ranked-venue papers
13as first author
1since 2021 · last 2025
0000-0003-4525-1949ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 12 first-authorDatabases, data management, data science and information retrieval · 29 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Information retrieval · 40% Data mining · 36% Machine learning and data management · 20%
Artificial intelligence
4 papers
Trustworthy machine learning · 89% Image recognition and object detection · 7% Probabilistic and Bayesian machine learning · 4%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
active learning
0.912025
Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025
Information retrieval › information filtering › technology-assisted review
stopping criteria
0.912025
Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025
Information retrieval › information filtering
technology-assisted review
0.912025
Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025
Data mining
pattern mining
0.432012
Different slopes for different folks: mining for exceptional regression models with cook's distance · KDD 2012
Efficient Algorithms for Finding Richer Subgroup Descriptions in Numeric and Nominal Data · ICDM 2012
Subgroup Discovery Meets Bayesian Networks -- An Exceptional Model Mining Approach · ICDM 2010
Data mining › pattern mining
subgroup discovery
0.332012
Efficient Algorithms for Finding Richer Subgroup Descriptions in Numeric and Nominal Data · ICDM 2012
Subgroup Discovery Meets Bayesian Networks -- An Exceptional Model Mining Approach · ICDM 2010
Different slopes for different folks: mining for exceptional regression models with cook's distance · KDD 2012
Data mining › pattern mining › subgroup discovery
exceptional model mining
0.322012
Different slopes for different folks: mining for exceptional regression models with cook's distance · KDD 2012
Subgroup Discovery Meets Bayesian Networks -- An Exceptional Model Mining Approach · ICDM 2010
Data mining › predictive modeling › classification › multiclass classification
ordinal classification
0.222011
On Generating All Optimal Monotone Classifications · ICDM 2011
Monotone Relabeling in Ordinal Classification · ICDM 2010
Data mining › predictive modeling
classification
0.222012
On Generating All Optimal Monotone Classifications · ICDM 2011
Efficient Algorithms for Finding Richer Subgroup Descriptions in Numeric and Nominal Data · ICDM 2012
Data models and query languages › query language design
pattern language
0.112012
Efficient Algorithms for Finding Richer Subgroup Descriptions in Numeric and Nominal Data · ICDM 2012
Data mining › predictive modeling › classification › interpretable classification
monotone classification
0.112011
On Generating All Optimal Monotone Classifications · ICDM 2011
Machine learning › Trustworthy machine learning
monotonicity
0.112010
Monotone Relabeling in Ordinal Classification · ICDM 2010
Machine learning › Trustworthy machine learning
interpretability
0.112008
Nonparametric Monotone Classification with MOCA · ICDM 2008
Machine learning › Trustworthy machine learning › interpretability
monotonic classification
0.112008
Nonparametric Monotone Classification with MOCA · ICDM 2008
Machine learning › Trustworthy machine learning
robustness
0.112005
Instability of Classifiers on Categorical Data · ICDM 2005
Data mining
anomaly detection
0.112005
Instability of Classifiers on Categorical Data · ICDM 2005
Computational complexity › counting complexity
#p-completeness
0.012011
On Generating All Optimal Monotone Classifications · ICDM 2011
Computational complexity
counting complexity
0.012011
On Generating All Optimal Monotone Classifications · ICDM 2011
Computer vision › Image recognition and object detection
ordinal classification
0.012008
Nonparametric Monotone Classification with MOCA · ICDM 2008
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.011996
Learning from Biased Data Using Mixture Models · KDD 1996
Computational finance and economics › credit risk
credit scoring
0.011995
Data Mining for Loan Evaluation at ABN AMRO: A Case Study · KDD 1995

Methods — techniques the papers use, named apart from their topics

ensemble · 1.7chao's population size estimator · 1.7active learning · 1.7uniform sampling · 0.2enumeration · 0.2regression model fitting · 0.1greedy search · 0.1cook's distance · 0.1convex hull · 0.1ROC space analysis · 0.1convex loss optimization · 0.1l1 loss minimization · 0.1interpolation · 0.1local instability analysis · 0.1mixture model · 0.0
YearPublicationVenuePosition
2025 Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review
abstract
Technology-Assisted Review aims to reduce the human effort required for screening processes such as abstract screening for Systematic Literature Reviews. Human reviewers label documents as relevant or irrelevant during this process, while the system incrementally updates a prediction model based on the reviewers’ previous decisions. After each model update, the system proposes new documents it deems relevant, to prioritize relevant documents over irrelevant ones. A stopping criterion is necessary to guide users in stopping the review process to minimize the number of missed relevant documents and the number of read irrelevant documents. In this article, we propose and evaluate a new ensemble-based Active Learning strategy and a stopping criterion based on Chao’s Population Size Estimator that estimates the prevalence of relevant documents in the dataset. Our simulation study demonstrates that this criterion performs well on several datasets and is compared to other methods presented in the literature.
Michiel P. Bron, Peter G. M. van der Heijden, A. J. Feelders, Arno Siebes
ACM Trans. Inf. Syst.3
2020 Tailored Graph Embeddings for Entity Alignment on Historical Data
abstract
In the domain of the Dutch cultural heritage various data sets describe different aspects of life during the Dutch Golden Age. These data sets, in the form of RDF graphs, use different standards and contain noise in the values of literal nodes, such as misspelled names and uncertainty in dates. The Golden Agents project aims at answering queries about the Dutch Golden ages using these distributed and independently maintained data sets. A problem in this project, among many other problems, is the identification of persons who occur in multiple data sets but under different URI's. This paper aims to solve this specific problem and generate a linkset, i.e. a set of pairs of URI's which are judged to represent the same person. We use domain knowledge in the application of an existing node context generation algorithm to serve as input for GloVe, an algorithm originally designed for embedding words. This embedding is then used to train a classifier on pairs of URI's which are known duplicates and non-duplicates. Using just the cosine similarity between URI-pairs in embedding space for prediction, we obtain a simple classifier with an F½-score of around 0.85, even when very few training examples are provided. On larger training sets, more complex classifiers are shown to reach an F½-score of up to 0.88.
Jurian Baas, Mehdi Dastani, A. J. Feelders
iiWAS3
2018 Data-driven fraud detection in international shipping
Ron Triepels, Hennie A. M. Daniels, A. J. Feelders
Expert Syst. Appl.3
2016 Exploiting monotonicity constraints to reduce label noise: An experimental evaluation
abstract
In some ordinal classification problems we know beforehand that the class label should be increasing (or decreasing) in the attributes. Such relations between class label and attributes are called monotone. We attempt to exploit such monotonicity constraints to reduce label noise. Noise may cause violations of the monotonicity constraint in the data set. In an attempt to reduce label noise, we make the data set monotone by relabeling data points. Through experiments on artificial data, we demonstrate that relabeling almost always produces an improved data set.
A. J. Feelders, Tijmen Kolkman
IJCNN1
2016 Exceptional Model Mining - Supervised descriptive local pattern mining with complex target concepts
Wouter Duivesteijn, A. J. Feelders, Arno J. Knobbe
Data Min. Knowl. Discov.2
2016 Learning from incomplete data in Bayesian networks with qualitative influences
Andrés R. Masegosa, A. J. Feelders, Linda C. van der Gaag
Int. J. Approx. Reason.2
2015 A Quantitative Comparison of Semantic Web Page Segmentation Approaches
Robert Kreuzer, Jurriaan Hage, A. J. Feelders
ICWE3
2015 Exceptional Model Mining with Tree-Constrained Gradient Ascent
abstract
Exceptional Model Mining (EMM) generalizes the well-known data mining task of Subgroup Discovery (SD). Given a model class of interest, the goal of EMM is to find subgroups of the data for which a model fitted to the subgroup deviates substantially from the global model. In both SD and EMM, heuristic search is often employed because exhaustive search of the subgroup description space is in general not feasible. We present a new heuristic search strategy for EMM called tree-constrained gradient ascent (TCGA). It is designed to exploit information about the influence of individual records on the quality of a subgroup, and guarantee at the same time that the subgroups can be described in the pattern language. Using the notion of soft subgroups, numerical optimization can be applied to find subgroups of high quality. We introduce a form of constrained gradient ascent that constructs the constraint simultaneously with the numerical optimization so as to guarantee that the subgroups can be described in the pattern language, while limiting the optimization as little as possible. We show how TCGA can be applied toEMMwith linear regression models. The proposed algorithm is evaluated on both synthetic and real world data sets. We compare the results to those obtained with beam search, a heuristic search algorithm commonly employed in SD and EMM.
Thomas E. Krak, A. J. Feelders
SDM2
2015 Efficient algorithms for finding optimal binary features in numeric and nominal labeled data
Michael Mampaey, Siegfried Nijssen, A. J. Feelders, Rob M. Konijn, Arno J. Knobbe
Knowl. Inf. Syst.3
2014 Real-time Adaptive Problem Detection in Poultry
abstract
Real-time identification of unexpected values upon monitoring the production parameters of egg laying hens is quite challenging, as the collected data includes natural variability in addition to chance fluctuation. We present an adaptive method for calculating residuals that reflect the latter type of fluctuation only, and thereby provide for more accurate detection of potential problems. We report on the application of our method to real-world poultry data.
Steven P. D. Woudenberg, Linda C. van der Gaag, A. J. Feelders, Armin Elbers
ECAI3
2014 Real-Time Adaptive Residual Calculation for Detecting Trend Deviations in Systems with Natural Variability
Steven P. D. Woudenberg, Linda C. van der Gaag, A. J. Feelders, Armin Elbers
IDA3
2014 Exploiting Monotonicity Constraints in Active Learning for Ordinal Classification
abstract
We consider ordinal classification and instance ranking problems where each attribute is known to have an increasing or decreasing relation with the class label or rank. For example, it stands to reason that the number of query terms occurring in a document has a positive influence on its relevance to the query. We aim to exploit such monotonicity constraints by using labeled attribute vectors to draw conclusions about the class labels of order related unlabeled ones. Assuming we have a pool of unlabeled attribute vectors, and an oracle that can be queried for class labels, the central problem is to choose a query point whose label is expected to provide the most information. We evaluate different query strategies by comparing the number of inferred labels after some limited number of queries, as well as by comparing the prediction errors of models trained on the points whose labels have been determined so far. We present an efficient algorithm to determine the query point preferred by the well-known active learning strategy generalized binary search. This algorithm can be applied to binary classification on incomplete matrix orders. For non-binary classification, we propose to include attribute vectors in the training set whose class labels have not been uniquely determined yet. We perform experiments on artificial and real data.
Pieter Soons, A. J. Feelders
SDM2
2012 Efficient Algorithms for Finding Richer Subgroup Descriptions in Numeric and Nominal Data
abstract
Subgroup discovery systems are concerned with finding interesting patterns in labeled data. How these systems deal with numeric and nominal data has a large impact on the quality of their results. In this paper, we consider two ways to extend the standard pattern language of subgroup discovery: using conditions that test for interval membership for numeric attributes, and value set membership for nominal attributes. We assume a greedy search setting, that is, iteratively refining a given subgroup, with respect to a (convex) quality measure. For numeric attributes, we propose an algorithm that finds the optimal interval in linear (rather than quadratic) time, with respect to the number of examples and split points. Similarly, for nominal attributes, we show that finding the optimal set of values can be achieved in linear (rather than exponential) time, with respect to the number of examples and the size of the domain of the attribute. These algorithms operate by only considering subgroup refinements that lie on a convex hull in ROC space, thus significantly narrowing down the search space. We further provide efficient algorithms specifically for the popular Weighted Relative Accuracy quality measure, taking advantage of some of its properties. Our algorithms are shown to perform well in practice, and furthermore provide additional expressive power leading to higher-quality results.
Michael Mampaey, Siegfried Nijssen, A. J. Feelders, Arno J. Knobbe
ICDM3
2012 Different slopes for different folks: mining for exceptional regression models with cook's distance
abstract
Exceptional Model Mining (EMM) is an exploratory data analysis technique that can be regarded as a generalization of subgroup discovery. In EMM we look for subgroups of the data for which a model fitted to the subgroup differs substantially from the same model fitted to the entire dataset. In this paper we develop methods to mine for exceptional regression models. We propose a measure for the exceptionality of regression models (Cook's distance), and explore the possibilities to avoid having to fit the regression model to each candidate subgroup. The algorithm is evaluated on a number of real life datasets. These datasets are also used to illustrate the results of the algorithm. We find interesting subgroups with deviating models on datasets from several different domains. We also show that under certain circumstances one can forego fitting regression models on up to 40% of the subgroups, and these 40% are the relatively expensive regression models to compute.
Wouter Duivesteijn, A. J. Feelders, Arno J. Knobbe
KDD2
2012 Probability estimation and a competence model for rule based e-tutoring systems
abstract
In this paper, we present a student model for rule based e-tutoring systems. This model describes both properties of rewrite rules (difficulty and discriminativity) and of students (start competence and learning speed). The model is an extension of the two-parameter logistic ogive function of Item Response Theory. We show that the model can be applied even to relatively small datasets. We gather data from students working on problems in the logic domain, and show that the model estimates of rule difficulty correspond well to expert opinions. We also show that the estimated start competence corresponds well to our expectations based on the previous experience of the students in the logic domain. We point out that this model can be used to inform students about their competence and learning, and teachers about the students and the difficulty and discriminativity of the rules.
Diederik M. Roijers, Johan Jeuring, A. J. Feelders
LAK3
2012 Active Learning with Monotonicity Constraints
abstract
In many applications of data mining it is known beforehand that the response variable should be increasing (or decreasing) in the attributes. We propose two algorithms to exploit such monotonicity constraints for active learning in ordinal classification in two different settings. The basis of our approach is the observation that if the class label of an object is given, then the monotonicity constraints may allow the labels of other objects to be inferred. For instance, from knowing that loan applicant a is rejected, it can be concluded that all applicants that score worse than a on all criteria should be rejected as well. We propose two heuristics to select good query points. These heuristics make a selection based on a point's potential to determine the labels of other points. The algorithms, each implemented with the proposed heuristics, are evaluated on artificial and real data sets to study their performance. We conclude that exploitation of monotonicity constraints can be very beneficial in active learning.
Nicola Barile, A. J. Feelders
SDM2
2011 Monotone Instance Ranking with mira
Nicola Barile, A. J. Feelders
Discovery Science2
2011 When Learning Naive Bayesian Classifiers Preserves Monotonicity
Barbara F. I. Pieters, Linda C. van der Gaag, A. J. Feelders
ECSQARU3
2011 On Generating All Optimal Monotone Classifications
abstract
In many applications of data mining one knows beforehand that the response variable should be monotone (either increasing or decreasing) in the attributes. In ordinal classification, changing the class labels of a data set (relabeling) so that the data becomes monotone, is useful for at least two reasons. Firstly, models trained on relabeled data tend to have better predictive performance than models trained on the original data. Secondly, relabeling is an important building block for the construction of monotone classifiers. However, optimal monotone relabelings are rarely unique, and so far an efficient algorithm to generate them all has been lacking. The main result of this paper is an efficient algorithm to produce the structure of all optimal monotone relabelings. We also show that counting the solutions is #P-complete and give algorithms for efficiently enumerating all solutions, as well as sampling uniformly from the set of solutions. Experiments show that relabeling non-monotone data can improve the predictive performance of models trained on that data.
Luite Stegeman, A. J. Feelders
ICDM2
2010 Subgroup Discovery Meets Bayesian Networks -- An Exceptional Model Mining Approach
abstract
Whenever a dataset has multiple discrete target variables, we want our algorithms to consider not only the variables themselves, but also the interdependencies between them. We propose to use these interdependencies to quantify the quality of subgroups, by integrating Bayesian networks with the Exceptional Model Mining framework. Within this framework, candidate subgroups are generated. For each candidate, we fit a Bayesian network on the target variables. Then we compare the network's structure to the structure of the Bayesian network fitted on the whole dataset. To perform this comparison, we define an edit distance-based distance metric that is appropriate for Bayesian networks. We show interesting subgroups that we experimentally found with our method on datasets from music theory, semantic scene classification, biology and zoogeography.
Wouter Duivesteijn, Arno J. Knobbe, A. J. Feelders, Matthijs van Leeuwen
ICDM3
2010 Monotone Relabeling in Ordinal Classification
abstract
In many applications of data mining we know beforehand that the response variable should be increasing (or decreasing) in the attributes. Such relations between response and attributes are called monotone. In this paper we present a new algorithm to compute an optimal monotone classification of a data set for convex loss functions. Moreover, we show how the algorithm can be extended to compute all optimal monotone classifications with little additional effort. Monotone relabeling is useful for at least two reasons. Firstly, models trained on relabeled data sets often have better predictive performance than models trained on the original data. Secondly, relabeling is an important building block for the construction of monotone classifiers. We apply the new algorithm to investigate the effect on the prediction error of relabeling the training sample for k nearest neighbour classification and classification trees. In contrast to previous work in this area, we consider all optimal monotone relabelings. The results show that, for small training samples, relabeling the training data results in significantly better predictive performance.
A. J. Feelders
ICDM1
2009 Isotonic Classification Trees
Rémon van de Kamp, A. J. Feelders, Nicola Barile
IDA2
2008 Nonparametric Monotone Classification with MOCA
abstract
We describe a monotone classification algorithm called MOCA that attempts to minimize the mean absolute prediction error for classification problems with ordered class labels.We first find a monotone classifier with minimum L1loss on the training sample, and then use a simple interpolation scheme to predict the class labels for attribute vectors not present in the training data.We compare MOCA to the ordinal stochastic dominance learner (OSDL), on artificial as well as real data sets. We show that MOCA often outperforms OSDL with respect to mean absolute prediction error.
Nicola Barile, A. J. Feelders
ICDM2
2008 Nearest Neighbour Classification with Monotonicity Constraints
Wouter Duivesteijn, A. J. Feelders
ECML/PKDD (1)2
2008 Exceptional Model Mining
Dennis Leman, A. J. Feelders, Arno J. Knobbe
ECML/PKDD (2)2
2008 An Experimental Comparison of Different Inclusion Relations in Frequent Tree Mining
Jeroen De Knijf, A. J. Feelders
Fundam. Informaticae2
2007 Parameter Learning for Bayesian Networks with Strict Qualitative Influences
A. J. Feelders, Robert van Straalen
IDA1
2007 A new parameter Learning Method for Bayesian Networks with Qualitative Influences
A. J. Feelders
UAI1
2006 Learning Bayesian network parameters under order constraints
A. J. Feelders, Linda C. van der Gaag
Int. J. Approx. Reason.1
2005 Instability of Classifiers on Categorical Data
abstract
In this paper we study the local behaviour of arbitrary classifiers using the instability of that classifier in a data point. Moreover, we introduce two algorithms. The first to find highly unstable points, the second to find islands of stability.
Arno Siebes, Muhammad Subianto, A. J. Feelders
ICDM3
2005 Bringing order into bayesian-network construction
abstract
Among the tasks involved in building a Bayesian network, obtaining the required probabilities is generally considered the most daunting. Available data collections are often too small to allow for estimating reliable probabilities. Most domain experts, on the other hand, consider assessing the numbers to be quite demanding. Qualitative probabilistic knowledge, however, is provided more easily by experts. We propose a method for obtaining probabilities, that uses qualitative expert knowledge to constrain the probabilities learned from a small data collection. A dedicated elicitation technique is designed to support the acquisition of the qualitative knowledge required for this purpose. We demonstrate the application of our method by quantifying part of a network in the field of classical swine fever.
Eveline M. Helsper, Linda C. van der Gaag, A. J. Feelders, Willie Loeffen, Petra L. Geenen, Armin Elbers
K-CAP3
2005 Learning Bayesian Network Parameters with Prior Knowledge about Context-Specific Qualitative Influences
A. J. Feelders, Linda C. van der Gaag
UAI1
2004 Monotonicity in Bayesian Networks
Linda C. van der Gaag, Hans L. Bodlaender, A. J. Feelders
UAI3
2003 Pruning for Monotone Classification Trees
A. J. Feelders, Martijn Pardoel
IDA1
2001 MAMBO: Discovering Association Rules Based on Conditional Independencies
Robert Castelo, A. J. Feelders, Arno Siebes
IDA2
2000 Prior Knowledge in Economic Applications of Data Mining
A. J. Feelders
PKDD1
2000 Methodological and practical aspects of data mining
A. J. Feelders, Hennie A. M. Daniels, Marcel Holsheimer
Inf. Manag.1
1999 Handling Missing Data in Trees: Surrogate Splits or Statistical Imputation
A. J. Feelders
PKDD1
1998 Mining in the Presence of Selectivity Bias and its Application to Reject Inference
A. J. Feelders, Soong Chang, Geoffrey J. McLachlan
KDD1
1996 Learning from Biased Data Using Mixture Models
A. J. Feelders
KDD1
1995 Data Mining for Loan Evaluation at ABN AMRO: A Case Study
A. J. Feelders, A. J. F. le Loux, J. W. van't Zand
KDD1
1992 Explanation and diagnosis in business assessment
abstract
An architecture for the diagnosis of a firm that combines quantitative methods with qualitative reasoning is presented. The system consists of a problem identification module and a diagnostic module. The problem identification module operates in quantitative mode to discover deviations from norm values for key variables. The diagnostic module uses both quantitative and qualitative relations to generate intuitively appealing explanations for the symptoms discovered. This architecture provides a convenient and flexible way of coding different kinds of knowledge and reasoning methods of human financial experts.>
Hennie A. M. Daniels, A. J. Feelders
IEEE Trans. Syst. Man Cybern.2