Peter G. M. van der Heijden

dblp:14/300 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-3345-096XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 67% Machine learning and data management · 33%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
active learning
0.912025
Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025
Information retrieval › information filtering › technology-assisted review
stopping criteria
0.912025
Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025
Information retrieval › information filtering
technology-assisted review
0.912025
Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025

Methods — techniques the papers use, named apart from their topics

ensemble · 1.7chao's population size estimator · 1.7active learning · 1.7
YearPublicationVenuePosition
2026 Identification of NMF by Choosing Maximum-Volume Basis Vectors
abstract
In nonnegative matrix factorization (NMF), minimum-volume-constrained NMF is a widely used framework for identifying the solution of NMF by making basis vectors as similar as possible. This typically induces sparsity in the coefficient matrix, with each row containing zero entries. Consequently, minimum-volume-constrained NMF may fail for highly mixed data, where such sparsity does not hold. Moreover, the estimated basis vectors in minimum-volume-constrained NMF may be difficult to interpret as they may be mixtures of the ground truth basis vectors. To address these limitations, in this paper we propose a new NMF framework, called maximum-volume-constrained NMF, which makes the basis vectors as distinct as possible. We further establish an identifiability theorem for maximum-volume-constrained NMF and provide an algorithm to estimate it. Experimental results demonstrate the effectiveness of the proposed method.
Qianqian Qi 0002, Zhongming Chen, Peter G. M. van der Heijden
IEEE Signal Process. Lett.3
2025 Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review
abstract
Technology-Assisted Review aims to reduce the human effort required for screening processes such as abstract screening for Systematic Literature Reviews. Human reviewers label documents as relevant or irrelevant during this process, while the system incrementally updates a prediction model based on the reviewers’ previous decisions. After each model update, the system proposes new documents it deems relevant, to prioritize relevant documents over irrelevant ones. A stopping criterion is necessary to guide users in stopping the review process to minimize the number of missed relevant documents and the number of read irrelevant documents. In this article, we propose and evaluate a new ensemble-based Active Learning strategy and a stopping criterion based on Chao’s Population Size Estimator that estimates the prevalence of relevant documents in the dataset. Our simulation study demonstrates that this criterion performs well on several datasets and is compared to other methods presented in the literature.
Michiel P. Bron, Peter G. M. van der Heijden, A. J. Feelders, Arno Siebes
ACM Trans. Inf. Syst.2
2024 Improving information retrieval through correspondence analysis instead of latent semantic analysis
abstract
Abstract The initial dimensions extracted by latent semantic analysis (LSA) of a document-term matrix have been shown to mainly display marginal effects, which are irrelevant for information retrieval. To improve the performance of LSA, usually the elements of the raw document-term matrix are weighted and the weighting exponent of singular values can be adjusted. An alternative information retrieval technique that ignores the marginal effects is correspondence analysis (CA). In this paper, the information retrieval performance of LSA and CA is empirically compared. Moreover, it is explored whether the two weightings also improve the performance of CA. The results for four empirical datasets show that CA always performs better than LSA. Weighting the elements of the raw data matrix can improve CA; however, it is data dependent and the improvement is small. Adjusting the singular value weighting exponent often improves the performance of CA; however, the extent of the improvement depends on the dataset and the number of dimensions.
Qianqian Qi 0002, David J. Hessen, Peter G. M. van der Heijden
J. Intell. Inf. Syst.3
2024 A comparison of latent semantic analysis and correspondence analysis of document-term matrices
abstract
Abstract Latent semantic analysis (LSA) and correspondence analysis (CA) are two techniques that use a singular value decomposition for dimensionality reduction. LSA has been extensively used to obtain low-dimensional representations that capture relationships among documents and terms. In this article, we present a theoretical analysis and comparison of the two techniques in the context of document-term matrices. We show that CA has some attractive properties as compared to LSA, for instance that effects of margins, that is, sums of row elements and column elements, arising from differing document lengths and term frequencies are effectively eliminated so that the CA solution is optimally suited to focus on relationships among documents and terms. A unifying framework is proposed that includes both CA and LSA as special cases. We empirically compare CA to various LSA-based methods on text categorization in English and authorship attribution on historical Dutch texts and find that CA performs significantly better. We also apply CA to a long-standing question regarding the authorship of the Dutch national anthem Wilhelmus and provide further support that it can be attributed to the author Datheen, among several contenders.
Qianqian Qi 0002, David J. Hessen, Tejaswini Deoskar, Peter G. M. van der Heijden
Nat. Lang. Eng.4
2020 ETM: Enrichment by topic modeling for automated clinical sentence classification to detect patients' disease history
abstract
Abstract Given the rapid rate at which text data are being digitally gathered in the medical domain, there is growing need for automated tools that can analyze clinical notes and classify their sentences in electronic health records (EHRs). This study uses EHR texts to detect patients’ disease history from clinical sentences. However, in EHRs, sentences are less topic-focused and shorter than that in general domain, which leads to the sparsity of co-occurrence patterns and the lack of semantic features. To tackle this challenge, current approaches for clinical sentence classification are dependent on external information to improve classification performance. However, this is implausible owing to a lack of universal medical dictionaries. This study proposes the ETM (enrichment by topic modeling) algorithm, based on latent Dirichlet allocation, to smoothen the semantic representations of short sentences. The ETM enriches text representation by incorporating probability distributions generated by an unsupervised algorithm into it. It considers the length of the original texts to enhance representation by using an internal knowledge acquisition procedure. When it comes to clinical predictive modeling, interpretability improves the acceptance of the model. Thus, for clinical sentence classification, the ETM approach employs an initial TFiDF (term frequency inverse document frequency) representation, where we use the support vector machine and neural network algorithms for the classification task. We conducted three sets of experiments on a data set consisting of clinical cardiovascular notes from the Netherlands to test the sentence classification performance of the proposed method in comparison with prevalent approaches. The results show that the proposed ETM approach outperformed state-of-the-art baselines.
Ayoub Bagheri, Arjan Sammani, Peter G. M. van der Heijden, Folkert W. Asselbergs, Daniel L. Oberski
J. Intell. Inf. Syst.3