VLDB 2026 Research / reviewers in the wild / expert
Peter G. M. van der Heijden
dblp:14/300
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-3345-096XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 67% Machine learning and data management · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
active learning |
0.9 | 1 | 2025 | Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025 |
Information retrieval › information filtering › technology-assisted review
stopping criteria |
0.9 | 1 | 2025 | Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025 |
Information retrieval › information filtering
technology-assisted review |
0.9 | 1 | 2025 | Using Chao's Estimator as a Stopping Criterion for Technology-Assisted Review · ACM Trans. Inf. Syst. 2025 |
Methods — techniques the papers use, named apart from their topics
ensemble · 1.7chao's population size estimator · 1.7active learning · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identification of NMF by Choosing Maximum-Volume Basis VectorsabstractIn nonnegative matrix factorization (NMF), minimum-volume-constrained NMF is a widely used framework for identifying the solution of NMF by making basis vectors as similar as possible. This typically induces sparsity in the coefficient matrix, with each row containing zero entries. Consequently, minimum-volume-constrained NMF may fail for highly mixed data, where such sparsity does not hold. Moreover, the estimated basis vectors in minimum-volume-constrained NMF may be difficult to interpret as they may be mixtures of the ground truth basis vectors. To address these limitations, in this paper we propose a new NMF framework, called maximum-volume-constrained NMF, which makes the basis vectors as distinct as possible. We further establish an identifiability theorem for maximum-volume-constrained NMF and provide an algorithm to estimate it. Experimental results demonstrate the effectiveness of the proposed method. Qianqian Qi 0002, Zhongming Chen, Peter G. M. van der Heijden |
IEEE Signal Process. Lett. | 3 |
| 2025 | Using Chao's Estimator as a Stopping Criterion for Technology-Assisted ReviewabstractTechnology-Assisted Review aims to reduce the human effort required for screening processes such as abstract screening for Systematic Literature Reviews. Human reviewers label documents as relevant or irrelevant during this process, while the system incrementally updates a prediction model based on the reviewers’ previous decisions. After each model update, the system proposes new documents it deems relevant, to prioritize relevant documents over irrelevant ones. A stopping criterion is necessary to guide users in stopping the review process to minimize the number of missed relevant documents and the number of read irrelevant documents. In this article, we propose and evaluate a new ensemble-based Active Learning strategy and a stopping criterion based on Chao’s Population Size Estimator that estimates the prevalence of relevant documents in the dataset. Our simulation study demonstrates that this criterion performs well on several datasets and is compared to other methods presented in the literature. Michiel P. Bron, Peter G. M. van der Heijden, A. J. Feelders, Arno Siebes |
ACM Trans. Inf. Syst. | 2 |
| 2024 | Improving information retrieval through correspondence analysis instead of latent semantic analysisabstractAbstract The initial dimensions extracted by latent semantic analysis (LSA) of a document-term matrix have been shown to mainly display marginal effects, which are irrelevant for information retrieval. To improve the performance of LSA, usually the elements of the raw document-term matrix are weighted and the weighting exponent of singular values can be adjusted. An alternative information retrieval technique that ignores the marginal effects is correspondence analysis (CA). In this paper, the information retrieval performance of LSA and CA is empirically compared. Moreover, it is explored whether the two weightings also improve the performance of CA. The results for four empirical datasets show that CA always performs better than LSA. Weighting the elements of the raw data matrix can improve CA; however, it is data dependent and the improvement is small. Adjusting the singular value weighting exponent often improves the performance of CA; however, the extent of the improvement depends on the dataset and the number of dimensions. Qianqian Qi 0002, David J. Hessen, Peter G. M. van der Heijden |
J. Intell. Inf. Syst. | 3 |
| 2024 | A comparison of latent semantic analysis and correspondence analysis of document-term matricesabstractAbstract Latent semantic analysis (LSA) and correspondence analysis (CA) are two techniques that use a singular value decomposition for dimensionality reduction. LSA has been extensively used to obtain low-dimensional representations that capture relationships among documents and terms. In this article, we present a theoretical analysis and comparison of the two techniques in the context of document-term matrices. We show that CA has some attractive properties as compared to LSA, for instance that effects of margins, that is, sums of row elements and column elements, arising from differing document lengths and term frequencies are effectively eliminated so that the CA solution is optimally suited to focus on relationships among documents and terms. A unifying framework is proposed that includes both CA and LSA as special cases. We empirically compare CA to various LSA-based methods on text categorization in English and authorship attribution on historical Dutch texts and find that CA performs significantly better. We also apply CA to a long-standing question regarding the authorship of the Dutch national anthem Wilhelmus and provide further support that it can be attributed to the author Datheen, among several contenders. Qianqian Qi 0002, David J. Hessen, Tejaswini Deoskar, Peter G. M. van der Heijden |
Nat. Lang. Eng. | 4 |
| 2020 | ETM: Enrichment by topic modeling for automated clinical sentence classification to detect patients' disease historyabstractAbstract Given the rapid rate at which text data are being digitally gathered in the medical domain, there is growing need for automated tools that can analyze clinical notes and classify their sentences in electronic health records (EHRs). This study uses EHR texts to detect patients’ disease history from clinical sentences. However, in EHRs, sentences are less topic-focused and shorter than that in general domain, which leads to the sparsity of co-occurrence patterns and the lack of semantic features. To tackle this challenge, current approaches for clinical sentence classification are dependent on external information to improve classification performance. However, this is implausible owing to a lack of universal medical dictionaries. This study proposes the ETM (enrichment by topic modeling) algorithm, based on latent Dirichlet allocation, to smoothen the semantic representations of short sentences. The ETM enriches text representation by incorporating probability distributions generated by an unsupervised algorithm into it. It considers the length of the original texts to enhance representation by using an internal knowledge acquisition procedure. When it comes to clinical predictive modeling, interpretability improves the acceptance of the model. Thus, for clinical sentence classification, the ETM approach employs an initial TFiDF (term frequency inverse document frequency) representation, where we use the support vector machine and neural network algorithms for the classification task. We conducted three sets of experiments on a data set consisting of clinical cardiovascular notes from the Netherlands to test the sentence classification performance of the proposed method in comparison with prevalent approaches. The results show that the proposed ETM approach outperformed state-of-the-art baselines. Ayoub Bagheri, Arjan Sammani, Peter G. M. van der Heijden, Folkert W. Asselbergs, Daniel L. Oberski |
J. Intell. Inf. Syst. | 3 |