Ruben van Heusden

dblp:327/3279 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-9204-9220ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Redacted text detection using neural image segmentation methods
abstract
The redaction of sensitive information in documents is common practice in specific types of organizations. This happens for example in court proceedings or in documents released under the Freedom of Information Act (FOIA). The ability to automatically detect when information has been redacted has several practical applications, such as the gathering of statistics on the amount of redaction present in documents, enabling a critical view on redaction practices. It can also be used to further investigate redactions, and whether or not the used techniques provide sufficient anonymization. The task is particularly challenging because of the large variety of redaction methods and techniques, from software for automatic redaction to manual redactions by pen. Any detection system must be robust to a large variety of inputs, as it will be run on many documents that might not even contain redactions. In this study, we evaluate two neural methods for the task, namely a Mask R-CNN model and a Mask2Former model, and compare them to a rule-based model based on optical character recognition and morphological operations. The best performing, the Mask R-CNN model, has a recall of .94 with a precision of .96 over a challenging data set containing several redaction types. Adding many pages without redaction barely lowers this score (precision drops to .90, recall drops to .92). The Mask2Former model is most robust to inputs without redactions, producing the least false positives of all models.
Ruben van Heusden, Kaj Meijer, Maarten Marx
Int. J. Document Anal. Recognit.1
2024 OpenPSS: An Open Page Stream Segmentation Benchmark
Ruben van Heusden, Jaap Kamps, Maarten Marx
TPDL (1)1
2024 Bcubed revisited: elements like me
abstract
Abstract BCubed is a mathematically clean, elegant and intuitively well behaved external performance metric for clustering tasks. BCubed compares a predicted clustering to a known ground truth clustering through elementwise precision and recall scores. For each element, the predicted and ground truth clusters containing the element are compared, and the mean over all elements is taken. We argue that BCubed overestimates performance, for the intuitive reason that the clustering gets credit for putting an element into its own cluster. This is repaired, and we investigate the repaired version, called “Elements Like Me (ELM)”. We extensively evaluate ELM from both a theoretical and empirical perspective, and conclude that it retains all of its positive properties, and yields a minimum zero score when it should. Synthetic experiments show that ELM can produce different rankings of predicted clusterings when compared to BCubed, and that the ELM scores are distributed with lower mean and a larger variance than BCubed.
Ruben van Heusden, Jaap Kamps, Maarten Marx
Discov. Comput.1
2024 A sharper definition of alignment for Panoptic Quality
abstract
The Panoptic Quality metric, developed by Kirillov et al. in 2019, makes object-level precision, recall and F1 measures available for evaluating image segmentation, and more generally any partitioning task, against a gold standard. Panoptic Quality is based on partial isomorphisms between hypothesized and true segmentations. Kirillov et al. desire that functions defining these one-to-one matchings should be simple, interpretable and effectively computable. They show that for t and h, true and hypothesized segments, the condition stating that there are more correct than wrongly predicted pixels, formalized as IoU(t,h)>.5 or equivalently as |t∩h|>.5|t∪h| has these properties. We show that a weaker function, requiring that more than half of the pixels in the hypothesized segment are in the true segment and vice-versa, formalized as |t∩h|>.5|t| and |t∩h|>.5|h|, is not only sufficient but also necessary. With a small proviso, every function defining a partial isomorphism satisfies this condition. We theoretically and empirically compare the two conditions.
Ruben van Heusden, Maarten Marx
Pattern Recognit. Lett.1
2023 Making PDFs Accessible for Visually Impaired Users (and Findable for Everybody Else)
Ruben van Heusden, Hazel Ling, Lars Nelissen, Maarten Marx
TPDL1
2023 Detection of Redacted Text in Legal Documents
Ruben van Heusden, Aron de Ruijter, Roderick Majoor, Maarten Marx
TPDL1