VLDB 2026 Research / reviewers in the wild / expert
Frank W. Takes
dblp:32/10440
· DBLP profile ↗
18ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-5468-1030ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Anonymization Problem in Social Networks
Rachel G. de Jong, Mark P. J. van der Loo, Frank W. Takes |
WAW | 3 |
| 2024 | Fair tree classifier using strong demographic parityabstractAbstract When dealing with sensitive data in automated data-driven decision-making, an important concern is to learn predictors with high performance towards a class label, whilst minimising for the discrimination towards any sensitive attribute, like gender or race, induced from biased data. Hybrid tree optimisation criteria have been proposed which combine classification performance and fairness. Although the threshold-free ROC-AUC is the standard for measuring classification model performance, current fair tree classification methods mainly optimise for a fixed threshold on the fairness metric. In this paper, we propose SCAFF—splitting criterion AUC for Fairness—a compound decision tree splitting criterion which combines the threshold-free strong demographic parity with ROC-AUC termed, easily applicable as an ensemble. Our method simultaneously leverages multiple sensitive attributes of which the values may be multicategorical, and is tunable with respect to the unavoidable performance-fairness trade-off. In our experiments, we demonstrate how SCAFF generates effective models with competitive performance and fairness with respect to binary, multicategorical, and multiple sensitive attributes. António Pereira Barata, Frank W. Takes, H. Jaap van den Herik, Cor J. Veenman |
Mach. Learn. | 2 |
| 2023 | CADENCE: Community-Aware Detection of Dynamic Network StatesabstractDynamic interaction data is often aggregated in a sequence of network snapshots before being employed in downstream analysis. The two common ways of defining network snapshots are i) a fixed time interval or ii) fixed number of interactions per snapshot. The choice of aggregation has a significant impact on subsequent analysis, and it is not trivial to select one approach over another for a given dataset. More importantly assuming snapshot regularity is data-agnostic and may be at odds with the underlying interaction dynamics. To address these challenges, we propose a method for community-aware detection of network states (CADENCE) based on the premise of stable interaction time-frames within network communities. We simultaneously detect network communities and partition the global interaction activity into scale-adaptive snapshots where the level of interaction within communities remains stable. We model a temporal network as a node-node-time tensor and use a structured canonical polyadic decomposition with a piece-wise constant temporal factor to iteratively identify communities and their activity levels. We demonstrate that transitions between network snapshots learned by CADENCE constitute network change points of better quality than those predicted by state-of-the-art network change point detectors. Furthermore, the network structure within individual snapshots reflects ground truth communities better than baselines for adaptive tensor granularity. Through a case study on a real-world Reddit dataset, we showcase the interpretability of CADENCE motivated snapshots as periods separated by significant events. Maxwell McNeil, Carolina Mattsson, Frank W. Takes, Petko Bogdanov |
SDM | 3 |
| 2020 | The eXPose Approach to Crosslier DetectionabstractTransit of wasteful materials within the European Union is highly regulated through a system of permits. Waste processing costs vary greatly depending on the waste category of a permit. Therefore, companies may have a financial incentive to allege transporting waste with erroneous categorisation. Our goal is to assist inspectors in selecting potentially manipulated permits for further investigation, making their task more effective and efficient. Due to data limitations, a supervised learning approach based on historical cases is not possible. Standard unsupervised approaches, such as outlier detection and data quality-assurance techniques, are not suited since we are interested in targeting non-random modifications in both category and category-correlated features. For this purpose we (1) introduce the concept of crosslier: an anomalous instance of a category which lies across other categories; (2) propose eXPose: a novel approach to crosslier detection based on supervised category modelling; and (3) present the crosslier diagram: a visualisation tool specifically designed for domain experts to easily assess crossliers. We compare eXPose against traditional outlier detection methods in various benchmark datasets with synthetic crossliers and show the superior performance of our method in targeting these instances. António Pereira Barata, Frank W. Takes, H. Jaap van den Herik, Cor J. Veenman |
ICPR | 2 |
| 2020 | Foreword to the special issue on mining actionable insights from social networks
Marcelo Gabriel Armentano, Ebrahim Bagheri, Frank W. Takes, Virginia Yannibelli |
Inf. Process. Manag. | 3 |
| 2020 | Foreword to the special issue on mining actionable insights from online user generated content
Marcelo Gabriel Armentano, Ebrahim Bagheri, Julia Kiseleva, Frank W. Takes |
Inf. Retr. J. | 4 |
| 2019 | Fast incremental computation of harmonic closeness centrality in directed weighted networksabstractThis paper proposes a novel approach to efficiently compute the exact closeness centrality values of all nodes in dynamically evolving directed and weighted networks. Closeness centrality is one of the most frequently used centrality measures in the field of social network analysis. It uses the total distance to all other nodes to determine node centrality. Previous work has addressed the problem of dynamically updating closeness centrality values for either undirected networks or only for the top-k nodes in terms of closeness centrality. Here, we propose a fast approach for exactly computing all closeness centrality values at each timestamp of directed and weighted evolving networks. Such networks are prevalent in many real-world situations. The main ingredients of our approach are a combination of work filtering methods and efficient incremental updates that avoid unnecessary recomputation. We tested the approach on several real-world datasets of dynamic small-world networks and found that we have mean speed-ups of about 33 times. In addition, the method is highly parallelizable. K. (Lynn) Putman, Hanjo D. Boekhout, Frank W. Takes |
ASONAM | 3 |
| 2019 | A Bayesian Approach for Accurate Classification-Based AggregatesabstractIn this paper, we study the accuracy of values aggregated over classes predicted by a classification algorithm. The problem is that the resulting aggregates (e.g., sums of a variable) are known to be biased. The bias can be large even for highly accurate classification algorithms, in particular when dealing with class-imbalanced data. To correct this bias, the algorithm's classification error rates have to be estimated. In this estimation, two issues arise when applying existing bias correction methods. First, inaccuracies in estimating classification error rates have to be taken into account. Second, impermissible estimates, such as a negative estimate for a positive value, have to be dismissed. We show that both issues are relevant in applications where the true labels are known only for a small set of data points. We propose a novel bias correction method using Bayesian inference. The novelty of our method is that it imposes constraints on the model parameters. We show that our method solves the problem of biased classification-based aggregates as well as the two issues above, in the general setting of multi-class classification. In the empirical evaluation, using a binary classifier on a real-world dataset of company tax returns, we show that our method outperforms existing methods in terms of mean squared error. Quinten Meertens, C. G. H. Diks, H. Jaap van den Herik, Frank W. Takes |
SDM | 4 |
| 2018 | Estimating Subgraph Generation Models to Understand Large Network FormationabstractRecently, a new network formation model was proposed: SUGM. Our research looks into a method to estimate the parameters of this model based on the subgraph census. Laurens Bogaardt, Frank W. Takes |
eScience | 2 |
| 2018 | Understanding Evolving Communities in Transnational Board Interlock NetworksabstractThe network structure of the corporate elite is well studied through board interlocks: the sharing of common directors between companies. Corporate networks, where nodes are companies and lines represent interlocks, model how corporations and the individuals therein exert power over others, gain access to information and in general interact within the global economy. Exploring corporate network structures using network analysis techniques has greatly improved our understanding of the global corporate system. Network studies have amongst others aided in unraveling the spread of corporate practice, the formation of a corporate elite, and the formation of business groups and elite transnatic-malization. We apply the proposed method on the evolving board interlock network over time, and relate the most significant changes in communities to other attributes, such as geographical location, allowing us to study the dynamics of the global corporate elite structure. Analysis of the evolving global corporate network benefits the debates on globalization and the balance of power between west and east, for example enabling a quantification of the gradual integration of new corporate powers such as China. Dafne E. van Kuppevelt, Frank W. Takes, Eelke M. Heemskerk |
eScience | 2 |
| 2018 | Accurate WiFi-Based Indoor Positioning with Continuous Location Sampling
Jesper E. van Engelen, J. J. van Lier, Frank W. Takes, Heike Trautmann |
ECML/PKDD (3) | 3 |
| 2018 | The effects of data quality on the analysis of corporate board interlock networks
Javier Garcia-Bernardo, Frank W. Takes |
Inf. Syst. | 2 |
| 2017 | Exploiting GPUs for Fast Force-Directed Visualization of Large-Scale NetworksabstractNetwork analysis software relies on graph layout algorithms to enable users to visually explore network data. Nowadays, networks easily consist of millions of nodes and edges, resulting in hours of computation time to obtain a readable graph layout on a typical workstation. Although these machines usually do not have a very large number of CPU cores, they can easily be equipped with Graphics Processing Units (GPUs), opening up the possibility of exploiting hundreds or even thousands of cores to counter the aforementioned computational challenges. In this paper we introduce a novel GPU framework for visualizing large real-world network data. The main focus is on a GPU implementation of force-directed graph layout algorithms, which are known to create high quality network visualizations. The proposed framework is used to parallelize the well-known ForceAtlas2 algorithm, which is widely used in many popular network analysis packages and toolkits. The different procedures and data structures of the algorithm are adjusted to the CUDA GPU architecture's specifics in terms of memory coalescing, shared memory usage and thread workload balance. To evaluate its performance, the GPU implementation is tested using a diverse set of 38 different large-scale real-world networks. This allows for a thorough characterization of the parallelizable components of both force-directed layout algorithms in general as well as the proposed GPU framework as a whole. Experiments demonstrate how the approach can efficiently process very large real-world networks, showing overall speedup factors between 40x and 123x compared to existing CPU implementations. In practice, this means that a network with 4 million nodes and 120 million edges can be visualized in 14 minutes rather than 9 hours. Govert G. Brinkmann, Kristian F. D. Rietveld, Frank W. Takes |
ICPP | 3 |
| 2016 | Explainable and Efficient Link Prediction in Real-World Network Data
Jesper E. van Engelen, Hanjo D. Boekhout, Frank W. Takes |
IDA | 3 |
| 2015 | Fast diameter and radius BFS-based computation in (weakly connected) real-world graphs: With an application to the six degrees of separation games
Michele Borassi, Pierluigi Crescenzi, Michel Habib, Walter A. Kosters, Andrea Marino 0001, Frank W. Takes |
Theor. Comput. Sci. | 6 |
| 2014 | Fast Diameter Computation of Large Sparse Graphs Using GPUsabstractIn this paper we propose a highly parallel GPU-based bounding algorithm for computing the exact diameter of large real-world sparse graphs. The diameter is defined as the length of the longest shortest path between vertices in the graph, and serves as a relevant property of all types of graphs that are nowadays frequently studied. Examples include social networks, webgraphs and routing networks. We verify the performance of our parallel approach on a set of large graphs comprised of millions of vertices, and using a CUDA GPU observe an increase in performance of up to 21.1x compared to a CPU algorithm using the same strategy. Based on these results, we provide a characterization of the types of graphs that are well-suited for traversal by means of our parallel diameter algorithm. We furthermore include a comparison of different GPU algorithms for single-source shortest path computations, which is not only a crucial step in computing the diameter, but also relevant in many other distance and neighborhood-based algorithms. Giso H. Dal, Walter A. Kosters, Frank W. Takes |
PDP | 3 |
| 2013 | Mining User-Generated Path Traversal Patterns in an Information NetworkabstractThis paper studies patterns occurring in user-generated click paths within the online encyclopedia Wikipedia. The click path data originates from over seven million goal-oriented clicks gathered from the Wiki Game, an online game in which the goal is to find a path between two given random Wikipedia articles. First we propose to use node-based path traversal patterns to derive a new measure of node centrality, arguing that a node is central if it proves useful in navigating through the network. A comparison with centrality measures from literature is provided, showing that users generally "know" only a relatively small portion of the network, which they employ frequently in finding their goal, and that this set of nodes differs significantly from the set of central nodes according to various centrality measures. Next, using the notion of sub graph centrality, we show that users are able to identify a small yet efficient portion of the graph that is useful for successfully completing their navigation goals. Frank W. Takes, Walter A. Kosters |
Web Intelligence | 1 |
| 2011 | Determining the diameter of small world networksabstractIn this paper we present a novel approach to determine the exact diameter (longest shortest path length) of large graphs, in particular of the nowadays frequently studied small world networks. Typical examples include social networks, gene networks, web graphs and internet topology networks. Due to complexity issues, the diameter is often calculated based on a sample of only a fraction of the nodes in the graph, or some approximation algorithm is applied. We instead propose an exact algorithm that uses various lower and upper bounds as well as effective node selection and pruning strategies in order to evaluate only the critical nodes which ultimately determine the diameter. We will show that our algorithm is able to quickly determine the exact diameter of various large datasets of small world networks with millions of nodes and hundreds of millions of links, whereas before only approximations could be given. Frank W. Takes, Walter A. Kosters |
CIKM | 1 |