VLDB 2026 Research / reviewers in the wild / expert
Ingo Scholtes
dblp:s/IngoScholtes
· DBLP profile ↗
26ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-2253-0216ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Network Science Meets AI: A Converging FrontierabstractThe convergence of network science and artificial intelligence (AI) represents a rich area of research, where both fields can mutually enhance one another.Network science offers a comprehensive framework to analyze and model complex relationships, while machine learning (ML) and AI provide powerful tools for recognizing patterns and making predictions from large datasets.Combining these two disciplines can advance the study of complex systems and lead to new innovations in data-driven research.This tutorial paper reviews fundamental concepts of network science, describes the current and promising research direction for bridging network science and AI, and summarizes the contributions that have been accepted for publication in the ESANN 2025 special session on the topic. Matteo Zignani, Fragkiskos D. Malliaros, Ingo Scholtes, Roberto Interdonato, Manuel Dileo |
ESANN | 3 |
| 2024 | The Map Equation Goes Neural: Mapping Network Flows with Graph Neural NetworksabstractCommunity detection is an essential tool for unsupervised data exploration and revealing the organisational structure of networked systems. With a long history in network science, community detection typically relies on objective functions, optimised with custom-tailored search algorithms, but often without leveraging recent advances in deep learning. Recently, first works have started incorporating such objectives into loss functions for deep graph clustering and pooling. We consider the map equation, a popular information-theoretic objective function for unsupervised community detection, and express it in differentiable tensor form for optimisation through gradient descent. Our formulation turns the map equation compatible with any neural network architecture, enables end-to-end learning, incorporates node features, and chooses the optimal number of clusters automatically, all without requiring explicit regularisation. Applied to unsupervised graph clustering tasks, we achieve competitive performance against state-of-the-art deep graph clustering baselines in synthetic and real-world datasets. Christopher Blöcker, Chester Tan, Ingo Scholtes |
NeurIPS | 3 |
| 2024 | Using Time-Aware Graph Neural Networks to Predict Temporal Centralities in Dynamic GraphsabstractNode centralities play a pivotal role in network science, social network analysis, and recommender systems.
In temporal data, static path-based centralities like closeness or betweenness can give misleading results about the true importance of nodes in a temporal graph. To address this issue, temporal generalizations of betweenness and closeness have been defined that are based on the shortest time-respecting paths between pairs of nodes. However, a major issue of those generalizations is that the calculation of such paths is computationally expensive.
Addressing this issue, we study the application of De Bruijn Graph Neural Networks (DBGNN), a time-aware graph neural network architecture, to predict temporal path-based centralities in time series data. We experimentally evaluate our approach in 13 temporal graphs from biological and social systems and show that it considerably improves the prediction of betweenness and closeness centrality compared to (i) a static Graph Convolutional Neural Network, (ii) an efficient sampling-based approximation technique for temporal betweenness, and (iii) two state-of-the-art time-aware graph learning techniques for dynamic graphs. Franziska Heeg, Ingo Scholtes |
NeurIPS | 2 |
| 2022 | Predicting Influential Higher-Order Patterns in Temporal Network DataabstractNetworks are frequently used to model complex systems comprised of interacting elements. While edges capture the topology of direct interactions, the true complexity of many systems originates from higher-order patterns in paths by which nodes can indirectly influence each other. Path data, representing ordered sequences of consecutive direct interactions, can be used to model these patterns. On the one hand, to avoid overfitting, such models should only consider those higher-order patterns for which the data provide sufficient statistical evidence. On the other hand, we hypothesise that network models, which capture only direct interactions, underfit higher-order patterns present in data. Consequently, both approaches are likely to misidentify influential nodes in complex networks. We contribute to this issue by proposing five centrality measures based on MOGen, a multi-order generative model that accounts for all indirect influences up to a maximum distance but disregards influences at higher distances. We compare MOGen-based centralities to equivalent measures for network models and path data in a prediction experiment where we aim to identify influential nodes in out-of-sample data. Our results show strong evidence supporting our hypothesis. MOGen consistently outperforms both the network model and path-based prediction. We further show that the performance difference between MOGen and the path-based approach disappears if we have sufficient observations, confirming that the error is due to overfitting. Christoph Gote, Vincenzo Perri, Ingo Scholtes |
ASONAM | 3 |
| 2022 | Big Data = Big Insights? Operationalising Brooks' Law in a Massive GitHub Data SetabstractMassive data from software repositories and collaboration tools are widely used to study social aspects in software development. One question that several recent works have addressed is how a software project's size and structure influence team productivity, a question famously considered in Brooks' law. Recent studies using massive repository data suggest that developers in larger teams tend to be less productive than smaller teams. Despite using similar methods and data, other studies argue for a positive linear or even super-linear relationship between team size and productivity, thus contesting the view of software economics that software projects are diseconomies of scale. Christoph Gote, Pavlin Mavrodiev, Frank Schweitzer, Ingo Scholtes |
ICSE | 4 |
| 2022 | Learning the Markov Order of Paths in GraphsabstractWe address the problem of learning the Markov order in categorical sequences that represent paths in a network, i.e., sequences of variable lengths where transitions between states are constrained to a known graph. Such data pose challenges for standard Markov order detection methods and demand modeling techniques that explicitly account for the graph constraint. Adopting a multi-order modeling framework for paths, we develop a Bayesian learning technique that (i) detects the correct Markov order more reliably than a competing method based on the likelihood ratio test, (ii) requires considerably less data than methods using AIC or BIC, and (iii) is robust against partial knowledge of the underlying constraints. We further show that a recently published method that uses a likelihood ratio test exhibits a tendency to overfit the true Markov order of paths, which is not the case for our Bayesian technique. Our method is important for data scientists analyzing patterns in categorical sequence data that are subject to (partially) known constraints, e.g. click stream data or other behavioral data on the Web, information propagation in social networks, mobility trajectories, or pathway data in bioinformatics. Addressing the key challenge of model selection, our work is also relevant for the growing body of research that emphasizes the need for higher-order models in network analysis. Luka V. Petrovic, Ingo Scholtes |
WWW | 2 |
| 2021 | Analysing Time-Stamped Co-Editing Networks in Software Development Teams using git2netabstractAbstract Data from software repositories have become an important foundation for the empirical study of software engineering processes. A recurring theme in the repository mining literature is the inference of developer networks capturing e.g. collaboration, coordination, or communication from the commit history of projects. Many works in this area studied networks ofco-authorshipof software artefacts, neglecting detailed information on code changes and code ownership available in software repositories. To address this issue, we introduce , a scalable software that facilitates the extraction of fine-grainedco-editing networksin large repositories. It uses text mining techniques to analyse the detailed history of textual modificationswithinfiles. We apply our tool in two case studies using repositories of multiple Open Source as well as a proprietary software project. Specifically, we use data on more than 1.2 million commits and more than 25,000 developers to test a hypothesis on the relation between developer productivity and co-editing patterns in software teams. We argue that opens up an important new source of high-resolution data on human collaboration patterns that can be used to advance theory in empirical software engineering, computational social science, and organisational studies. Christoph Gote, Ingo Scholtes, Frank Schweitzer |
Empir. Softw. Eng. | 2 |
| 2020 | HOTVis: Higher-Order Time-Aware Visualisation of Dynamic Graphs
Vincenzo Perri, Ingo Scholtes |
GD | 2 |
| 2020 | HYPA: Efficient Detection of Path Anomalies in Time Series Data on NetworksabstractThe unsupervised detection of anomalies in time series data has important applications in user behavioral modeling, fraud detection, and cybersecurity. Anomaly detection has, in fact, been extensively studied in categorical sequences. However, we often have access to time series data that represent paths through networks. Examples include transaction sequences in financial networks, click streams of users in networks of cross-referenced documents, or travel itineraries in transportation networks. To reliably detect anomalies, we must account for the fact that such data contain a large number of independent observations of paths constrained by a graph topology. Moreover, the heterogeneity of real systems rules out frequency-based anomaly detection techniques, which do not account for highly skewed edge and degree statistics. To address this problem, we introduce HYPA, a novel framework for the unsupervised detection of anomalies in large corpora of variable-length temporal paths in a graph. HYPA provides an efficient analytical method to detect paths with anomalous frequencies that result from nodes being traversed in unexpected chronological order. Timothy LaRock, Vahan Nanumyan, Ingo Scholtes, Giona Casiraghi, Tina Eliassi-Rad, Frank Schweitzer |
SDM | 3 |
| 2019 | git2net: mining time-stamped co-editing networks from large git repositoriesabstractData from software repositories have become an important foundation for the empirical study of software engineering processes. A recurring theme in the repository mining literature is the inference of developer networks capturing e.g. collaboration, coordination, or communication, from the commit history of projects. Most of the studied networks are based on the co-authorship of software artefacts defined at the level of files, modules, or packages. While this approach has led to insights into the social aspects of software development, it neglects detailed information on code changes and code ownership, e.g. which exact lines of code have been authored by which developers, that is contained in the commit log of software projects. Addressing this issue, we introduce git2net, a scalable python software that facilitates the extraction of fine-grained co-editing networks in large git repositories. It uses text mining techniques to analyse the detailed history of textual modifications within files. This information allows us to construct directed, weighted, and time-stamped networks, where a link signifies that one developer has edited a block of source code originally written by another developer. Our tool is applied in case studies of an Open Source and a commercial software project. We argue that it opens up a massive new source of high-resolution data on human collaboration patterns. Christoph Gote, Ingo Scholtes, Frank Schweitzer |
MSR | 2 |
| 2017 | When is a Network a Network?: Multi-Order Graphical Model Selection in Pathways and Temporal NetworksabstractWe introduce a framework for the modeling of sequential data capturing pathways of varying lengths observed in a network. Such data are important, e.g., when studying click streams in the Web, travel patterns in transportation systems, information cascades in social networks, biological pathways, or time-stamped social interactions. While it is common to apply graph analytics and network analysis to such data, recent works have shown that temporal correlations can invalidate the results of such methods. This raises a fundamental question: When is a network abstraction of sequential data justified?Addressing this open question, we propose a framework that combines Markov chains of multiple, higher orders into a multi-layer graphical model that captures temporal correlations in pathways at multiple length scales simultaneously. We develop a model selection technique to infer the optimal number of layers of such a model and show that it outperforms baseline Markov order detection techniques. An application to eight real-world data sets on pathways and temporal networks shows that it allows to infer graphical models that capture both topological and temporal characteristics of such data. Our work highlights fallacies of network abstractions and provides a principled answer to the open question when they are justified. Generalizing network representations to multi-order graphical models, it opens perspectives for new data mining and knowledge discovery algorithms. Ingo Scholtes |
KDD | 1 |
| 2016 | From Aristotle to Ringelmann: a large-scale analysis of team productivity and coordination in Open Source Software projects
Ingo Scholtes, Pavlin Mavrodiev, Frank Schweitzer |
Empir. Softw. Eng. | 1 |
| 2013 | Categorizing bugs with social networks: a case study on four open source software communitiesabstractEfficient bug triaging procedures are an important precondition for successful collaborative software engineering projects. Triaging bugs can become a laborious task particularly in open source software (OSS) projects with a large base of comparably inexperienced part-time contributors. In this paper, we propose an efficient and practical method to identify valid bug reports which a) refer to an actual software bug, b) are not duplicates and c) contain enough information to be processed right away. Our classification is based on nine measures to quantify the social embeddedness of bug reporters in the collaboration network. We demonstrate its applicability in a case study, using a comprehensive data set of more than 700, 000 bug reports obtained from the Bugzilla installation of four major OSS communities, for a period of more than ten years. For those projects that exhibit the lowest fraction of valid bug reports, we find that the bug reporters' position in the collaboration network is a strong indicator for the quality of bug reports. Based on this finding, we develop an automated classification scheme that can easily be integrated into bug tracking platforms and analyze its performance in the considered OSS communities. A support vector machine (SVM) to identify valid bug reports based on the nine measures yields a precision of up to 90.3% with an associated recall of 38.9%. With this, we significantly improve the results obtained in previous case studies for an automated early identification of bugs that are eventually fixed. Furthermore, our study highlights the potential of using quantitative measures of social organization in collaborative software engineering. It also opens a broad perspective for the integration of social awareness in the design of support infrastructures. Marcelo Serrano Zanetti, Ingo Scholtes, Claudio J. Tessone, Frank Schweitzer |
ICSE | 2 |
| 2012 | Hierarchical Consensus Formation Reduces The Influence Of Opinion BiasabstractWe study the role of hierarchical structures in a simple model of collective consensus formation based on the bounded confidence model with continuous individual opinions. For the particular variation of this model considered in this paper, we assume that a bias towards an extreme opinion is introduced whenever two individuals interact and form a common decision. As a simple proxy for hierarchical social structures, we introduce a two-step decision making process in which in the second step groups of like-minded individuals are replaced by representatives once they have reached local consensus, and the representatives in turn form a collective decision in a downstream process. We find that the introduction of such a hierarchical decision making structure can improve consensus formation, in the sense that the eventual collective opinion is closer to the true average of individual opinions than without it. In particular, we numerically study how the size of groups of like-minded individuals being represented by delegate individuals affects the impact of the bias on the final population-wide consensus. These results are of interest for the design of organisational policies and the optimisation of hierarchical structures in the context of group decision making. Nicolas Perony, René Pfitzner, Ingo Scholtes, Claudio J. Tessone, Frank Schweitzer |
ECMS | 3 |
| 2012 | A Tunable Mechanism for Identifying Trusted Nodes in Large Scale Distributed NetworksabstractIn this paper, we propose a simple randomized protocol for identifying trusted nodes based on personalized trust in large scale distributed networks. The problem of identifying trusted nodes, based on personalized trust, in a large network setting stems from the huge computation and message overhead involved in exhaustively calculating and propagating the trust estimates by the remote nodes. However, in any practical scenario, nodes generally communicate with a small subset of nodes and thus exhaustively estimating the trust of all the nodes can lead to huge resource consumption. In contrast, our mechanism can be tuned to locate a desired subset of trusted nodes, based on the allowable overhead, with respect to a particular user. The mechanism is based on a simple exchange of random walk messages and nodes counting the number of times they are being hit by random walkers of nodes in their neighborhood. Simulation results to analyze the effectiveness of the algorithm show that using the proposed algorithm, nodes identify the top trusted nodes in the network with a very high probability by exploring only around 45% of the total nodes, and in turn generates nearly 90% less overhead as compared to an exhaustive trust estimation mechanism, named TrustWebRank. Finally, we provide a measure of the global trustworthiness of a node; simulation results indicate that the measures generated using our mechanism differ by only around 0.6% as compared to TrustWebRank. Joydeep Chandra, Ingo Scholtes, Niloy Ganguly, Frank Schweitzer |
TrustCom | 2 |
| 2011 | TweetGames: A framework for twitter-based collaborative social online gamesabstractToday social networks and microblogging services attract much attention and their importance and pervasion is constantly increasing. This trend is fostered by the increasing prevalence of mobile internet devices, which enable users to be online every time and everywhere. A popular kind of applicatio Markus Esch, Aleksandrina Kovacheva, Ingo Scholtes, Steffen Rothkugel |
CollaborateCom | 3 |
| 2010 | Resilience and Multicast Aspects of the Structured Network Overlay GP3abstractToday Massive Multiuser Virtual Environments (MMVEs) are for the most part realized in a centralized fashion. While this provides advantages in terms of manageability and controllability to the providers, it is not feasible for a global-scale scenario like a 3D Web. In order to provide such a global-scale virtual environment it is indispensable to utilize P2P technologies. Within the HyperVerse project, we have developed and recently presented a two-tier Peer-to-Peer (P2P) architecture that incorporates a loosely structured P2P overlay of user peers and a highly structured overlay of server machines, constituting a reliable backbone service. For the interconnection of this backbone servers we proposed a structured P2P overlay named GP3. It has been shown, that this overlay allows efficient join and node lookup operations and allows for the reference locality in virtual environments. This paper studies the GP3 algorithm with respect to multicast aspects as well as resilience against node failures. Markus Esch, Ingo Scholtes |
CISIS | 2 |
| 2010 | A self-organized resource allocation scheme for decentralized distributed virtual environmentsabstractWith the growth of Massively Multiuser Virtual Environments (MMVEs) and increasingly interactive social net working platforms, it is widely accepted that their convergence renders today's centralized hosting approaches impracticable. To handle virtual environments of such massive scale, decentralize Jean Botev, Ingo Scholtes |
CollaborateCom | 2 |
| 2009 | P2P-Based Avatar Interaction in Massive Multiuser Virtual EnvironmentsabstractThe idea of the 3D Web as a global scale Distributed Virtual Environment (DVE) currently is very popular and a lot of research work is done in this field. In the course of the HyperVerse project we have developed a two-tier Peer-To-Peer (P2P) architecture as basic infrastructure for a federated, open and scalable 3D Web. Our approach relies on a concept that incorporates a loosely-structured P2P network overlay of user clients and a highly-structured overlay that connects a federation of reliable server machines constituting a reliable backbone service. A central problem when intending to build a massive virtual online environment is how avatar interaction as well as tracking and provision of avatar positions can be realized in a globally scalable manner. This paper presents a hybrid avatar interaction scheme developed for HyperVerse that incorporates the user clients and the backbone service for avatar position tracking. The backbone is used as reliable fall-back service, while the avatar tracking is handled in a pure P2P fashion whenever possible. This way backbone overload is prevented in a self-organizing manner whenever client density tends to overburden the backbone infrastructure. Markus Esch, Jean Botev, Hermann Schloss, Ingo Scholtes |
CISIS | 4 |
| 2009 | A scale-free and self-organized P2P overlay for massive multiuser virtual environmentsabstractMassive Multiuser Virtual Environments have recently grown popular, and commercial virtual online worlds like Second Life or World of Warcraft attract a lot of attention. In this context, for the research on distributed systems, especially the idea of a 3D Web as a global scale virtual environment i Markus Esch, Ingo Scholtes |
CollaborateCom | 2 |
| 2009 | Evaluation of the HyperVerse Avatar Management Scheme Based on the Analysis of Second Life TracesabstractMassive multiuser virtual environments (MMVEs) and the idea of a global scale 3D Web have grown popular in recent years. While commercial precursors of such environments for the most part rely on centralized client/server architectures, it is commonly accepted that a global scale virtual online world can only be realized in a distributed fashion. Within the HyperVerse project, we have developed and recently presented a two-tier Peer-to-Peer (P2P) architecture that incorporates a loosely structured P2P overlay of user peers and a highly structured overlay of server machines constituting a reliable backbone service. In such a distributed environment, an essential question is how avatars are tracked and interconnected in order to allow mutual rendering and interaction. We have previously proposed a hybrid avatar management scheme that utilizes the backbone service for avatar tracking if necessary, but handles tracking in a P2P fashion when peers can track each other to reduce the backbone load. This paper presents a detailed performance analysis of this algorithm under a realistic scenario, using traces from a large scale MMVE called Second Life. Moreover this paper presents and evaluates an optimization for the hybrid avatar tracking scheme that can be utilized under a weaker condition. Markus Esch, Wei Tsang Ooi, Ingo Scholtes |
ICPADS | 3 |
| 2008 | Using Epidemic Hoarding to Minimize Load Delays in P2P Distributed Virtual Environments
Ingo Scholtes, Jean Botev, Markus Esch, Hermann Schloss, Peter Sturm 0001 |
CollaborateCom | 1 |
| 2008 | GP3 - A Distributed Grid-Based Spatial Index Infrastructure for Massive Multiuser Virtual EnvironmentsabstractMassive Multiuser Virtual Environments (MMVEs) and especially the idea of a ''3D Web'' as a combination of a MMVE and today's WWW currently attracts a lot of attention. The realization of such a vision on a global scale though poses severe technical challenges to the underlying network infrastructure. It is generally accepted that such a global scale scenario can only be realized in a distributed fashion. The HyperVerse project aims at the provision of a federated global scalable infrastructure for such a ''3D Web'' scenario. We propose a two-tier Peer-To-Peer infrastructure that combines a loosely structured overlay network of user clients with a highly-structured overlay network of reliable so-called public servers constituting the backbone of our architecture. This paper presents the Grid-based plane Partitioning Protocol (GP3), a structured peer-to-peer overlay network for the interconnection of the public servers that realizes a spatial index in order to allow fast location based queries. Markus Esch, Jean Botev, Hermann Schloss, Ingo Scholtes |
ICPADS | 4 |
| 2007 | The SEMPA prototype - using XAML and web services for rich interactive Peer-to-Peer applicationsabstractCurrently one can see a surge of evolving technologies supporting the creation of Rich Internet Applications. All of these approaches however traditionally address the classic Client/Server scenario in which a dedicated Web Server acts as application provider for thin clients. This paper argues, that there are many scenarios in which a Peer-to-Peer approach enabling every application/process to export parts of its graphical user interface in a lightweight and interoperable way would be more desirable. This work-in-progress paper presents the lightweight middleware prototype SEMPA, proving that these scenarios can be supported by combining the readily available technologies eXtensible Application Markup Language (XAML) and Web Services. Markus Esch, Hermann Schloss, Ingo Scholtes |
CollaborateCom | 3 |
| 2006 | Web Service Interface Syndication and its Application to a Collaborative Weblog SearchabstractThis paper presents the idea of Web service interface syndication - a scheme for the collaborative creation of overlay networks based on common Web service interfaces. Rather than requiring complex management of dynamic peers we picture a setting of interconnected static Web applications forming syndicates and offering services in a self-organized fashion. Among other application scenarios, this paper will present vicinitySearch, a collaborative Weblog search service which has been created in order to demonstrate the concept. This service offers a remarkable added-value to the Weblog community, making use of existing collaborative structures and the interoperability provided by XML Web services. Apart from the general concept of Web service interface syndication, the paper presents an implementation of the vicinitySearch service that has been done in terms of a Wordpress plugin Ingo Scholtes, Daniel Görgen, Patrick Gratz |
CollaborateCom | 1 |
| 2006 | Web Federates - Towards a Middleware for Highly Scalable Peer-to-Peer Services
Ingo Scholtes, Peter Sturm 0001 |
WEBIST (1) | 1 |