VLDB 2026 Research / reviewers in the wild / expert
Frank Schweitzer
dblp:67/4271
· DBLP profile ↗
18ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0003-1551-6491ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Helping a Friend or Supporting a Cause? Disentangling Active and Passive Cosponsorship in the U.S. CongressabstractGiuseppe Russo, Christoph Gote, Laurence Brandenberger, Sophia Schlosser, Frank Schweitzer. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Giuseppe Russo 0001, Christoph Gote, Laurence Brandenberger, Sophia Schlosser, Frank Schweitzer |
ACL (1) | 5 |
| 2022 | Big Data = Big Insights? Operationalising Brooks' Law in a Massive GitHub Data SetabstractMassive data from software repositories and collaboration tools are widely used to study social aspects in software development. One question that several recent works have addressed is how a software project's size and structure influence team productivity, a question famously considered in Brooks' law. Recent studies using massive repository data suggest that developers in larger teams tend to be less productive than smaller teams. Despite using similar methods and data, other studies argue for a positive linear or even super-linear relationship between team size and productivity, thus contesting the view of software economics that software projects are diseconomies of scale. Christoph Gote, Pavlin Mavrodiev, Frank Schweitzer, Ingo Scholtes |
ICSE | 3 |
| 2021 | Analysing Time-Stamped Co-Editing Networks in Software Development Teams using git2netabstractAbstract Data from software repositories have become an important foundation for the empirical study of software engineering processes. A recurring theme in the repository mining literature is the inference of developer networks capturing e.g. collaboration, coordination, or communication from the commit history of projects. Many works in this area studied networks ofco-authorshipof software artefacts, neglecting detailed information on code changes and code ownership available in software repositories. To address this issue, we introduce , a scalable software that facilitates the extraction of fine-grainedco-editing networksin large repositories. It uses text mining techniques to analyse the detailed history of textual modificationswithinfiles. We apply our tool in two case studies using repositories of multiple Open Source as well as a proprietary software project. Specifically, we use data on more than 1.2 million commits and more than 25,000 developers to test a hypothesis on the relation between developer productivity and co-editing patterns in software teams. We argue that opens up an important new source of high-resolution data on human collaboration patterns that can be used to advance theory in empirical software engineering, computational social science, and organisational studies. Christoph Gote, Ingo Scholtes, Frank Schweitzer |
Empir. Softw. Eng. | 3 |
| 2020 | HYPA: Efficient Detection of Path Anomalies in Time Series Data on NetworksabstractThe unsupervised detection of anomalies in time series data has important applications in user behavioral modeling, fraud detection, and cybersecurity. Anomaly detection has, in fact, been extensively studied in categorical sequences. However, we often have access to time series data that represent paths through networks. Examples include transaction sequences in financial networks, click streams of users in networks of cross-referenced documents, or travel itineraries in transportation networks. To reliably detect anomalies, we must account for the fact that such data contain a large number of independent observations of paths constrained by a graph topology. Moreover, the heterogeneity of real systems rules out frequency-based anomaly detection techniques, which do not account for highly skewed edge and degree statistics. To address this problem, we introduce HYPA, a novel framework for the unsupervised detection of anomalies in large corpora of variable-length temporal paths in a graph. HYPA provides an efficient analytical method to detect paths with anomalous frequencies that result from nodes being traversed in unexpected chronological order. Timothy LaRock, Vahan Nanumyan, Ingo Scholtes, Giona Casiraghi, Tina Eliassi-Rad, Frank Schweitzer |
SDM | 6 |
| 2019 | Quantifying triadic closure in multi-edge social networksabstractIn social networks, edges often form closed triangles or triads. Standard approaches to measuring triadic closure, however, fail for multi-edge networks, because they do not consider that triads can be formed by edges of different multiplicity. We propose a novel measure of triadic closure for multi-edge networks based on a shared partner statistic and demonstrate that this measure can detect meaningful closure in synthetic and empirical multi-edge networks, where conventional approaches fail. This work is a cornerstone in driving inferential network analyses from the analysis of binary networks towards the analyses of multi-edge and weighted networks, which offer a more realistic representation of social interactions and relations. Laurence Brandenberger, Giona Casiraghi, Vahan Nanumyan, Frank Schweitzer |
ASONAM | 4 |
| 2019 | git2net: mining time-stamped co-editing networks from large git repositoriesabstractData from software repositories have become an important foundation for the empirical study of software engineering processes. A recurring theme in the repository mining literature is the inference of developer networks capturing e.g. collaboration, coordination, or communication, from the commit history of projects. Most of the studied networks are based on the co-authorship of software artefacts defined at the level of files, modules, or packages. While this approach has led to insights into the social aspects of software development, it neglects detailed information on code changes and code ownership, e.g. which exact lines of code have been authored by which developers, that is contained in the commit log of software projects. Addressing this issue, we introduce git2net, a scalable python software that facilitates the extraction of fine-grained co-editing networks in large git repositories. It uses text mining techniques to analyse the detailed history of textual modifications within files. This information allows us to construct directed, weighted, and time-stamped networks, where a link signifies that one developer has edited a block of source code originally written by another developer. Our tool is applied in case studies of an Open Source and a commercial software project. We argue that it opens up a massive new source of high-resolution data on human collaboration patterns. Christoph Gote, Ingo Scholtes, Frank Schweitzer |
MSR | 3 |
| 2016 | From Aristotle to Ringelmann: a large-scale analysis of team productivity and coordination in Open Source Software projects
Ingo Scholtes, Pavlin Mavrodiev, Frank Schweitzer |
Empir. Softw. Eng. | 3 |
| 2013 | Categorizing bugs with social networks: a case study on four open source software communitiesabstractEfficient bug triaging procedures are an important precondition for successful collaborative software engineering projects. Triaging bugs can become a laborious task particularly in open source software (OSS) projects with a large base of comparably inexperienced part-time contributors. In this paper, we propose an efficient and practical method to identify valid bug reports which a) refer to an actual software bug, b) are not duplicates and c) contain enough information to be processed right away. Our classification is based on nine measures to quantify the social embeddedness of bug reporters in the collaboration network. We demonstrate its applicability in a case study, using a comprehensive data set of more than 700, 000 bug reports obtained from the Bugzilla installation of four major OSS communities, for a period of more than ten years. For those projects that exhibit the lowest fraction of valid bug reports, we find that the bug reporters' position in the collaboration network is a strong indicator for the quality of bug reports. Based on this finding, we develop an automated classification scheme that can easily be integrated into bug tracking platforms and analyze its performance in the considered OSS communities. A support vector machine (SVM) to identify valid bug reports based on the nine measures yields a precision of up to 90.3% with an associated recall of 38.9%. With this, we significantly improve the results obtained in previous case studies for an automated early identification of bugs that are eventually fixed. Furthermore, our study highlights the potential of using quantitative measures of social organization in collaborative software engineering. It also opens a broad perspective for the integration of social awareness in the design of support infrastructures. Marcelo Serrano Zanetti, Ingo Scholtes, Claudio J. Tessone, Frank Schweitzer |
ICSE | 4 |
| 2012 | Hierarchical Consensus Formation Reduces The Influence Of Opinion BiasabstractWe study the role of hierarchical structures in a simple model of collective consensus formation based on the bounded confidence model with continuous individual opinions. For the particular variation of this model considered in this paper, we assume that a bias towards an extreme opinion is introduced whenever two individuals interact and form a common decision. As a simple proxy for hierarchical social structures, we introduce a two-step decision making process in which in the second step groups of like-minded individuals are replaced by representatives once they have reached local consensus, and the representatives in turn form a collective decision in a downstream process. We find that the introduction of such a hierarchical decision making structure can improve consensus formation, in the sense that the eventual collective opinion is closer to the true average of individual opinions than without it. In particular, we numerically study how the size of groups of like-minded individuals being represented by delegate individuals affects the impact of the bias on the final population-wide consensus. These results are of interest for the design of organisational policies and the optimisation of hierarchical structures in the context of group decision making. Nicolas Perony, René Pfitzner, Ingo Scholtes, Claudio J. Tessone, Frank Schweitzer |
ECMS | 5 |
| 2012 | Emotional Divergence Influences Information Spreading in Twitter
René Pfitzner, Antonios Garas, Frank Schweitzer |
ICWSM | 3 |
| 2012 | A Tunable Mechanism for Identifying Trusted Nodes in Large Scale Distributed NetworksabstractIn this paper, we propose a simple randomized protocol for identifying trusted nodes based on personalized trust in large scale distributed networks. The problem of identifying trusted nodes, based on personalized trust, in a large network setting stems from the huge computation and message overhead involved in exhaustively calculating and propagating the trust estimates by the remote nodes. However, in any practical scenario, nodes generally communicate with a small subset of nodes and thus exhaustively estimating the trust of all the nodes can lead to huge resource consumption. In contrast, our mechanism can be tuned to locate a desired subset of trusted nodes, based on the allowable overhead, with respect to a particular user. The mechanism is based on a simple exchange of random walk messages and nodes counting the number of times they are being hit by random walkers of nodes in their neighborhood. Simulation results to analyze the effectiveness of the algorithm show that using the proposed algorithm, nodes identify the top trusted nodes in the network with a very high probability by exploring only around 45% of the total nodes, and in turn generates nearly 90% less overhead as compared to an exhaustive trust estimation mechanism, named TrustWebRank. Finally, we provide a measure of the global trustworthiness of a node; simulation results indicate that the measures generated using our mechanism differ by only around 0.6% as compared to TrustWebRank. Joydeep Chandra, Ingo Scholtes, Niloy Ganguly, Frank Schweitzer |
TrustCom | 4 |
| 2012 | How Random Is Social Behaviour? Disentangling Social Complexity through the Study of a Wild House Mouse PopulationabstractOut of all the complex phenomena displayed in the behaviour of animal groups, many are thought to be emergent properties of rather simple decisions at the individual level. Some of these phenomena may also be explained by random processes only. Here we investigate to what extent the interaction dynamics of a population of wild house mice (Mus domesticus) in their natural environment can be explained by a simple stochastic model. We first introduce the notion of perceptual landscape, a novel tool used here to describe the utilisation of space by the mouse colony based on the sampling of individuals in discrete locations. We then implement the behavioural assumptions of the perceptual landscape in a multi-agent simulation to verify their accuracy in the reproduction of observed social patterns. We find that many high-level features--with the exception of territoriality--of our behavioural dataset can be accounted for at the population level through the use of this simplified representation. Our findings underline the potential importance of random factors in the apparent complexity of the mice's social structure. These results resonate in the general context of adaptive behaviour versus elementary environmental interactions. Nicolas Perony, Claudio J. Tessone, Barbara König 0002, Frank Schweitzer |
PLoS Comput. Biol. | 4 |
| 2012 | The Link between Dependency and Cochange: Empirical EvidenceabstractWe investigate the relationship between class dependency and change propagation (cochange) in software written in Java. On the one hand, we find a strong correlation between dependency and cochange. Furthermore, we provide empirical evidence for the propagation of change along paths of dependency. These findings support the often alleged role of dependencies as propagators of change. On the other hand, we find that approximately half of all dependencies are never involved in cochanges and that the vast majority of cochanges pertain to only a small percentage of dependencies. This means that inferring the cochange characteristics of a software architecture solely from its dependency structure results in a severely distorted approximation of cochange characteristics. Any metric which uses dependencies alone to pass judgment on the evolvability of a piece of Java software is thus unreliable. As a consequence, we suggest to always take both the change characteristics and the dependency structure into account when evaluating software architecture. Markus M. Geipel, Frank Schweitzer |
IEEE Trans. Software Eng. | 2 |
| 2009 | Personalised and dynamic trust in social networksabstractWe propose a novel trust metric for social networks which is suitable for application to recommender systems. It is personalised and dynamic, and allows to compute the indirect trust between two agents which are not neighbours based on the direct trust between agents that are neighbours. In analogy to some personalised versions of PageRank, this metric makes use of the concept of feedback centrality and overcomes some of the limitations of other trust metrics. In particular, it does not neglect cycles and other patterns characterising social networks, as some other algorithms do. In order to apply the metric to recommender systems, we propose a way to make trust dynamic over time. We show by means of analytical approximations and computer simulations that the metric has the desired properties. Finally, we carry out an empirical validation on a dataset crawled from an Internet community and compare the performance of a recommender system using our metric to one using collaborative filtering. Frank Edward Walter, Stefano Battiston, Frank Schweitzer |
RecSys | 3 |
| 2009 | Software change dynamics: evidence from 35 java projectsabstractIn this paper we investigate the relationship between class dependency and change propagation in Java software. By analyzing 35 large Open Source Java projects, we find that in the majority of the projects more than half of the dependencies are never involved in change propagation. Furthermore, our analysis shows that only a few dependencies are transmitting the majority of change propagation events. An additional analysis reveals that this concentration cannot be explained by the different ages of the dependencies. The conclusion is that the dependency structure alone is a poor measure for the change dynamics. This contrasts with current literature. Markus M. Geipel, Frank Schweitzer |
ESEC/SIGSOFT FSE | 2 |
| 2008 | A model of a trust-based recommendation system on a social network
Frank Edward Walter, Stefano Battiston, Frank Schweitzer |
Auton. Agents Multi Agent Syst. | 3 |
| 1997 | Optimization of Road Networks Using Evolutionary StrategiesabstractA road network usually has to fulfill two requirements: (i) it should as far as possible provide direct connections between nodes to avoid large detours; and (ii) the costs for road construction and maintenance, which are assumed proportional to the total length of the roads, should be low. The optimal solution is a compromise between these contradictory demands, which in our model can be weighted by a parameter. The road optimization problem belongs to the class of frustrated optimization problems. In this paper, a special class of evolutionary strategies, such as the Boltzmann and Darwin and mixed strategies, are applied to find differently optimized solutions (graphs of varying density) for the road network, depending on the degree of frustration. We show that the optimization process occurs on two different time scales. In the asymptotic limit, a fixed relation between the mean connection distance (detour) and the total length (costs) of the network exists that defines a range of possible compromises. Furthermore, we investigate the density of states, which describes the number of solutions with a certain fitness value in the stationary regime. We find that the network problem belongs to a class of optimization problems in which more effort in optimization certainly yields better solutions. An analytical approximation for the relation between effort and improvement is derived. Frank Schweitzer, Helge Rosé, Werner Ebeling, Olaf Weiss |
Evol. Comput. | 1 |
| 1996 | Network Optimization Using Evolutionary Strategies
Frank Schweitzer, Werner Ebeling, Helge Rosé, Olaf Weiss |
PPSN | 1 |