EDBT 2026 Demo / reviewers in the wild / expert
Andreas Spitz
dblp:135/5777
· DBLP profile ↗
20ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-5282-6133ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 15 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-supervised Data Augmentation for Text Classification in Low-Data Settings
Deyu Ding, Mengying Wang 0004, Andreas Spitz |
LREC | 3 |
| 2026 | R.U.Psycho? A Framework for Robust Unified Psychometric Testing of Language Models
Julian Schelb, Orr Borin, David García 0001, Andreas Spitz |
LREC | 4 |
| 2025 | Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language ModelsabstractPrompt-based language models like GPT4 and LLaMa have been used for a wide variety of use cases such as simulating agents, searching for information, or for content analysis.For all of these applications and others, political biases in these models can affect their performance.Several researchers have attempted to study political bias in language models using evaluation suites based on surveys, such as the Political Compass Test (PCT), often finding a particular leaning favored by these models.However, there is some variation in the exact prompting techniques, leading to diverging findings, and most research relies on constrained-answer settings to extract model responses.Moreover, the Political Compass Test is not a scientifically valid survey instrument.In this work, we contribute a political bias measured informed by political science theory, building on survey design principles to test a wide variety of input prompts, while taking into account prompt sensitivity.We then prompt 11 different open and commercial models, differentiating between instruction-tuned and non-instructiontuned models, and automatically classify their political stances from 88,110 responses.Leveraging this dataset, we compute political bias profiles across different prompt variations and find that while PCT exaggerates bias in certain models like GPT3.5, measures of political bias are often unstable, but generally more leftleaning for instruction-tuned models.Code and data are available on GitHub 1 . Mats Faulborn, Indira Sen, Max Pellert, Andreas Spitz, David García 0001 |
ACL (1) | 4 |
| 2023 | Quotatives Indicate Decline in Objectivity in U.S. Political NewsabstractAccording to journalistic standards, direct quotes should be attributed to sources with objective quotatives such as ``said'' and ``told,'' since nonobjective quotatives, e.g., ``argued'' and ``insisted,'' would influence the readers' perception of the quote and the quoted person. In this paper, we analyze the adherence to this journalistic norm to study trends in objectivity in political news across U.S. outlets of different ideological leanings. We ask: 1) How has the usage of nonobjective quotatives evolved? 2) How do news outlets use nonobjective quotatives when covering politicians of different parties? To answer these questions, we developed a dependency-parsing-based method to extract quotatives and applied it to Quotebank, a web-scale corpus of attributed quotes, obtaining nearly 7 million quotes, each enriched with the quoted speaker's political party and the ideological leaning of the outlet that published the quote. We find that, while partisan outlets are the ones that most often use nonobjective quotatives, between 2013 and 2020, the outlets that increased their usage of nonobjective quotatives the most were ``moderate'' centrist news outlets (around 0.6 percentage points, or 20% in relative percentage over seven years). Further, we find that outlets use nonobjective quotatives more often when quoting politicians of the opposing ideology (e.g., left-leaning outlets quoting Republicans) and that this ``quotative bias'' is rising at a swift pace, increasing up to 0.5 percentage points, or 25% in relative percentage, per year. These findings suggest an overall decline in journalistic objectivity in U.S. political news. Tiancheng Hu, Manoel Horta Ribeiro, Robert West 0001, Andreas Spitz |
ICWSM | 4 |
| 2022 | Quote Erat Demonstrandum: A Web Interface for Exploring the Quotebank CorpusabstractThe use of attributed quotes is the most direct and least filtered pathway of information propagation in news. Consequently, quotes play a central role in the conception, reception, and analysis of news stories. Since quotes provide a more direct window into a speaker's mind than regular reporting, they are a valuable resource for journalists and researchers alike. While substantial research efforts have been devoted to methods for the automated extraction of quotes from news and their attribution to speakers, few comprehensive corpora of attributed quotes from contemporary sources are available to the public. Here, we present an adaptive web interface for searching Quotebank, a massive collection of quotes from the news, which we make available at https://quotebank.dlab.tools. Vuk Vukovic, Akhil Arora 0001, Huan-Cheng Chang, Andreas Spitz, Robert West 0001 |
SIGIR | 4 |
| 2022 | ${\sf DeepNC}$DeepNC: Deep Generative Network CompletionabstractMost network data are collected from partially observable networks with both missing nodes and missing edges, for example, due to limited resources and privacy settings specified by users on social media. Thus, it stands to reason that inferring the missing parts of the networks by performing network completion should precede downstream applications. However, despite this need, the recovery of missing nodes and edges in such incomplete networks is an insufficiently explored problem due to the modeling difficulty, which is much more challenging than link prediction that only infers missing edges. In this paper, we present DeepNC, a novel method for inferring the missing parts of a network based on a deep generative model of graphs. Specifically, our method first learns a likelihood over edges via an autoregressive generative model, and then identifies the graph that maximizes the learned likelihood conditioned on the observable graph topology. Moreover, we propose a computationally efficient [Formula: see text] algorithm that consecutively finds individual nodes that maximize the probability in each node generation step, as well as an enhanced version using the expectation-maximization algorithm. The runtime complexities of both algorithms are shown to be almost linear in the number of nodes in the network. We empirically demonstrate the superiority of DeepNC over state-of-the-art network completion approaches. Cong Tran, Won-Yong Shin, Andreas Spitz, Michael Gertz 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Community Detection in Partially Observable Social NetworksabstractThe discovery of community structures in social networks has gained significant attention since it is a fundamental problem in understanding the networks’ topology and functions. However, most social network data are collected from partially observable networks with both missing nodes and edges . In this article, we address a new problem of detecting overlapping community structures in the context of such an incomplete network, where communities in the network are allowed to overlap since nodes belong to multiple communities at once. To solve this problem, we introduce KroMFac , a new framework that conducts community detection via regularized nonnegative matrix factorization (NMF) based on the Kronecker graph model. Specifically, from an inferred Kronecker generative parameter matrix, we first estimate the missing part of the network. As our major contribution to the proposed framework, to improve community detection accuracy, we then characterize and select influential nodes (which tend to have high degrees) by ranking, and add them to the existing graph. Finally, we uncover the community structures by solving the regularized NMF-aided optimization problem in terms of maximizing the likelihood of the underlying graph. Furthermore, adopting normalized mutual information (NMI), we empirically show superiority of our KroMFac approach over two baseline schemes by using both synthetic and real-world networks. Cong Tran, Won-Yong Shin, Andreas Spitz |
ACM Trans. Knowl. Discov. Data | 3 |
| 2021 | Quotebank: A Corpus of Quotations from a Decade of NewsabstractWe present Quotebank, an open corpus of 178 million quotations attributed to the speakers who uttered them, extracted from 162 million English news articles published between 2008 and 2020. In order to produce this Web-scale corpus, while at the same time benefiting from the performance of modern neural models, we introduce Quobert, a minimally supervised framework for extracting and attributing quotations from massive corpora. Quobert avoids the necessity of manually labeled input and instead exploits the redundancy of the corpus by bootstrapping from a single seed pattern to extract training data for fine-tuning a BERT-based model. Quobert is language- and corpus agnostic and correctly attributes 86.9% of quotations in our experiments. Quotebank and Quobert are publicly available at https://doi.org/10.5281/zenodo.4277311. Timoté Vaucher, Andreas Spitz, Michele Catasta, Robert West 0001 |
WSDM | 2 |
| 2021 | Interventions for Softening Can Lead to Hardening of Opinions: Evidence from a Randomized Controlled TrialabstractMotivated by the goal of designing interventions for softening polarized opinions on the Web, and building on results from psychology, we hypothesized that people would be moved more easily towards opposing opinions when the latter were voiced by a celebrity they like, rather than by a celebrity they dislike. We tested this hypothesis in a survey-based randomized controlled trial in which we exposed respondents to opinions that were randomly assigned to one of four spokespersons each: a disagreeing but liked celebrity, a disagreeing and disliked celebrity, a disagreeing expert, and an agreeing but disliked celebrity. After the treatment, we measured changes in the respondents’ opinions, empathy towards the spokespersons, and use of affective language. Andreas Spitz, Ahmad Abu-Akel, Robert West 0001 |
WWW | 1 |
| 2020 | A Versatile Hypergraph Model for Document CollectionsabstractEfficiently and effectively representing large collections of text is of central importance to information retrieval tasks such as summarization and search. Since models for these tasks frequently rely on an implicit graph structure of the documents or their contents, graph-based document representations are naturally appealing. For tasks that consider the joint occurrence of words or entities, however, existing document representations often fall short in capturing cooccurrences of higher order, higher multiplicity, or at varying proximity levels. Furthermore, while numerous applications benefit from structured knowledge sources, external data sources are rarely considered as integral parts of existing document models. Andreas Spitz, Dennis Aumiller, Bálint Soproni, Michael Gertz 0001 |
SSDBM | 1 |
| 2019 | Word Embeddings for Entity-Annotated Texts
Satya Almasian, Andreas Spitz, Michael Gertz 0001 |
ECIR (1) | 2 |
| 2019 | Retrieving Multi-Entity Associations: An Evaluation of Combination Modes for Word EmbeddingsabstractWord embeddings have gained significant attention as learnable representations of semantic relations between words, and have been shown to improve upon the results of traditional word representations. However, little effort has been devoted to using embeddings for the retrieval of entity associations beyond pairwise relations. In this paper, we use popular embedding methods to train vector representations of an entity-annotated news corpus, and evaluate their performance for the task of predicting entity participation in news events versus a traditional word cooccurrence network as a baseline. To support queries for events with multiple participating entities, we test a number of combination modes for the embedding vectors. While we find that even the best combination modes for word embeddings do not quite reach the performance of the full cooccurrence network, especially for rare entities, we observe that different embedding methods model different types of relations, thereby indicating the potential for ensemble methods. Gloria Feher, Andreas Spitz, Michael Gertz 0001 |
SIGIR | 2 |
| 2019 | TopExNet: Entity-Centric Network Topic Exploration in News StreamsabstractThe recent introduction of entity-centric implicit network representations of unstructured text offers novel ways for exploring entity relations in document collections and streams efficiently and interactively. Here, we present TopExNet as a tool for exploring entity-centric network topics in streams of news articles. The application is available as a web service at https://topexnet.ifi.uni-heidelberg.de. Andreas Spitz, Satya Almasian, Michael Gertz 0001 |
WSDM | 1 |
| 2018 | Entity-Centric Topic Extraction and Exploration: A Network-Based Approach
Andreas Spitz, Michael Gertz 0001 |
ECIR | 1 |
| 2018 | Efficient anti-community detection in complex networksabstractModeling the relations between the components of complex systems as networks of vertices and edges is a commonly used method in many scientific disciplines that serves to obtain a deeper understanding of the systems themselves. In particular, the detection of densely connected communities in these networks is frequently used to identify functionally related components, such as social circles in networks of personal relations or interactions between agents in biological networks. Traditionally, communities are considered to have a high density of internal connections, combined with a low density of external edges between different communities. However, not all naturally occurring communities in complex networks are characterized by this notion of structural equivalence, such as groups of energy states with shared quantum numbers in networks of spectral line transitions. In this paper, we focus on this inverse task of detecting anti-communities that are characterized by an exceptionally low density of internal connections and a high density of external connections. While anti-communities have been discussed in the literature for anecdotal applications or as a modification of traditional community detection, no rigorous investigation of algorithms for the problem has been presented. To this end, we introduce and discuss a broad range of possible approaches and evaluate them with regard to efficiency and effectiveness on a range of real-world and synthetic networks. Furthermore, we show that the presence of a community and anti-community structure are not mutually exclusive, and that even networks with a strong traditional community structure may also contain anti-communities. Sebastian Lackner, Andreas Spitz, Matthias Weidemüller, Michael Gertz 0001 |
SSDBM | 2 |
| 2016 | Terms over LOAD: Leveraging Named Entities for Cross-Document Extraction and Summarization of EventsabstractReal world events, such as historic incidents, typically contain both spatial and temporal aspects and involve a specific group of persons. This is reflected in the descriptions of events in textual sources, which contain mentions of named entities and dates. Given a large collection of documents, however, such descriptions may be incomplete in a single document, or spread across multiple documents. In these cases, it is beneficial to leverage partial information about the entities that are involved in an event to extract missing information. In this paper, we introduce the LOAD model for cross-document event extraction in large-scale document collections. The graph-based model relies on co-occurrences of named entities belonging to the classes locations, organizations, actors, and dates and puts them in the context of surrounding terms. As such, the model allows for efficient queries and can be updated incrementally in negligible time to reflect changes to the underlying document collection. We discuss the versatility of this approach for event summarization, the completion of partial event information, and the extraction of descriptions for named entities and dates. We create and provide a LOAD graph for the documents in the English Wikipedia from named entities extracted by state-of-the-art NER tools. Based on an evaluation set of historic data that include summaries of diverse events, we evaluate the resulting graph. We find that the model not only allows for near real-time retrieval of information from the underlying document collection, but also provides a comprehensive framework for browsing and summarizing event data. Andreas Spitz, Michael Gertz 0001 |
SIGIR | 1 |
| 2015 | Exploiting Phase Transitions for the Efficient Sampling of the Fixed Degree Sequence ModelabstractReal-world network data is often very noisy and contains erroneous or missing edges. These superfluous and missing edges can be identified statistically by assessing the number of common neighbors of the two incident nodes. To evaluate whether this number of common neighbors, the so called co-occurrence, is statistically significant, a comparison with the expected co-occurrence in a suitable random graph model is required. For networks with a skewed degree distribution, including most real-world networks, it is known that the fixed degree sequence model, which maintains the degrees of nodes, is favourable over using simplified graph models that are based on an independence assumption. However, the use of a fixed degree sequence model requires sampling from the space of all graphs with the given degree sequence and measuring the co-occurrence of each pair of nodes in each of the samples, since there is no known closed formula for this statistic. While there exist log-linear approaches such as Markov chain Monte Carlo sampling, the computational complexity still depends on the length of the Markov chain and the number of samples, which is significant in large-scale networks. In this article, we show based on ground truth data that there are various phase transition-like tipping points that enable us to choose a comparatively low number of samples and to reduce the length of the Markov chains without reducing the quality of the significance test. As a result, the computational effort can be reduced by an order of magnitudes. Christian Brugger, André Lucas Chinazzo, Alexandre Flores John, Christian de Schryver, Norbert Wehn, Andreas Spitz, Katharina A. Zweig |
ASONAM | 6 |
| 2015 | Beyond Friendships and Followers: The Wikipedia Social NetworkabstractMost traditional social networks rely on explicitly given relations between users, their friends and followers. In this paper, we go beyond well structured data repositories and create a person-centric network from unstructured text -- the Wikipedia Social Network. To identify persons in Wikipedia, we make use of interwiki links, Wikipedia categories and person related information available in Wikidata. From the co-occurrences of persons on a Wikipedia page we construct a large-scale person-centric network and provide a weighting scheme for the relationship of two persons based on the distances of their mentions within the text. We extract key characteristics of the network such as centrality, clustering coefficient and component sizes for which we find values that are typical for social networks. Using state-of-the-art algorithms for community detection in massive networks, we identify interesting communities and evaluate them against Wikipedia categories. The Wikipedia social network developed this way provides an important source for future social analysis tasks. Johanna Geiß, Andreas Spitz, Michael Gertz 0001 |
ASONAM | 2 |
| 2015 | Breaking the News: Extracting the Sparse Citation Network Backbone of Online News ArticlesabstractNetworks of online news articles and blog posts are some of the most commonly used data sets in network science. As a result, they have become a vital piece of network analysis and are used for the evaluation of algorithms that work on large networks, or serve as examples in the analysis of information diffusion and propagation. Similarly, scientific citation networks are part of the bedrock upon which much of modern network analysis is built and have been studied for decades. In this paper, we show that the backbone inherent to networks of online news articles shares significant structural similarities to scientific citation networks once the noise of spurious links is stripped away. We present a data set of news articles that, while it is extremely sparse and lightweight, still contains information relevant to the propagation of information in mass media and is remarkably similar to scientific citation networks, thus opening the door to the use of established methodologies from scientometrics and bibliometrics in the analysis of online news propagation. Andreas Spitz, Michael Gertz 0001 |
ASONAM | 1 |
| 2013 | SICOP: identifying significant co-interaction patternsabstractSUMMARY: Interactions between various types of molecules that regulate crucial cellular processes are extensively investigated by high-throughput experiments and require dedicated computational methods for the analysis of the resulting data. In many cases, these data can be represented as a bipartite graph because it describes interactions between elements of two different types such as the influence of different experimental conditions on cellular variables or the direct interaction between receptors and their activators/inhibitors. One of the major challenges in the analysis of such noisy datasets is the statistical evaluation of the relationship between any two elements of the same type. Here, we present SICOP (significant co-interaction patterns), an implementation of a method that provides such an evaluation based on the number of their common interaction partners, their so-called co-interaction. This general network analytic method, proved successful in diverse fields, provides a framework for assessing the significance of this relationship by comparison with the expected co-interaction in a suitable null model of the same bipartite graph. SICOP takes into consideration up to two distinct types of interactions such as up- or downregulation. The tool is written in Java and accepts several common input formats and supports different output formats, facilitating further analysis and visualization. Its key features include a user-friendly interface, easy installation and platform independence. AVAILABILITY: The software is open source and available at cna.cs.uni-kl.de/SICOP under the terms of the GNU General Public Licence (version 3 or later). Andreas Spitz, Katharina A. Zweig, Ágnes Horvát |
Bioinform. | 1 |