EDBT 2026 Demo / reviewers in the wild / expert
Markus Strohmaier
dblp:01/6659
· DBLP profile ↗
52ranked-venue papers in the field
6as first author
7since 2021 · last 2025
0000-0002-5485-5720ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 32 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 12 (3 first)Data Mining & Knowledge Discovery · 6 (1 first)Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language ModelsabstractThis paper proposes temporally aligned Large Language Models (LLMs) as a tool for longitudinal analysis of social media data. We fine-tune Temporal Adapters for Llama 3 8B on full timelines from a panel of British Twitter users and extract longitudinal aggregates of emotions and attitudes with established questionnaires. We focus our analysis on the beginning of the COVID-19 pandemic that had a strong impact on public opinion and collective emotions. We validate our estimates against representative British survey data and find strong positive and significant correlations for several collective emotions. The estimates obtained are robust across multiple training seeds and prompt formulations, and in line with collective emotions extracted using a traditional classification model trained on labeled data. We demonstrate the flexibility of our method on questions of public opinion for which no pre-trained classifier is available. Our work extends the analysis of affect in LLMs to a longitudinal setting through Temporal Adapters. It enables flexible and new approaches to the longitudinal analysis of social media data. Georg Ahnert, Max Pellert, David García 0001, Markus Strohmaier |
ICWSM | 4 |
| 2024 | Exploring Global Gender Gaps in the Blockchain Domain: Insights from LinkedIn Advertising DataabstractBlockchain technology has gained widespread attention through Bitcoin, but the blockchain domain is still striving to increase gender diversity and widely assess skills gaps by gender. There is limited awareness of women’s participation in blockchain, prompting this study to assess and explore global gender gaps in interests, skills, and professions within the field. By analyzing gender-disaggregated data from LinkedIn’s advertisement platform, we reveal that women are significantly underrepresented in blockchain compared to men, with the gender gap being even more pronounced than in the broader IT sector. This study delves into the volume, velocity, variety, veracity, and value that LinkedIn Ad data offers to assess gender gaps in the blockchain domain at a global level. Reham Al Tamime, Markus Strohmaier, Ingmar Weber |
IEEE Big Data | 2 |
| 2023 | Neighborhood Structure Configuration ModelsabstractWe develop a new method to efficiently sample synthetic networks that preserve the d-hop neighborhood structure of a given network for any given d. The proposed algorithm trades off the diversity in network samples against the depth of the neighborhood structure that is preserved. Our key innovation is to employ a colored Configuration Model with colors derived from iterations of the so-called Color Refinement algorithm. We prove that with increasing iterations the preserved structural information increases: the generated synthetic networks and the original network become more and more similar, and are eventually indistinguishable in terms of centrality measures such as PageRank, HITS, Katz centrality and eigenvector centrality. Our work enables to efficiently generate samples with a precisely controlled similarity to the original network, especially for large networks. Felix I. Stamm, Michael Scholkemper, Michael T. Schaub, Markus Strohmaier |
WWW | 4 |
| 2022 | Adversarial Inter-Group Link Injection Degrades the Fairness of Graph Neural NetworksabstractWe present evidence for the existence and effectiveness of adversarial attacks on graph neural networks (GNNs) that aim to degrade fairness. These attacks can disadvantage a particular subgroup of nodes in GNN-based node classification, where nodes of the underlying network have sensitive attributes, such as race or gender. We conduct qualitative and experimental analyses explaining how adversarial link injection impairs the fairness of GNN predictions. For example, an attacker can compromise the fairness of GNN-based node classification by injecting adversarial links between nodes belonging to opposite subgroups and opposite class labels. Our experiments on empirical datasets demonstrate that adversarial fairness attacks can significantly degrade the fairness of GNN predictions (attacks are effective) with a low perturbation rate (attacks are efficient) and without a significant drop in accuracy (attacks are deceptive). This work demonstrates the vulnerability of GNN models to adversarial fairness attacks. We hope our findings raise awareness about this issue in our community and lay a foundation for the future development of GNN models that are more robust to such attacks. Hussain Hussain, Sandipan Sikdar, Denis Helic, Elisabeth Lex, Markus Strohmaier, Roman Kern |
ICDM | 6 |
| 2021 | Global Gender Differences in Wikipedia Readership
Isaac L. Johnson, Florian Lemmerich, Diego Sáez-Trumper, Robert West 0001, Markus Strohmaier, Leila Zia |
ICWSM | 5 |
| 2021 | Sudden Attention Shifts on Wikipedia During the COVID-19 Crisis
Manoel Horta Ribeiro, Kristina Gligoric, Maxime Peyrard, Florian Lemmerich, Markus Strohmaier, Robert West 0001 |
ICWSM | 5 |
| 2021 | Redescription Model MiningabstractThis paper introduces Redescription Model Mining, a novel approach to identify interpretable patterns across two datasets that share only a subset of attributes and have no common instances. In particular, Redescription Model Mining aims to find pairs of describable data subsets -- one for each dataset -- that induce similar exceptional models with respect to a prespecified model class. To achieve this, we combine two previously separate research areas: Exceptional Model Mining and Redescription Mining. For this new problem setting, we develop interestingness measures to select promising patterns, propose efficient algorithms, and demonstrate their potential on synthetic and real-world data. Uncovered patterns can hint at common underlying phenomena that manifest themselves across datasets, enabling the discovery of possible associations between (combinations of) attributes that do not appear in the same dataset. Felix I. Stamm, Martin Becker 0003, Markus Strohmaier, Florian Lemmerich |
KDD | 3 |
| 2020 | Detecting Different Forms of Semantic Shift in Word Embeddings via Paradigmatic and Syntagmatic Association Changes
Anna Wegmann, Florian Lemmerich, Markus Strohmaier |
ISWC (1) | 3 |
| 2020 | The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word EmbeddingsabstractWe introduce ‘POLAR’ — a framework that adds interpretability to pre-trained word embeddings via the adoption of semantic differentials. Semantic differentials are a psychometric construct for measuring the semantics of a word by analysing its position on a scale between two polar opposites (e.g., cold – hot, soft – hard). The core idea of our approach is to transform existing, pre-trained word embeddings via semantic differentials to a new “polar” space with interpretable dimensions defined by such polar opposites. Our framework also allows for selecting the most discriminative dimensions from a set of polar dimensions provided by an oracle, i.e., an external source. We demonstrate the effectiveness of our framework by deploying it to various downstream tasks, in which our interpretable word embeddings achieve a performance that is comparable to the original word embeddings. We also show that the interpretable dimensions selected by our framework align with human judgement. Together, these results demonstrate that interpretability can be added to word embeddings without compromising performance. Our work is relevant for researchers and engineers interested in interpreting pre-trained word embeddings. Binny Mathew, Sandipan Sikdar, Florian Lemmerich, Markus Strohmaier |
WWW | 4 |
| 2019 | HopRank: How Semantic Structure Influences Teleportation in PageRank (A Case Study on BioPortal)abstractThis paper introduces HopRank, an algorithm for modeling human navigation on semantic networks. HopRank leverages the assumption that users know or can see the whole structure of the network. Therefore, besides following links, they also follow nodes at certain distances (i.e., k-hop neighborhoods), and not at random as suggested by PageRank, which assumes only links are known or visible. We observe such preference towards k-hop neighborhoods on BioPortal, one of the leading repositories of biomedical ontologies on the Web. In general, users navigate within the vicinity of a concept. But they also “jump” to distant concepts less frequently. We fit our model on 11 ontologies using the transition matrix of clickstreams, and show that semantic structure can influence teleportation in PageRank. This suggests that users-to some extent-utilize knowledge about the underlying structure of ontologies, and leverage it to reach certain pieces of information. Our results help the development and improvement of user interfaces for ontology exploration. Lisette Espin Noboa, Florian Lemmerich, Simon Walk, Markus Strohmaier, Mark A. Musen |
WWW | 4 |
| 2019 | Self- and Cross-Excitation in Stack Exchange Question & Answer CommunitiesabstractIn this paper, we quantify the impact of self- and cross-excitation on the temporal development of user activity in Stack Exchange Question & Answer (Q&A) communities. We study differences in user excitation between growing and declining Stack Exchange communities, and between those dedicated to STEM and humanities topics by leveraging Hawkes processes. We find that growing communities exhibit early stage, high cross-excitation by a small core of power users reacting to the community as a whole, and strong long-term self-excitation in general and cross-excitation by casual users in particular, suggesting community openness towards less active users. Further, we observe that communities in the humanities exhibit long-term power user cross-excitation, whereas in STEM communities activity is more evenly distributed towards casual user self-excitation. We validate our findings via permutation tests and quantify the impact of these excitation effects with a range of prediction experiments. Our work enables researchers to quantitatively assess the evolution and activity potential of Q&A communities. Simon Walk, Roman Kern, Markus Strohmaier, Denis Helic |
WWW | 4 |
| 2019 | Processing social media in real-time
Damiano Spina, Arkaitz Zubiaga, Amit P. Sheth, Markus Strohmaier |
Inf. Process. Manag. | 4 |
| 2018 | (Don't) Mention the War: A Comparison of Wikipedia and Britannica Articles on National HistoriesabstractIn this paper we present a large-scale quantitative comparison between expert- and crowdsourced writing of history by analysing articles from the English Wikipedia and Britannica. In order to quantify attention to particular periods, we extract mentioned year numbers and utilise them to study historical timelines of nations stretched over the last thousand years. By combining this temporal analysis with lexical analysis of both encyclopedic corpora we can identify distinctive historiographic points of view in each encyclopedia. We find that Britannica focuses on social and cultural phenomena, e.g. religion, as well as the geographical characteristics of states, while Wikipedia puts emphasis on political aspects, concentrating on wars and violent conflicts, and events of high popularity. Finally, both encyclopedias exhibit characteristics of English Academic prose, with Britannica being slightly less readable compared to Wikipedia, according to several readability scores. Anna Samoilenko, Florian Lemmerich, Maria Zens, Mohsen Jadidi, Mathieu Génois, Markus Strohmaier |
WWW | 6 |
| 2017 | Analysing Timelines of National Histories Across Wikipedia Editions: A Comparative Computational Approach
Anna Samoilenko, Florian Lemmerich, Katrin Weller, Maria Zens, Markus Strohmaier |
ICWSM | 5 |
| 2017 | Comparing Hypotheses About Sequential Data: A Bayesian Approach and Its Applications
Florian Lemmerich, Philipp Singer, Martin Becker 0003, Lisette Espin Noboa, Dimitar Dimitrov 0002, Denis Helic, Andreas Hotho, Markus Strohmaier |
ECML/PKDD (3) | 8 |
| 2017 | What Makes a Link Successful on Wikipedia?abstractWhile a plethora of hypertext links exist on the Web, only a small amount of them are regularly clicked. Starting from this observation, we set out to study large-scale click data from Wikipedia in order to understand what makes a link successful. We systematically analyze effects of link properties on the popularity of links. By utilizing mixed-effects hurdle models supplemented with descriptive insights, we find evidence of user preference towards links leading to the periphery of the network, towards links leading to semantically similar articles, and towards links in the top and left-side of the screen. We integrate these findings as Bayesian priors into a navigational Markov chain model and by doing so successfully improve the model fits. We further adapt and improve the well-known classic PageRank algorithm that assumes random navigation by accounting for observed navigational preferences of users in a weighted variation. This work facilitates understanding navigational click behavior and thus can contribute to improving link structures and algorithms utilizing these structures. Dimitar Dimitrov 0002, Philipp Singer, Florian Lemmerich, Markus Strohmaier |
WWW | 4 |
| 2017 | Why We Read WikipediaabstractWikipedia is one of the most popular sites on the Web, with millions of users relying on it to satisfy a broad range of information needs every day. Although it is crucial to understand what exactly these needs are in order to be able to meet them, little is currently known about why users visit Wikipedia. The goal of this paper is to fill this gap by combining a survey of Wikipedia readers with a log-based analysis of user activity. Based on an initial series of user surveys, we build a taxonomy of Wikipedia use cases along several dimensions, capturing users' motivations to visit Wikipedia, the depth of knowledge they are seeking, and their knowledge of the topic of interest prior to visiting Wikipedia. Then, we quantify the prevalence of these use cases via a large-scale user survey conducted on live Wikipedia with almost 30,000 responses. Our analyses highlight the variety of factors driving users to Wikipedia, such as current events, media coverage of a topic, personal curiosity, work or school assignments, or boredom. Finally, we match survey responses to the respondents' digital traces in Wikipedia's server logs, enabling the discovery of behavioral patterns associated with specific use cases. For instance, we observe long and fast-paced page sequences across topics for users who are bored or exploring randomly, whereas those using Wikipedia for work or school spend more time on individual articles focused on topics such as science. Our findings advance our understanding of reader motivations and behavior on Wikipedia and can have implications for developers aiming to improve Wikipedia's user experience, editors striving to cater to their readers' needs, third-party services (such as search engines) providing access to Wikipedia content, and researchers aiming to build tools such as recommendation engines. Philipp Singer, Florian Lemmerich, Robert West 0001, Leila Zia, Ellery Wulczyn, Markus Strohmaier, Jure Leskovec |
WWW | 6 |
| 2017 | Sampling from Social Networks with AttributesabstractSampling from large networks represents a fundamental challenge for social network research. In this paper, we explore the sensitivity of different sampling techniques (node sampling, edge sampling, random walk sampling, and snowball sampling) on social networks with attributes. We consider the special case of networks (i) where we have one attribute with two values (e.g., male and female in the case of gender), (ii) where the size of the two groups is unequal (e.g., a male majority and a female minority), and (iii) where nodes with the same or different attribute value attract or repel each other (i.e., homophilic or heterophilic behavior). We evaluate the different sampling techniques with respect to conserving the position of nodes and the visibility of groups in such networks. Experiments are conducted both on synthetic and empirical social networks. Our results provide evidence that different network sampling techniques are highly sensitive with regard to capturing the expected centrality of nodes, and that their accuracy depends on relative group size differences and on the level of homophily that can be observed in the network. We conclude that uninformed sampling from social networks with attributes thus can significantly impair the ability of researchers to draw valid conclusions about the centrality of nodes and the visibility or invisibility of groups in social networks. Claudia Wagner 0001, Philipp Singer, Fariba Karimi 0001, Jürgen Pfeffer, Markus Strohmaier |
WWW | 5 |
| 2017 | How Users Explore Ontologies on the Web: A Study of NCBO's BioPortal Usage LogsabstractOntologies in the biomedical domain are numerous, highly specialized and very expensive to develop. Thus, a crucial prerequisite for ontology adoption and reuse is effective support for exploring and finding existing ontologies. Towards that goal, the National Center for Biomedical Ontology (NCBO) has developed BioPortal---an online repository containing more than 500 biomedical ontologies. In 2016, BioPortal represents one of the largest portals for exploration of semantic biomedical vocabularies and terminologies, which is used by many researchers and practitioners. While usage of this portal is high, we know very little about how exactly users search and explore ontologies and what kind of usage patterns or user groups exist in the first place. Deeper insights into user behavior on such portals can provide valuable information to devise strategies for a better support of users in exploring and finding existing ontologies, and thereby enable better ontology reuse. To that end, we study and group users according to their browsing behavior on BioPortal and use data mining techniques to characterize and compare exploration strategies across ontologies. In particular, we were able to identify seven distinct browsing types, all relying on different functionality provided by BioPortal. For example, Search Explorers extensively use the search functionality while Ontology Tree Explorers mainly rely on the class hierarchy for exploring ontologies. Further, we show that specific characteristics of ontologies influence the way users explore and interact with the website. Our results may guide the development of more user-oriented systems for ontology exploration on the Web. Simon Walk, Lisette Espin Noboa, Denis Helic, Markus Strohmaier, Mark A. Musen |
WWW | 4 |
| 2017 | MixedTrails: Bayesian hypothesis comparison on heterogeneous sequential data
Martin Becker 0003, Florian Lemmerich, Philipp Singer, Markus Strohmaier, Andreas Hotho |
Data Min. Knowl. Discov. | 4 |
| 2017 | A Bayesian Method for Comparing Hypotheses About Human TrailsabstractWhen users interact with the Web today, they leave sequential digital trails on a massive scale. Examples of such human trails include Web navigation, sequences of online restaurant reviews, or online music play lists. Understanding the factors that drive the production of these trails can be useful, for example, for improving underlying network structures, predicting user clicks, or enhancing recommendations. In this work, we present a method called HypTrails for comparing a set of hypotheses about human trails on the Web, where hypotheses represent beliefs about transitions between states. Our method utilizes Markov chain models with Bayesian inference. The main idea is to incorporate hypotheses as informative Dirichlet priors and to calculate the evidence of the data under them. For eliciting Dirichlet priors from hypotheses, we present an adaption of the so-called (trial) roulette method, and to compare the relative plausibility of hypotheses, we employ Bayes factors. We demonstrate the general mechanics and applicability of HypTrails by performing experiments with (i) synthetic trails for which we control the mechanisms that have produced them and (ii) empirical trails stemming from different domains including Web site navigation, business reviews, and online music played. Our work expands the repertoire of methods available for studying human trails. Philipp Singer, Denis Helic, Andreas Hotho, Markus Strohmaier |
ACM Trans. Web | 4 |
| 2016 | Mining Subgroups with Exceptional Transition BehaviorabstractWe present a new method for detecting interpretable subgroups with exceptional transition behavior in sequential data. Identifying such patterns has many potential applications, e.g., for studying human mobility or analyzing the behavior of internet users. To tackle this task, we employ exceptional model mining, which is a general approach for identifying interpretable data subsets that exhibit unusual interactions between a set of target attributes with respect to a certain model class. Although exceptional model mining provides a well-suited framework for our problem, previously investigated model classes cannot capture transition behavior. To that end, we introduce first-order Markov chains as a novel model class for exceptional model mining and present a new interestingness measure that quantifies the exceptionality of transition subgroups. The measure compares the distance between the Markov transition matrix of a subgroup and the respective matrix of the entire data with the distance of random dataset samples. In addition, our method can be adapted to find subgroups that match or contradict given transition hypotheses. We demonstrate that our method is consistently able to recover subgroups with exceptional transition models from synthetic data and illustrate its potential in two application examples. Our work is relevant for researchers and practitioners interested in detecting exceptional transition behavior in sequential data. Florian Lemmerich, Martin Becker 0003, Philipp Singer, Denis Helic, Andreas Hotho, Markus Strohmaier |
KDD | 6 |
| 2016 | Steering the Random Surfer on Directed WebgraphsabstractEver since the inception of the Web website administrators have tried to steer user browsing behavior for a variety of reasons. For example, to be able to provide the most relevant information, for offering specific products, or to increase revenue from advertisements. One common approach to steer or bias the browsing behavior of users is to influence the link selection process by, for example, highlighting or repositioning links on a website. In this paper, we present a methodology for (i) expressing such navigational biases based on the random surfer model, and for (ii) measuring the consequences of the implemented biases. By adopting a model-based approach we are able to perform a wide range of experiments on seven empirical datasets. Our analyses allows us to gain novel insights into the consequences of navigational biases. Further, we unveil that navigational biases may have significant effects on the browsing processes of users and their typical whereabouts on a website. The first contribution of our work is the formalization of an approach to analyze consequences of navigational biases on the browsing dynamics and visit probabilities of specific pages of a website. Second, we apply this approach to analyze several empirical datasets and improve our understanding of the effects of different biases on real-world websites. In particular, we find that on webgraphs - contrary to undirected networks - typical biases always increase the certainty of the random surfer when selecting a link. Further, we observe significant side effects of biases, which indicate that for practical settings website administrators might need to carefully balance the desired outcomes against undesirable side effects. Florian Geigl, Simon Walk, Markus Strohmaier, Denis Helic |
WI | 3 |
| 2016 | The QWERTY Effect on the Web: How Typing Shapes the Meaning of Words in Online Human-Computer InteractionabstractThe QWERTY effect postulates that the keyboard layout influences word meanings by linking positivity to the use of the right hand and negativity to the use of the left hand. For example, previous research has established that words with more right hand letters are rated more positively than words with more left hand letters by human subjects in small scale experiments. In this paper, we perform large scale investigations of the QWERTY effect on the web. Using data from eleven web platforms related to products, movies, books, and videos, we conduct observational tests whether a hand-meaning relationship can be found in text interpretations by web users. Furthermore, we investigate whether writing text on the web exhibits the QWERTY effect as well, by analyzing the relationship between the text of online reviews and their star ratings in four additional datasets. Overall, we find robust evidence for the QWERTY effect both at the point of text interpretation (decoding) and at the point of text creation (encoding). We also find under which conditions the effect might not hold. Our findings have implications for any algorithmic method aiming to evaluate the meaning of words on the web, including for example semantic or sentiment analysis, and show the existence of "dactilar onomatopoeias" that shape the dynamics of word-meaning associations. To the best of our knowledge, this is the first work to reveal the extent to which the QWERTY effect exists in large scale human-computer interaction on the web. David García 0001, Markus Strohmaier |
WWW | 2 |
| 2016 | What Users Actually Do in a Social Tagging System: A Study of User Behavior in BibSonomyabstractSocial tagging systems have established themselves as an important part in today’s Web and have attracted the interest of our research community in a variety of investigations. Henceforth, several aspects of social tagging systems have been discussed and assumptions have emerged on which our community builds their work. Yet, testing such assumptions has been difficult due to the absence of suitable usage data in the past. In this work, we thoroughly investigate and evaluate four aspects about tagging systems, covering social interaction, retrieval of posted resources, the importance of the three different types of entities, users, resources, and tags, as well as connections between these entities’ popularity in posted and in requested content. For that purpose, we examine live server log data gathered from the real-world, public social tagging system BibSonomy. Our empirical results paint a mixed picture about the four aspects. Although typical assumptions hold to a certain extent for some, other aspects need to be reflected in a very critical light. Our observations have implications for the understanding of social tagging systems and the way they are used on the Web. We make the dataset used in this work available to other researchers. Stephan Doerfel, Daniel Zoller, Philipp Singer, Thomas Niebler, Andreas Hotho, Markus Strohmaier |
ACM Trans. Web | 6 |
| 2016 | Activity Dynamics in Collaboration NetworksabstractMany online collaboration networks struggle to gain user activity and become self-sustaining due to the ramp-up problem or dwindling activity within the system. Prominent examples include online encyclopedias such as (Semantic) MediaWikis, Question and Answering portals such as StackOverflow, and many others. Only a small fraction of these systems manage to reach self-sustaining activity, a level of activity that prevents the system from reverting to a nonactive state. In this article, we model and analyze activity dynamics in synthetic and empirical collaboration networks. Our approach is based on two opposing and well-studied principles: (i) without incentives, users tend to lose interest to contribute and thus, systems become inactive, and (ii) people are susceptible to actions taken by their peers (social or peer influence). With the activity dynamics model that we introduce in this article we can represent typical situations of such collaboration networks. For example, activity in a collaborative network, without external impulses or investments, will vanish over time, eventually rendering the system inactive. However, by appropriately manipulating the activity dynamics and/or the underlying collaboration networks, we can jump-start a previously inactive system and advance it toward an active state. To be able to do so, we first describe our model and its underlying mechanisms. We then provide illustrative examples of empirical datasets and characterize the barrier that has to be breached by a system before it can become self-sustaining in terms of critical mass and activity dynamics. Additionally, we expand on this empirical illustration and introduce a new metricp—theActivity Momentum—to assess the activity robustness of collaboration networks. Simon Walk, Denis Helic, Florian Geigl, Markus Strohmaier |
ACM Trans. Web | 4 |
| 2015 | Voting Behaviour and Power in Online Democracy: A Study of LiquidFeedback in Germany's Pirate Party
Christoph Carl Kling, Jérôme Kunegis, Heinrich Hartmann, Markus Strohmaier, Steffen Staab |
ICWSM | 4 |
| 2015 | It's a Man's Wikipedia? Assessing Gender Inequality in an Online Encyclopedia
Claudia Wagner 0001, David García 0001, Mohsen Jadidi, Markus Strohmaier |
ICWSM | 4 |
| 2015 | Associating Intent with Sentiment in Weblogs
Mark Kröll, Markus Strohmaier |
NLDB | 2 |
| 2015 | Understanding How Users Edit Ontologies: Comparing Hypotheses About Four Real-World Projects
Simon Walk, Philipp Singer, Lisette Espin Noboa, Tania Tudorache, Mark A. Musen, Markus Strohmaier |
ISWC (1) | 6 |
| 2015 | HypTrails: A Bayesian Approach for Comparing Hypotheses About Human Trails on the WebabstractWhen users interact with the Web today, they leave sequential digital trails on a massive scale. Examples of such human trails include Web navigation, sequences of online restaurant reviews, or online music play lists. Understanding the factors that drive the production of these trails can be useful for e.g., improving underlying network structures, predicting user clicks or enhancing recommendations. In this work, we present a general approach called HypTrails for comparing a set of hypotheses about human trails on the Web, where hypotheses represent beliefs about transitions between states. Our approach utilizes Markov chain models with Bayesian inference. The main idea is to incorporate hypotheses as informative Dirichlet priors and to leverage the sensitivity of Bayes factors on the prior for comparing hypotheses with each other. For eliciting Dirichlet priors from hypotheses, we present an adaption of the so-called (trial) roulette method. We demonstrate the general mechanics and applicability of HypTrails by performing experiments with (i) synthetic trails for which we control the mechanisms that have produced them and (ii) empirical trails stemming from different domains including website navigation, business reviews and online music played. Our work expands the repertoire of methods available for studying human trails on the Web. Philipp Singer, Denis Helic, Andreas Hotho, Markus Strohmaier |
WWW | 4 |
| 2014 | Sequential Action Patterns in Collaborative Ontology-Engineering Projects: A Case-Study in the Biomedical DomainabstractWithin the last few years the importance of collaborative ontology-engineering projects, especially in the biomedical domain, has drastically increased. This recent trend is a direct consequence of the growing complexity of these structured data representations, which no single individual is able to handle anymore. For example, the World Health Organization is currently actively developing the next revision of the International Classification of Diseases (ICD), using an OWL-based core for data representation and Web 2.0 technologies to augment collaboration. This new revision of ICD consists of roughly 50,000 diseases and causes of death and is used in many countries around the world to encode patient history, to compile health-related statistics and spendings. Hence, it is crucial for practitioners to better understand and steer the underlying processes of how users collaboratively edit an ontology. Particularly, generating predictive models is a pressing issue as these models may be leveraged for generating recommendations in collaborative ontology-engineering projects and to determine the implications of potential actions on the ontology and community. Simon Walk, Philipp Singer, Markus Strohmaier |
CIKM | 3 |
| 2014 | When Politicians Talk: Assessing Online Conversational Practices of Political Parties on Twitter
Haiko Lietz, Claudia Wagner 0001, Arnim Bleier, Markus Strohmaier |
ICWSM | 4 |
| 2014 | Semantic stability in social tagging streamsabstractOne potential disadvantage of social tagging systems is that due to the lack of a centralized vocabulary, a crowd of users may never manage to reach a consensus on the description of resources (e.g., books, users or songs) on the Web. Yet, previous research has provided interesting evidence that the tag distributions of resources may become semantically stable over time as more and more users tag them. At the same time, previous work has raised an array of new questions such as: (i) How can we assess the semantic stability of social tagging systems in a robust and methodical way? (ii) Does semantic stabilization of tags vary across different social tagging systems and ultimately, (iii) what are the factors that can explain semantic stabilization in such systems? In this work we tackle these questions by (i) presenting a novel and robust method which overcomes a number of limitations in existing methods, (ii) empirically investigating semantic stabilization processes in a wide range of social tagging systems with distinct domains and properties and (iii) detecting potential causes for semantic stabilization, specifically imitation behavior, shared background knowledge and intrinsic properties of natural language. Our results show that tagging streams which are generated by a combination of imitation dynamics and shared background knowledge exhibit faster and higher semantic stability than tagging streams which are generated via imitation dynamics or natural language phenomena alone. Claudia Wagner 0001, Philipp Singer, Markus Strohmaier, Bernardo A. Huberman |
WWW | 3 |
| 2013 | How Tagging Pragmatics Influence Tag Sense Discovery in Social Annotation Systems
Thomas Niebler, Philipp Singer, Dominik Benz, Christian Körner, Markus Strohmaier, Andreas Hotho |
ECIR | 5 |
| 2013 | Measuring the Topical Specificity of Online Communities
Matthew Rowe 0001, Claudia Wagner 0001, Markus Strohmaier, Harith Alani |
ESWC | 3 |
| 2013 | The Wisdom of the Audience: An Empirical Study of Social Semantics in Twitter Streams
Claudia Wagner 0001, Philipp Singer, Lisa Posch, Markus Strohmaier |
ESWC | 4 |
| 2013 | Computing Semantic Relatedness from Human Navigational Paths: A Case Study on WikipediaabstractIn this article, the authors present a novel approach for computing semantic relatedness and conduct a large-scale study of it on Wikipedia. Unlike existing semantic analysis methods that utilize Wikipedia’s content or link structure, the authors propose to use human navigational paths on Wikipedia for this task. The authors obtain 1.8 million human navigational paths from a semi-controlled navigation experiment – a Wikipedia-based navigation game, in which users are required to find short paths between two articles in a given Wikipedia article network. The authors’ results are intriguing: They suggest that (i) semantic relatedness computed from human navigational paths may be more precise than semantic relatedness computed from Wikipedia’s plain link structure alone and (ii) that not all navigational paths are equally useful. Intelligent selection based on path characteristics can improve accuracy. The authors’ work makes an argument for expanding the existing arsenal of data sources for calculating semantic relatedness and to consider the utility of human navigational paths for this task. Philipp Singer, Thomas Niebler, Markus Strohmaier, Andreas Hotho |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2013 | PragmatiX: An Interactive Tool for Visualizing the Creation Process Behind Collaboratively Engineered OntologiesabstractWith the emergence of tools for collaborative ontology engineering, more and more data about the creation process behind collaborative construction of ontologies is becoming available. Today, collaborative ontology engineering tools such as Collaborative Protégé offer rich and structured logs of changes, thereby opening up new challenges and opportunities to study and analyze the creation of collaboratively constructed ontologies. While there exists a plethora of visualization tools for ontologies, they have primarily been built to visualize aspects of the final product (the ontology) and not the collaborative processes behind construction (e.g. the changes made by contributors over time). To the best of the authors’ knowledge, there exists no ontology visualization tool today that focuses primarily on visualizing the history behind collaboratively constructed ontologies. Since the ontology engineering processes can influence the quality of the final ontology, they believe that visualizing process data represents an important stepping-stone towards better understanding of managing the collaborative construction of ontologies in the future. In this application paper, the authors present a tool – PragmatiX – which taps into structured change logs provided by tools such as Collaborative Protégé to visualize various pragmatic aspects of collaborative ontology engineering. The tool is aimed at managers and leaders of collaborative ontology engineering projects to help them in monitoring progress, in exploring issues and problems, and in tracking quality-related issues such as overrides and coordination among contributors. The paper makes the following contributions: (i) They present PragmatiX, a tool for visualizing the creation process behind collaboratively constructed ontologies (ii) the authors illustrate the functionality and generality of the tool by applying it to structured logs of changes of two large collaborative ontology-engineering projects and (iii) they conduct a heuristic evaluation of the tool with domain experts to uncover early design challenges and opportunities for improvement. Finally, the authors hope that this work sparks a new line of research on visualization tools for collaborative ontology engineering projects. Simon Walk, Jan Pöschko, Markus Strohmaier, Keith Andrews, Tania Tudorache, Natasha F. Noy, Csongor Nyulas, Mark A. Musen |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2013 | How ontologies are made: Studying the hidden social dynamics behind collaborative ontology engineering projects
Markus Strohmaier, Simon Walk, Jan Pöschko, Daniel Lamprecht, Tania Tudorache, Csongor Nyulas, Mark A. Musen, Natasha F. Noy |
J. Web Semant. | 1 |
| 2012 | What Catches Your Attention? An Empirical Study of Attention Patterns in Community Forums
Claudia Wagner 0001, Matthew Rowe 0001, Markus Strohmaier, Harith Alani |
ICWSM | 3 |
| 2012 | Acquiring knowledge about human goals from Search Query Logs
Markus Strohmaier, Mark Kröll |
Inf. Process. Manag. | 1 |
| 2012 | Evaluation of Folksonomy Induction AlgorithmsabstractAlgorithms for constructing hierarchical structures from user-generated metadata have caught the interest of the academic community in recent years. In social tagging systems, the output of these algorithms is usually referred to as folksonomies (from folk-generated taxonomies). Evaluation of folksonomies and folksonomy induction algorithms is a challenging issue complicated by the lack of golden standards, lack of comprehensive methods and tools as well as a lack of research and empirical/simulation studies applying these methods. In this article, we report results from a broad comparative study of state-of-the-art folksonomy induction algorithms that we have applied and evaluated in the context of five social tagging systems. In addition to adopting semantic evaluation techniques, we present and adopt a new technique that can be used to evaluate the usefulness of folksonomies for navigation . Our work sheds new light on the properties and characteristics of state-of-the-art folksonomy induction algorithms and introduces a new pragmatic approach to folksonomy evaluation, while at the same time identifying some important limitations and challenges of folksonomy evaluation. Our results show that folksonomy induction algorithms specifically developed to capture intuitions of social tagging systems outperform traditional hierarchical clustering techniques. To the best of our knowledge, this work represents the largest and most comprehensive evaluation study of state-of-the-art folksonomy induction algorithms to date. Markus Strohmaier, Denis Helic, Dominik Benz, Christian Körner, Roman Kern |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | Understanding why users tag: A survey of tagging motivation literature and results from an empirical studyabstractWhile recent progress has been achieved in understanding the structure and dynamics of social tagging systems, we know little about the underlying user motivations for tagging, and how they influence resulting folksonomies and tags. This paper addresses three issues related to this question. (1) What distinctions of user motivations are identified by previous research, and in what ways are the motivations of users amenable to quantitative analysis? (2) To what extent does tagging motivation vary across different social tagging systems? (3) How does variability in user motivation influence resulting tags and folksonomies? In this paper, we present measures to detect whether a tagger is primarily motivated by categorizing or describing resources, and apply these measures to datasets from seven different tagging systems. Our results show that (a) users’ motivation for tagging varies not only across, but also within tagging systems, and that (b) tag agreement among users who are motivated by categorizing resources is significantly lower than among users who are motivated by describing resources. Our findings are relevant for (1) the development of tag-based user interfaces, (2) the analysis of tag semantics and (3) the design of search algorithms for social tagging systems. Markus Strohmaier, Christian Körner, Roman Kern |
J. Web Semant. | 1 |
| 2011 | Building directories for social tagging systemsabstractToday, a number of algorithms exist for constructing tag hierarchies from social tagging data. While these algorithms were designed with ontological goals in mind, we know very little about their properties from an information retrieval perspective, such as whether these tag hierarchies support efficient navigation in social tagging systems. The aim of this paper is to investigate the usefulness of such tag hierarchies (sometimes also called folksonomies - from folk-generated taxonomy) as directories that aid navigation in social tagging systems. To this end, we simulate navigation of directories as decentralized search on a network of tags using Kleinberg's model. In this model, a tag hierarchy can be applied as background knowledge for decentralized search. By constraining the visibility of nodes in the directories we aim to mimic typical constraints imposed by a practical user interface (UI), such as limiting the number of displayed subcategories or related categories. Our experiments on five different social tagging datasets show that existing tag hierarchy algorithms can support navigation in theory, but our results also demonstrate that they face tremendous challenges when user interface (UI) restrictions are taken into account. Based on this observation, we introduce a new algorithm that constructs efficiently navigable directories on our datasets. The results are relevant for engineers and scientists aiming to improve navigability of social tagging systems. Denis Helic, Markus Strohmaier |
CIKM | 2 |
| 2011 | One Tag to Bind Them All: Measuring Term Abstractness in Social Metadata
Dominik Benz, Christian Körner, Andreas Hotho, Gerd Stumme, Markus Strohmaier |
ESWC (2) | 5 |
| 2011 | Automatically Constructing Concept Hierarchies of Health-Related Human Goals
Mark Kröll, Yusuke Fukazawa, Jun Ota 0001, Markus Strohmaier |
KSEM | 4 |
| 2011 | Pragmatic evaluation of folksonomiesabstractRecently, a number of algorithms have been proposed to obtain hierarchical structures - so-called folksonomies - from social tagging data. Work on these algorithms is in part driven by a belief that folksonomies are useful for tasks such as: (a) Navigating social tagging systems and (b) Acquiring semantic relationships between tags. While the promises and pitfalls of the latter have been studied to some extent, we know very little about the extent to which folksonomies are pragmatically useful for navigating social tagging systems. This paper sets out to address this gap by presenting and applying a pragmatic framework for evaluating folksonomies. We model exploratory navigation of a tagging system as decentralized search on a network of tags. Evaluation is based on the fact that the performance of a decentralized search algorithm depends on the quality of the background knowledge used. The key idea of our approach is to use hierarchical structures learned by folksonomy algorithm as background knowledge for decentralized search. Utilizing decentralized search on tag networks in combination with different folksonomies as hierarchical background knowledge allows us to evaluate navigational tasks in social tagging systems. Our experiments with four state-of-the-art folksonomy algorithms on five different social tagging datasets reveal that existing folksonomy algorithms exhibit significant, previously undiscovered, differences with regard to their utility for navigation. Our results are relevant for engineers aiming to improve navigability of social tagging systems and for scientists aiming to evaluate different folksonomy algorithms from a pragmatic perspective. Denis Helic, Markus Strohmaier, Christoph Trattner, Markus Muhr, Kristina Lerman |
WWW | 2 |
| 2010 | Why do Users Tag? Detecting Users' Motivation for Tagging in Social Tagging Systems
Markus Strohmaier, Christian Körner, Roman Kern |
ICWSM | 1 |
| 2010 | Stop thinking, start tagging: tag semantics emerge from collaborative verbosityabstractRecent research provides evidence for the presence of emergent semantics in collaborative tagging systems. While several methods have been proposed, little is known about the factors that influence the evolution of semantic structures in these systems. A natural hypothesis is that the quality of the emergent semantics depends on the pragmatics of tagging: Users with certain usage patterns might contribute more to the resulting semantics than others. In this work, we propose several measures which enable a pragmatic differentiation of taggers by their degree of contribution to emerging semantic structures. We distinguish between categorizers, who typically use a small set of tags as a replacement for hierarchical classification schemes, and describers, who are annotating resources with a wealth of freely associated, descriptive keywords. To study our hypothesis, we apply semantic similarity measures to 64 different partitions of a real-world and large-scale folksonomy containing different ratios of categorizers and describers. Our results not only show that "verbose" taggers are most useful for the emergence of tag semantics, but also that a subset containing only 40% of the most 'verbose' taggers can produce results that match and even outperform the semantic precision obtained from the whole dataset. Moreover, the results suggest that there exists a causal link between the pragmatics of tagging and resulting emergent semantics. This work is relevant for designers and analysts of tagging systems interested (i) in fostering the semantic development of their platforms, (ii) in identifying users introducing "semantic noise", and (iii) in learning ontologies. Christian Körner, Dominik Benz, Andreas Hotho, Markus Strohmaier, Gerd Stumme |
WWW | 4 |
| 2009 | Analyzing human intentions in natural language textabstractIn this paper, we introduce the idea of Intent Analysis, which is to create a profile of the goals and intentions present in textual content. Intent Analysis, similar to Sentiment Analysis, represents a type of document classification that differs from traditional topic categorization by focusing on classification by intent. We investigate the extent to which the automatic analysis of human intentions in text is feasible and report our preliminary results, and discuss potential applications. In addition, we present results from a study that focused on evaluating intent profiles generated from transcripts of American presidential candidate speeches in 2008. Mark Kröll, Markus Strohmaier |
K-CAP | 2 |
| 2009 | Studying databases of intentions: do search query logs capture knowledge about common human goals?abstractAccess to knowledge about common human goals has been found critical for realizing the vision of intelligent agents acting upon user intent on the web. Yet, the acquisition of knowledge about common human goals represents a major challenge. In a departure from existing approaches, this paper investigates a novel resource for knowledge acquisition: The utilization of search query logs for this task. By relating goals contained in search query logs with goals contained in existing commonsense knowledge bases such as ConceptNet, we aim to shed light on the usefulness of search query logs for capturing knowledge about common human goals. The main contribution of this paper consists of an empirical study comparing common human goals contained in two large search query logs (AOL and Microsoft Research) with goals contained in the commonsense knowledge base ConceptNet. The paper sketches ways how goals from search query logs could be used to address the goal acquisition and goal coverage problem related to common-sense knowledge bases. Markus Strohmaier, Mark Kröll |
K-CAP | 1 |