VLDB 2026 Research / reviewers in the wild / expert
Tim Weninger
dblp:73/2015
· DBLP profile ↗
42ranked-venue papers in the field
7as first author
13since 2021 · last 2025
0000-0003-3164-2615ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 16 (4 first)Data Mining & Knowledge Discovery · 15 (2 first)Database Systems & Data Management · 5 (1 first)Big Data, Cloud & Distributed Data Systems · 5Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Modeling Information Narrative Evolution on Telegram During the Russia-Ukraine WarabstractFollowing the Russian Federation's full-scale invasion of Ukraine in February 2022, a multitude of information narratives emerged within both pro-Russian and pro-Ukrainian communities online. As the conflict progresses, so too do the information narratives, constantly adapting and influencing local and global community perceptions and attitudes. This dynamic nature of the evolving information environment (IE) underscores a critical need to fully discern how narratives evolve and affect online communities. Existing research, however, often fails to capture information narrative evolution, overlooking both the fluid nature of narratives and the internal mechanisms that drive their evolution. Recognizing this, we introduce a novel approach designed to both model narrative evolution and uncover the underlying mechanisms driving them. In this work we perform a comparative discourse analysis across communities on Telegram covering the initial three months following the invasion. First, we uncover substantial disparities in narratives and perceptions between pro-Russian and pro-Ukrainian communities. Then, we probe deeper into prevalent narratives of each group, identifying key themes and examining the underlying mechanisms fueling their evolution. Finally, we explore influences and factors that may shape the development and spread of narratives. Patrick Gérard, Svitlana Volkova, Louis Penafiel, Kristina Lerman, Tim Weninger |
ICWSM | 5 |
| 2025 | Fear and Loathing on the Frontline: Decoding the Language of Othering by Russia-Ukraine War BloggersabstractOthering—the process of portraying an outgroup as fundamentally different and inferior—often escalates into framing the outgroup as an existential threat, thereby legitimizing exclusion and violence. Throughout history, othering has played a central role in conflicts, from genocides in Nazi Germany and Rwanda to contemporary hostility toward migrants in the US and Europe. Traditional computational methods, such as those used for hate speech detection, frequently overlook the subtle, context-dependent nature of othering language, limiting their effectiveness in real-time detection and analysis. Our work addresses these limitations through three key contributions: (1) a computational framework that combines sociological theory with large language models (LLMs) to identify and analyze othering language, (2) an in-depth examination of othering discourse dynamics, focusing on attention patterns and its interplay with moral framing, and (3) a rapid domain adaptation enabling robust analysis across different platforms and contexts. We apply our framework to a large corpus of Telegram messages from Russo-Ukrainian war bloggers and political discourse on Gab, revealing several previously unquantified patterns: othering rhetoric surges during crises, often intertwines with moralized language, and escalates during critical periods. Our findings demonstrate that this approach not only surpasses existing hate and fear speech detection methods but also offers actionable insights for anticipating and mitigating threats to social cohesion in conflict-prone environments. Patrick Gérard, Tim Weninger, Kristina Lerman |
ICWSM | 2 |
| 2025 | SCHENO: Measuring Schema vs. Noise in GraphsabstractReal-world data is typically a noisy manifestation of a core pattern (schema), and the purpose of data mining algorithms is to uncover that pattern, thereby splitting (i.e.decomposing) the data into schema and noise. We introduce SCHENO, a principled evaluation metric for the goodness of a schema-noise decomposition of a graph. SCHENO captures how schematic the schema is, how noisy the noise is, and how well the combination of the two represent the original graph data. We visually demonstrate what this metric prioritizes in small graphs, then show that if SCHENO is used as the fitness function for a simple optimization strategy, we can uncover a wide variety of patterns. Finally, we evaluate several well-known graph mining algorithms with this metric; we find that although they produce patterns, those patterns are not always the best representation of the input data. Justus Hibshman, Adnan Hoq, Tim Weninger |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Navigating the Post-API DilemmaabstractRecent decisions to discontinue access to social media APIs are having detrimental effects on Internet research and the field of computational social science as a whole. This lack of access to data has been dubbed the Post-API era of Internet research. Fortunately, popular search engines have the means to crawl, capture, and surface social media data on their Search Engine Results Pages (SERP) if provided the proper search query, and may provide a solution to this dilemma. In the present work we ask: does SERP provide a complete and unbiased sample of social media data? Is SERP a viable alternative to direct API-access? To answer these questions, we perform a comparative analysis between (Google) SERP results and nonsampled data from Reddit and Twitter/X. We find that SERP results are highly biased in favor of popular posts; against political, pornographic, and vulgar posts; are more positive in their sentiment; and have large topical gaps. Overall, we conclude that SERP is not a viable alternative to social media API access. Amrit Poudel, Tim Weninger |
WWW | 2 |
| 2023 | Truth Social DatasetabstractFormally announced to the public following former President Donald Trump’s bans and suspensions from mainstream social networks in early 2022 following his role in the January 6 Capitol Riots, Truth Social was launched as an ``alternative'' social media platform that claims to be a refuge for free speech, offering a platform for those disaffected by the content moderation policies of then existing, mainstream social networks. The subsequent rise of Truth Social has been driven largely by hard-line supporters of the former president as well as those affected by the content moderation of other social networks. These distinct qualities combined with the its status as the main mouthpiece of the former president positions Truth Social as a particularly influential social media platform and give rise to several research questions. However, outside of a handful of news reports, little is known about the new social media platform partially due to a lack of well-curated data. In the current work, we describe a dataset of over 823,000 posts to Truth Social and and social network with over 454,000 distinct users. In addition to the dataset itself, we also present some basic analysis of its content, certain temporal features, and its network. Patrick Gérard, Nicholas Botzer, Tim Weninger |
ICWSM | 3 |
| 2023 | Entity graphs for exploring online discourse
Nicholas Botzer, Tim Weninger |
Knowl. Inf. Syst. | 2 |
| 2023 | The Infinity Mirror Test for Graph ModelsabstractGraph models, like other machine learning models, have implicit and explicit biases built-in, which often impact performance in nontrivial ways. The model’s faithfulness is often measured by comparing the newly generated graph against the source graph using any number of graph properties. Therefore, differences in the size or topology of the generated graph indicate a loss in the model. Yet, in many systems, errors encoded in loss functions are subtle and not well understood. In the present work, we introduce theInfinity Mirrortest for analyzing the robustness of graph models. This straightforward stress test works by repeatedly fitting a model to its outputs. A hypothetically perfect graph model would have no deviation from the source graph; however, a model’s implicit biases and assumptions are exaggerated by the Infinity Mirror test, exposing potential previously obscured issues. Through an analysis of thousands of experiments on synthetic and real-world graphs, we show that several conventional graph models degenerate in exciting and informative ways. We believe that the observed degenerative patterns are clues to the future development of better graph models. Satyaki Sikdar, Daniel Gonzalez 0001, Trenton Ford, Tim Weninger |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Subreddit Links Drive Community Creation and User Engagement on Reddit
Rachel Krohn, Tim Weninger |
ICWSM | 2 |
| 2022 | 17th International Workshop on Mining and Learning with Graphs (MLG)abstractThe 17th International Workshop on Mining and Learning with Graphs (MLG) is held in Washington DC, USA on August 15, 2022 and is co-located with the Eighth International Workshop on Deep Learning on Graphs (DLG) as part of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. This workshop is a forum for exchanging ideas and methods for mining and learning with graphs, developing new common understandings of the problems at hand, sharing of data sets where applicable, and leveraging existing knowledge from different disciplines. In doing so, we aim to better understand the overarching principles and the limitations of our current methods, and to inspire research on new algorithms and techniques for mining and learning with graphs. Topics of interest include, but are not limited to, graph mining, statistical relational learning, social network analysis, and network science. The target audience spans researchers and practitioners across academia, government, and industry. Shobeir Fakhraei, Tim Weninger, Neil Shah, Sami Abu-El-Haija, Saurabh Verma, Tara Safavi |
KDD | 2 |
| 2022 | Attributed Graph Modeling with Vertex Replacement GrammarsabstractRecent work at the intersection of formal language theory and graph theory has explored graph grammars for graph modeling. However, existing models and formalisms can only operate on homogeneous (i.e., untyped or unattributed) graphs. We relax this restriction and introduce the Attributed Vertex Replacement Grammar (AVRG), which can be efficiently extracted from heterogeneous (i.e., typed, colored, or attributed) graphs. Unlike current state-of-the-art methods, which train enormous models over complicated deep neural architectures, the AVRG model is unsupervised and interpretable. It is based on context-free string grammars and works by encoding graph rewriting rules into a graph grammar containing graphlets and instructions on how they fit together. We show that the AVRG can encode succinct models of input graphs yet faithfully preserve their structure and assortativity properties. Experiments on large real-world datasets show that graphs generated from the AVRG model exhibit substructures and attribute configurations that match those found in the input networks. Satyaki Sikdar, Neil Shah, Tim Weninger |
WSDM | 3 |
| 2021 | Automatic Discovery of Political Meme Genres with Diverse Appearances
William Theisen, Joel Brogan, Pamela Bilo Thomas, Daniel Moreira, Pascal Phoa, Tim Weninger, Walter J. Scheirer |
ICWSM | 6 |
| 2021 | Joint Subgraph-to-Subgraph Transitions: Generalizing Triadic Closure for Powerful and Interpretable Graph ModelingabstractWe generalize triadic closure, along with previous generalizations of triadic closure, under an intuitive umbrella generalization: the Subgraph-to-Subgraph Transition (SST). We present algorithms and code to model graph evolution in terms of collections of these SSTs. We then use the SST framework to create link prediction models for both static and temporal, directed and undirected graphs which produce highly interpretable results. Quantitatively, our models match out-of-the-box performance of state of the art graph neural network models, thereby validating the correctness and meaningfulness of our interpretable results. Justus Hibshman, Daniel Gonzalez 0001, Satyaki Sikdar, Tim Weninger |
WSDM | 4 |
| 2021 | Reddit entity linking dataset
Nicholas Botzer, Yifan Ding 0001, Tim Weninger |
Inf. Process. Manag. | 3 |
| 2019 | Dynamics of team library adoptions: an exploration of GitHub commit logsabstractWhen a group of people strives to understand new information, struggle ensues as various ideas compete for attention. Steep learning curves are surmounted as teams learn together. To understand how these team dynamics play out in software development, we explore Git logs, which provide a complete change history of software repositories. In these repositories, we observe code additions, which represent successfully implemented ideas, and code deletions, which represent ideas that have failed or been superseded. By examining the patterns between these commit types, we can begin to understand how teams adopt new information. We specifically study what happens after a software library is adopted by a project, i.e., when a library is used for the first time in the project. We find that a variety of factors, including team size, library popularity, and prevalence on Stack Overflow are associated with how quickly teams learn and successfully adopt new software libraries. Pamela Bilo Thomas, Rachel Krohn, Tim Weninger |
ASONAM | 3 |
| 2019 | Preserving Composition and Crystal Structures of Chemical Compounds in Atomic EmbeddingabstractWe develop a new representation learning method in the chemistry domain. Given a large set of compounds of inorganic crystals, the extraction model learns the embeddings of atoms so that the predictive model can place them into the periodic table correctly. Our method preserves not only the compounds' compositions but also their crystal structures. Experiments demonstrate the effectiveness of the proposed method, compared to the state-of-the-art method (in PNAS 2018). Yifan Ding 0001, Daheng Wang, Tim Weninger, Meng Jiang 0001 |
IEEE BigData | 3 |
| 2019 | Towards Interpretable Graph Modeling with Vertex Replacement GrammarsabstractAn enormous amount of real-world data exists in the form of graphs. Oftentimes, interesting patterns that describe the complex dynamics of these graphs are captured in the form of frequently reoccurring substructures. Recent work at the intersection of formal language theory and graph theory has explored the use of graph grammars for graph modeling and pattern mining. However, existing formulations do not extract meaningful and easily interpretable patterns from the data. The present work addresses this limitation by extracting a special type of vertex replacement grammar, which we call a KT grammar, according to the Minimum Description Length (MDL) heuristic. In experiments on synthetic and real-world datasets, we show that KT-grammars can be efficiently extracted from a graph and that these grammars encode meaningful patterns that represent the dynamics of the real-world system. Justus Hibshman, Satyaki Sikdar, Tim Weninger |
IEEE BigData | 3 |
| 2019 | Modelling Online Comment Threads from their StartabstractThe social Web is a widely used platform for online discussion. Across social media, users can start discussions by posting a topical image, url, or message. Upon seeing this initial post, other users may add their own comments to the post, or to another user's comment. The resulting online discourse produces a comment thread, which constitutes an enormous portion of modern online communication. Comment threads are often viewed as trees: nodes represent the post and its comments, while directed edges represent reply-to relationships. The goal of the present work is to predict the size and shape of these comment threads. Existing models do this by observing the first several comments and then fitting a predictive model. However, most comment threads are relatively small, and waiting for data to materialize runs counter to the goal of the prediction task. We therefore introduce the Comment Thread Prediction Model (CTPM) that accurately predicts the size and shape of a comment thread using only the text of the initial post, allowing for the prediction of new posts without observable comments. We find that the CTPM significantly outperforms existing models and competitive baselines on thousands of Reddit discussions from nine varied subreddits, particularly for new posts. Rachel Krohn, Tim Weninger |
IEEE BigData | 2 |
| 2019 | Representation Learning in Heterogeneous Professional Social Networks with Ambiguous Social ConnectionsabstractNetwork representations have been shown to improve performance within a variety of tasks, including classification, clustering, and link prediction. However, most models either focus on moderate-sized, homogeneous networks or require a significant amount of auxiliary input to be provided by the user. Moreover, few works have studied network representations in real-world heterogeneous social networks with ambiguous social connections and are often incomplete. In the present work, we investigate the problem of learning low-dimensional node representations in heterogeneous professional social networks (HPSNs), which are incomplete and have ambiguous social connections. We present a general heterogeneous network representation learning model called Star2Vec that learns entity and person embeddings jointly using a social connection strength-aware biased random walk combined with a node-structure expansion function. Experiments on LinkedIn's Economic Graph and publicly available snapshots of Facebook's network show that Star2Vec outperforms existing methods on members' industry and social circle classification, skill and title clustering, and member-entity link predictions. We also conducted large-scale case studies to demonstrate practical applications of the Star2Vec embeddings trained on LinkedIn's Economic Graph such as next career move, alternative career suggestions, and general entity similarity searches. Baoxu Shi, Jaewon Yang, Tim Weninger, How Jing, Qi He 0002 |
IEEE BigData | 3 |
| 2019 | Modeling Graphs with Vertex Replacement GrammarsabstractOne of the principal goals of graph modeling is to capture the building blocks of network data in order to study various physical and natural phenomena. Recent work at the intersection of formal language theory and graph theory has explored the use of graph grammars for graph modeling. However, existing graph grammar formalisms, like Hyperedge Replacement Grammars, can only operate on small tree-like graphs. The present work relaxes this restriction by revising a different graph grammar formalism called Vertex Replacement Grammars (VRGs). We show that a variant of the VRG called Clustering-based Node Replacement Grammar (CNRG) can be efficiently extracted from many hierarchical clusterings of a graph. We show that CNRGs encode a succinct model of the graph, yet faithfully preserves the structure of the original graph. In experiments on large real-world datasets, we show that graphs generated from the CNRG model exhibit a diverse range of properties that are similar to those found in the original networks. Satyaki Sikdar, Justus Hibshman, Tim Weninger |
ICDM | 3 |
| 2018 | How Humans Versus Bots React to Deceptive and Trusted News Sources: A Case Study of Active UsersabstractSociety`s reliance on social media as a primary source of news has spawned a renewed focus on the spread of misinformation. In this work, we identify the differences in how social media accounts identified as bots react to news sources of varying credibility, regardless of the veracity of the content those sources have shared. We analyze bot and human responses annotated using a fine-grained model that labels responses as being an answer, appreciation, agreement, disagreement, an elaboration, humor, or a negative reaction. We present key findings of our analysis into the prevalence of bots, the variety and speed of bot and human reactions, and the disparity in authorship of reaction tweets between these two sub-populations. We observe that bots are responsible for 9-15% of the reactions to sources of any given type but comprise only 7-10% of accounts responsible for reaction-tweets; trusted news sources have the highest proportion of humans who reacted; bots respond with significantly shorter delays than humans when posting answer-reactions in response to sources identified as propaganda. Finally, we report significantly different inequality levels in reaction rates for accounts identified as bots vs not. Maria Glenski, Tim Weninger, Svitlana Volkova |
ASONAM | 2 |
| 2018 | Synchronous Hyperedge Replacement Graph Grammars
Corey Pennycuff, Satyaki Sikdar, Catalina Vajiac, David Chiang 0001, Tim Weninger |
ICGT | 5 |
| 2018 | HeteroNAM: International Workshop on Heterogeneous Networks Analysis and MiningabstractThe first International Workshop on Heterogeneous Networks Analysis and Mining is held in Los Angeles, California, USA on February 9th, 2018 and is co-located with the 11th ACM International Conference on Web Search and Data Mining. The goal of this workshop is to bring together computing researchers and practitioners to address challenges in the mining and analysis of real-world heterogeneous networks. This workshop has an exciting program that spans a number of subareas including: graph mining, learning from structured data, statistical relational learning, and network science in general. The program includes six invited speakers, lively discussion on emerging topics, and presentations of several original research papers. Shobeir Fakhraei, Yanen Li, Yizhou Sun, Tim Weninger |
WSDM | 4 |
| 2017 | Predicting User-Interactions on RedditabstractIn order to keep up with the demand of curating the deluge of crowd-sourced content, social media platforms leverage user interaction feedback to make decisions about which content to display, highlight, and hide. User interactions such as likes, votes, clicks, and views are assumed to be a proxy of a content's quality, popularity, or news-worthiness. In this paper we ask: how predictable are the interactions of a user on social media? To answer this question we recorded the clicking, browsing, and voting behavior of 186 Reddit users over a year. We present interesting descriptive statistics about their combined 339,270 interactions, and we find that relatively simple models are able to predict users' individual browse- or vote-interactions with reasonable accuracy. Maria Glenski, Tim Weninger |
ASONAM | 2 |
| 2017 | Rating Effects on Social News Posts and CommentsabstractAt a time when information seekers first turn to digital sources for news and opinion, it is critical that we understand the role that social media plays in human behavior. This is especially true when information consumers also act as information producers and editors through their online activity. In order to better understand the effects that editorial ratings have on online human behavior, we report the results of a two large-scale in vivo experiments in social media. We find that small, random rating manipulations on social media posts and comments created significant changes in downstream ratings, resulting in significantly different final outcomes. We found positive herding effects for positive treatments on posts, increasing the final rating by 11.02% on average, but not for positive treatments on comments. Contrary to the results of related work, we found negative herding effects for negative treatments on posts and comments, decreasing the final ratings, on average, of posts by 5.15% and of comments by 37.4%. Compared to the control group, the probability of reaching a high rating ( ⩾ 2,000) for posts is increased by 24.6% when posts receive the positive treatment and for comments it is decreased by 46.6% when comments receive the negative treatment. Maria Glenski, Tim Weninger |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | Growing Graphs from Hyperedge Replacement Graph GrammarsabstractDiscovering the underlying structures present in large real world graphs is a fundamental scientific problem. In this paper we show that a graph's clique tree can be used to extract a hyperedge replacement grammar. If we store an ordering from the extraction process, the extracted graph grammar is guaranteed to generate an isomorphic copy of the original graph. Or, a stochastic application of the graph grammar rules can be used to quickly create random graphs. In experiments on large real world networks, we show that random graphs, generated from extracted graph grammars, exhibit a wide range of properties that are very similar to the original graphs. In addition to graph properties like degree or eigenvector centrality, what a graph ``looks like'' ultimately depends on small details in local graph substructures that are difficult to define at a global level. We show that our generative graph model is able to preserve these local substructures when generating new graphs and performs well on new and difficult tests of model robustness. Salvador Aguiñaga, Rodrigo Palácios, David Chiang 0001, Tim Weninger |
CIKM | 4 |
| 2016 | Scalable models for computing hierarchies in information networks
Baoxu Shi, Tim Weninger |
Knowl. Inf. Syst. | 2 |
| 2015 | Concept hierarchies and human navigationabstractWe are confronted with massive amounts of information at every turn. In order to efficiently reason about knowledge and information, humans have evolved efficient strategies for organizing complex concepts in order to form connections between and recall information. This behavior can be observed and codified when people search for objects within digital information networks. Current models of search behavior exhibit unnecessary or extraneous complexity. Minimal or simple modifications to well established algorithms yield valid models of human navigation by exploring hierarchical information inherent in networks. We explore and validate a new model of how humans navigate an information networks. To that end, we present a new path finding algorithm that approximates human navigation by leveraging the categorical classification of the nodes within the network. We compare our new model, CatPath, to existing graph distance measures when possible and show that the category paths are largely correlated with traces of human navigation. Salvador Aguiñaga, Aditya Nambiar, Zuozhu Liu, Tim Weninger |
IEEE BigData | 4 |
| 2013 | An exploration of discussion threads in social news sites: a case study of the Reddit communityabstractSocial news and content aggregation Web sites have become massive repositories of valuable knowledge on a diverse range of topics. Millions of Web-users are able to leverage these platforms to submit, view and discuss nearly anything. The users themselves exclusively curate the content with an intricate system of submissions, voting and discussion. Furthermore, the data on social news Web sites is extremely well organized by its user-base, which opens the door for opportunities to leverage this data for other purposes just like Wikipedia data has been used for many other purposes. In this paper we study a popular social news Web site called Reddit. Our investigation looks at the dynamics of its discussion threads, and asks two main questions: (1) to what extent do discussion threads resemble a topical hierarchy? and (2) Can discussion threads be used to enhance Web search? We show interesting results for these questions on a very large snapshot several sub-communities of the Reddit Web site. Finally, we discuss the implications of these results and suggest ways by which social news Web site's can be used to perform other tasks. Tim Weninger, Xihao Avi Zhu, Jiawei Han 0001 |
ASONAM | 1 |
| 2013 | Research-insight: providing insight on research by publication network analysisabstractA database contains rich, inter-related, multi-typed data and information, forming one or a set of gigantic, intercon- nected, heterogeneous information networks. Much knowl- edge can be derived from such information networks if we systematically develop an effective and scalable database-oriented information network analysis technology. In this system demo, we take a computer science research publica- tion network as an example, which is an information net- work derived from an integration of DBLP, other web-based information about researchers, and partially available cita- tion data, and construct a Research-Insight system in order to demonstrate the power of database-oriented information network analysis. We show that nontrivial research insight can be obtained from such analysis, including (1) ranking, clustering, classification and similarity search of researchers, terms and venues for research subfields and themes, (2) recommending good researchers and good research papers to read or cite when conducting research on certain topics (3) predicting potential collaborators for certain theme-oriented research, and (4) predicting advisor-advisee rela- tionships and affiliation history based on historical research publications. Although some of these functions have been studied in recent research, effective and scalable realization of such functions in large networks still poses challenging research problems. Moreover, some function are our on- going research tasks. By integrating these functionalities, Research-Insight may not only provide with us insightful rec- ommendations in CS research but also help us gain insight on how to perform effective data mining in large databases. Fangbo Tao, Xiao Yu 0007, Kin Hou Lei, George Brova, Jiawei Han 0001, Rucha Kanade, Yizhou Sun, Chi Wang 0001, Tim Weninger |
SIGMOD Conference | 11 |
| 2013 | Exploring structure and content on the web: extraction and integration of the semi-structured webabstractIn this tutorial we view the World Wide Web as a type of massive, decentralized database. At present, this "Web database" is presented in a manner largely devoid of any consistent meaning or schema. That is not to say that Web-data lacks an underlying organization; in fact, most Web content is generated from an underlying schema-bound, or otherwise structured database. Information extraction is generally concerned with the reconciliation of unstructured or semi-structured Web content with the neatly structured database paradigm. With this Web-database in hand, researchers and practitioners have recently begun developing mechanisms which return structured results in response to an unstructured query. These new developments are a product of (1) record, list and table extraction from large numbers of semi-structured Web pages, (2) integration of these disparate extraction results into a consistent form, and (3) analysis of the newly extracted and integrated Web data. Tim Weninger, Jiawei Han 0001 |
WSDM | 1 |
| 2013 | The parallel path framework for entity discovery on the webabstractIt has been a dream of the database and Web communities to reconcile the unstructured nature of the World Wide Web with the neat, structured schemas of the database paradigm. Even though databases are currently used to generate Web content in some sites, the schemas of these databases are rarely consistent across a domain. This makes the comparison and aggregation of information from different domains difficult. We aim to make an important step towards resolving this disparity by using the structural and relational information on the Web to (1) extract Web lists, (2) find entity-pages, (3) map entity-pages to a database, and (4) extract attributes of the entities. Specifically, given a Web site and an entity-page (e.g., university department and faculty member home page) we seek to find all of the entity-pages of the same type (e.g., all faculty members in the department), as well as attributes of the specific entities (e.g., their phone numbers, email addresses, office numbers). To do this, we propose a Web structure mining method which grows parallel paths through the Web graph and DOM trees and propagates relevant attribute information forward. We show that by utilizing these parallel paths we can efficiently discover entity-pages and attributes. Finally, we demonstrate the accuracy of our method with a large case study. Tim Weninger, Thomas J. Johnston, Jiawei Han 0001 |
ACM Trans. Web | 1 |
| 2012 | Document-topic hierarchies from document graphsabstractTopic taxonomies present a multi-level view of a document collection, where general topics live towards the top of the taxonomy and more specific topics live towards the bottom. Topic taxonomies allow users to quickly drill down into their topic of interest to find documents. We show that hierarchies of documents, where documents live at the inner nodes of the hierarchy-tree can also be inferred by combining document text with inter-document links. We present a Bayesian generative model by which an explicit hierarchy of documents is created. Experiments on three document-graph data sets shows that the generated document hierarchies are able to fit the observed data, and that the levels in the constructed document hierarchy represent practical groupings. Tim Weninger, Yonatan Bisk, Jiawei Han 0001 |
CIKM | 1 |
| 2011 | Authorship classification: a discriminative syntactic tree mining approachabstractIn the past, there have been dozens of studies on automatic authorship classification, and many of these studies concluded that the writing style is one of the best indicators for original authorship. From among the hundreds of features which were developed, syntactic features were best able to reflect an author's writing style. However, due to the high computational complexity for extracting and computing syntactic features, only simple variations of basic syntactic features such as function words, POS(Part of Speech) tags, and rewrite rules were considered. In this paper, we propose a new feature set of k-embedded-edge subtree patterns that holds more syntactic information than previous feature sets. We also propose a novel approach to directly mining them from a given set of syntactic trees. We show that this approach reduces the computational burden of using complex syntactic structures as the feature set. Comprehensive experiments on real-world datasets demonstrate that our approach is reliable and more accurate than previous studies. Sangkyum Kim, Hyungsul Kim, Tim Weninger, Jiawei Han 0001, Hyun Duk Kim |
SIGIR | 3 |
| 2011 | WINACS: construction and analysis of web-based computer science information networksabstractWINACS (Web-based Information Network Analysis for Computer Science) is a project that incorporates many recent, exciting developments in data sciences to construct a Web-based computer science information network and to discover, retrieve, rank, cluster, and analyze such an information network. With the rapid development of the Web, huge amounts of information are available in the form of Web documents, structures, and links. It has been a dream of the database and Web communities to harvest such information and reconcile the unstructured nature of the Web with the neat, semi-structured schemas of the database paradigm. Taking computer science as a dedicated domain, WINACS first discovers related Web entity structures, and then constructs a heterogeneous computer science information network in order to rank, cluster and analyze this network and support intelligent and analytical queries. Tim Weninger, Marina Danilevsky, Fabio Fumarola, Joshua M. Hailpern, Jiawei Han 0001, Thomas J. Johnston, Surya Kallumadi, Hyungsul Kim, Zhijin Li, David McCloskey, Yizhou Sun, Nathan E. TeGrotenhuis, Chi Wang 0001, Xiao Yu 0007 |
SIGMOD Conference | 1 |
| 2011 | Mining Flipping Correlations from Large Datasets with TaxonomiesabstractIn this paper we introduce a new type of pattern -- a flipping correlation pattern. The flipping patterns are obtained from contrasting the correlations between items at different levels of abstraction. They represent surprising correlations, both positive and negative, which are specific for a given abstraction level, and which "flip" from positive to negative and vice versa when items are generalized to a higher level of abstraction. We design an efficient algorithm for finding flipping correlations, the Flipper algorithm, which outperforms naïve pattern mining methods by several orders of magnitude. We apply Flipper to real-life datasets and show that the discovered patterns are non-redundant, surprising and actionable. Flipper finds strong contrasting correlations in itemsets with low-to-medium support, while existing techniques cannot handle the pattern discovery in this frequency range. Marina Barsky, Sangkyum Kim, Tim Weninger, Jiawei Han 0001 |
Proc. VLDB Endow. | 3 |
| 2010 | A Unified Framework for Link Recommendation Using Random WalksabstractThe phenomenal success of social networking sites, such as Facebook, Twitter and LinkedIn, has revolutionized the way people communicate. This paradigm has attracted the attention of researchers that wish to study the corresponding social and technological problems. Link recommendation is a critical task that not only helps increase the linkage inside the network and also improves the user experience. In an effective link recommendation algorithm it is essential to identify the factors that influence link creation. This paper enumerates several of these intuitive criteria and proposes an approach which satisfies these factors. This approach estimates link relevance by using random walk algorithm on an augmented social graph with both attribute and structure information. The global and local influences of the attributes are leveraged in the framework as well. Other than link recommendation, our framework can also rank the attributes in the network. Experiments on DBLP and IMDB data sets demonstrate that our method outperforms state-of-the-art methods for link recommendation. Zhijun Yin, Manish Gupta 0001, Tim Weninger, Jiawei Han 0001 |
ASONAM | 3 |
| 2010 | Mapping web pages to database records via link pathsabstractIn this paper we propose a new knowledge management task which aims to map Web pages to their corresponding records in a structured database. For example, the DBLP database contains records for many computer scientists, and most of these persons have public Web pages; if we can map the database record with the appropriate Web page then the new information could be used to further describe the person's database record. To accomplish this goal we employ link paths which contain anchor texts from multiple paths through the Web ending at the Web page in question. We hypothesize that the information from these link paths can be used to generate an accurate Web page to database record mapping. Experiments on two large, real world data sets, DBLP and IMDB for the structured data and computer science faculty members' Web pages and official movie homepages for the Web page data, show that our method does provide an accurate mapping. Finally, we conclude by issuing a call for further research on this promising new task. Tim Weninger, Fabio Fumarola, Jiawei Han 0001, Donato Malerba |
CIKM | 1 |
| 2010 | NDPMine: Efficiently Mining Discriminative Numerical Features for Pattern-Based Classification
Hyungsul Kim, Sangkyum Kim, Tim Weninger, Jiawei Han 0001, Tarek F. Abdelzaher |
ECML/PKDD (2) | 3 |
| 2010 | Entity relation discovery from web tables and linksabstractThe World-Wide Web consists not only of a huge number of unstructured texts, but also a vast amount of valuable structured data. Web tables [2] are a typical type of structured information that are pervasive on the web, and Web-scale methods that automatically extract web tables have been studied extensively [1]. Many powerful systems (e.g.OCTOPUS [4], Mesa [3]) use extracted web tables as a fundamental component. Cindy Xide Lin, Bo Zhao 0001, Tim Weninger, Jiawei Han 0001, Bing Liu 0001 |
WWW | 3 |
| 2010 | CETR: content extraction via tag ratiosabstractWe present Content Extraction via Tag Ratios (CETR) - a method to extract content text from diverse webpages by using the HTML document's tag ratios. We describe how to compute tag ratios on a line-by-line basis and then cluster the resulting histogram into content and non-content areas. Initially, we find that the tag ratio histogram is not easily clustered because of its one-dimensionality; therefore we extend the original approach in order to model the data in two dimensions. Next, we present a tailored clustering technique which operates on the two-dimensional model, and then evaluate our approach against a large set of alternative methods using standard accuracy, precision and recall metrics on a large and varied Web corpus. Finally, we show that, in most cases, CETR achieves better content extraction performance than existing methods, especially across varying web domains, languages and styles. Tim Weninger, William H. Hsu, Jiawei Han 0001 |
WWW | 1 |
| 2010 | LINKREC: a unified framework for link recommendation with user attributes and graph structureabstractWith the phenomenal success of networking sites (e.g., Facebook, Twitter and LinkedIn), social networks have drawn substantial attention. On online social networking sites, link recommendation is a critical task that not only helps improve user experience but also plays an essential role in network growth. In this paper we propose several link recommendation criteria, based on both user attributes and graph structure. To discover the candidates that satisfy these criteria, link relevance is estimated using a random walk algorithm on an augmented social graph with both attribute and structure information. The global and local influence of the attributes is leveraged in the framework as well. Besides link recommendation, our framework can also rank attributes in a social network. Experiments on DBLP and IMDB data sets demonstrate that our method outperforms state-of-the-art methods based on network structure and node attribute information for link recommendation. Zhijun Yin, Manish Gupta 0001, Tim Weninger, Jiawei Han 0001 |
WWW | 3 |
| 2007 | Structural Link Analysis from User Profiles and Friends Networks: A Feature Construction Approach
William H. Hsu, Joseph P. Lancaster, Martin S. R. Paradesi, Tim Weninger |
ICWSM | 4 |