Filippo Menczer

dblp:79/3056 · DBLP profile ↗
← Back
40ranked-venue papers in the field
3as first author
10since 2021 · last 2025
0000-0003-4384-2876ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 30 (3 first)Data Mining & Knowledge Discovery · 9Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 Coordinated Reply Attacks in Influence Operations: Characterization and Detection
abstract
Coordinated reply attacks are a tactic observed in online influence operations and other coordinated campaigns to support or harass targeted individuals, or influence them or their followers. Despite its potential to influence the public, past studies have yet to analyze or provide a methodology to detect this tactic. In this study, we characterize coordinated reply attacks in the context of influence operations on Twitter. Our analysis reveals that the primary targets of these attacks are influential people such as journalists, news media, state officials, and politicians. We propose two supervised machine-learning models, one to classify tweets to determine whether they are targeted by a reply attack, and one to classify accounts that reply to a targeted tweet to determine whether they are part of a coordinated attack. The classifiers achieve AUC scores of 0.88 and 0.97, respectively. These results indicate that accounts involved in reply attacks can be detected, and the targeted accounts themselves can serve as sensors for influence operation detection.
Manita Pote, Tugrulcan Elmas, Alessandro Flammini, Filippo Menczer
ICWSM4
2025 Labeled Datasets for Research on Information Operations
abstract
Social media platforms have become a hub for political activities and discussions, democratizing participation in these endeavors. However, they have also become an incubator for manipulation campaigns, like information operations (IOs). Some social media platforms have released datasets related to such IOs originating from different countries. However, we lack comprehensive control data that can enable the development of IO detection methods. To bridge this gap, we present new labeled datasets about 26 campaigns, which contain both IO posts verified by a social media platform and over 13M posts by 303k accounts that discussed similar topics in the same time frames (control data). The datasets will facilitate the study of narratives, network interactions, and engagement strategies employed by coordinated accounts across various campaigns and countries. By comparing these coordinated accounts against organic ones, researchers can develop and benchmark IO detection algorithms.
Ozgur Can Seckin, Manita Pote, Alexander C. Nwala, Lake Yin, Luca Luceri, Alessandro Flammini, Filippo Menczer
ICWSM7
2024 The Dawn of Decentralized Social Media: An Exploration of Bluesky's Public Opening
Erfan Samieyan Sahneh, Gianluca Nogara, Matthew DeVerna, Nick Liu, Luca Luceri, Filippo Menczer, Francesco Pierri 0002, Silvia Giordano
ASONAM (1)6
2023 A Multi-Platform Collection of Social Media Posts about the 2022 U.S. Midterm Elections
abstract
Social media are utilized by millions of citizens to discuss important political issues. Politicians use these platforms to connect with the public and broadcast policy positions. Therefore, data from social media has enabled many studies of political discussion. While most analyses are limited to data from individual platforms, people are embedded in a larger information ecosystem spanning multiple social networks. Here we describe and provide access to the Indiana University 2022 U.S. Midterms Multi-Platform Social Media Dataset (MEIU22), a collection of social media posts from Twitter, Facebook, Instagram, Reddit, and 4chan. MEIU22 links to posts about the midterm elections based on a comprehensive list of keywords and tracks the social media accounts of 1,011 candidates from October 1 to December 25, 2022. We also publish the source code of our pipeline to enable similar multi-platform research projects.
Rachith Aiyappa, Matthew DeVerna, Manita Pote, Bao Tran Truong, Wanying Zhao, David Axelrod, Aria Pessianzadeh, Zoher Kachwala, Munjung Kim, Ozgur Can Seckin, Minsuk Kim, Sunny Gandhi, Amrutha Manikonda, Francesco Pierri 0002, Filippo Menczer, Kai-Cheng Yang
ICWSM15
2022 Ukraine as a Political Tool in Facebook Sponsored Content
abstract
This work provides an investigation of the narratives in Facebook sponsored content about the Russian invasion of Ukraine. Advertisers used Ukraine to promote narratives related to their goals. Different types of advertisers spent their money on varying demographics, and some reached more viewers while spending less money.
Oliver Melbourne Allen, Filippo Menczer
ASONAM2
2022 Manipulating Twitter through Deletions
Christopher Torres-Lugo, Manita Pote, Alexander C. Nwala, Filippo Menczer
ICWSM4
2022 The Manufacture of Partisan Echo Chambers by Follow Train Abuse on Twitter
Christopher Torres-Lugo, Kai-Cheng Yang, Filippo Menczer
ICWSM3
2022 Do Recommender Systems Make Social Media More Susceptible to Misinformation Spreaders?
abstract
Recommender systems are central to online information consumption and user-decision processes, as they help users find relevant information and establish new social relationships. However, recommenders could also (unintendedly) help propagate misinformation and increase the social influence of the spreading it. In this context, we study the impact of friend recommender systems on the social influence of misinformation spreaders on Twitter. To this end, we applied several user recommenders to a COVID-19 misinformation data collection. Then, we explore what-if scenarios to simulate changes in user misinformation spreading behaviour as an effect of the interactions in the recommended network. Our study shows that recommenders can indeed affect how misinformation spreaders interact with other users and influence them.
Antonela Tommasel, Filippo Menczer
RecSys2
2021 CoVaxxy: A Collection of English-Language Twitter Posts About COVID-19 Vaccines
Matthew DeVerna, Francesco Pierri 0002, Bao Tran Truong, John Bollenbacher, David Axelrod, Niklas Loynes, Christopher Torres-Lugo, Kai-Cheng Yang, Filippo Menczer, John Bryden
ICWSM9
2021 Uncovering Coordinated Networks on Social Media: Methods and Case Studies
Diogo Pacheco, Pik-Mai Hui, Christopher Torres-Lugo, Bao Tran Truong, Alessandro Flammini, Filippo Menczer
ICWSM6
2020 Detection of Novel Social Bots by Ensembles of Specialized Classifiers
abstract
Malicious actors create inauthentic social media accounts controlled in part by algorithms, known as social bots, to disseminate misinformation and agitate online discussion. While researchers have developed sophisticated methods to detect abuse, novel bots with diverse behaviors evade detection. We show that different types of bots are characterized by different behavioral features. As a result, supervised learning techniques suffer severe performance deterioration when attempting to detect behaviors not observed in the training data. Moreover, tuning these models to recognize novel bots requires retraining with a significant amount of new annotations, which are expensive to obtain. To address these issues, we propose a new supervised learning method that trains classifiers specialized for each class of bots and combines their decisions through the maximum rule. The ensemble of specialized classifiers (ESC) can better generalize, leading to an average improvement of 56% in F1 score for unseen accounts across datasets. Furthermore, novel bot behaviors are learned with fewer labeled examples during retraining. We deployed ESC in the newest version of Botometer, a popular tool to detect social bots in the wild, with a cross-validation AUC of 0.99.
Mohsen Sayyadiharikandeh, Onur Varol, Kai-Cheng Yang, Alessandro Flammini, Filippo Menczer
CIKM5
2020 BotSlayer: DIY Real-Time Influence Campaign Detection
Pik-Mai Hui, Kai-Cheng Yang, Christopher Torres-Lugo, Filippo Menczer
ICWSM4
2020 4 Reasons Why Social Media Make Us Vulnerable to Manipulation
abstract
As social media become major channels for the diffusion of news and information, it becomes critical to understand how the complex interplay between cognitive, social, and algorithmic biases triggered by our reliance on online social networks makes us vulnerable to manipulation and disinformation. This talk overviews ongoing network analytics, modeling, and machine learning efforts to study the viral spread of misinformation and to develop tools for countering the online manipulation of opinions.
Filippo Menczer
RecSys1
2019 Quantifying Biases in Online Information Exposure
abstract
Our consumption of online information is mediated by filtering, ranking, and recommendation algorithms that introduce unintentional biases as they attempt to deliver relevant and engaging content. It has been suggested that our reliance on online technologies such as search engines and social media may limit exposure to diverse points of view and make us vulnerable to manipulation by disinformation. In this article, we mine a massive data set of web traffic to quantify two kinds of bias: (i) homogeneity bias, which is the tendency to consume content from a narrow set of information sources, and (ii) popularity bias, which is the selective exposure to content from top sites. Our analysis reveals different bias levels across several widely used web platforms. Search exposes users to a diverse set of sources, while social media traffic tends to exhibit high popularity and homogeneity bias. When we focus our analysis on traffic to news sites, we find higher levels of popularity bias, with smaller differences across applications. Overall, our results quantify the extent to which our choices of online systems confine us inside “social bubbles.”
Dimitar Nikolov, Mounia Lalmas-Roelleke, Alessandro Flammini, Filippo Menczer
J. Assoc. Inf. Sci. Technol.4
2018 The Hoaxy Misinformation and Fact-Checking Diffusion Network
Pik-Mai Hui, Chengcheng Shao, Alessandro Flammini, Filippo Menczer, Giovanni Luca Ciampaglia
ICWSM4
2018 Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data
abstract
Machine learning methods are used to discover complex nonlinear relationships in biological and medical data. However, sophisticated learning models are computationally unfeasible for data with millions of features. Here, we introduce the first feature selection method for nonlinear learning problems that can scale up to large, ultra-high dimensional biological data. More specifically, we scale up the novel Hilbert-Schmidt Independence Criterion Lasso (HSIC Lasso) to handle millions of features with tens of thousand samples. The proposed method is guaranteed to find an optimal subset of maximally predictive features with minimal redundancy, yielding higher predictive power and improved interpretability. Its effectiveness is demonstrated through applications to classify phenotypes based on module expression in human prostate cancer patients and to detect enzymes among protein structures. We achieve high accuracy with as few as 20 out of one million features-a dimensionality reduction of 99.998 percent. Our algorithm can be implemented on commodity cloud computing platforms. The dramatic reduction of features may lead to the ubiquitous deployment of sophisticated prediction models in mobile health care applications.
Makoto Yamada, Jiliang Tang, Jose Lugo-Martinez, Ermin Hodzic, Raunak Shrestha, Avishek Saha, Hua Ouyang, Dawei Yin 0001, Hiroshi Mamitsuka, Süleyman Cenk Sahinalp, Predrag Radivojac, Filippo Menczer, Yi Chang 0001
IEEE Trans. Knowl. Data Eng.12
2017 Finding Streams in Knowledge Graphs to Support Fact Checking
abstract
The volume of information generated online makes it impossible to manually fact-check all claims. Computational approaches for fact checking may be the key to help mitigate the risks of massive misinformation spread. Such approaches can be designed to not only be scalable and effective at assessing veracity of dubious claims, but also to boost a human fact checker's productivity by surfacing relevant facts and patterns to aid their analysis. We present a novel, unsupervised network-flow based approach to determine the truthfulness of a statement of fact expressed in the form of a triple. We view a knowledge graph of background information about real-world entities as a flow network, and show that computational fact checking then amounts to finding a "knowledge stream" connecting the subject and object of the triple. Evaluation on a range of real-world and hand-crafted datasets of facts reveals that this network-flow model can be very effective in discerning true statements from false ones, outperforming existing algorithms on many test cases. Moreover, the model is expressive in its ability to automatically discover several useful patterns and surface relevant facts that may help a human fact checker.
Prashant Shiralkar, Alessandro Flammini, Filippo Menczer, Giovanni Luca Ciampaglia
ICDM3
2017 Online Human-Bot Interactions: Detection, Estimation, and Characterization
Onur Varol, Emilio Ferrara, Clayton A. Davis, Filippo Menczer, Alessandro Flammini
ICWSM4
2016 Detection of Promoted Social Media Campaigns
Emilio Ferrara, Onur Varol, Filippo Menczer, Alessandro Flammini
ICWSM3
2016 Mining for Topics to Suggest Knowledge Model Extensions
abstract
Electronic concept maps, interlinked with other concept maps and multimedia resources, can provide rich knowledge models to capture and share human knowledge. This article presents and evaluates methods to support experts as they extend existing knowledge models, by suggesting new context-relevant topics mined from Web search engines. The task of generating topics to support knowledge model extension raises two research questions: first, how to extract topic descriptors and discriminators from concept maps; and second, how to use these topic descriptors and discriminators to identify candidate topics on the Web with the right balance of novelty and relevance. To address these questions, this article first develops the theoretical framework required for a “topic suggester” to aid information search in the context of a knowledge model under construction. It then presents and evaluates algorithms based on this framework and applied in E xtender , an implemented tool for topic suggestion. E xtender has been developed and tested within CmapTools, a widely used system for supporting knowledge modeling using concept maps. However, the generality of the algorithms makes them applicable to a broad class of knowledge modeling systems, and to Web search in general.
Carlos M. Lorenzetti, Ana Gabriela Maguitman, David B. Leake, Filippo Menczer, Thomas Reichherzer
ACM Trans. Knowl. Discov. Data4
2014 Predicting Successful Memes Using Network and Community Structure
Lilian Weng, Filippo Menczer, Yong-Yeol Ahn
ICWSM2
2013 Clustering memes in social media
abstract
The increasing pervasiveness of social media creates new opportunities to study human social behavior, while challenging our capability to analyze their massive data streams. One of the emerging tasks is to distinguish between different kinds of activities, for example engineered misinformation campaigns versus spontaneous communication. Such detection problems require a formal definition of meme, or unit of information that can spread from person to person through the social network. Once a meme is identified, supervised learning methods can be applied to classify different types of communication. The appropriate granularity of a meme, however, is hardly captured from existing entities such as tags and keywords. Here we present a framework for the novel task of detecting memes by clustering messages from large streams of social data. We evaluate various similarity measures that leverage content, metadata, network features, and their combinations. We also explore the idea of pre-clustering on the basis of existing entities. A systematic evaluation is carried out using a manually curated dataset as ground truth. Our analysis shows that pre-clustering and a combination of heterogeneous features yield the best trade-off between number of clusters and their quality, demonstrating that a simple combination based on pairwise maximization of similarity is as effective as a non-trivial optimization of parameters. Our approach is fully automatic, unsupervised, and scalable for real-time detection of memes in streaming data.
Emilio Ferrara, Mohsen JafariAsbagh, Onur Varol, Vahed Qazvinian, Filippo Menczer, Alessandro Flammini
ASONAM5
2013 The role of information diffusion in the evolution of social networks
abstract
Every day millions of users are connected through online social networks, generating a rich trove of data that allows us to study the mechanisms behind human interactions. Triadic closure has been treated as the major mechanism for creating social links: if Alice follows Bob and Bob follows Charlie, Alice will follow Charlie. Here we present an analysis of longitudinal micro-blogging data, revealing a more nuanced view of the strategies employed by users when expanding their social circles. While the network structure affects the spread of information among users, the network is in turn shaped by this communication activity. This suggests a link creation mechanism whereby Alice is more likely to follow Charlie after seeing many messages by Charlie. We characterize users with a set of parameters associated with different link creation strategies, estimated by a Maximum-Likelihood approach. Triadic closure does have a strong effect on link formation, but shortcuts based on traffic are another key factor in interpreting network evolution. However, individual strategies for following other users are highly heterogeneous. Link creation behaviors can be summarized by classifying users in different categories with distinct structural and behavioral characteristics. Users who are popular, active, and influential tend to create traffic-based shortcuts, making the information diffusion process more efficient in the network.
Lilian Weng, Jacob Ratkiewicz, Nicola Perra, Bruno Gonçalves, Carlos Castillo 0001, Francesco Bonchi, Rossano Schifanella, Filippo Menczer, Alessandro Flammini
KDD8
2013 Ambiguous author query detection using crowdsourced digital library annotations
Jasleen Kaur 0002, Lino Possamai, Filippo Menczer
Inf. Process. Manag.4
2012 Friendship prediction and homophily in social media
abstract
Social media have attracted considerable attention because their open-ended nature allows users to create lightweight semantic scaffolding to organize and share content. To date, the interplay of the social and topical components of social media has been only partially explored. Here, we study the presence of homophily in three systems that combine tagging social media with online social networks. We find a substantial level of topical similarity among users who are close to each other in the social network. We introduce a null model that preserves user activity while removing local correlations, allowing us to disentangle the actual local similarity between users from statistical effects due to the assortative mixing of user activity and centrality in the social network. This analysis suggests that users with similar interests are more likely to be friends, and therefore topical similarity measures among users based solely on their annotation metadata should be predictive of social links. We test this hypothesis on several datasets, confirming that social networks constructed from topical similarity capture actual friendship accurately. When combined with topological features, topical similarity achieves a link prediction accuracy of about 92%.
Luca Maria Aiello, Alain Barrat, Rossano Schifanella, Ciro Cattuto, Benjamin Markines, Filippo Menczer
ACM Trans. Web6
2011 Behavior-driven clustering of queries into topics
abstract
Categorization of web-search queries in semantically coherent topics is a crucial task to understand the interest trends of search engine users and, therefore, to provide more intelligent personalization services. Query clustering usually relies on lexical and clickthrough data, while the information originating from the user actions in submitting their queries is currently neglected. In particular, the intent that drives users to submit their requests is an important element for meaningful aggregation of queries. We propose a new intent-centric notion of topical query clusters and we define a query clustering technique that differs from existing algorithms in both methodology and nature of the resulting clusters. Our method extracts topics from the query log by merging missions, i.e., activity fragments that express a coherent user intent, on the basis of their topical affinity. Our approach works in a bottom-up way, without any a-priori knowledge of topical categorization, and produces good quality topics compared to state-of-the-art clustering techniques. It can also summarize topically-coherent missions that occur far away from each other, thus enabling a more compact user profiling on a topical basis. Furthermore, such a topical user profiling discriminates the stream of activity of a particular user from the activity of others, with a potential to predict future user search activity.
Luca Maria Aiello, Debora Donato, Umut Ozertem, Filippo Menczer
CIKM4
2011 Political Polarization on Twitter
Michael D. Conover, Jacob Ratkiewicz, Matthew R. Francisco, Bruno Gonçalves, Filippo Menczer, Alessandro Flammini
ICWSM5
2011 Detecting and Tracking Political Abuse in Social Media
Jacob Ratkiewicz, Michael D. Conover, Mark R. Meiss, Bruno Gonçalves, Alessandro Flammini, Filippo Menczer
ICWSM6
2010 Folks in Folksonomies: social link prediction from shared metadata
abstract
Web 2.0 applications have attracted a considerable amount of attention because their open-ended nature allows users to create lightweight semantic scaffolding to organize and share content. To date, the interplay of the social and semantic components of social media has been only partially explored. Here we focus on Flickr and Last.fm, two social media systems in which we can relate the tagging activity of the users with an explicit representation of their social network. We show that a substantial level of local lexical and topical alignment is observable among users who lie close to each other in the social network. We introduce a null model that preserves user activity while removing local correlations, allowing us to disentangle the actual local alignment between users from statistical effects due to the assortative mixing of user activity and centrality in the social network. This analysis suggests that users with similar topical interests are more likely to be friends, and therefore semantic similarity measures among users based solely on their annotation metadata should be predictive of social links. We test this hypothesis on the Last.fm data set, confirming that the social network constructed from semantic similarity captures actual friendship more accurately than Last.fm's suggestions based on listening patterns.
Rossano Schifanella, Alain Barrat, Ciro Cattuto, Benjamin Markines, Filippo Menczer
WSDM5
2009 Incentives for social annotation
abstract
The effectiveness of community-driven annotation, such as social bookmarking, depends on user participation. Since the participation of many users is motivated by selfish reasons, an effective way to encourage participation is to create useful or entertaining applications. We demo two such tools -- a browser extension and a game.
Heather Roinestad, John Burgoon, Benjamin Markines, Filippo Menczer
SIGIR4
2009 Evaluating similarity measures for emergent semantics of social tagging
abstract
Social bookmarking systems are becoming increasingly important data sources for bootstrapping and maintaining Semantic Web applications. Their emergent information structures have become known as folksonomies. A key question for harvesting semantics from these systems is how to extend and adapt traditional notions of similarity to folksonomies, and which measures are best suited for applications such as community detection, navigation support, semantic search, user profiling and ontology learning. Here we build an evaluation framework to compare various general folksonomy-based similarity measures, which are derived from several established information-theoretic, statistical, and practical measures. Our framework deals generally and symmetrically with users, tags, and resources. For evaluation purposes we focus on similarity between tags and between resources and consider different methods to aggregate annotations across users. After comparing the ability of several tag similarity measures to predict user-created tag relations, we provide an external grounding by user-validated semantic proxies based on WordNet and the Open Directory Project. We also investigate the issue of scalability. We find that mutual information with distributional micro-aggregation across users yields the highest accuracy, but is not scalable; per-user projection with collaborative aggregation provides the best scalable approach via incremental computations. The results are consistent across resource and tag similarity.
Benjamin Markines, Ciro Cattuto, Filippo Menczer, Dominik Benz, Andreas Hotho, Gerd Stumme
WWW3
2008 Ranking web sites with real user traffic
abstract
We analyze the traffic-weighted Web host graph obtained from a large sample of real Web users over about seven months. A number of interesting structural properties are revealed by this complex dynamic network, some in line with the well-studied boolean link host graph and others pointing to important differences. We find that while search is directly involved in a surprisingly small fraction of user clicks, it leads to a much larger fraction of all sites visited. The temporal traffic patterns display strong regularities, with a large portion of future requests being statistically predictable by past ones. Given the importance of topological measures such as PageRank in modeling user navigation, as well as their role in ranking sites for Web search, we use the traffic data to validate the PageRank random surfing model. The ranking obtained by the actual frequency with which a site is visited by users differs significantly from that approximated by the uniform surfing/teleportation behavior modeled by PageRank, especially for the most important sites. To interpret this finding, we consider each of the fundamental assumptions underlying PageRank and show how each is violated by actual user behavior
Mark R. Meiss, Filippo Menczer, Santo Fortunato, Alessandro Flammini, Alessandro Vespignani
WSDM2
2007 Introduction to the special topic section on mining Web resources for enhancing information retrieval
abstract
Abstract The amount of information on the Web has been expanding at an enormous pace. There are a variety of Web documents in different genres, such as news, reports, reviews. Traditionally, the information displayed on Web sites has been static. Recently, there are many Web sites offering content that is dynamically generated and frequently updated. It is also common for Web sites to contain information in different languages since many countries adopt more than one language. Moreover, content may exist in multimedia formats including text, images, video, and audio.
Wai Lam, Christopher C. Yang, Filippo Menczer
J. Assoc. Inf. Sci. Technol.3
2005 Algorithmic detection of semantic similarity
abstract
Automatic extraction of semantic information from text and links in Web pages is key to improving the quality of search results. However, the assessment of automatic semantic measures is limited by the coverage of user studies, which do not scale with the size, heterogeneity, and growth of the Web. Here we propose to leverage human-generated metadata --- namely topical directories --- to measure semantic relationships among massive numbers of pairs of Web pages or topics. The Open Directory Project classifies millions of URLs in a topical ontology, providing a rich source from which semantic relationships between Web pages can be derived. While semantic similarity measures based on taxonomies (trees) are well studied, the design of well-founded similarity measures for objects stored in the nodes of arbitrary ontologies (graphs) is an open problem. This paper defines an information-theoretic measure of semantic similarity that exploits both the hierarchical and non-hierarchical structure of an ontology. An experimental study shows that this measure improves significantly on the traditional taxonomy-based approach. This novel measure allows us to address the general question of how text and link analyses can be combined to derive measures of relevance that are in good agreement with semantic similarity. Surprisingly, the traditional use of text similarity turns out to be ineffective for relevance ranking.
Ana Gabriela Maguitman, Filippo Menczer, Heather Roinestad, Alessandro Vespignani
WWW2
2005 On the lack of typical behavior in the global Web traffic network
abstract
We offer the first large-scale analysis of Web traffic based on network flow data. Using data collected on the Internet2 network, we constructed a weighted bipartite clientserver host graph containing more than 18 × 10^6 vertices and 68 × 10^6 edges valued by relative traffic flows. When considered as a traffic map of the World-Wide Web, the generated graph provides valuable information on the statistical patterns that characterize the global information flow on the Web. Statistical analysis shows that client-server connections and traffic flows exhibit heavy-tailed probability distributions lacking any typical scale. In particular, the absence of an intrinsic average in some of the distributions implies the absence of a prototypical scale appropriate for server design, Web-centric network design, or traffic modeling. The inspection of the amount of traffic handled by clients and servers and their number of connections highlights non-trivial correlations between information flow and patterns of connectivity as well as the presence of anomalous statistical patterns related to the behavior of users on the Web. The results presented here may impact considerably the modeling, scalability analysis, and behavioral study of Web applications.
Mark R. Meiss, Filippo Menczer, Alessandro Vespignani
WWW2
2005 A General Evaluation Framework for Topical Crawlers
Padmini Srinivasan, Filippo Menczer, Gautam Pant
Inf. Retr.2
2004 Dynamic extraction topic descriptors and discriminators: towards automatic context-based topic search
abstract
Effective knowledge management may require going beyond initial knowledge capture, to support decisions about how to extend previously-captured knowledge. Electronic concept maps, interlinked with other concept maps and multimedia resources, can provide rich knowledge models for human knowledge capture and sharing. This paper presents research on methods for supporting experts as they extend these knowledge models, by searching the Web for new context-relevant topics as candidates for inclusion. This topic search problem presents two challenges: First, how to formulate queries to seek topics that reflect the context of the current knowledge model, and, second, how to identify candidate topics with the right balance of novelty and relevance. More generally, this problem raises the broad question of the interaction of topic information from the local analysis space (a collected set of documents) and the global search space (the Web). The paper develops a framework for understanding this interaction, and proposes and evaluates techniques for addressing the query formation and topic identification questions by dynamically extracting topic descriptors and discriminators from a knowledge model, to characterize information needs for retrieval and filtering of relevant material. Using these techniques, we have developed a support tool that starts from a knowledge model under construction and automatically produces a set of suggestions for topics to include, proactively supporting users as they extend knowledge models.
Ana Gabriela Maguitman, David B. Leake, Thomas Reichherzer, Filippo Menczer
CIKM4
2004 Lexical and semantic clustering by Web links
abstract
Abstract Recent Web‐searching and ‐mining tools are combining text and link analysis to improve ranking and crawling algorithms. The central assumption behind such approaches is that there is a correlation between the graph structure of the Web and the text and meaning of pages. Here I formalize and empirically evaluate two general conjectures drawing connections from link information to lexical and semantic Web content. The link‐content conjecture states that a page is similar to the pages that link to it, and the link‐cluster conjecture that pages about the same topic are clustered together. These conjectures are often simply assumed to hold, and Web search tools are built on such assumptions. The present quantitative confirmation sheds light on the connection between the success of the latest Web‐mining techniques and the small world topology of the Web, with encouraging implications for the design of better crawling algorithms.
Filippo Menczer
J. Assoc. Inf. Sci. Technol.1
2001 Evaluating Topic-Driven Web Crawlers
abstract
Due to limited bandwidth, storage, and computational resources, and to the dynamic nature of the Web, search engines cannot index every Web page, and even the covered portion of the Web cannot be monitored continuously for changes. Therefore it is essential to develop effective crawling strategies to prioritize the pages to be indexed. The issue is even more important for topic-specific search engines, where crawlers must make additional decisions based on the relevance of visited pages. However, it is difficult to evaluate alternative crawling strategies because relevant sets are unknown and the search space is changing. We propose three different methods to evaluate crawling strategies. We apply the proposed metrics to compare three topic-driven crawling algorithms based on similarity ranking, link analysis, and adaptive agents.
Filippo Menczer, Gautam Pant, Padmini Srinivasan, Miguel E. Ruiz
SIGIR1
2000 Feature selection in unsupervised learning via evolutionary search
abstract
Feature subset selection is an important problem in knowl- edge discovery, not only for the insight gained from deter- mining relevant modeling variables but also for the improved understandability, scalability, and possibly, accuracy of the resulting models. In this paper we consider the problem of feature selection for unsupervised learning. A number of heuristic criteria can be used to estimate the quality of clusters built from a given featuresubset. Rather than combining such criteria, we use ELSA, an evolutionary lo- cal selection algorithm that maintains a diverse population of solutions that approximate the Pareto front in a multi- dimensional objectiv espace. Each evolved solution repre- sents a feature subset and a number of clusters; a standard K-means algorithm is applied to form the given n umber of clusters based on the selected features. Preliminary results on both real and synthetic data show promise in finding Pareto-optimal solutions through which we can identify the significant features and the correct number of clusters.
YongSeog Kim, W. Nick Street, Filippo Menczer
KDD3