EDBT 2026 Demo / reviewers in the wild / expert
Emilio Ferrara
dblp:38/8773
· DBLP profile ↗
32ranked-venue papers in the field
5as first author
13since 2021 · last 2026
0000-0002-1942-2831ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 21 (3 first)Data Mining & Knowledge Discovery · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Platform Narrative Prediction: Leveraging Platform-Invariant Discourse Networks
Patrick Gerard, Luca Luceri, Leonardo Blas, Emilio Ferrara |
WWW | 4 |
| 2026 | Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information OperationsabstractGenerative agents are rapidly advancing in sophistication, raising urgent questions about how they might coordinate when deployed in online ecosystems. This is particularly consequential in information operations (IOs), influence campaigns that aim to manipulate public opinion on social media. While traditional IOs have been orchestrated by human operators and relied on manually crafted tactics, agentic AI promises to make campaigns more automated, adaptive, and difficult to detect. This work presents the first systematic study of emergent coordination among generative agents in simulated IO campaigns. Using generative agent-based modeling, we instantiate IO and organic agents in a simulated environment and evaluate coordination across operational regimes, from simple goal alignment to team knowledge and collective decision-making. As operational regimes become more structured, IO networks become denser and more clustered, interactions more reciprocal and positive, narratives more homogeneous, amplification more synchronized, and hashtag adoption faster and more sustained. Remarkably, simply revealing to agents which other agents share their goals can produce coordination levels nearly equivalent to those achieved through explicit deliberation and collective voting. Overall, we show that generative agents, even without human guidance, can reproduce coordination strategies characteristic of real-world IOs, underscoring the societal risks posed by increasingly automated, self-organizing IOs. Gian Marco Orlando, Jinyi Ye, Valerio La Gatta, Mahdi Saeedi, Vincenzo Moscato, Emilio Ferrara, Luca Luceri |
WWW | 6 |
| 2026 | Information suppression in large language models: Auditing, quantifying, and characterizing censorship in DeepSeek
Peiran Qiu, Emilio Ferrara |
Inf. Sci. | 3 |
| 2025 | The Susceptibility Paradox in Online Social InfluenceabstractUnderstanding susceptibility to online influence is crucial for mitigating the spread of misinformation and protecting vulnerable audiences. This paper investigates susceptibility to influence within social networks, focusing on the differential effects of influence-driven versus spontaneous behaviors on user content adoption. Our analysis reveals that influence-driven adoption exhibits high homophily, indicating that individuals prone to influence often connect with similarly susceptible peers, thereby reinforcing peer influence dynamics, whereas spontaneous adoption shows significant but lower homophily. Additionally, we extend the Generalized Friendship Paradox to influence-driven behaviors, demonstrating that users' friends are generally more susceptible to influence than the users themselves, de facto establishing the notion of Susceptibility Paradox in online social influence. This pattern does not hold for spontaneous behaviors, where friends exhibit fewer spontaneous adoptions. We find that susceptibility to influence can be predicted using friends' susceptibility alone, while predicting spontaneous adoption requires additional features, such as user metadata. These findings highlight the complex interplay between user engagement and characteristics in spontaneous content adoption. Our results provide new insights into social influence mechanisms and offer implications for designing more effective moderation strategies to protect vulnerable audiences. Luca Luceri, Jinyi Ye, Julie Jiang, Emilio Ferrara |
ICWSM | 4 |
| 2025 | Exposing Cross-Platform Coordinated Inauthentic Activity in the Run-Up to the 2024 U.S. ElectionabstractCoordinated information operations remain a persistent challenge on social media, despite platform efforts to curb them. While previous research has primarily focused on identifying these operations within individual platforms, this study shows that coordination frequently transcends platform boundaries. Leveraging newly collected data of online conversations related to the 2024 U.S. Election across 𝕏 (formerly, Twitter), Facebook, and Telegram, we construct similarity networks to detect coordinated communities exhibiting suspicious sharing behaviors within and across platforms. Proposing an advanced coordination detection model, we reveal evidence of potential foreign interference, with Russian-affiliated media being systematically promoted across Telegram and 𝕏. Our analysis also uncovers substantial intra- and cross-platform coordinated inauthentic activity, driving the spread of highly partisan, low-credibility, and conspiratorial content. These findings highlight the urgent need for regulatory measures that extend beyond individual platforms to effectively address the growing challenge of cross-platform coordinated influence campaigns. Federico Cinus, Marco Minici, Luca Luceri, Emilio Ferrara |
WWW | 4 |
| 2024 | Unmasking the Web of Deceit: Uncovering Coordinated Activity to Expose Information Operations on Twitter
Luca Luceri, Valeria Pantè, Keith Burghardt, Emilio Ferrara |
WWW | 4 |
| 2024 | Susceptibility to Unreliable Information Sources: Swift Adoption with Minimal ExposureabstractMisinformation proliferation on social media platforms is a pervasive threat to the integrity of online public discourse. Genuine users, susceptible to others' influence, often unknowingly engage with, endorse, and re-share questionable pieces of information, collectively amplifying the spread of misinformation. In this study, we introduce an empirical framework to investigate users' susceptibility to influence when exposed to unreliable and reliable information sources. Leveraging two datasets on political and public health discussions on Twitter, we analyze the impact of exposure on the adoption of information sources, examining how the reliability of the source modulates this relationship. Our findings provide evidence that increased exposure augments the likelihood of adoption. Users tend to adopt low-credibility sources with fewer exposures than high-credibility sources, a trend that persists even among non-partisan users. Furthermore, the number of exposures needed for adoption varies based on the source credibility, with extreme ends of the spectrum (very high or low credibility) requiring fewer exposures for adoption. Additionally, we reveal that the adoption of information sources often mirrors users' prior exposure to sources with comparable credibility levels. Our research offers critical insights for mitigating the endorsement of misinformation by vulnerable users, offering a framework to study the dynamics of content exposure and adoption on social media platforms. Jinyi Ye, Luca Luceri, Julie Jiang, Emilio Ferrara |
WWW | 4 |
| 2023 | Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War between Ukraine and RussiaabstractOn February 24, 2022, Russia invaded Ukraine. In the days that followed, reports kept flooding in from laymen to news anchors of a conflict quickly escalating into war. Russia faced immediate backlash and condemnation from the world at large. While the war continues to contribute to an ongoing humanitarian and refugee crisis in Ukraine, a second battlefield has emerged in the online space, both in the use of social media to garner support for both sides of the conflict and also in the context of information warfare. In this paper, we present a collection of nearly half a billion tweets, from February 22, 2022, through January 8, 2023, that we are publishing for the wider research community to use. This dataset can be found at https://github.com/echen102/ukraine-russia. Our preliminary analysis on a subset of our dataset already shows evidence of public engagement with Russian state-sponsored media and other domains that are known to push unreliable information towards the beginning of the war; the former saw a spike in activity on the day of the Russian invasion, while the other saw spikes in engagement within the first month of the war. Our hope is that this public dataset can help the research community to further understand the ever-evolving role that social media plays in information dissemination, influence campaigns, grassroots mobilization, and much more, during a time of conflict. Emily Chen, Emilio Ferrara |
ICWSM | 2 |
| 2023 | Retweet-BERT: Political Leaning Detection Using Language Features and Information Diffusion on Social NetworksabstractEstimating the political leanings of social media users is a challenging and ever more pressing problem given the increase in social media consumption. We introduce Retweet-BERT, a simple and scalable model to estimate the political leanings of Twitter users. Retweet-BERT leverages the retweet network structure and the language used in users' profile descriptions. Our assumptions stem from patterns of networks and linguistics homophily among people who share similar ideologies. Retweet-BERT demonstrates competitive performance against other state-of-the-art baselines, achieving 96%-97% macro-F1 on two recent Twitter datasets (a COVID-19 dataset and a 2020 United States presidential elections dataset). We also perform manual validation to validate the performance of Retweet-BERT on users not in the training data. Finally, in a case study of COVID-19, we illustrate the presence of political echo chambers on Twitter and show that it exists primarily among right-leaning users. Our code is open-sourced and our data is publicly available. Julie Jiang, Xiang Ren 0001, Emilio Ferrara |
ICWSM | 3 |
| 2023 | Identifying and Characterizing Behavioral Classes of Radicalization within the QAnon Conspiracy on TwitterabstractSocial media provide a fertile ground where conspiracy theories and radical ideas can flourish, reach broad audiences, and sometimes lead to hate or violence beyond the online world itself. QAnon represents a notable example of a political conspiracy that started out on social media but turned mainstream, in part due to public endorsement by influential political figures. Nowadays, QAnon conspiracies often appear in the news, are part of political rhetoric, and are espoused by significant swaths of people in the United States. It is therefore crucial to understand how such a conspiracy took root online, and what led so many social media users to adopt its ideas. In this work, we propose a framework that exploits both social interaction and content signals to uncover evidence of user radicalization or support for QAnon. Leveraging a large dataset of 240M tweets collected in the run-up to the 2020 US Presidential election, we define and validate a multivariate metric of radicalization. We use that to separate users in distinct, naturally-emerging, classes of behaviors associated with radicalization processes, from self-declared QAnon supporters to hyper-active conspiracy promoters. We also analyze the impact of Twitter's moderation policies on the interactions among different classes: we discover aspects of moderation that succeed, yielding a substantial reduction in the endorsement received by hyperactive QAnon accounts. But we also uncover where moderation fails, showing how QAnon content amplifiers are not deterred or affected by the Twitter intervention. Our findings refine our understanding of online radicalization processes, reveal effective and ineffective aspects of moderation, and call for the need to further investigate the role social media play in the spread of conspiracies. Emily L. Wang, Luca Luceri, Francesco Pierri 0002, Emilio Ferrara |
ICWSM | 4 |
| 2022 | Characterizing Online Engagement with Disinformation and Conspiracies in the 2020 U.S. Presidential Election
Karishma Sharma, Emilio Ferrara, Yan Liu 0002 |
ICWSM | 2 |
| 2022 | Construction of Large-Scale Misinformation Labeled Datasets from Social Media Discourse using Label RefinementabstractMalicious accounts spreading misinformation has led to widespread false and misleading narratives in recent times, especially during the COVID-19 pandemic, and social media platforms struggle to eliminate these contents rapidly. This is because adapting to new domains requires human intensive fact-checking that is slow and difficult to scale. To address this challenge, we propose to leverage news-source credibility labels as weak labels for social media posts and propose model-guided refinement of labels to construct large-scale, diverse misinformation labeled datasets in new domains. The weak labels can be inaccurate at the article or social media post level where the stance of the user does not align with the news source or article credibility. We propose a framework to use a detection model self-trained on the initial weak labels with uncertainty sampling based on entropy in predictions of the model to identify potentially inaccurate labels and correct for them using self-supervision or relabeling. The framework will incorporate social context of the post in terms of the community of its associated user for surfacing inaccurate labels towards building a large-scale dataset with minimum human effort. To provide labeled datasets with distinction of misleading narratives where information might be missing significant context or has inaccurate ancillary details, the proposed framework will use the few labeled samples as class prototypes to separate high confidence samples into false, unproven, mixture, mostly false, mostly true, true, and debunk information. The approach is demonstrated for providing a large-scale misinformation dataset on COVID-19 vaccines. Karishma Sharma, Emilio Ferrara, Yan Liu 0002 |
WWW | 2 |
| 2021 | Identifying Coordinated Accounts on Social Media through Hidden Influence and Group BehavioursabstractDisinformation campaigns on social media, involving coordinated activities from malicious accounts towards manipulating public opinion, have become increasingly prevalent. Existing approaches to detect coordinated accounts either make very strict assumptions about coordinated behaviours, or require part of the malicious accounts in the coordinated group to be revealed in order to detect the rest. To address these drawbacks, we propose a generative model, AMDN-HAGE (Attentive Mixture Density Network with Hidden Account Group Estimation) which jointly models account activities and hidden group behaviours based on Temporal Point Processes (TPP) and Gaussian Mixture Model (GMM), to capture inherent characteristics of coordination which is, accounts that coordinate must strongly influence each other's activities, and collectively appear anomalous from normal accounts. To address the challenges of optimizing the proposed model, we provide a bilevel optimization algorithm with theoretical guarantee on convergence. We verified the effectiveness of the proposed method and training algorithm on real-world social network data collected from Twitter related to coordinated campaigns from Russia's Internet Research Agency targeting the 2016 U.S. Presidential Elections, and to identify coordinated campaigns related to the COVID-19 pandemic. Leveraging the learned model, we find that the average influence between coordinated account pairs is the highest. On COVID-19, we found coordinated group spreading anti-vaccination, anti-masks conspiracies that suggest the pandemic is a hoax and political scam. Karishma Sharma, Emilio Ferrara, Yan Liu 0002 |
KDD | 3 |
| 2020 | Fair Class Balancing: Enhancing Model Fairness without Observing Sensitive AttributesabstractMachine learning models are at the foundation of modern society. Accounts of unfair models penalizing subgroups of a population have been reported in domains including law enforcement, job screening, etc. Unfairness can spur from biases in the training data, as well as from class imbalance, i.e., when a sensitive group's data is not sufficiently represented. Under such settings, balancing techniques are commonly used to achieve better prediction performance, but their effects on model fairness are largely unknown. In this paper, we first illustrate the extent to which common balancing techniques exacerbate unfairness in real-world data. Then, we propose a new method, called fair class balancing, that allows to enhance model fairness without using any information about sensitive attributes. We show that our method can achieve accurate prediction performance while concurrently improving fairness. Shen Yan 0007, Hsien-Te Kao, Emilio Ferrara |
CIKM | 3 |
| 2020 | ReCOVery: A Multimodal Repository for COVID-19 News Credibility ResearchabstractFirst identified in Wuhan, China, in December 2019, the outbreak of COVID-19 has been declared as a global emergency in January, and a pandemic in March 2020 by the World Health Organization (WHO). Along with this pandemic, we are also experiencing an "infodemic" of information with low credibility such as fake news and conspiracies. In this work, we present ReCOVery, a repository designed and constructed to facilitate research on combating such information regarding COVID-19. We first broadly search and investigate ~2,000 news publishers, from which 60 are identified with extreme [high or low] levels of credibility. By inheriting the credibility of the media on which they were published, a total of 2,029 news articles on coronavirus, published from January to May 2020, are collected in the repository, along with 140,820 tweets that reveal how these news articles have spread on the Twitter social network. The repository provides multimodal information of news articles on coronavirus, including textual, visual, temporal, and network information. The way that news credibility is obtained allows a trade-off between dataset scalability and label accuracy. Extensive experiments are conducted to present data statistics and distributions, as well as to provide baseline performances for predicting news credibility so that future methods can be compared. Our repository is available at http://coronavirus-fakenews.com. Xinyi Zhou 0001, Apurva Mulay, Emilio Ferrara, Reza Zafarani |
CIKM | 3 |
| 2020 | Detecting Troll Behavior via Inverse Reinforcement Learning: A Case Study of Russian Trolls in the 2016 US Election
Luca Luceri, Silvia Giordano, Emilio Ferrara |
ICWSM | 3 |
| 2019 | Effects of Network Structure on Subjective Preference DiversityabstractDifferent social media environments enable different levels of community connectedness, which in turn affects the information that a user is exposed to. In this study, we create an agent-based model to investigate how the different levels of connectedness as well as network structure affects a group's diversity of opinions. In the model, agents are tasked with “liking” or “disliking” a set of objects. At each turn, each agent sees the most popular object amongst the agents that they are connected to. There are two main findings: 1) low network connectivity leads to more diversity amongst agents, 2) a complete network leads information popularity distribution to be more skewed, 3) the more random a network is, the more skewed information popularity distribution is. These findings suggest that online platforms that either create well-connected communities or present aggregated information of users (such as Billboard rankings) may lead to homogeneity in the subjective preference of the users over time. Anne Lin, Andrés Abeliuk, Emilio Ferrara |
IEEE BigData | 3 |
| 2019 | Linguistic Cues to Deception: Identifying Political Trolls on Social Media
Aseel Addawood, Adam Badawy, Kristina Lerman, Emilio Ferrara |
ICWSM | 4 |
| 2018 | Analyzing the Digital Traces of Political Manipulation: The 2016 Russian Interference Twitter CampaignabstractUntil recently, social media was seen to promote democratic discourse on social and political issues. However, this powerful communication platform has come under scrutiny for allowing hostile actors to exploit online discussions in an attempt to manipulate public opinion. A case in point is the ongoing U.S. Congress investigation of Russian interference in the 2016 U.S. election campaign, with Russia accused of, among other things, using trolls (malicious accounts created for the purpose of manipulation) and bots (automated accounts) to spread misinformation and politically biased information. In this study, we explore the effects of this manipulation campaign, taking a closer look at users who re-shared the posts produced on Twitter by the Russian troll accounts publicly disclosed by U.S. Congress investigation. We collected a dataset with over 43 million elections-related posts shared on Twitter between September 16 and November 9, 2016 by about 5.7 million distinct users. This dataset includes accounts associated with the identified Russian trolls. We use label propagation to infer the users' ideology based on the news sources they shared, to classify a large number of them as liberal or conservative with precision and recall above 90%. Conservatives retweeted Russian trolls significantly more often than liberals and produced 36 times more tweets. Additionally, most of the troll content originated in, and was shared by users from Southern states. Using state-of-the-art bot detection techniques, we estimated that about 4.9% and 6.2% of liberal and conservative users respectively were bots. Text analysis on the content shared by trolls reveals that they had a mostly conservative, pro-Trump agenda. Although an ideologically broad swath of Twitter users were exposed to Russian trolls in the period leading up to the 2016 U.S. Presidential election, it was mainly conservatives who helped amplify their message. Adam Badawy, Emilio Ferrara, Kristina Lerman |
ASONAM | 2 |
| 2018 | Social Bots for Online Public Health InterventionsabstractAccording to the Center for Disease Control and Prevention, hundreds of thousands initiate smoking each year, and millions live with smoking-related diseases in the United States. Many tobacco users discuss their opinions, habits and preferences on social media. This work conceptualizes a framework for targeted health interventions to inform tobacco users about the consequences of tobacco use. We designed a Twitter bot named Notobot (short for No-Tobacco Bot) that leverages machine learning to identify users posting pro-tobacco tweets and select individualized interventions to curb their tobacco use. We searched the Twitter feed for tobacco-related keywords and phrases, and trained a convolutional neural network using over 4,000 tweets manually labeled as either pro-tobacco or not pro-tobacco. This model achieved a 90% accuracy rate on the training set and 74% on test data. Users posting protobacco tweets were matched with former smokers with similar interests who posted anti-tobacco tweets. Algorithmic matching, leveraging the power of peer influence, allows for the systematic delivery of personalized interventions based on real anti-tobacco tweets from former smokers. Experimental evaluation suggested that our system would perform well if deployed. Ashok Deb, Anuja Majmundar, Sungyong Seo, Akira Matsui, Rajat Tandon, Shen Yan 0007, Jon-Patrick Allem, Emilio Ferrara |
ASONAM | 8 |
| 2018 | Deep neural networks for bot detection
Sneha Reddy Kudugunta, Emilio Ferrara |
Inf. Sci. | 2 |
| 2017 | Dynamics of Content Quality in Collaborative Knowledge Production
Emilio Ferrara, Nazanin Alipourfard, Keith Burghardt, Chiranth Gopal, Kristina Lerman |
ICWSM | 1 |
| 2017 | Online Human-Bot Interactions: Detection, Estimation, and Characterization
Onur Varol, Emilio Ferrara, Clayton A. Davis, Filippo Menczer, Alessandro Flammini |
ICWSM | 2 |
| 2017 | Contagion dynamics of extremist propaganda in social networks
Emilio Ferrara |
Inf. Sci. | 1 |
| 2016 | Detection of Promoted Social Media Campaigns
Emilio Ferrara, Onur Varol, Filippo Menczer, Alessandro Flammini |
ICWSM | 1 |
| 2016 | Latent Space Model for Multi-Modal Social DataabstractWith the emergence of social networking services, researchers enjoy the increasing availability of large-scale heterogenous datasets capturing online user interactions and behaviors. Traditional analysis of techno-social systems data has focused mainly on describing either the dynamics of social interactions, or the attributes and behaviors of the users. However, overwhelming empirical evidence suggests that the two dimensions affect one another, and therefore they should be jointly modeled and analyzed in a multi-modal framework. The benefits of such an approach include the ability to build better predictive models, leveraging social network information as well as user behavioral signals. To this purpose, here we propose the Constrained Latent Space Model (CLSM), a generalized framework that combines Mixed Membership Stochastic Blockmodels (MMSB) and Latent Dirichlet Allocation (LDA) incorporating a constraint that forces the latent space to concurrently describe the multiple data modalities. We derive an efficient inference algorithm based on Variational Expectation Maximization that has a computational cost linear in the size of the network, thus making it feasible to analyze massive social datasets. We validate the proposed framework on two problems: prediction of social interactions from user attributes and behaviors, and behavior prediction exploiting network information. We perform experiments with a variety of multi-modal social systems, spanning location-based social networks (Gowalla), social media services (Instagram, Orkut), e-commerce and review sites (Amazon, Ciao), and finally citation networks (Cora). The results indicate significant improvement in prediction accuracy over state of the art methods, and demonstrate the flexibility of the proposed approach for addressing a variety of different learning problems commonly occurring with multi-modal social data. Yoon-Sik Cho, Greg Ver Steeg, Emilio Ferrara, Aram Galstyan |
WWW | 3 |
| 2016 | Network structure and resilience of Mafia syndicates
Santa Agreste, Salvatore Catanese, Pasquale De Meo, Emilio Ferrara, Giacomo Fiumara |
Inf. Sci. | 4 |
| 2013 | Clustering memes in social mediaabstractThe increasing pervasiveness of social media creates new opportunities to study human social behavior, while challenging our capability to analyze their massive data streams. One of the emerging tasks is to distinguish between different kinds of activities, for example engineered misinformation campaigns versus spontaneous communication. Such detection problems require a formal definition of meme, or unit of information that can spread from person to person through the social network. Once a meme is identified, supervised learning methods can be applied to classify different types of communication. The appropriate granularity of a meme, however, is hardly captured from existing entities such as tags and keywords. Here we present a framework for the novel task of detecting memes by clustering messages from large streams of social data. We evaluate various similarity measures that leverage content, metadata, network features, and their combinations. We also explore the idea of pre-clustering on the basis of existing entities. A systematic evaluation is carried out using a manually curated dataset as ground truth. Our analysis shows that pre-clustering and a combination of heterogeneous features yield the best trade-off between number of clusters and their quality, demonstrating that a simple combination based on pairwise maximization of similarity is as effective as a non-trivial optimization of parameters. Our approach is fully automatic, unsupervised, and scalable for real-time detection of memes in streaming data. Emilio Ferrara, Mohsen JafariAsbagh, Onur Varol, Vahed Qazvinian, Filippo Menczer, Alessandro Flammini |
ASONAM | 1 |
| 2013 | Enhancing community detection using a network weighting strategy
Pasquale De Meo, Emilio Ferrara, Giacomo Fiumara, Alessandro Provetti |
Inf. Sci. | 2 |
| 2013 | Scientific impact evaluation and the effect of self-citations: Mitigating the bias by discounting the h-indexabstractIn this article, we propose a measure to assess scientific impact that discounts self‐citations and does not require any prior knowledge of their distribution among publications. This index can be applied to both researchers and journals. In particular, we show that it fills the gap of the h‐index and similar measures that do not take into account the effect of self‐citations for authors or journals impact evaluation. We provide 2 real‐world examples: First, we evaluate the research impact of the most productive scholars in computer science (according to DBLP Computer Science Bibliography, Universität Trier, Trier, Germany); then we revisit the impact of the journals ranked in the Computer Science Applications section of the SCImago Journal & Country Rank ranking service (Consejo Superior de Investigaciones Científicas, University of Granada, Extremadura, Madrid, Spain). We observe how self‐citations, in many cases, affect the rankings obtained according to different measures (including h‐index and ch‐index), and show how the proposed measure mitigates this effect. Emilio Ferrara, Alfonso E. Romero |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2013 | Analyzing user behavior across social sharing environmentsabstractIn this work we present an in-depth analysis of the user behaviors on different Social Sharing systems. We consider three popular platforms, Flickr, Delicious and StumbleUpon, and, by combining techniques from social network analysis with techniques from semantic analysis, we characterize the tagging behavior as well as the tendency to create friendship relationships of the users of these platforms. The aim of our investigation is to see if (and how) the features and goals of a given Social Sharing system reflect on the behavior of its users and, moreover, if there exists a correlation between the social and tagging behavior of the users. We report our findings in terms of the characteristics of user profiles according to three different dimensions: (i) intensity of user activities, (ii) tag-based characteristics of user profiles, and (iii) semantic characteristics of user profiles. Pasquale De Meo, Emilio Ferrara, Fabian Abel, Lora Aroyo, Geert-Jan Houben |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | Effective retrieval of resources in folksonomies using a new tag similarity measureabstractSocial (or folksonomic) tagging has become a very popular way to describe content within Web 2.0 websites. However, as tags are informally defined, continually changing, and ungoverned, it has often been criticised for lowering, rather than increasing, the efficiency of searching. To address this issue, a variety of approaches have been proposed that recommend users what tags to use, both when labeling and when looking for resources. These techniques work well in dense folksonomies, but they fail to do so when tag usage exhibits a power law distribution, as it often happens in real-life folksonomies. To tackle this issue, we propose an approach that induces the creation of a dense folksonomy, in a fully automatic and transparent way: when users label resources, an innovative tag similarity metric is deployed, so to enrich the chosen tag set with related tags already present in the folksonomy. The proposed metric, which represents the core of our approach, is based on the mutual reinforcement principle. Our experimental evaluation proves that the accuracy and coverage of searches guaranteed by our metric are higher than those achieved by applying classical metrics. Giovanni Quattrone, Licia Capra, Pasquale De Meo, Emilio Ferrara, Domenico Ursino |
CIKM | 4 |