EDBT 2026 Demo / reviewers in the wild / expert
Jussara M. Almeida
dblp:34/5480 · also Jussara Marques de Almeida
· DBLP profile ↗
55ranked-venue papers in the field
0as first author
10since 2021 · last 2025
0000-0001-9142-2919ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 34Data Mining & Knowledge Discovery · 11Database Systems & Data Management · 5Knowledge Engineering, Semantic Web & Information Systems · 3Big Data, Cloud & Distributed Data Systems · 1Business Process & Enterprise Data · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Join the Chat: How Curiosity Sparks Participation in Telegram GroupsabstractThis study delves into the mechanisms that spark user curiosity driving active engagement within public Telegram groups. By analyzing approximately 6 million messages from 29,196 users across 409 groups, we identify and quantify the key factors that stimulate users to actively participate (i.e., send messages) in group discussions. These factors include social influence, novelty, complexity, uncertainty, and conflict, all measured through metrics derived from message sequences and user participation over time. After clustering the messages, we apply explainability techniques to assign meaningful labels to the clusters. This approach uncovers macro categories representing distinct curiosity stimulation profiles, each characterized by a unique combination of various stimuli. Social influence from peers and influencers drives engagement for some users, while for others, rare media types or a diverse range of senders and media sparks curiosity. Analyzing patterns, we found that user curiosity stimuli are mostly stable, but, as the time between the initial message increases, curiosity occasionally shifts. A graph-based analysis of influence networks reveals that users motivated by direct social influence tend to occupy more peripheral positions, while those who are not stimulated by any specific factors are often more central, potentially acting as initiators and conversation catalysts. These findings contribute to understanding information dissemination and spread processes on social media networks, potentially contributing to more effective communication strategies. Giordano Paoletti, Jussara M. Almeida, Luca Vassio, Marcos André Gonçalves, Marco Mellia |
ICWSM | 2 |
| 2025 | The Silent Signature: Behavior-Based User Exposure in Mobility DataabstractMobility is a fundamental aspect of human life, and mobility data offers valuable insights into user behavior. Yet, this data also exposes users to privacy risks given pattern unicity in their trajectories, i.e., the singularity in the displacements made by users. Existing strategies to quantify such user exposure either focus only on the sequences of places visited by each user, as the widely used uniqueness measure, or are tied to specific attack models. We here introduce MoBES, a novel, scalable, customizable and highly interpretable measure of user exposure in mobility data. MoBES leverages multiple existing metrics to build a multi-dimensional space, which in turn is used to capture each user's mobility signature behavior. MoBES quantifies user exposure based on how distinct a user's signature is from her neighbors in the defined metric space. As such, MoBES is designed to be a fundamental expression of user behavior, and not tied to any specific attack model. We evaluate MoBES on a real mobility dataset, showing that it effectively captures user exposure within the behavioral metric space. We also compare MoBES with the uniqueness measure, showing that MoBES is able to uncover users who, even though visiting the same places as others in the crowd, are still at risk of exposure due to the unicity of their mobility behavior. Lucas G. S. Félix, Josiane Kouam, Aline Carneiro Viana, Nadjib Achir, Jussara M. Almeida |
MDM | 5 |
| 2025 | CoDÆN: Benchmarks and Comparison of Evolutionary Community Detection Algorithms for Dynamic NetworksabstractWeb data are often modelled as complex networks in which entities interact and form communities. Nevertheless, web data evolves over time, and network communities change alongside it. This makes Community Detection (CD) in dynamic graphs a relevant problem, calling for evolutionary CD algorithms. The choice and evaluation of such algorithm performance is challenging because of the lack of a comprehensive set of benchmarks and specific metrics. To address these challenges, we propose CoDÆN—Community Detection Algorithms in Evolving Networks—a benchmarking framework for evolutionary CD algorithms in dynamic networks, that we offer as open source to the community. CoDÆN allows us to generate synthetic community-structured graphs with known ground truth and design evolving scenarios combining nine basic graph transformations that modify edges, nodes, and communities. We propose three complementary metrics (i.e., Correctness, Delay, and Stability) to compare evolutionary CD algorithms. Armed with CoDÆN, we consider three evolutionary modularity-based CD approaches, dissecting their performance to gauge the trade-off between the stability of the communities and their correctness. Next, we compare the algorithms in real Web-oriented datasets, confirming such a trade-off. Our findings reveal that algorithms that introduce memory in the graph maximise stability but add delay when abrupt changes occur. Conversely, algorithms that introduce memory by initialising the CD algorithms with the previous solution fail to identify the split and birth of new communities. These observations underscore the value of CoDÆN in facilitating the study and comparison of alternative evolutionary community detection algorithms. Giordano Paoletti, Luca Gioacchini, Marco Mellia, Luca Vassio, Jussara M. Almeida |
ACM Trans. Web | 5 |
| 2024 | Beauty or Beast: Human Behavioral Insights and Learning Power of Federated Mobility PredictionabstractMobility patterns are inherently linked to human nature (e.g., individual variability, temporal dynamics, behavioral factors, curiosity, social interaction), making mobility prediction a multifaceted and challenging problem that requires sophisticated models and comprehensive data. Machine learning (ML) models excel at predicting the location a person will be at the next time interval, but they often raise privacy concerns. To address these privacy issues while maintaining the benefits of ML models, Federated Learning (FL) offers a distributed framework that enables collaborative training of human mobility prediction models without requiring the sharing of highly sensitive location data. However, in the domain of FL for individual mobility prediction, prior work lacks a thorough understanding of the many factors that may impact the performance of FL-based prediction models. In this work, we provide a comprehensive study of the impact of various aspects related to human behavior, data characteristics, ML algorithmic solutions, and FL architectural structuring. We quantify the impact of such factors on effectiveness (accuracy) and efficiency (execution time, memory, and energy usage) of the prediction, revealing that, ignoring these factors lead to misleading result interpretation, and acknowledging them empowers both effectiveness and efficiency results. João Paulo Esper, Aline Carneiro Viana, Jussara M. Almeida |
SIGSPATIAL/GIS | 3 |
| 2024 | Unraveling User Coordination on Telegram: A Comprehensive Analysis of Political Mobilization during the 2022 Brazilian Presidential ElectionabstractSocial media has gained importance as a channel to influence people's behavior and decisions, affecting not only the online world but also real-life (offline) events. This is especially evident in Brazil, where platforms like Telegram have been instrumental in disseminating political content rapidly and widely. However, the potential coordinated use of Telegram for promoting specific political narratives at critical times, such as the 2022 Brazilian elections, remains an area that requires further investigation. This study aims to investigate this phenomenon, focusing on the first and second rounds of voting and the January 8th riots. To this end, we conducted a comprehensive analysis of 620,000 messages from 256 Telegram groups, focusing on the dynamics of message dissemination and user interactions. Using network backbone extraction and text analysis methods, we identified key users who may be orchestrating the distribution of content. Our findings suggest that these individuals play a central role in the network's topology, relaying messages to a broader audience on dominant topics of discussion that reflect Brazil's political landscape during this turbulent period. This study not only highlights the growing influence of messaging apps on political mobilization but also contributes to our understanding of digital communication strategies in modern electoral contexts, emphasizing the need for further research in this field. Otavio R. Venâncio, Carlos Henrique Gomes Ferreira, Jussara M. Almeida, Ana Paula Couto da Silva |
ICWSM | 3 |
| 2022 | Uncovering Coordinated Communities on Twitter During the 2020 U.S. ElectionabstractA large volume of content related to claims of election fraud, often associated with hate speech and extremism, was reported on Twitter during the 2020 US election, with evidence that coordinated efforts took place to promote such content on the platform. In response, Twitter announced the suspension of thousands of user accounts allegedly involved in such actions. Motivated by these events, we here propose a novel network-based approach to uncover evidence of coordination in a set of user interactions. Our approach is designed to address the challenges incurred by the often sheer volume of noisy edges in the network (i.e., edges that are unrelated to coordination) and the effects of data sampling. To that end, it exploits the joint use of two network backbone extraction techniques, namely Disparity Filter and Neighborhood Overlap, to reveal strongly tied groups of users (here referred to as communities) exhibiting repeatedly common behavior, consistent with coordination. We employ our strategy to a large dataset of tweets related to the aforementioned fraud claims, in which users were labeled as suspended, deleted or active, according to their accounts status after the election. Our findings reveal well-structured communities, with strong evidence of coordination to promote (i.e., retweet) the aforementioned fraud claims. Moreover, many of those communities are formed not only by suspended and deleted users, but also by users who, despite exhibiting very similar sharing patterns, remained active in the platform. This observation suggests that a significant number of users who were potentially involved in the coordination efforts went unnoticed by the platform, and possibly remained actively spreading this content on the system. Renan Saldanha Linhares, Jose Martins da Rosa, Carlos Henrique Gomes Ferreira, Fabricio Murai, Gabriel Peres Nobre, Jussara M. Almeida |
ASONAM | 6 |
| 2022 | A hierarchical network-oriented analysis of user participation in misinformation spread on WhatsApp
Gabriel Peres Nobre, Carlos Henrique Gomes Ferreira, Jussara M. Almeida |
Inf. Process. Manag. | 3 |
| 2021 | Analyzing topic attention in online small groupsabstractAttention is a scarce resource disputed by algorithms and people on the Internet. This competition for attention is part of online spaces especially online small groups where there is a limited number of individuals interacting with each other using text and media content that is not controlled by algorithms or human curators. In these groups, as certain participants and piece of content can catch the collective attention, a question that naturally arises is: how to analyze topic attention in online small groups? In this paper, we propose a methodology aimed at answering this question. Our proposal consists of sets of analyses over topical (obtained from topic analysis) transition graphs for characterizing attention allocation, permanence and shifting as well as participant role characterization during discussions in online small groups. We experimented with our methodology using WhatsApp groups as a case study. Among other results, we identified and characterized abrupt and smooth topic transitions as well as patterns of participant activity related to certain topics. Josemar Alves Caetano, Jussara M. Almeida, Marcos André Gonçalves, Wagner Meira Jr., Humberto Torres Marques-Neto, Virgílio A. F. Almeida |
ASONAM | 2 |
| 2021 | On the cost-effectiveness of neural and non-neural approaches and representations for text classification: A comprehensive comparative study
Washington Cunha, Vítor Mangaravite, Christian Gomes, Sérgio D. Canuto, Elaine Resende, Cecilia Nascimento, Felipe Viegas, Celso França, Wellington Santos Martins, Jussara M. Almeida, Thierson Couto, Leonardo Rocha 0001, Marcos André Gonçalves |
Inf. Process. Manag. | 10 |
| 2021 | Characterizing client usage patterns and service demand for car-sharing systems
Victor Aquiles Alencar, Felipe Rooke, Michele Cocca, Luca Vassio, Jussara M. Almeida, Alex Borges Vieira |
Inf. Syst. | 5 |
| 2020 | A Dataset of Fact-Checked Images Shared on WhatsApp During the Brazilian and Indian Elections
Julio C. S. Reis, Philipe F. Melo, Venkata Rama Kiran Garimella, Jussara M. Almeida, Dean Eckles, Fabrício Benevenuto |
ICWSM | 4 |
| 2020 | Analyzing the Use of Audio Messages in WhatsApp GroupsabstractWhatsApp is a free messaging app with more than one billion active monthly users which has become one of the main communication platforms in many countries, including Saudi Arabia, Germany, and Brazil. In addition to allowing the direct exchange of messages among pairs of users, the app also enables group conversations, where multiple people can interact with one another. A number of recent studies have shown that WhatsApp groups play an important role as an information dissemination platform, especially during important social mobilization events. In this paper, we build upon those prior efforts by taking a first look into the use of audio messages in WhatsApp groups, a type of content that is becoming increasingly important in the platform. We present a methodology to analyze audio messages shared in WhatsApp groups, characterizing content properties (e.g, topics and language characteristics), their propagation dynamics and the impact of different types of audios (e.g., speech versus music) on such dynamics. Alexandre Maros, Jussara M. Almeida, Fabrício Benevenuto, Marisa A. Vasconcelos |
WWW | 2 |
| 2020 | "Fixing the curse of the bad product descriptions" - Search-boosted tag recommendation for E-commerce products
Fabiano Muniz Belém, Rodrigo M. Silva, Cláudio M. V. de Andrade, Gabriel Person, Felipe Mingote, Raphael Ballet, Helton Alponti, Henrique P. de Oliveira, Jussara M. Almeida, Marcos André Gonçalves |
Inf. Process. Manag. | 9 |
| 2020 | Fine-grained tourism prediction: Impact of social and environmental features
Amir Khatibi, Fabiano Muniz Belém, Ana Paula Couto da Silva, Jussara M. Almeida, Marcos André Gonçalves |
Inf. Process. Manag. | 4 |
| 2019 | Analyzing and modeling user curiosity in online content consumption: a LastFM case studyabstractCuriosity is a natural trait of human behavior. When we take into account the time we spend consuming content online, it is expected that at least a fraction of that time was driven by curious behavior. Aiming at understanding how curiosity drives online information consumption, we here propose a model that captures user curiosity relying on several stimulus metrics. Our model relies on the well-established Wundt's curve from psychology and is based on metrics capturing Novelty, Complexity and Uncertainty as key stimuli driving one's curiosity. As a case study, we apply our model on a dataset of online music consumption from LastFM. We found that there are four main types of user behaviors in terms of how the curiosity stimulus metrics drive the user accesses to online music. These are characterized based on the diversity in the songs, artists and musical genres accessed. Alexandre M. Sousa, Jussara M. Almeida, Flavio Figueiredo |
ASONAM | 2 |
| 2019 | Deciphering Predictability Limits in Human MobilityabstractHuman mobility has been studied from different perspectives. One approach addresses predictability, deriving theoretical limits on the accuracy that any prediction model can achieve in a given dataset. This approach focuses on the inherent nature and fundamental patterns of human behavior captured in the dataset, filtering out factors that depend on the specificities of the prediction method adopted. In this paper, we revisit the state-of-the-art method for estimating the predictability of a person's mobility, which, despite being widely adopted, suffers from low interpretability and disregards external factors that have been suggested to improve predictability estimation, notably the use of contextual information (e.g., weather, day of the week, and time of the day). We also conduct a thorough analysis of how this widely used method works, by looking into two different measures (one proposed by us) which are easier to understand and, as shown, capture reasonably well the effects of the original technique. Additionally, we investigate strategies to incorporate different types of contextual information into predictability estimates, and show that the benefits vary depending on the underlying prediction task. Finally, we propose and evaluate alternative estimates of predictability which, while being much easier to interpret, provide comparable results to the state-of-the-art. Douglas do Couto Teixeira, Aline Carneiro Viana, Mário S. Alvim, Jussara M. Almeida |
SIGSPATIAL/GIS | 4 |
| 2019 | WhatsApp Monitor: A Fact-Checking System for WhatsApp
Philipe F. Melo, Johnnatan Messias, Gustavo Resende, Venkata Rama Kiran Garimella, Jussara M. Almeida, Fabrício Benevenuto |
ICWSM | 5 |
| 2019 | (Mis)Information Dissemination in WhatsApp: Gathering, Analyzing and CountermeasuresabstractWhatsApp has revolutionized the way people communicate and interact. It is not only cheaper than the traditional Short Message Service (SMS) communication but it also brings a new form of mobile communication: the group chats. Such groups are great forums for collective discussions on a variety of topics. In particular, in events of great social mobilization, such as strikes and electoral campaigns, WhatsApp group chats are very attractive as they facilitate information exchange among interested people. Yet, recent events have raised concerns about the spreading of misinformation in WhatsApp. In this work, we analyze information dissemination within WhatsApp, focusing on publicly accessible political-oriented groups, collecting all shared messages during major social events in Brazil: a national truck drivers' strike and the Brazilian presidential campaign. We analyze the types of content shared within such groups as well as the network structures that emerge from user interactions within and cross-groups. We then deepen our analysis by identifying the presence of misinformation among the shared images using labels provided by journalists and by a proposed automatic procedure based on Google searches. We identify the most important sources of the fake images and analyze how they propagate across WhatsApp groups and from/to other Web platforms. Gustavo Resende, Philipe F. Melo, Hugo Sousa, Johnnatan Messias, Marisa A. Vasconcelos, Jussara M. Almeida, Fabrício Benevenuto |
WWW | 6 |
| 2019 | Exploiting syntactic and neighbourhood attributes to address cold start in tag recommendation
Fabiano Muniz Belém, André G. Heringer, Jussara M. Almeida, Marcos André Gonçalves |
Inf. Process. Manag. | 3 |
| 2018 | Characterizing Politically Engaged Users' Behavior During the 2016 US Presidential CampaignabstractPolitical campaigns have frequently used the online social network as an important environment to exhibit the candidate ideas, their activities, and their electoral plans if elected. Some users are more politically engaged than others. As an example, we can observe intense political debates, especially during major campaigns on Twitter. In such context, this paper presents a characterization of politically engaged user groups on Twitter during the 2016 US Presidential Campaign. Using a rich dataset with 23 million tweets, 115 thousand user profiles and their contact network collected from January 2016 to November 2016, we identified four politically engaged user groups: advocates for both main candidates, political bots, and regular users. We present a characterization of how Twitter users behave during a political campaign through the language patterns analysis of tweets, which users receive more popularity during the campaign and how tweets from each candidate may have affected their mood variation, as expressed by the messages they share. Josemar Alves Caetano, Jussara M. Almeida, Humberto Torres Marques-Neto |
ASONAM | 2 |
| 2017 | Mining and modeling web trajectories from passive tracesabstractIn modern web, users contact lots of services, identified by the domain name of the server. The temporal sequence and transitions of visited domains form a trajectory of the user on the web. In this work, we analyze 4 weeks of such trajectories, extracted from logs collected in our university network, and mine them via big data and machine learning methodologies to extract the interests of users. Our goal is to create a model of such trajectories and find similarities so to observe peculiarity of users' browsing. Thanks to the model, we propose a methodology to automatically group together the trajectories of single users and/or communities into highly descriptive environments which in turn allow the analyst to identify the topic of interest. We propose an automatic way to highlight differences in terms of popularity and content of environments. Lastly, we analyze the transition among environments, showing how people in smaller communities, e.g., in the same department, have a much more homogeneous behaviour than people at large, e.g., in the university. Luca Vassio, Marco Mellia, Flavio Figueiredo, Ana Paula Couto da Silva, Jussara M. Almeida |
IEEE BigData | 5 |
| 2017 | A large-scale study of cultural differences using urban data about eating and drinking preferences
Thiago H. Silva 0001, Pedro O. S. Vaz de Melo, Jussara M. Almeida, Mirco Musolesi, Antonio Alfredo Ferreira Loureiro |
Inf. Syst. | 3 |
| 2017 | A survey on tag recommendation methodsabstractTags (keywords freely assigned by users to describe web content) have become highly popular on Web 2.0 applications, because of the strong stimuli and easiness for users to create and describe their own content. This increase in tag popularity has led to a vast literature on tag recommendation methods. These methods aim at assisting users in the tagging process, possibly increasing the quality of the generated tags and, consequently, improving the quality of the information retrieval (IR) services that rely on tags as data sources. Regardless of the numerous and diversified previous studies on tag recommendation, to our knowledge, no previous work has summarized and organized them into a single survey article. In this article, we propose a taxonomy for tag recommendation methods, classifying them according to the target of the recommendations, their objectives, exploited data sources, and underlying techniques. Moreover, we provide a critical overview of these methods, pointing out their advantages and disadvantages. Finally, we describe the main open challenges related to the field, such as tag ambiguity, cold start, and evaluation issues. Fabiano Muniz Belém, Jussara M. Almeida, Marcos André Gonçalves |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Dissecting a Scholar Popularity Ranking into Different Knowledge Areas
Gabriel Pacheco, Pablo Figueira, Jussara M. Almeida, Marcos André Gonçalves |
TPDL | 3 |
| 2016 | TribeFlow: Mining & Predicting User TrajectoriesabstractWhich song will Smith listen to next? Which restaurant will Alice go to tomorrow? Which product will John click next? These applications have in common the prediction of user trajectories that are in a constant state of flux over a hidden network (e.g. website links, geographic location). Moreover, what users are doing now may be unrelated to what they will be doing in an hour from now. Mindful of these challenges we propose TribeFlow, a method designed to cope with the complex challenges of learning personalized predictive models of non-stationary, transient, and time-heterogeneous user trajectories. TribeFlow is a general method that can perform next product recommendation, next song recommendation, next location prediction, and general arbitrary-length user trajectory prediction without domain-specific knowledge. TribeFlow is more accurate and up to 413x faster than top competitors. Flavio Figueiredo, Bruno Ribeiro 0001, Jussara M. Almeida, Christos Faloutsos |
WWW | 3 |
| 2016 | TrendLearner: Early prediction of popularity trends of user generated content
Flavio Figueiredo, Jussara M. Almeida, Marcos André Gonçalves, Fabrício Benevenuto |
Inf. Sci. | 2 |
| 2016 | On cold start for associative tag recommendationabstractTag recommendation strategies that exploit term co‐occurrence patterns with tags previously assigned to the target object have consistently produced state‐of‐the‐art results. However, such techniques work only for objects with previously assigned tags. Here we focus on tag recommendation for objects with no tags, a variation of the well‐known \textit{cold start} problem. We start by evaluating state‐of‐the‐art co‐occurrence based methods in cold start. Our results show that the effectiveness of these methods suffers in this situation. Moreover, we show that employing various automatic filtering strategies to generate an initial tag set that enables the use of co‐occurrence patterns produces only marginal improvements. We then propose a new approach that exploits both positive and negative user feedback to iteratively select input tags along with a genetic programming strategy to learn the recommendation function. Our experimental results indicate that extending the methods to include user relevance feedback leads to gains in precision of up to 58% over the best baseline in cold start scenarios and gains of up to 43% over the best baseline in objects that contain some initial tags (i.e., no cold start). We also show that our best relevance‐feedback‐driven strategy performs well even in scenarios that lack user cooperation (i.e., users may refuse to provide feedback) and user reliability (i.e., users may provide the wrong feedback). Eder Ferreira Martins, Fabiano Muniz Belém, Jussara M. Almeida, Marcos André Gonçalves |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | A quantitative analysis of the temporal effects on automatic text classificationabstractAutomatic text classification (TC) continues to be a relevant research topic and several TC algorithms have been proposed. However, the majority of TC algorithms assume that the underlying data distribution does not change over time. In this work, we are concerned with the challenges imposed by the temporal dynamics observed in textual data sets. We provide evidence of the existence of temporal effects in three textual data sets, reflected by variations observed over time in the class distribution, in the pairwise class similarities, and in the relationships between terms and classes. We then quantify, using a series of full factorial design experiments, the impact of these effects on four well‐known TC algorithms. We show that these temporal effects affect each analyzed data set differently and that they restrict the performance of each considered TC algorithm to different extents. The reported quantitative analyses, which are the original contributions of this article, provide valuable new insights to better understand the behavior of TC algorithms when faced with nonstatic (temporal) data distributions and highlight important requirements for the proposal of more accurate classification models. Thiago Salles, Leonardo Rocha 0001, Marcos André Gonçalves, Jussara M. Almeida, Fernando Mourão, Wagner Meira Jr., Felipe Viegas |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2016 | Beyond Relevance: Explicitly Promoting Novelty and Diversity in Tag RecommendationabstractThe design and evaluation of tag recommendation methods has historically focused on maximizing the relevance of the suggested tags for a given object, such as a movie or a song. However, relevance by itself may not be enough to guarantee recommendation usefulness. Promoting novelty and diversity in tag recommendation not only increases the chances that the user will select “some” of the recommended tags but also promotes complementary information (i.e., tags), which helps to cover multiple aspects or topics related to the target object. Previous work has addressed the tag recommendation problem by exploiting at most two of the following aspects: (1) relevance, (2) explicit topic diversity, and (3) novelty. In contrast, here we tackle these three aspects conjointly, by introducing two new tag recommendation methods that cover all three aspects of the problem at different levels. Our first method, called Random Forest with topic-related attributes , or RF t , extends a relevance-driven tag recommender based on the Random Forest ( RF ) learning-to-rank method by including new tag attributes to capture the extent to which a candidate tag is related to the topics of the target object. This solution captures topic diversity as well as novelty at the attribute level while aiming at maximizing relevance in its objective function. Our second method, called Explicit Tag Recommendation Diversifier with Novelty Promotion , or xTReND , reranks the recommendations provided by any tag recommender to jointly promote relevance, novelty, and topic diversity. We use RF t as a basic recommender applied before the reranking, thus building a solution that addresses the problem at both attribute and objective levels. Furthermore, to enable the use of our solutions on applications in which category information is unavailable, we investigate the suitability of using latent Dirichlet allocation (LDA) to automatically generate topics for objects. We evaluate all tag recommendation approaches using real data from five popular Web 2.0 applications. Our results show that RF t greatly outperforms the relevance-driven RF baseline in diversity while producing gains in relevance as well. We also find that our new xTReND reranker obtains considerable gains in both novelty and relevance when compared to that same baseline while keeping the same relevance levels. Furthermore, compared to our previous reranker method, xTReD , which does not consider novelty, xTReND is also quite effective, improving the novelty of the recommended tags while keeping similar relevance and diversity levels in most datasets and scenarios. Comparing our two new proposals, we find that xTReND considerably outperforms RF t in terms of novelty and diversity with only small losses (under 4%) in relevance. Overall, considering the trade-off among relevance, novelty, and diversity, our results demonstrate the superiority of xTReND over the baselines and the proposed alternative, RF t . Finally, the use of automatically generated latent topics as an alternative to manually labeled categories also provides significant improvements, which greatly enhances the applicability of our solutions to applications where the latter is not available. Fabiano Muniz Belém, Carolina S. Batista, Rodrygo L. T. Santos, Jussara M. Almeida, Marcos André Gonçalves |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Twitter Population Sample Bias and its impact on predictive outcomes: a case study on electionsabstractIn the past years a lot of effort has been spent analyzing online social network data to understand how the world reality is reflected in the "virtual" world. Twitter is by far the network most used in these studies, given its policy of public data availability. However, a big discussion is still on on whether the data available is enough to make user characterization or event outcomes prediction, and what are the pitfalls people do not usually account for. In this direction, we propose a new methodology for drawing representative samples from Twitter data, which is divided into four phases: (i) user filtering, (ii) user demographic characterization, (iii) user sampling, and (iv) event prediction. The methodology is tested into a common scenario in Twitter event outcome prediction: elections. The methodology was tested with municipality elections from six different Brazilian cities, and compared to official election results. Results show it is worth further investigating the topic, but that a very hight number of messages is required to match real data distributions. Renato Miranda, Jussara M. Almeida, Gisele L. Pappa |
ASONAM | 2 |
| 2015 | On the Impact of Academic Factors on Scholar Popularity: A Cross-Area Study
Pablo Figueira, Gabriel Pacheco, Jussara M. Almeida, Marcos André Gonçalves |
TPDL | 3 |
| 2015 | A genetic programming framework to schedule webpage updates
Aécio S. R. Santos, Cristiano R. de Carvalho, Jussara M. Almeida, Edleno Silva de Moura, Altigran S. da Silva, Nivio Ziviani |
Inf. Retr. J. | 3 |
| 2015 | Predicting the popularity of micro-reviews: A Foursquare case study
Marisa A. Vasconcelos, Jussara M. Almeida, Marcos André Gonçalves |
Inf. Sci. | 2 |
| 2014 | On the Effectiveness of Concern Metrics to Detect Code Smells: An Empirical Study
Juliana Padilha, Juliana Alves Pereira, Eduardo Figueiredo 0001, Jussara M. Almeida, Alessandro F. Garcia 0001, Cláudio Sant'Anna |
CAiSE | 4 |
| 2014 | You Are What You Eat (and Drink): Identifying Cultural Boundaries by Analyzing Food and Drink Habits in Foursquare
Thiago H. Silva 0001, Pedro O. S. Vaz de Melo, Jussara M. Almeida, Mirco Musolesi, Antonio Alfredo Ferreira Loureiro |
ICWSM | 3 |
| 2014 | Revisit Behavior in Social Media: The Phoenix-R Model and Discoveries
Flavio Figueiredo, Jussara M. Almeida, Yasuko Matsubara, Bruno Ribeiro 0001, Christos Faloutsos |
ECML/PKDD (1) | 2 |
| 2014 | Personalized and object-centered tag recommendation methods for Web 2.0 applications
Fabiano Muniz Belém, Eder Ferreira Martins, Jussara M. Almeida, Marcos André Gonçalves |
Inf. Process. Manag. | 3 |
| 2013 | Exploiting Novelty and Diversity in Tag Recommendation
Fabiano Muniz Belém, Eder Ferreira Martins, Jussara M. Almeida, Marcos André Gonçalves |
ECIR | 3 |
| 2013 | Topic diversity in tag recommendationabstractTag recommendation approaches have historically focused on maximizing the relevance of the recommended tags for a given object, such as a movie or a song. Nevertheless, different users may be interested in the same object for different reasons---for instance, the Star Wars movies may appeal to both adventure as well as to fantasy movie fans. In this situation, a sensible strategy is to provide a user with diverse recommendations of how to tag the object. In this paper, we address the problem of recommending relevant and diverse tags as a ranking problem. In particular, we propose a novel tag recommendation approach that explicitly takes into account the possible topics (e.g., categories) underlying an object in order to promote tags with high coverage and low redundancy with respect to these topics. We thoroughly evaluate our proposed approach using data collected from two popular Web 2.0 applications, namely, LastFM and MovieLens. Our experimental results attest the effectiveness of our approach at promoting more relevant and diverse tags in contrast to state-of-the-art relevance-based methods as well as a recently proposed method that takes both relevance and diversity into account. Fabiano Muniz Belém, Rodrygo L. T. Santos, Jussara M. Almeida, Marcos André Gonçalves |
RecSys | 3 |
| 2013 | Learning to Schedule Webpage Updates Using Genetic Programming
Aécio S. R. Santos, Nivio Ziviani, Jussara M. Almeida, Cristiano R. de Carvalho, Edleno Silva de Moura, Altigran S. da Silva |
SPIRE | 3 |
| 2013 | Using early view patterns to predict the popularity of youtube videosabstractPredicting Web content popularity is an important task for supporting the design and evaluation of a wide range of systems, from targeted advertising to effective search and recommendation services. We here present two simple models for predicting the future popularity of Web content based on historical information given by early popularity measures. Our approach is validated on datasets consisting of videos from the widely used YouTube video-sharing portal. Our experimental results show that, compared to a state-of-the-art baseline model, our proposed models lead to significant decreases in relative squared errors, reaching up to 20% reduction on average, and larger reductions (of up to 71%) for videos that experience a high peak in popularity in their early days followed by a sharp decrease in popularity. Henrique Pinto, Jussara M. Almeida, Marcos André Gonçalves |
WSDM | 2 |
| 2013 | Assessing the quality of textual features in social media
Flavio Figueiredo, Henrique Pinto, Fabiano Muniz Belém, Jussara M. Almeida, Marcos André Gonçalves, David Fernandes de Oliveira, Edleno Silva de Moura |
Inf. Process. Manag. | 4 |
| 2012 | Automatic query expansion based on tag recommendationabstractWe here propose a new method for expanding entity related queries that automatically filters, weights and ranks candidate expasion terms extracted from Wikipedia articles related to the original query. Our method is based on state-of-the-art tag recommendation methods that exploit heuristic metrics to estimate the descriptive capacity of a given term. Originally proposed for the context of tags, we here apply these recommendation methods to weight and rank terms extracted from multiple fields of Wikipedia articles according to their relevance for the article. We evaluate our method comparing it against three state-of-the-art baselines in three collections. Our results indicate that our method outperforms all baselines in all collections, with relative gains in MAP of up to 14% against the best ones. Vitor Campos de Oliveira, Guilherme de C. M. Gomes, Fabiano Muniz Belém, Wladmir Cardoso Brandão, Jussara M. Almeida, Nivio Ziviani, Marcos André Gonçalves |
CIKM | 5 |
| 2012 | Automatic Vandalism Detection in Wikipedia with Active Associative Classification
Maria I. M. Sumbana, Marcos André Gonçalves, Rodrigo Silva Oliveira, Jussara M. Almeida, Adriano Veloso |
TPDL | 4 |
| 2012 | Tips, dones and todos: uncovering user profiles in foursquareabstractOnline Location Based Social Networks (LBSNs), which combine social network features with geographic information sharing, are becoming increasingly popular. One such application is Foursquare, which doubled its user population in less than six months. Among other features, Foursquare allows users to leave tips (i.e., reviews or recommendations) at specific venues as well as to give feedback on previously posted tips by adding them to their to-do lists or marking them as done. In this paper, we analyze how Foursquare users exploit these three features - tips, dones and to-dos - uncovering different behavior profiles. Our study reveals the existence of very active and influential users, some of which are famous businesses and brands, that seem engaged in posting tips at a large variety of venues while also receiving a great amount of user feedback on them. We also provide evidence of spamming, showing the existence of users that post tips whose contents are unrelated to the nature or domain of the venue where the tips were left. Marisa A. Vasconcelos, Saulo M. R. Ricci, Jussara M. Almeida, Fabrício Benevenuto, Virgílio A. F. Almeida |
WSDM | 3 |
| 2012 | A tool for generating synthetic authorship records for evaluating author name disambiguation methods
Anderson A. Ferreira, Marcos André Gonçalves, Jussara M. Almeida, Alberto H. F. Laender, Adriano Veloso |
Inf. Sci. | 3 |
| 2011 | Associative tag recommendation exploiting multiple textual featuresabstractThis work addresses the task of recommending relevant tags to a target object by jointly exploiting three dimensions of the problem: (i) term co-occurrence with tags pre-assigned to the target object, (ii) terms extracted from multiple textual features, and (iii) several metrics of tag relevance. In particular, we propose several new heuristic methods, which extend state-of-the-art strategies by including new metrics that try to capture how accurately a candidate term describes the object's content. We also exploit two learning-to-rank (L2R) techniques, namely RankSVM and Genetic Programming, for the task of generating ranking functions that combine multiple metrics to accurately estimate the relevance of a tag to a given object. We evaluate all proposed methods in various scenarios for three popular Web 2.0 applications, namely, LastFM, YouTube and YahooVideo. We found that our new heuristics greatly outperform the methods on which they are based, producing gains in precision of up to 181%, as well as another state-of-the-art technique, with improvements in precision of up to 40% over the best baseline in any scenario. Further improvements can also be achieved with the new L2R strategies, which have the additional advantage of being quite flexible and extensible to exploit other aspects of the tag recommendation problem. Fabiano Muniz Belém, Eder Ferreira Martins, Tatiana Pontes, Jussara M. Almeida, Marcos André Gonçalves |
SIGIR | 4 |
| 2011 | GreenMeter: a tool for assessing the quality and recommending tags for web 2.0 applicationsabstractWe present GreenMeter, a tool for assessing the quality and recommending tags for Web 2.0 content. Its goal is to improve tag quality and the effectiveness of various information services (e.g., search, content recommendation) that rely on tags as data sources. We demonstrate an implementation of GreenMeter for the popular Last.fm application. Saulo M. R. Ricci, Dilson Almeida Guimarães, Fabiano Muniz Belém, Jussara M. Almeida, Marcos André Gonçalves, Raquel Oliveira Prates |
SIGIR | 4 |
| 2011 | The tube over time: characterizing popularity growth of youtube videosabstractUnderstanding content popularity growth is of great importance to Internet service providers, content creators and online marketers. In this work, we characterize the growth patterns of video popularity on the currently most popular video sharing application, namely YouTube. Using newly provided data by the application, we analyze how the popularity of individual videos evolves since the video's upload time. Moreover, addressing a key aspect that has been mostly overlooked by previous work, we characterize the types of the referrers that most often attracted users to each video, aiming at shedding some light into the mechanisms (e.g., searching or external linking) that often drive users towards a video, and thus contribute to popularity growth. Our analyses are performed separately for three video datasets, namely, videos that appear in the YouTube top lists, videos removed from the system due to copyright violation, and videos selected according to random queries submitted to YouTube's search engine. Our results show that popularity growth patterns depend on the video dataset. In particular, copyright protected videos tend to get most of their views much earlier in their lifetimes, often exhibiting a popularity growth characterized by a viral epidemic-like propagation process. In contrast, videos in the top lists tend to experience sudden significant bursts of popularity. We also show that not only search but also other YouTube internal mechanisms play important roles to attract users to videos in all three datasets. Flavio Figueiredo, Fabrício Benevenuto, Jussara M. Almeida |
WSDM | 3 |
| 2010 | Exploiting co-occurrence and information quality metrics to recommend tags in web 2.0 applicationsabstractThis work addresses the task of recommending high quality tags by exploiting not only previously assigned tags, but also terms extracted from other textual features (e.g., title and description) associated with the target object.To estimate the quality of a candidate tag recommendation, we use several metrics related to both tag co-occurrence and information quality. We also propose a heuristic function to combine the metrics to produce a final ranking of the recommended tags. We evaluate our heuristic function in various scenarios, for three popular Web 2.0 applications. Our experimental results indicate that our heuristic function significantly outperforms two state-of-the-art tag recommendation algorithms. Fabiano Muniz Belém, Eder Ferreira Martins, Jussara M. Almeida, Marcos André Gonçalves, Gisele L. Pappa |
CIKM | 3 |
| 2010 | Demand-Driven Tag Recommendation
Guilherme Vale Menezes, Jussara M. Almeida, Fabiano Muniz Belém, Marcos André Gonçalves, Anísio Lacerda, Edleno Silva de Moura, Gisele L. Pappa, Adriano Veloso, Nivio Ziviani |
ECML/PKDD (2) | 2 |
| 2009 | Evidence of quality of textual features on the web 2.0abstractThe growth of popularity of Web 2.0 applications greatly increased the amount of social media content available on the Internet. However, the unsupervised, user-oriented nature of this source of information, and thus, its potential lack of quality, have posed a challenge to information retrieval (IR) services. Previous work focuses mostly only on tags, although a consensus about its effectiveness as supporting information for IR services has not yet been reached. Moreover, other textual features of the Web 2.0 are generally overseen by previous research. Flavio Figueiredo, Fabiano Muniz Belém, Henrique Pinto, Jussara M. Almeida, Marcos André Gonçalves, David Fernandes de Oliveira, Edleno Silva de Moura, Marco Cristo |
CIKM | 4 |
| 2009 | Detecting spammers and content promoters in online video social networksabstractA number of online video social networks, out of which YouTube is the most popular, provides features that allow users to post a video as a response to a discussion topic. These features open opportunities for users to introduce polluted content, or simply pollution, into the system. For instance, spammers may post an unrelated video as response to a popular one aiming at increasing the likelihood of the response being viewed by a larger number of users. Moreover, opportunistic users--promoters--may try to gain visibility to a specific video by posting a large number of (potentially unrelated) responses to boost the rank of the responded video, making it appear in the top lists maintained by the system. Content pollution may jeopardize the trust of users on the system, thus compromising its success in promoting social interactions. In spite of that, the available literature is very limited in providing a deep understanding of this problem. Fabrício Benevenuto, Virgílio A. F. Almeida, Jussara M. Almeida, Marcos André Gonçalves |
SIGIR | 4 |
| 2007 | Traffic Characteristics and Communication Patterns in Blogosphere
Fernando Duarte, Bernardo Mattos, Azer Bestavros, Virgílio A. F. Almeida, Jussara M. Almeida |
ICWSM | 5 |
| 2004 | Analyzing client interactivity in streaming mediaabstractThis paper provides an extensive analysis of pre-stored streaming media workloads, focusing on the client interactive behavior. We analyze four workloads that fall into three different domains, namely, education, entertainment video and entertainment audio. Our main goals are: (a) to identify qualitative similarities and differences in the typical client behavior for the three workload classes and (b) to provide data for generating realistic synthetic workloads. Cristiano P. Costa, Ítalo S. Cunha, Alex Borges Vieira, Claudiney Vander Ramos, Marcus Vinicius de Melo Rocha, Jussara M. Almeida, Berthier A. Ribeiro-Neto |
WWW | 6 |