VLDB 2026 Research / reviewers in the wild / expert
Carlos Castillo 0001
dblp:c/CarlosCastillo1
· DBLP profile ↗
81ranked-venue papers in the field
8as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 49 (6 first)Data Mining & Knowledge Discovery · 19Database Systems & Data Management · 6 (1 first)Other / Interdisciplinary · 6 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Does fair ranking lead to fair recruitment outcomes? A study of interventions, interfaces, and interactionsabstract• Fair exposure ≠ fair outcomes. Visibility in rankings does not guarantee equitable shortlisting in online recruitment. • Human and task factors matter. Recruiter behavior, task design, and candidate cues influence fairness beyond algorithms, as we show in user studies. • Design implications. Our findings translate into practical implications for model evaluations, interfaces, and recruiter practices to support equity. Personnel recruitment is increasingly mediated by Applicant Tracking Systems (ATS), which rank candidates for job positions, making them a central decision-support tool in modern Human Resources (HR) processes. Often framed as an information retrieval (IR) problem, the ranking of candidates in ATS is typically driven by relevance to the job position, with algorithms sorting applicants according to a set of predefined criteria. In recent years, fairness-aware ranking methods have emerged to mitigate the risk of indirect discrimination, where the ordering of candidates may inadvertently favor one demographic group over another. These approaches are inspired by browsing models developed for web search and aim to balance candidate exposure based on protected characteristics. However, ATS in recruitment introduce unique challenges due to their high-stakes nature and the decision-making context in which they operate. In this paper, we present a series of user studies that explore the disconnect between fair exposure and fair outcomes in candidate shortlisting. We focus on how factors such as task design (e.g., how recruiters interact with candidate lists), individual representations of candidates (e.g., national origin cues), and ranking order influence both position bias and demographic balance. Our findings show that while demographic balance may be achieved in terms of ranking visibility, this does not necessarily translate to fair outcomes in terms of who gets shortlisted. Through a crowdsourced experiment and in-depth interviews with recruiters, we identify key task-level, individual, and ranking factors that mediate these effects. We conclude that fairness in ATS rankings is contingent not only on algorithmic design but also on the shortlisting tasks they support, as well as the interfaces, strategies, and assumptions that recruiters use when interacting with candidate lists. Based on these insights, we provide implications for the design of algorithms, interfaces, and recruitment processes that support fairer and more equitable recruitment outcomes. Alessandro Fabris, Clara Rus, Jorge Saldivar, Anna Gatzioura, Asia J. Biega, Carlos Castillo 0001 |
Inf. Process. Manag. | 6 |
| 2024 | Assessing the Impact of Music Recommendation Diversity on Listeners: A Longitudinal StudyabstractWe present the results of a 12-week longitudinal user study wherein the participants, 110 subjects from Southern Europe, received on a daily basis Electronic Music (EM) diversified recommendations. By analyzing their explicit and implicit feedback, we show that exposure to specific levels of music recommendation diversity may be responsible for long-term impacts on listeners’ attitudes. In particular, we highlight the function of diversity in increasing the openness in listening to EM, a music genre not particularly known or liked by the participants previous to their participation in the study. Moreover, we demonstrate that recommendations may help listeners in removing positive and negative attachments towards EM, deconstructing pre-existing implicit associations but also stereotypes associated with this music. In addition, our results show the significant influence that recommendation diversity has in generating curiosity in listeners. Lorenzo Porcaro, Emilia Gómez, Carlos Castillo 0001 |
Trans. Recomm. Syst. | 3 |
| 2023 | Disparity, Inequality, and Accuracy Tradeoffs in Graph Neural Networks for Node ClassificationabstractGraph neural networks (GNNs) are increasingly used in critical human applications for predicting node labels in attributed graphs. Their ability to aggregate features from nodes' neighbors for accurate classification also has the capacity to exacerbate existing biases in data or to introduce new ones towards members from protected demographic groups. Thus, it is imperative to quantify how GNNs may be biased and to what extent their harmful effects may be mitigated. To this end, we propose two new GNN-agnostic interventions namely, (i) PFR-AX which decreases the separability between nodes in protected and non-protected groups, and (ii) PostProcess which updates model predictions based on a blackbox policy to minimize differences between error rates across demographic groups. Through a large set of experiments on four datasets, we frame the efficacies of our approaches (and three variants) in terms of their algorithmic fairness-accuracy tradeoff and bench- mark our results against three strong baseline interventions on three state-of-the-art GNN models. Our results show that no single intervention offers a universally optimal tradeoff, but PFR-AX and PostProcess provide granular control and improve model confidence when correctly predicting positive outcomes for nodes in protected groups. Arpit Merchant, Carlos Castillo 0001 |
CIKM | 2 |
| 2023 | SciLander: Mapping the Scientific News LandscapeabstractThe COVID-19 pandemic has fueled the spread of misinformation on social media and the Web as a whole. The phenomenon dubbed `infodemic' has taken the challenges of information veracity and trust to new heights by massively introducing seemingly scientific and technical elements into misleading content. Despite the existing body of work on modeling and predicting misinformation, the coverage of very complex scientific topics with inherent uncertainty and an evolving set of findings, such as COVID-19, provides many new challenges that are not easily solved by existing tools. To address these issues, we introduce SciLander, a method for learning representations of news sources reporting on science-based topics. We extract four heterogeneous indicators for the sources; two generic indicators that capture (1) the copying of news stories between sources, and (2) the use of the same terms to mean different things (semantic shift), and two scientific indicators that capture (1) the usage of jargon and (2) the stance towards specific citations. We use these indicators as signals of source agreement, sampling pairs of positive (similar) and negative (dissimilar) samples, and combine them in a unified framework to train unsupervised news source embeddings with a triplet margin loss objective. We evaluate our method on a novel COVID-19 dataset containing nearly 1M news articles from 500 sources spanning a period of 18 months since the beginning of the pandemic in 2020. Our results show that the features learned by our model outperform state-of-the-art baseline methods on the task of news veracity classification. Furthermore, a clustering analysis suggests that the learned representations encode information about the reliability, political leaning, and partisanship bias of these sources. Maurício Gruppi, Panayiotis Smeros, Sibel Adali, Carlos Castillo 0001, Karl Aberer |
ICWSM | 4 |
| 2022 | Diversity in the Music Listening Experience: Insights from Focus Group InterviewsabstractMusic listening in today’s digital spaces is highly characterized by the availability of huge music catalogues, accessible by people all over the world. In this scenario, recommender systems are designed to guide listeners in finding tracks and artists that best fit their requests, having therefore the power to influence the diversity of the music they listen to. Albeit several works have proposed new techniques for developing diversity-aware recommendations, little is known about how people perceive diversity while interacting with music recommendations. In this study, we interview several listeners about the role that diversity plays in their listening experience, trying to get a better understanding of how they interact with music recommendations. We recruit the listeners among the participants of a previous quantitative study, where they were confronted with the notion of diversity when asked to identify, from a series of electronic music lists, the most diverse ones according to their beliefs. As a follow-up, in this qualitative study we carry out semi-structured interviews to understand how listeners may assess the diversity of a music list and to investigate their experiences with music recommendation diversity. We report here our main findings on 1) what can influence the diversity assessment of tracks and artists’ music lists, and 2) which factors can characterize listeners’ interaction with music recommendation diversity. Lorenzo Porcaro, Emilia Gómez, Carlos Castillo 0001 |
CHIIR | 3 |
| 2022 | Exposure Inequality in People Recommender Systems: The Long-Term Effects
Francesco Fabbri, Maria Luisa Croci, Francesco Bonchi, Carlos Castillo 0001 |
ICWSM | 4 |
| 2022 | Rewiring What-to-Watch-Next Recommendations to Reduce Radicalization PathwaysabstractRecommender systems typically suggest to users content similar to what they consumed in the past. If a user happens to be exposed to strongly polarized content, she might subsequently receive recommendations which may steer her towards more and more radicalized content, eventually being trapped in what we call a “radicalization pathway”. In this paper, we study the problem of mitigating radicalization pathways using a graph-based approach. Specifically, we model the set of recommendations of a “what-to-watch-next” recommender as a d-regular directed graph where nodes correspond to content items, links to recommendations, and paths to possible user sessions. Francesco Fabbri, Yanhao Wang 0001, Francesco Bonchi, Carlos Castillo 0001, Michael Mathioudakis |
WWW | 4 |
| 2022 | Fair Top-k Ranking with multiple protected groups
Meike Zehlike, Tom Sühr, Ricardo Baeza-Yates, Francesco Bonchi, Carlos Castillo 0001, Sara Hajian |
Inf. Process. Manag. | 5 |
| 2021 | SciClops: Detecting and Contextualizing Scientific Claims for Assisting Manual Fact-CheckingabstractThis paper describes SciClops, a method to help combat online scientific misinformation. Although automated fact-checking methods have gained significant attention recently, they require pre-existing ground-truth evidence, which, in the scientific context, is sparse and scattered across a constantly-evolving scientific literature. Existing methods do not exploit this literature, which can effectively contextualize and combat science-related fallacies. Furthermore, these methods rarely require human intervention, which is essential for the convoluted and critical domain of scientific misinformation. Panayiotis Smeros, Carlos Castillo 0001, Karl Aberer |
CIKM | 2 |
| 2021 | Learning to Classify Morals and Conventions: Artificial Intelligence in Terms of the Economics of Convention
David Solans, Christopher Tauchmann, Aideen Farrell, Karolin Kappler, Hans-Hendrik Huber, Carlos Castillo 0001, Kristian Kersting |
ICWSM | 6 |
| 2020 | The Effect of Homophily on Disparate Visibility of Minorities in People Recommender Systems
Francesco Fabbri, Francesco Bonchi, Ludovico Boratto, Carlos Castillo 0001 |
ICWSM | 4 |
| 2020 | Poisoning Attacks on Algorithmic Fairness
David Solans, Battista Biggio, Carlos Castillo 0001 |
ECML/PKDD (1) | 3 |
| 2020 | Reducing Disparate Exposure in Ranking: A Learning To Rank ApproachabstractRanked search results have become the main mechanism by which we find content, products, places, and people online. Thus their ordering contributes not only to the satisfaction of the searcher, but also to career and business opportunities, educational placement, and even social success of those being ranked. Researchers have become increasingly concerned with systematic biases in data-driven ranking models, and various post-processing methods have been proposed to mitigate discrimination and inequality of opportunity. This approach, however, has the disadvantage that it still allows an unfair ranking model to be trained. Meike Zehlike, Carlos Castillo 0001 |
WWW | 2 |
| 2020 | SciLens News Platform: A System for Real-Time Evaluation of News ArticlesabstractWe demonstrate the SciLens News Platform, a novel system for evaluating the quality of news articles. The SciLens News Platform automatically collects contextual information about news articles in real-time and provides quality indicators about their validity and trustworthiness. These quality indicators derive from i) social media discussions regarding news articles, showcasing the reach and stance towards these articles, and ii) their content and their referenced sources, showcasing the journalistic foundations of these articles. Furthermore, the platform enables domain-experts to review articles and rate the quality of news sources. This augmented view of news articles, which combines automatically extracted indicators and domain-expert reviews, has provably helped the platform users to have a better consensus about the quality of the underlying articles. The platform is built in a distributed and robust fashion and runs operationally handling daily thousands of news articles. We evaluate the SciLens News Platform on the emerging topic of COVID-19 where we highlight the discrepancies between low and high-quality news outlets based on three axes, namely their newsroom activity, evidence seeking and social engagement. A live demonstration of the platform can be found here: http://scilens.epfl.ch. Angelika Romanou, Panayiotis Smeros, Carlos Castillo 0001, Karl Aberer |
Proc. VLDB Endow. | 3 |
| 2019 | Modeling human annotation errors to design bias-aware systems for social stream processingabstractHigh-quality human annotations are necessary to create effective machine learning systems for social media. Low-quality human annotations indirectly contribute to the creation of inaccurate or biased learning systems. We show that human annotation quality is dependent on the ordering of instances shown to annotators (referred as 'annotation schedule'), and can be improved by local changes in the instance ordering provided to the annotators, yielding a more accurate annotation of the data stream for efficient real-time social media analytics. Rahul Pandey, Carlos Castillo 0001, Hemant Purohit |
ASONAM | 2 |
| 2019 | SciLens: Evaluating the Quality of Scientific News Articles Using Social Media and Scientific Literature IndicatorsabstractThis paper describes, develops, and validates SciLens, a method to evaluate the quality of scientific news articles. The starting point for our work are structured methodologies that define a series of quality aspects for manually evaluating news. Based on these aspects, we describe a series of indicators of news quality. According to our experiments, these indicators help non-experts evaluate more accurately the quality of a scientific news article, compared to non-experts that do not have access to these indicators. Furthermore, SciLens can also be used to produce a completely automated quality score for an article, which agrees more with expert evaluators than manual evaluations done by non-experts. One of the main elements of SciLens is the focus on both content and context of articles, where context is provided by (1) explicit and implicit references on the article to scientific literature, and (2) reactions in social media referencing the article. We show that both contextual elements can be valuable sources of information for determining article quality. The validation of SciLens, done through a combination of expert and non-expert annotation, demonstrates its effectiveness for both semi-automatic and automatic quality evaluation of scientific news. Panayiotis Smeros, Carlos Castillo 0001, Karl Aberer |
WWW | 2 |
| 2018 | Social-EOC: Serviceability Model to Rank Social Media Requests for Emergency Operation CentersabstractThe public expects a prompt response from emergency services to address requests for help posted on social media. However, the information overload of social media experienced by these organizations, coupled with their limited human resources, challenges them to timely identify and prioritize critical requests. This is particularly acute in crisis situations where any delay may have a severe impact on the effectiveness of the response. While social media has been extensively studied during crises, there is limited work on formally characterizing serviceable help requests and automatically prioritizing them for a timely response. In this paper, we present a formal model of serviceability called Social-EOC (Social Emergency Operations Center), which describes the elements of a serviceable message posted in social media that can be expressed as a request. We also describe a system for the discovery and ranking of highly serviceable requests, based on the proposed serviceability model. We validate the model for emergency services, by performing an evaluation based on real-world data from six crises, with ground truth provided by emergency management practitioners. Our experiments demonstrate that features based on the serviceability model improve the performance of discovering and ranking (nDCG up to 25%) service requests over different baselines. In the light of these experiments, the application of the serviceability model could reduce the cognitive load on emergency operation center personnel, in filtering and ranking public requests at scale. Hemant Purohit, Carlos Castillo 0001, Muhammad Imran 0002, Rahul Pandey |
ASONAM | 2 |
| 2018 | EviDense: A Graph-Based Method for Finding Unique High-Impact Events with Succinct Keyword-Based Descriptions
Oana Balalau, Carlos Castillo 0001, Mauro Sozio |
ICWSM | 2 |
| 2018 | The Effect of Extremist Violence on Hateful Speech Online
Alexandra Olteanu, Carlos Castillo 0001, Jeremy Boy, Kush R. Varshney |
ICWSM | 2 |
| 2018 | Algorithms for Hiring and Outsourcing in the Online Labor MarketabstractAlthough freelancing work has grown substantially in recent years, in part facilitated by a number of online labor marketplaces, %(e.g., Guru, Freelancer, Amazon Mechanical Turk), traditional forms of "in-sourcing" work continue being the dominant form of employment. % in most companies. This means that, at least for the time being, freelancing and salaried employment will continue to co-exist. In this paper, we provide algorithms for outsourcing and hiring workers in a general setting, where workers form a team and contribute different skills to perform a task. We call this model team formation with outsourcing. In our model, tasks arrive in an online fashion: neither the number nor the composition of the tasks are known a-priori. At any point in time, there is a team of hired workers who receive a fixed salary independently of the work they perform. This team is dynamic: new members can be hired and existing members can be fired, at some cost. Additionally, some parts of the arriving tasks can be outsourced and thus completed by non-team members, at a premium. Our contribution is an efficient online cost-minimizing algorithm for hiring and firing team members and outsourcing tasks. We present theoretical bounds obtained using a primal--dual scheme proving that our algorithms have logarithmic competitive approximation ratio. We complement these results with experiments using semi-synthetic datasets based on actual task requirements and worker skills from three large online labor marketplaces. Aris Anagnostopoulos, Carlos Castillo 0001, Adriano Fazzone, Stefano Leonardi 0001, Evimaria Terzi |
KDD | 2 |
| 2018 | Ranking of Social Media Alerts with Workload Bounds in Emergency Operation CentersabstractExtensive research on social media usage during emergencies has shown its value to provide life-saving information, if a mechanism is in place to filter and prioritize messages. Existing ranking systems can provide a baseline for selecting which updates or alerts to push to emergency responders. However, prior research has not investigated in depth how many and how often should these updates be generated, considering a given bound on the workload for a user due to the limited budget of attention in this stressful work environment. This paper presents a novel problem and a model to quantify the relationship between the performance metrics of ranking systems (e.g., recall, NDCG) and the bounds on the user workload. We then synthesize an alert-based ranking system that enforces these bounds to avoid overwhelming end-users. We propose a Pareto optimal algorithm for ranking selection that adaptively determines the preference of top-k ranking and user workload over time. We demonstrate the applicability of this approach for Emergency Operation Centers (EOCs) by performing an evaluation based on real world data from six crisis events. We analyze the trade-off between recall and workload recommendation across periodic and realtime settings. Our experiments demonstrate that the proposed ranking selection approach can improve the efficiency of monitoring social media requests while optimizing the need for user attention. Hemant Purohit, Carlos Castillo 0001, Muhammad Imran 0002, Rahul Pandey |
WI | 2 |
| 2018 | The 5th International Workshop on Social Web for Disaster Management(SWDM'18): Collective Sensing, Trust, and Resilience in Global CrisesabstractDuring large-scale emergencies such as natural and man-made disasters, a massive amount of information is posted by the public in social media. Collecting, aggregating, and presenting this information to stakeholders can be extremely challenging, particularly if an understanding of the "big picture»» is sought. This international workshop, the fifth in the series, is a key venue for researchers and practitioners to discuss research challenges and technical issues around the usage of social media in disaster management. Workshop»s website: https://sites.google.com/site/swdm2018/ Yu-Ru Lin, Carlos Castillo 0001, Jie Yin 0001 |
WSDM | 2 |
| 2018 | A Critical Review of Online Social Data: Biases, Methodological Pitfalls, and Ethical BoundariesabstractOnline social data like user-generated content, expressed or implicit relations among people, and behavioral traces are at the core of many popular web applications and platforms, driving the research agenda of researchers in both academia and industry. The promises of social data are many, including the understanding of "what the world thinks»» about a social issue, brand, product, celebrity, or other entity, as well as enabling better decision-making in a variety of fields including public policy, healthcare, and economics. However, many academics and practitioners are increasingly warning against the naive usage of social data. They highlight that there are biases and inaccuracies occurring at the source of the data, but also introduced during data processing pipeline; there are methodological limitations and pitfalls, as well as ethical boundaries and unexpected outcomes that are often overlooked. Such an overlook can lead to wrong or inappropriate results that can be consequential. Alexandra Olteanu, Emre Kiciman, Carlos Castillo 0001 |
WSDM | 3 |
| 2017 | Efficient Document Filtering Using Vector Space Topic Expansion and Pattern-Mining: The Case of Event Detection in MicropostsabstractAutomatically extracting information from social media is challenging given that social content is often noisy, ambiguous, and inconsistent. However, as many stories break on social channels first before being picked up by mainstream media, developing methods to better handle social content is of utmost importance. In this paper, we propose a robust and effective approach to automatically identify microposts related to a specific topic defined by a small sample of reference documents. Our framework extracts clusters of semantically similar microposts that overlap with the reference documents, by extracting combinations of key features that define those clusters through frequent pattern mining. This allows us to construct compact and interpretable representations of the topic, dramatically decreasing the computational burden compared to classical clustering and k-NN-based machine learning techniques and producing highly-competitive results even with small training sets (less than 1'000 training objects). Our method is efficient and scales gracefully with large sets of incoming microposts. We experimentally validate our approach on a large corpus of over 60M microposts, showing that it significantly outperforms state-of-the-art techniques. Julia Proskurnia, Ruslan Mavlyutov, Carlos Castillo 0001, Karl Aberer, Philippe Cudré-Mauroux |
CIKM | 3 |
| 2017 | FA*IR: A Fair Top-k Ranking AlgorithmabstractIn this work, we define and solve the Fair Top-k Ranking problem, in which we want to determine a subset of k candidates from a large pool of n » k candidates, maximizing utility (i.e., select the "best" candidates) subject to group fairness criteria. Meike Zehlike, Francesco Bonchi, Carlos Castillo 0001, Sara Hajian, Mohamed Megahed, Ricardo Baeza-Yates |
CIKM | 3 |
| 2017 | Predicting the Success of Online Petitions Leveraging Multidimensional Time-SeriesabstractApplying classical time-series analysis techniques to online content is challenging, as web data tends to have data quality issues and is often incomplete, noisy, or poorly aligned. In this paper, we tackle the problem of predicting the evolution of a time series of user activity on the web in a manner that is both accurate and interpretable, using related time series to produce a more accurate prediction. We test our methods in the context of predicting signatures for online petitions using data from thousands of petitions posted on The Petition Site - one of the largest platforms of its kind. We observe that the success of these petitions is driven by a number of factors, including promotion through social media channels and on the front page of the petitions platform. We propose an interpretable model that incorporates seasonality, aging effects, self-excitation, and external effects. The interpretability of the model is important for understanding the elements that drives the activity of an online content. We show through an extensive empirical evaluation that our model is significantly better at predicting the outcome of a petition than state-of-the-art techniques. Julia Proskurnia, Przemyslaw A. Grabowicz, Ryota Kobayashi, Carlos Castillo 0001, Philippe Cudré-Mauroux, Karl Aberer |
WWW | 4 |
| 2017 | Story-focused reading in online news and its potential for user engagementabstractWe study the news reading behavior of several hundred thousand users on 65 highly visited news sites. We focus on a specific phenomenon: users reading several articles related to a particular news development, which we call story‐focused reading. Our goal is to understand the effect of story‐focused reading on user engagement and how news sites can support this phenomenon. We found that most users focus on stories that interest them and that even casual news readers engage in story‐focused reading. During story‐focused reading, users spend more time reading and a larger number of news sites are involved. In addition, readers employ different strategies to find articles related to a story. We also analyze how news sites promote story‐focused reading by looking at how they link their articles to related content published by them, or by other sources. The results show that providing links to related content leads to a higher engagement of the users, and that this is the case even for links to external sites. We also show that the performance of links can be affected by their type, their position, and how many of them are present within an article. Janette Lehmann, Carlos Castillo 0001, Mounia Lalmas-Roelleke, Ricardo Baeza-Yates |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | The Fourth International Workshop on Social Web for Disaster Management (SWDM 2016)abstractThe proliferation of social media platforms together with the wide adoption of smartphone devices has transformed how we communicate and share news. During large-scale emergencies, such as natural disasters or armed attacks, victims, responders, and volunteers increasingly use social media to post situation updates and to request and offer help. The use of social media for emergency and disaster response has been a prominent application of information and knowledge management techniques in recent years. There are a number of challenges associated with near real-time processing of vast volumes of information in a way that makes sense for people directly affected, for volunteer organizations, and for official emergency response agencies. As massive amount of messages posted by users are transformed into semi-structured records via information extraction and natural language processing techniques, there is a growing need for developing advanced techniques to aggregate this large-scale data to gain an understanding of the ``big picture'' of an emergency, and to detect and predict how a disaster could develop. This workshop seeks to provide a platform for the exchange of ideas, identification of important problems, and discovery of possible synergies. It will enable interesting discussions and encouraged collaboration between various disciplines, and information and knowledge management approaches is the core of this workshop. Carlos Castillo 0001, Fernando Diaz 0001, Yu-Ru Lin, Jie Yin 0001 |
CIKM | 1 |
| 2016 | A Robust Framework for Classifying Evolving Document Streams in an Expert-Machine-Crowd SettingabstractAn emerging challenge in the online classification of social media data streams is to keep the categories used for classification up-to-date. In this paper, we propose an innovative framework based on an Expert-Machine-Crowd (EMC) triad to help categorize items by continuously identifying novel concepts in heterogeneous data streams often riddled with outliers. We unify constrained clustering and outlier detection by formulating a novel optimization problem: COD-Means. We design an algorithm to solve the COD-Means problem and show that COD-Means will not only help detect novel categories but also seamlessly discover human annotation errors and improve the overall quality of the categorization process. Experiments on diverse real data sets demonstrate that our approach is both effective and efficient. Muhammad Imran 0002, Sanjay Chawla, Carlos Castillo 0001 |
ICDM | 3 |
| 2016 | Algorithmic Bias: From Discrimination Discovery to Fairness-aware Data MiningabstractAlgorithms and decision making based on Big Data have become pervasive in all aspects of our daily lives lives (offline and online), as they have become essential tools in personal finance, health care, hiring, housing, education, and policies. It is therefore of societal and ethical importance to ask whether these algorithms can be discriminative on grounds such as gender, ethnicity, or health status. It turns out that the answer is positive: for instance, recent studies in the context of online advertising show that ads for high-income jobs are presented to men much more often than to women [Datta et al., 2015]; and ads for arrest records are significantly more likely to show up on searches for distinctively black names [Sweeney, 2013]. This algorithmic bias exists even when there is no discrimination intention in the developer of the algorithm. Sometimes it may be inherent to the data sources used (software making decisions based on data can reflect, or even amplify, the results of historical discrimination), but even when the sensitive attributes have been suppressed from the input, a well trained machine learning algorithm may still discriminate on the basis of such sensitive attributes because of correlations existing in the data. These considerations call for the development of data mining systems which are discrimination-conscious by-design. This is a novel and challenging research area for the data mining community. Sara Hajian, Francesco Bonchi, Carlos Castillo 0001 |
KDD | 3 |
| 2016 | Overview of the Special Issue on Trust and Veracity of Information in Social Mediaabstractresearch-article Share on Overview of the Special Issue on Trust and Veracity of Information in Social Media Authors: Symeon Papadopoulos Centre for Research and Technology Hellas; Thessaloniki, Greece Centre for Research and Technology Hellas; Thessaloniki, GreeceView Profile , Kalina Bontcheva University of Sheffield, Sheffield, UK University of Sheffield, Sheffield, UKView Profile , Eva Jaho Athens Technology Center, Athens, Greece Athens Technology Center, Athens, GreeceView Profile , Mihai Lupu Vienna University of Technology, Vienna, Austria Vienna University of Technology, Vienna, AustriaView Profile , Carlos Castillo Sapienza University of Rome, Rome, Italy Sapienza University of Rome, Rome, ItalyView Profile Authors Info & Claims ACM Transactions on Information SystemsVolume 34Issue 3May 2016 Article No.: 14pp 1–5https://doi.org/10.1145/2870630Published:11 April 2016Publication History 20citation1,578DownloadsMetricsTotal Citations20Total Downloads1,578Last 12 Months57Last 6 weeks12 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Symeon Papadopoulos, Kalina Bontcheva, Eva Jaho, Mihai Lupu, Carlos Castillo 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2015 | Comparing Events Coverage in Online News and Social Media: The Case of Climate Change
Alexandra Olteanu, Carlos Castillo 0001, Nicholas Diakopoulos, Karl Aberer |
ICWSM | 2 |
| 2014 | Ranking item features by mining online user-item interactionsabstractWe assume a database of items in which each item is described by a set of attributes, some of which could be multi-valued. We refer to each of the distinct attribute values as a feature. We also assume that we have information about the interactions (such as visits or likes) between a set of users and those items. In our paper, we would like to rank the features of an item using user-item interactions. For instance, if the items are movies, features could be actors, directors or genres, and user-item interaction could be user liking the movie. These information could be used to identify the most important actors for each movie. While users are drawn to an item due to a subset of its features, a user-item interaction only provides an expression of user preference over the entire item, and not its component features. We design algorithms to rank the features of an item depending on whether interaction information is available at aggregated or individual level granularity and extend them to rank composite features (set of features). Our algorithms are based on constrained least squares, network flow and non-trivial adaptations to non-negative matrix factorization. We evaluate our algorithms using both real-world and synthetic datasets. Sofiane Abbar, Habibur Rahman 0001, Saravanan Thirumuruganathan, Carlos Castillo 0001, Gautam Das 0001 |
ICDE | 4 |
| 2014 | CrisisLex: A Lexicon for Collecting and Filtering Microblogged Communications in Crises
Alexandra Olteanu, Carlos Castillo 0001, Fernando Diaz 0001, Sarah Vieweg |
ICWSM | 2 |
| 2014 | Composite Retrieval of Diverse and Complementary BundlesabstractUsers are often faced with the problem of finding complementary items that together achieve a single common goal (e.g., a starter kit for a novice astronomer, a collection of question/answers related to low-carb nutrition, a set of places to visit on holidays). In this paper, we argue that for some application scenarios returning item bundles is more appropriate than ranked lists. Thus we define composite retrieval as the problem of finding$k$bundles of complementary items. Beyond complementarity of items, the bundles must be valid w.r.t. a given budget, and the answer set of$k$bundles must exhibit diversity. We formally define the problem and show that in its general form is${\bf NP}$-hard and that also the special cases in which each bundle is formed by only one item, or only one bundle is sought, are hard. Our characterization however suggests how to adopt a two-phase approach (Produce-and-Choose, or PAC) in which we first produce many valid bundles, and then we choose$k$among them. For the first phase we devise two ad-hoc clustering algorithms, while for the second phase we adapt heuristics with approximation guarantees for a related problem. We also devise another approach which is based on first finding a$k$-clustering and then selecting a valid bundle from each of the produced clusters (Cluster-and-Pick, or CAP). We compare experimentally the proposed methods on two real-world data sets: the first data set is given by a sample of touristic attractions in 10 large European cities, while the second is a large database of user-generated restaurant reviews from Yahoo! Local. Our experiments show that when diversity is highly important, CAP is the best option, while when diversity is less important, a PAC approach constructing bundles around randomly chosen pivots, is better. Sihem Amer-Yahia, Francesco Bonchi, Carlos Castillo 0001, Esteban Feuerstein, Isabel Méndez-Díaz, Paula Zabala |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Social media news communities: gatekeeping, coverage, and statement biasabstractWe examine biases in online news sources and social media communities around them. To that end, we introduce unsupervised methods considering three types of biases: selection or ``gatekeeping'' bias, coverage bias, and statement bias, characterizing each one through a series of metrics. Our results, obtained by analyzing 80 international news sources during a two-week period, show that biases are subtle but observable, and follow geographical boundaries more closely than political ones. We also demonstrate how these biases are to some extent amplified by social media. Diego Sáez-Trumper, Carlos Castillo 0001, Mounia Lalmas-Roelleke |
CIKM | 2 |
| 2013 | Transient News Crowds in Social Media
Janette Lehmann, Carlos Castillo 0001, Mounia Lalmas-Roelleke, Ethan Zuckerman |
ICWSM | 2 |
| 2013 | Measuring and Summarizing Movement in Microblog Postings
Eduardo J. Ruiz, Vagelis Hristidis, Carlos Castillo 0001, Aristides Gionis |
ICWSM | 3 |
| 2013 | The role of information diffusion in the evolution of social networksabstractEvery day millions of users are connected through online social networks, generating a rich trove of data that allows us to study the mechanisms behind human interactions. Triadic closure has been treated as the major mechanism for creating social links: if Alice follows Bob and Bob follows Charlie, Alice will follow Charlie. Here we present an analysis of longitudinal micro-blogging data, revealing a more nuanced view of the strategies employed by users when expanding their social circles. While the network structure affects the spread of information among users, the network is in turn shaped by this communication activity. This suggests a link creation mechanism whereby Alice is more likely to follow Charlie after seeing many messages by Charlie. We characterize users with a set of parameters associated with different link creation strategies, estimated by a Maximum-Likelihood approach. Triadic closure does have a strong effect on link formation, but shortcuts based on traffic are another key factor in interpreting network evolution. However, individual strategies for following other users are highly heterogeneous. Link creation behaviors can be summarized by classifying users in different categories with distinct structural and behavioral characteristics. Users who are popular, active, and influential tend to create traffic-based shortcuts, making the information diffusion process more efficient in the network. Lilian Weng, Jacob Ratkiewicz, Nicola Perra, Bruno Gonçalves, Carlos Castillo 0001, Francesco Bonchi, Rossano Schifanella, Filippo Menczer, Alessandro Flammini |
KDD | 5 |
| 2013 | Online matching of web content to closed captions in IntoNowabstractIntoNow is a mobile application that provides a second-screen experience to television viewers. IntoNow uses the microphone of the companion device to sample the audio coming from the TV set, and compares it against a database of TV shows in order to identify the program being watched. Carlos Castillo 0001, Gianmarco De Francisci Morales, Ajay Shekhawat |
SIGIR | 1 |
| 2013 | Meme ranking to maximize posts virality in microblogging platforms
Francesco Bonchi, Carlos Castillo 0001, Dino Ienco |
J. Intell. Inf. Syst. | 2 |
| 2012 | Mining search behavior and user-generated content: presentation at the industrial session - EDBT/ICDT 2012abstractIn the first part of this presentation, we will overview two systems that gather and display intelligence from search behavior: "Yahoo! Search Clues" and "Yahoo! Political Insights". Carlos Castillo 0001 |
EDBT | 1 |
| 2012 | Correlating financial time series with micro-blogging activityabstractWe study the problem of correlating micro-blogging activity with stock-market events, defined as changes in the price and traded volume of stocks. Specifically, we collect messages related to a number of companies, and we search for correlations between stock-market events for those companies and features extracted from the micro-blogging messages. The features we extract can be categorized in two groups. Features in the first group measure the overall activity in the micro-blogging platform, such as number of posts, number of re-posts, and so on. Features in the second group measure properties of an induced interaction graph, for instance, the number of connected components, statistics on the degree distribution, and other graph-based properties. Eduardo J. Ruiz, Vagelis Hristidis, Carlos Castillo 0001, Aristides Gionis, Alejandro Jaimes |
WSDM | 3 |
| 2012 | Online team formation in social networksabstractWe study the problem of online team formation. We consider a setting in which people possess different skills and compatibility among potential team members is modeled by a social network. A sequence of tasks arrives in an online fashion, and each task requires a specific set of skills. The goal is to form a new team upon arrival of each task, so that (i) each team possesses all skills required by the task, (ii) each team has small communication overhead, and (iii) the workload of performing the tasks is balanced among people in the fairest possible way. Aris Anagnostopoulos, Luca Becchetti, Carlos Castillo 0001, Aristides Gionis, Stefano Leonardi 0001 |
WWW | 3 |
| 2011 | Sparsification of influence networksabstractWe present Spine, an efficient algorithm for finding the "backbone" of an influence network. Given a social graph and a log of past propagations, we build an instance of the independent-cascade model that describes the propagations. We aim at reducing the complexity of that model, while preserving most of its accuracy in describing the data. Michael Mathioudakis, Francesco Bonchi, Carlos Castillo 0001, Aristides Gionis, Antti Ukkonen |
KDD | 3 |
| 2011 | Information credibility on twitterabstractWe analyze the information credibility of news propagated through Twitter, a popular microblogging service. Previous research has shown that most of the messages posted on Twitter are truthful, but the service is also used to spread misinformation and false rumors, often unintentionally. Carlos Castillo 0001, Marcelo Mendoza, Barbara Poblete |
WWW | 1 |
| 2011 | Query reformulation mining: models, patterns, and applications
Paolo Boldi, Francesco Bonchi, Carlos Castillo 0001, Sebastiano Vigna |
Inf. Retr. | 3 |
| 2011 | Social Network Analysis and Mining for Business ApplicationsabstractSocial network analysis has gained significant attention in recent years, largely due to the success of online social networking and media-sharing sites, and the consequent availability of a wealth of social network data. In spite of the growing interest, however, there is little understanding of the potential business applications of mining social networks. While there is a large body of research on different problems and methods for social network mining, there is a gap between the techniques developed by the research community and their deployment in real-world applications. Therefore the potential business impact of these techniques is still largely unexplored. In this article we use a business process classification framework to put the research topics in a business context and provide an overview of what we consider key problems and techniques in social network analysis and mining from the perspective of business applications. In particular, we discuss data acquisition and preparation, trust, expertise, community structure, network dynamics, and information propagation. In each case we present a brief overview of the problem, describe state-of-the art approaches, discuss business application examples, and map each of the topics to a business process classification framework. In addition, we provide insights on prospective business applications, challenges, and future research directions. The main contribution of this article is to provide a state-of-the-art overview of current techniques while providing a critical perspective on business applications of social network analysis and mining. Francesco Bonchi, Carlos Castillo 0001, Aristides Gionis, Alejandro Jaimes |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2010 | Power in unity: forming teams in large-scale community systemsabstractThe internet has enabled the collaboration of groups at a scale that was unseen before. A key problem for large collaboration groups is to be able to allocate tasks effectively. An effective task assignment method should consider both how fit teams are for each job as well as how fair the assignment is to team members, in terms that no one should be overloaded or unfairly singled out. The assignment has to be done automatically or semi-automatically given that it is difficult and time-consuming to keep track of the skills and the workload of each person. Obviously the method to do this assignment must also be computationally efficient. Aris Anagnostopoulos, Luca Becchetti, Carlos Castillo 0001, Aristides Gionis, Stefano Leonardi 0001 |
CIKM | 3 |
| 2010 | Query similarity by projecting the query-flow graphabstractDefining a measure of similarity between queries is an interesting and difficult problem. A reliable query-similarity measure can be used in a variety of applications such as query recommendation, query expansion, and advertising. Ilaria Bordino, Carlos Castillo 0001, Debora Donato, Aristides Gionis |
SIGIR | 2 |
| 2010 | The demographics of web searchabstractHow does the web search behavior of "rich" and "poor" people differ? Do men and women tend to click on difffferent results for the same query? What are some queries almost exclusively issued by African Americans? These are some of the questions we address in this study. Ingmar Weber, Carlos Castillo 0001 |
SIGIR | 2 |
| 2010 | The Effects of Query Bursts on Web SearchabstractA query burst is a period of heightened interest of users on a topic which yields a higher frequency of the search queries related to it. In this paper we examine the behavior of search engine users during a query burst, compared to before and after this period. The purpose of this study is to get insights about how search engines and content providers should respond to a query burst. We analyze one year of web-search logs, looking at query bursts from two perspectives. First, we adopt the user's perspective describing changes in user's effort and interest while searching. Second, we look at the burst from the general content providers' view, answering the question of under which conditions a content provider should ``ride'' a wave of increased interest to obtain a significant share of clicks. Ilija Subasic, Carlos Castillo 0001 |
Web Intelligence | 2 |
| 2010 | An optimization framework for query recommendationabstractQuery recommendation is an integral part of modern search engines. The goal of query recommendation is to facilitate users while searching for information. Query recommendation also allows users to explore concepts related to their information needs. Aris Anagnostopoulos, Luca Becchetti, Carlos Castillo 0001, Aristides Gionis |
WSDM | 3 |
| 2010 | Efficient algorithms for large-scale local triangle countingabstractIn this article, we study the problem of approximate local triangle counting in large graphs. Namely, given a large graph G =( V,E ) we want to estimate as accurately as possible the number of triangles incident to every node v ∈ V in the graph. We consider the question both for undirected and directed graphs. The problem of computing the global number of triangles in a graph has been considered before, but to our knowledge this is the first contribution that addresses the problem of approximate local triangle counting with a focus on the efficiency issues arising in massive graphs and that also considers the directed case. The distribution of the local number of triangles and the related local clustering coefficient can be used in many interesting applications. For example, we show that the measures we compute can help detect the presence of spamming activity in large-scale Web graphs, as well as to provide useful features for content quality assessment in social networks. For computing the local number of triangles (undirected and directed), we propose two approximation algorithms, which are based on the idea of min-wise independent permutations [Broder et al. 1998]. Our algorithms operate in a semi-streaming fashion, using O (| V |) space in main memory and performing O (log | V |) sequential scans over the edges of the graph. The first algorithm we describe in this article also uses O (| E |) space of external memory during computation, while the second algorithm uses only main memory. We present the theoretical analysis as well as experimental results on large graphs, demonstrating the practical efficiency of our approach. Luca Becchetti, Paolo Boldi, Carlos Castillo 0001, Aristides Gionis |
ACM Trans. Knowl. Discov. Data | 3 |
| 2009 | Aging effects on query flow graphs for query suggestionabstractWorld Wide Web content continuously grows in size and importance. Furthermore, users ask Web search engines to satisfy increasingly disparate information needs. New techniques and tools are constantly developed aimed at assisting users in the interaction with the Web search engine. Query recommender systems suggesting interesting queries to users are an example of such tools. Most query recommendation techniques are based on the knowledge of the behaviors of past users of the search engine recorded in query logs. Ranieri Baraglia, Carlos Castillo 0001, Debora Donato, Franco Maria Nardini, Raffaele Perego 0001, Fabrizio Silvestri |
CIKM | 2 |
| 2009 | Voting in social networksabstractA voting system is a set of rules that a community adopts to take collective decisions. In this paper we study voting systems for a particular kind of community: electronically mediated social networks. In particular, we focus on delegative democracy (a.k.a. proxy voting) that has recently received increased interest for its ability to combine the benefits of direct and representative systems, and that seems also perfectly suited for electronically mediated social networks. In such a context, we consider a voting system in which users can only express their preference for one among the people they are explicitly connected with, and this preference can be propagated transitively, using an attenuation factor. We present this system and we study its properties. We also take into consideration the problem of missing votes, which is particularly relevant in online networks, as some recent case shows. Our experiments on real-world networks provide interesting insight into the significance and stability of the results obtained with the suggested voting system. Paolo Boldi, Francesco Bonchi, Carlos Castillo 0001, Sebastiano Vigna |
CIKM | 3 |
| 2009 | Fast shortest path distance estimation in large networksabstractIn this paper we study approximate landmark-based methods for point-to-point distance estimation in very large networks. These methods involve selecting a subset of nodes as landmarks and computing offline the distances from each node in the graph to those landmarks. At runtime, when the distance between a pair of nodes is needed, it can be estimated quickly by combining the precomputed distances. We prove that selecting the optimal set of landmarks is an NP-hard problem, and thus heuristic solutions need to be employed. We therefore explore theoretical insights to devise a variety of simple methods that scale well in very large networks. The efficiency of the suggested techniques is tested experimentally using five real-world graphs having millions of edges. While theoretical bounds support the claim that random landmarks work well in practice, our extensive experimentation shows that smart landmark selection can yield dramatically more accurate results: for a given target accuracy, our methods require as much as 250 times less space than selecting landmarks at random. In addition, we demonstrate that at a very small accuracy loss our techniques are several orders of magnitude faster than the state-of-the-art exact methods. Finally, we study an application of our methods to the task of social search in large graphs. Michalis Potamias, Francesco Bonchi, Carlos Castillo 0001, Aristides Gionis |
CIKM | 3 |
| 2009 | Taxonomy-Driven Lumping for Sequence Mining
Francesco Bonchi, Carlos Castillo 0001, Debora Donato, Aristides Gionis |
ECML/PKDD (1) | 2 |
| 2009 | The Geographical Life of SearchabstractThis article describes a geographical study on the usage of a search engine, focusing on the traffic details at the level of countries and continents. The main objective is to understand from a geographic point of view, how the needs of the users are satisfied, taking into account the geographic location of the host in which the search originates, and the host that contains the Web page that was selected by the user in the answers. Our results confirm that the Web is a cultural mirror of society and shed light on the implicit social network behind search. These results are also useful as input for the design of distributed search engines. Ricardo Baeza-Yates, Christian Middleton, Carlos Castillo 0001 |
Web Intelligence | 3 |
| 2009 | From "Dango" to "Japanese Cakes": Query Reformulation Models and PatternsabstractUnderstanding query reformulation patterns is a key step towards next generation web search engines: it can help improving users' web-search experience by predicting their intent, and thus helping them to locate information more effectively. As a step in this direction, we build an accurate model for classifying user query reformulations into broad classes (generalization, specialization, error correction or parallel move), achieving 92\% accuracy. We apply the model to automatically label two large query logs, creating annotated query-flow graphs. We study the resulting reformulation patterns, finding results consistent with previous studies done on smaller manually annotated datasets, and discovering new interesting patterns, including connections between reformulation types and topical categories. Finally, applying our findings to a third query log that is publicly available for research purposes, we demonstrate that our reformulation classifier leads to improved recommendations in a query recommendation system. Paolo Boldi, Francesco Bonchi, Carlos Castillo 0001, Sebastiano Vigna |
Web Intelligence | 3 |
| 2009 | Taxonomy-driven lumping for sequence mining
Francesco Bonchi, Carlos Castillo 0001, Debora Donato, Aristides Gionis |
Data Min. Knowl. Discov. | 2 |
| 2008 | The query-flow graph: model and applicationsabstractQuery logs record the queries and the actions of the users of engines, and as such they contain valuable information about the interests, the preferences, and the behavior of the users, as well as their implicit feedback to engine results. Mining the wealth of information available in the query logs has many important applications including query-log analysis, user profiling and personalization, advertising, query recommendation, and more.In this paper we introduce the query-flow graph, a graph representation of the interesting knowledge about latent querying behavior. Intuitively, in the query-flow graph a directed edge from query qi to query qj means that the two queries are likely to be part of the same search mission. Any path over the query-flow graph may be seen as a searching behavior, whose likelihood is given by the strength of the edges along the path.The query-flow graph is an outcome of query-log mining and, at the same time, a useful tool for it. We propose a methodology that builds such a graph by mining time and textual information as well as aggregating queries from different users. Using this approach we build a real-world query-flow graph from a large-scale query log and we demonstrate its utility in concrete applications, namely, finding logical sessions, and query recommendation. We believe, however, that the usefulness of the query-flow graph goes beyond these two applications. Paolo Boldi, Francesco Bonchi, Carlos Castillo 0001, Debora Donato, Aristides Gionis, Sebastiano Vigna |
CIKM | 3 |
| 2008 | Dr. Searcher and Mr. Browser: a unified hyperlink-click graphabstractWe introduce a unified graph representation of the Web, which includes both structural and usage information. We model this graph using a simple union of the Web's hyperlink and click graphs. The hyperlink graph expresses link structure among Web pages, while the click graph is a bipartite graph of queries and documents denoting users' searching behavior extracted from a search engine's query log. Barbara Poblete, Carlos Castillo 0001, Aristides Gionis |
CIKM | 2 |
| 2008 | Searching the wikipedia with contextual informationabstractWe propose a framework for searching the Wikipedia with contextual information. Our framework extends the typical keyword search, by considering queries of the type (q,p), where q is a set of terms (as in classical Web search), and p is a source Wikipedia document. The query terms q represent the information that the user is interested in finding, and the document p provides the context of the query. The task is to rank other documents in Wikipedia with respect to their relevance to the query terms q given the context document p. By associating a context to the query terms, the search results of a search initiated in a particular page can be made more relevant. Antti Ukkonen, Carlos Castillo 0001, Debora Donato, Aristides Gionis |
CIKM | 2 |
| 2008 | Efficient semi-streaming algorithms for local triangle counting in massive graphsabstractIn this paper we study the problem of local triangle counting in large graphs. Namely, given a large graph G = (V;E) we want to estimate as accurately as possible the number of triangles incident to every node υ ∈ V in the graph. The problem of computing the global number of triangles in a graph has been considered before, but to our knowledge this is the first paper that addresses the problem of local triangle counting with a focus on the efficiency issues arising in massive graphs. The distribution of the local number of triangles and the related local clustering coefficient can be used in many interesting applications. For example, we show that the measures we compute can help to detect the presence of spamming activity in large-scale Web graphs, as well as to provide useful features to assess content quality in social networks. Luca Becchetti, Paolo Boldi, Carlos Castillo 0001, Aristides Gionis |
KDD | 3 |
| 2008 | Topical query decompositionabstractWe introduce the problem of query decomposition, where we are given a query and a document retrieval system, and we want to produce a small set of queries whose union of resulting documents corresponds approximately to that of the original query. Ideally, these queries should represent coherent, conceptually well-separated topics. Francesco Bonchi, Carlos Castillo 0001, Debora Donato, Aristides Gionis |
KDD | 2 |
| 2008 | Finding high-quality content in social mediaabstractThe quality of user-generated content varies drastically from excellent to abuse and spam. As the availability of such content increases, the task of identifying high-quality content sites based on user contributions --social media sites -- becomes increasingly important. Social media in general exhibit a rich variety of information sources: in addition to the content itself, there is a wide array of non-content information available, such as links between items and explicit quality ratings from members of the community. In this paper we investigate methods for exploiting such community feedback to automatically identify high quality content. As a test case, we focus on Yahoo! Answers, a large community question/answering portal that is particularly rich in the amount and types of content and social interactions available in it. We introduce a general classification framework for combining the evidence from different sources of information, that can be tuned automatically for a given social media type and quality definition. In particular, for the community question/answering domain, we show that our system is able to separate high-quality items from the rest with an accuracy close to that of humans Eugene Agichtein, Carlos Castillo 0001, Debora Donato, Aristides Gionis, Gilad Mishne |
WSDM | 2 |
| 2008 | Fourth international workshop on adversarial information retrieval on the web (AIRWeb 2008)abstractAdversarial IR in general, and search engine spam, in particular, are engaging research topics with a real-world impact for Web users, advertisers and publishers. The AIRWeb workshop will bring researchers and practitioners in these areas together, to present and discuss state-of-the-art techniques as well as real-world experiences. Given the continued growth in search engine spam creation and detection efforts, we expect interest in this AIRWeb to surpass that of the previous three editions of the workshop (held jointly with WWW 2005, SIGIR 2006, and WWW 2007 respectively). Carlos Castillo 0001, Kumar Chellapilla, Dennis Fetterly |
WWW | 1 |
| 2008 | Link analysis for Web spam detectionabstractWe propose link-based techniques for automatic detection of Web spam, a term referring to pages which use deceptive techniques to obtain undeservedly high scores in search engines. The use of Web spam is widespread and difficult to solve, mostly due to the large size of the Web which means that, in practice, many algorithms are infeasible. We perform a statistical analysis of a large collection of Web pages. In particular, we compute statistics of the links in the vicinity of every Web page applying rank propagation and probabilistic counting over the entire Web graph in a scalable way. These statistical features are used to build Web spam classifiers which only consider the link structure of the Web, regardless of page contents. We then present a study of the performance of each of the classifiers alone, as well as their combined performance, by testing them over a large collection of Web link spam. After tenfold cross-validation, our best classifiers have a performance comparable to that of state-of-the-art spam classifiers that use content attributes, but are orthogonal to content-based methods. Luca Becchetti, Carlos Castillo 0001, Debora Donato, Ricardo Baeza-Yates, Stefano Leonardi 0001 |
ACM Trans. Web | 2 |
| 2008 | The Juxtaposed approximate PageRank method for robust PageRank approximation in a peer-to-peer web search networkabstractWe present Juxtaposed approximate PageRank (JXP), a distributed algorithm for computing PageRank-style authority scores of Web pages on a peer-to-peer (P2P) network. Unlike previous algorithms, JXP allows peers to have overlapping content and requires no a priori knowledge of other peers’ content. Our algorithm combines locally computed authority scores with information obtained from other peers by means of random meetings among the peers in the network. This computation is based on a Markov-chain state-lumping technique, and iteratively approximates global authority scores. The algorithm scales with the number of peers in the network and we show that the JXP scores converge to the true PageRank scores that one would obtain with a centralized algorithm. Finally, we show how to deal with misbehaving peers by extending JXP with a reputation model. Josiane Xavier Parreira, Carlos Castillo 0001, Debora Donato, Sebastian Michel 0001, Gerhard Weikum |
VLDB J. | 2 |
| 2007 | Challenges on Distributed Web RetrievalabstractIn the ocean of Web data, Web search engines are the primary way to access content. As the data is on the order of petabytes, current search engines are very large centralized systems based on replicated clusters. Web data, however, is always evolving. The number of Web sites continues to grow rapidly and there are currently more than 20 billion indexed pages. In the near future, centralized systems are likely to become ineffective against such a load, thus suggesting the need of fully distributed search engines. Such engines need to achieve the following goals: high quality answers, fast response time, high query throughput, and scalability. In this paper we survey and organize recent research results, outlining the main challenges of designing a distributed Web retrieval system. Ricardo Baeza-Yates, Carlos Castillo 0001, Flavio Paiva Junqueira, Vassilis Plachouras, Fabrizio Silvestri |
ICDE | 2 |
| 2007 | Know your neighbors: web spam detection using the web topologyabstractWeb spam can significantly deteriorate the quality of search engine results. Thus there is a large incentive for commercial search engines to detect spam pages efficiently and accurately. In this paper we present a spam detection system that combines link-based and content-based features, and uses the topology of the Web graph by exploiting the link dependencies among the Web pages. We find that linked hosts tend to belong to the same class: either both are spam or both are non-spam. We demonstrate three methods of incorporating the Web graph topology into the predictions obtained by our base classifier: (i) clustering the host graph, and assigning the label of all hosts in the cluster by majority vote, (ii) propagating the predicted labels to neighboring hosts, and (iii) using the predicted labels of neighboring hosts as new features and retraining the classifier. The result is an accurate system for detecting Web spam, tested on a large and public dataset, using algorithms that can be applied in practice to large-scale Web data. Carlos Castillo 0001, Debora Donato, Aristides Gionis, Vanessa Murdock 0001, Fabrizio Silvestri |
SIGIR | 1 |
| 2007 | Estimating Number of Citations Using Author Reputation
Carlos Castillo 0001, Debora Donato, Aristides Gionis |
SPIRE | 1 |
| 2006 | Generalizing PageRank: damping functions for link-based ranking algorithmsabstractThis paper introduces a family of link-based ranking algorithms that propagate page importance through links. In these algorithms there is a damping function that decreases with distance, so a direct link implies more endorsement than a link through a long path. PageRank is the most widely known ranking function of this family.The main objective of this paper is to determine whether this family of ranking techniques has some interest per se, and how different choices for the damping function impact on rank quality and on convergence speed. Even though our results suggest that PageRank can be approximated with other simpler forms of rankings that may be computed more efficiently, our focus is of more speculative nature, in that it aims at separating the kernel of PageRank, that is, link-based importance propagation, from the way propagation decays over paths.We focus on three damping functions, having linear, exponential, and hyperbolic decay on the lengths of the paths. The exponential decay corresponds to PageRank, and the other functions are new. Our presentation includes algorithms, analysis, comparisons and experiments that study their behavior under different parameters in real Web graph data.Among other results, we show how to calculate a linear approximation that induces a page ordering that is almost identical to PageRank's using a fixed small number of iterations; comparisons were performed using Kendall's τ on large domain datasets. Ricardo Baeza-Yates, Paolo Boldi, Carlos Castillo 0001 |
SIGIR | 3 |
| 2006 | Temporal Analysis of the WikigraphabstractWikipedia is an online encyclopedia, available in more than 100 languages and comprising over 1 million articles in its English version. If we consider each Wikipedia article as a node and each hyperlink between articles as an arc we have a "Wikigraph", a graph that represents the link structure of Wikipedia. The Wikigraph differs from other Web graphs studied in the literature by the fact that there are explicit timestamps associated with each node's events. This allows us to do a detailed analysis of the Wikipedia evolution over time. In the first part of this study we characterize this evolution in terms of users, editions and articles; in the second part, we depict the temporal evolution of several topological properties of the Wikigraph. The insights obtained from the Wikigraphs can be applied to large Web graphs from which the temporal data is usually not available. Luciana S. Buriol, Carlos Castillo 0001, Debora Donato, Stefano Leonardi 0001, Stefano Millozzi |
Web Intelligence | 2 |
| 2006 | A Memory-Efficient Strategy for Exploring the WebabstractSearch engines rely on Web crawlers to create an index of the Web. Web crawlers explore the Web downloading pages and finding links to new pages to be explored. At any given moment, there are a number of pages waiting to be downloaded in the crawler queue. We study the growth of this queue of pending pages during a crawl of a large subset of the Web. In a normal breadth-first crawler, the queue quickly grows very large. We present a strategy for managing the pending queue that reduces its maximum size by 50% while preserving the coverage and quality of the pages visited. This can be applied to general purpose Web crawlers as well as topic-specific crawling, peer-to-peer search, on-demand Web crawling, and other environments in which memory usage has to be kept to a minimum. Carlos Castillo 0001, Alberto Nelli, Alessandro Panconesi |
Web Intelligence | 1 |
| 2006 | Testing google ionterfaces modified for the blindabstractWe present the results of a research project focus on improving the usability of web search tools for blind users who interact via screen reader and voice synthesizer. In the first stage of our study, we proposed eight specific guidelines for simplifying this interaction with search engines. Next, we evaluated these criteria by applying them to Google UIs, re-implementing the simple search and the result page. Finally, we prepared the environment for a remote test with 12 totally blind users. The results highlight how Google interfaces could be improved in order to simplify interaction for the blind. Patrizia Andronico, Marina Buzzi, Barbara Leporini, Carlos Castillo 0001 |
WWW | 4 |
| 2006 | Relationship between web links and tradeabstractWe report on observations on Web characterization studies that suggest that the amount of Web links among sites under different country-code top-level domains is related to the amount of trade between the corresponding countries. Ricardo Baeza-Yates, Carlos Castillo 0001 |
WWW | 2 |
| 2006 | The distribution of pageRank follows a power-law only for particular values of the damping factorabstractThe empirical distribution of PageRank in a large sample of Web pages does not follow a powerlaw except for particular choices of the damping factor. The tail, comprising 5%10 % of the nodes, always follows a power law, but the distribution for the remaining 90%95 % of pages varies. This was observed in several Web samples having from 1 to 50 million pages with damping factors from 0.1 to 0.9 and the resulting behavior was very similar, specially if we restrict the sample to the main strongly connected component. DoublePareto Model As in [Mitzenmacher 2003], we can fit the powerlaw to the body and the tail of the distribution separately, as in the figure on the right. Given that the minimum value of PageRank is the baseline probability (1-α / N) if we assume the following: (i) The exponent of the body is equal to the exponent of the tail (ii)The distribution is normalized to 1, as in the case of PageRank (iii) N>> 1... we can show that the intersection point of the first figure, 1-F(1/N) does Conjecture not depend on N: 1-F(1/N) ≃ (1-α) θ-1 In a graph with powerlaw exponent θ for the indegree, calculating PageRank with: If we accept the hypothesis in [Pandurangan et al. 2002], that is, the powerlaw exponent for the distribution of the tail of the PageRank values is the same as for the indegree of pages, then: α = 1 θ – 1 In our collection of 1 million nodes with θ=2.2 and α=0.85, the Yields a powerlaw over the entire range of values value predicted for 1F(1/N) is 0.10 and the observed 0.12 In the WebBase collection of 130 million documents with θ=2.07 and α=0.85 the predicted value is 0.13, and the 0.16 Examples: θ=2.1, then α=0.90 yields a powerlaw θ=2.2, then α=0.83 yields a powerlawLognormal + Baseline Model Let X be a random variable distributed according to a lognormal distribution. We propose the following model for the PageRank distribution: X + (1-α)/N To obtain the lognormal parameters, we fit this to the distribution of PageRank with α=0.99. Then using the same parameters, we can fit the PageRank distribution obtained with other damping factors with high precision. Fit with our model: Fit with a powerlaw: Luca Becchetti, Carlos Castillo 0001 |
WWW | 2 |
| 2002 | Web Structure, Dynamics and Page Quality
Ricardo Baeza-Yates, Felipe Saint-Jean, Carlos Castillo 0001 |
SPIRE | 3 |
| 2001 | Relating Web Characteristics with Link Based Web Page RankingabstractIn the last years, several techniques based in link analysis have been proposed and used in search engines to rank Web pages. As links are generated by humans, link based ranking seems to give better results than traditional automatic techniques such as word based ranking. However, no studies have been done about their real impact. In this paper we extend global page ranking techniques to Web site ranking, and do a first experimental analysis of link ranking regarding the structure and dynamics of the Web. Ricardo Baeza-Yates, Carlos Castillo 0001 |
SPIRE | 2 |