Christo Wilson

dblp:79/5135 · DBLP profile ↗
← Back
25ranked-venue papers in the field
1as first author
7since 2021 · last 2026
0000-0002-5268-004XORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 23 (1 first)Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2026 Determinants and Effects of Buy Box Suppression on Amazon
abstract
98% of sales on Amazon.com flow through the Buy Box. However, Amazon sometimes decides not to feature any offers for a product and removes the Buy Box from the product page -- a situation known as Buy Box suppression. Suppression may have severe consequences on individual sellers and might affect competition if it disciplines prices across online marketplaces. This paper studies suppression using a new, high-frequency dataset that tracks 17,754 products on Amazon.com at hourly intervals for eight weeks in 2024, along with corresponding offers from 59,301 unique competitors. We open-source both the dataset and collection code. We find that the primary reason for suppression is changes to the offers for a product on Amazon. Lower competitor prices only modestly increase the probability of suppression. Suppression has substantial impacts on sellers: products immediately fall sharply in search, ad placements completely disappear within 12 hours, and sales ranks are 12% worse after 48 hours. In contrast, we find no evidence of impacts on competitor prices in the short run. Our results highlight how Buy Box suppression acts as an important and underexplored mechanism of platform governance.
Jeffrey L. Gleason, Christo Wilson
WWW3
2026 Inferring Users' Demographics and Sensitive Interests Using the Topics API
Athicha Srivirote, Muhammad Abu Bakar Aziz, Jeffrey L. Gleason, Desheng Hu, Christo Wilson
WWW5
2024 Search Engine Revenue from Navigational and Brand Advertising
abstract
Keyword advertising on general web search engines is a multi-billion dollar business. Keyword advertising turns contentious, however, when businesses target ads against their competitors' brand names---a practice known as "competitive poaching." To stave off poaching, companies defensively bid on ads for their own brand names. Google, in particular, has faced lawsuits and regulatory scrutiny since it altered its policies in 2004 to allow poaching. In this study, we investigate the sources of advertising revenue earned by Google, Bing, and DuckDuckGo by examining ad impressions, clicks, and revenue on navigational and brand searches. Using logs of searches performed by a representative panel of US residents, we estimate that ads on these searches account for 28--36% of Google's search revenue, while Bing earns even more. We also find that the effectiveness of these ads for advertisers varies. We conclude by discussing the implications of our findings for advertisers and regulators.
Jeffrey L. Gleason, Alice Koeninger, Desheng Hu, Jessica Teurn, Yakov Bart, Samsun Knight, Ronald E. Robertson, Christo Wilson
ICWSM8
2024 Market or Markets? Investigating Google Search's Market Shares under Vertical Segmentation
abstract
Is Google Search a monopoly with gatekeeping power? Regulators from the US, UK, and Europe have argued that it is based on the assumption that Google Search dominates the market for horizontal (a.k.a. “general”) web search. Google disputes this, claiming that competition extends to all vertical (a.k.a. “specialized”) search engines, and that under this market definition it does not have monopoly power. In this study we present the first analysis of Google Search’s market share under vertical segmentation of online search. We leverage observational trace data collected from a panel of US residents that includes their web browsing history and copies of the Google Search Engine Result Pages they were shown.We observe that participants’ search sessions begin at Google greater than 50% of the time in 24 out of 30 vertical market segments (which comprise almost all of our participants’ searches). Our results inform the consequential and ongoing debates about the market power of Google Search and the conceptualization of online markets in general.
Desheng Hu, Jeffrey L. Gleason, Muhammad Abu Bakar Aziz, Alice Koeninger, Nikolas Guggenberger, Ronald E. Robertson, Christo Wilson
ICWSM7
2024 Perceptions in Pixels: Analyzing Perceived Gender and Skin Tone in Real-world Image Search Results
abstract
The results returned by image search engines have the power to shape peoples' perceptions about social groups. Existing work on image search engines leverages hand-selected queries for occupations like "doctor" and "engineer" to quantify racial and gender bias in search results. We complement this work by analyzing peoples' real-world image search queries and measuring the distributions of perceived gender, skin tone, and age in their results. We collect 54,070 unique image search queries and analyze 1,481 open-ended people queries (i.e. not queries for named entities) from a representative sample of 643 US residents. For each query, we analyze the top 15 results returned on both Google and Bing Images.
Jeffrey L. Gleason, Avijit Ghosh, Ronald E. Robertson, Christo Wilson
WWW4
2023 Google the Gatekeeper: How Search Components Affect Clicks and Attention
abstract
The contemporary Google Search Engine Results Page (SERP) supplements classic blue hyperlinks with complex components. These components produce tensions between searchers, 3rd-party websites, and Google itself over clicks and attention. In this study, we examine 12 SERP components from two categories: (1) extracted results (e.g., featured-snippets) and (2) Google Services (e.g., shopping-ads) to determine their effect on peoples’ behavior. We measure behavior with two variables: (1) click- through rate (CTR) to Google’s own domains versus 3rd-party domains and (2) time spent on the SERP. We apply causal inference methods to an ecologically valid trace dataset comprising 477,485 SERPs from 1,756 participants. We find that multiple components substantially increase CTR to Google domains, while others decrease CTR and increase time on the SERP. These findings may inform efforts to regulate the design of powerful intermediary platforms like Google.
Jeffrey L. Gleason, Desheng Hu, Ronald E. Robertson, Christo Wilson
ICWSM4
2021 When Fair Ranking Meets Uncertain Inference
abstract
Existing fair ranking systems, especially those designed to be demographically fair, assume that accurate demographic information about individuals is available to the ranking algorithm. In practice, however, this assumption may not hold --- in real-world contexts like ranking job applicants or credit seekers, social and legal barriers may prevent algorithm operators from collecting peoples' demographic information. In these cases, algorithm operators may attempt to infer peoples' demographics and then supply these inferences as inputs to the ranking algorithm.
Avijit Ghosh, Ritam Dutt, Christo Wilson
SIGIR3
2020 Modeling and Measuring Expressed (Dis)belief in (Mis)information
Shan Jiang 0008, Miriam J. Metzger, Andrew J. Flanagin, Christo Wilson
ICWSM4
2019 Bringing the kid back into YouTube kids: detecting inappropriate content on video streaming platforms
abstract
With the advent of child-centric content-sharing platforms, such as YouTube Kids, thousands of children, from all age groups are consuming gigabytes of content on a daily basis. With PBS Kids, Disney Jr. and countless others joining in the fray, this consumption of video data stands to grow further in quantity and diversity. However, it has been observed increasingly that content unsuitable for children often slips through the cracks and lands on such platforms. To investigate this phenomenon in more detail, we collect a first of its kind dataset of inappropriate videos hosted on such children-focused apps and platforms. Alarmingly, our study finds that there is a noticeable percentage of such videos currently being watched by kids with some inappropriate videos having millions of views already. To address this problem, we develop a deep learning architecture that can flag such videos and report them. Our results show that the proposed system can be successfully applied to various types of animations, cartoons and CGI videos to detect any inappropriate content within them.
Rashid Tahir, Mohammad Hammas Saeed, Shiza Ali, Fareed Zaffar, Christo Wilson
ASONAM6
2019 Bias Misperceived: The Role of Partisanship and Misinformation in YouTube Comment Moderation
Shan Jiang 0008, Ronald E. Robertson, Christo Wilson
ICWSM3
2019 Auditing the Partisanship of Google Search Snippets
abstract
The text snippets presented in web search results provide users with a slice of page content that they can quickly scan to help inform their click decisions. However, little is known about how these snippets are generated or how they relate to a user's search query. Motivated by the growing body of evidence suggesting that search engine rankings can influence undecided voters, we conducted an algorithm audit of the political partisanship of Google Search snippets relative to the webpages they are extracted from. To accomplish this, we constructed lexicon of partisan cues to measure partisanship and construct a set of left- and right-leaning search queries. Then, we collected a large dataset of Search Engine Results Pages (SERPs) by running our partisan queries and their autocomplete suggestions on Google Search. After using our lexicon to score the machine-coded partisanship of snippets and webpages, we found that Google Search's snippets generally amplify partisanship, and that this effect is robust across different types of webpages, query topics, and partisan (left- and right-leaning) queries.
Desheng Hu, Shan Jiang 0008, Ronald E. Robertson, Christo Wilson
WWW4
2018 On Ridesharing Competition and Accessibility: Evidence from Uber, Lyft, and Taxi
abstract
Ridesharing services such as Uber and Lyft have become an important part of the Vehicle For Hire (VFH) market, which used to be dominated by taxis. Unfortunately, ridesharing services are not required to share data like taxi services, which has made it challenging to compare the competitive dynamics of these services, or assess their impact on cities. In this paper, we comprehensively compare Uber, Lyft, and taxis with respect to key market features (supply, demand, price, and wait time) in San Francisco and New York City. Based on point pattern statistics, we develop novel statistical techniques to validate our measurement methods. Using spatial lag models, we investigate the accessibility of VFH services, and find that transportation infrastructure and socio-economic features have substantial effects on VFH market features.
Shan Jiang 0008, Alan Mislove, Christo Wilson
WWW4
2018 Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages
abstract
Search engines are a primary means through which people obtain information in today»s connected world. Yet, apart from the search engine companies themselves, little is known about how their algorithms filter, rank, and present the web to users. This question is especially pertinent with respect to political queries, given growing concerns about filter bubbles, and the recent finding that bias or favoritism in search rankings can influence voting behavior. In this study, we conduct a targeted algorithm audit of Google Search using a dynamic set of political queries. We designed a Chrome extension to survey participants and collect the Search Engine Results Pages (SERPs) and autocomplete suggestions that they would have been exposed to while searching our set of political queries during the month after Donald Trump»s Presidential inauguration. Using this data, we found significant differences in the composition and personalization of politically-related SERPs by query type, subjects» characteristics, and date.
Ronald E. Robertson, David Lazer, Christo Wilson
WWW3
2017 Clickstream User Behavior Models
abstract
The next generation of Internet services is driven by users and user-generated content. The complex nature of user behavior makes it highly challenging to manage and secure online services. On one hand, service providers cannot effectively prevent attackers from creating large numbers of fake identities to disseminate unwanted content (e.g., spam). On the other hand, abusive behavior from real users also poses significant threats (e.g., cyberbullying). In this article, we propose clickstream models to characterize user behavior in large online services. By analyzing clickstream traces (i.e., sequences of click events from users), we seek to achieve two goals: (1) detection: to capture distinct user groups for the detection of malicious accounts, and (2) understanding: to extract semantic information from user groups to understand the captured behavior. To achieve these goals, we build two related systems. The first one is a semisupervised system to detect malicious user accounts (Sybils). The core idea is to build a clickstream similarity graph where each node is a user and an edge captures the similarity of two users’ clickstreams. Based on this graph, we propose a coloring scheme to identify groups of malicious accounts without relying on a large labeled dataset. We validate the system using ground-truth clickstream traces of 16,000 real and Sybil users from Renren, a large Chinese social network. The second system is an unsupervised system that aims to capture and understand the fine-grained user behavior. Instead of binary classification (malicious or benign), this model identifies the natural groups of user behavior and automatically extracts features to interpret their semantic meanings. Applying this system to Renren and another online social network, Whisper (100K users), we help service providers identify unexpected user behaviors and even predict users’ future actions. Both systems received positive feedback from our industrial collaborators including Renren, LinkedIn, and Whisper after testing on their internal clickstream data.
Gang Wang 0011, Xinyi Zhang 0003, Shiliang Tang, Christo Wilson, Haitao Zheng 0001, Ben Y. Zhao
ACM Trans. Web4
2016 An Empirical Analysis of Algorithmic Pricing on Amazon Marketplace
abstract
The rise of e-commerce has unlocked practical applications for algorithmic pricing (also called dynamic pricing algorithms), where sellers set prices using computer algorithms. Travel websites and large, well known e-retailers have already adopted algorithmic pricing strategies, but the tools and techniques are now available to small-scale sellers as well.
Alan Mislove, Christo Wilson
WWW3
2016 MapWatch: Detecting and Monitoring International Border Personalization on Online Maps
abstract
Maps have long played a crucial role in enabling people to conceptualize and navigate the world around them. However, maps also encode the world-views of their creators. Disputed international borders are one example of this: governments may mandate that cartographers produce maps that conform to their view of a territorial dispute. Today, online maps maintained by private corporations have become the norm. However, these new maps are still subject to old debates. Companies like Google and Bing resolve these disputes by localizing their maps to meet government requirements and user preferences, i.e., users in different locations are shown maps with different international boundaries. We argue that this non-transparent personalization of maps may exacerbate nationalistic disputes by promoting divergent views of geopolitical realities.
Gary Soeller, Karrie Karahalios, Christian Sandvig, Christo Wilson
WWW4
2015 Uncovering User Interaction Dynamics in Online Social Networks
Zhi Yang 0001, Jilong Xue, Christo Wilson, Ben Y. Zhao, Yafei Dai
ICWSM3
2015 Penny for Your Thoughts: Searching for the 50 Cent Party on Sina Weibo
Christo Wilson
ICWSM3
2014 Of Pins and Tweets: Investigating How Users Behave Across Image- and Text-Based Social Networks
Raphael Ottoni, Diego B. Las Casas, João Paulo Pesce, Wagner Meira Jr., Christo Wilson, Alan Mislove, Virgílio A. F. Almeida
ICWSM5
2014 Uncovering social network Sybils in the wild
abstract
Sybil accounts are fake identities created to unfairly increase the power or resources of a single malicious user. Researchers have long known about the existence of Sybil accounts in online communities such as file-sharing systems, but they have not been able to perform large-scale measurements to detect them or measure their activities. In this article, we describe our efforts to detect, characterize, and understand Sybil account activity in the Renren Online Social Network (OSN). We use ground truth provided by Renren Inc. to build measurement-based Sybil detectors and deploy them on Renren to detect more than 100,000 Sybil accounts. Using our full dataset of 650,000 Sybils, we examine several aspects of Sybil behavior. First, we study their link creation behavior and find that contrary to prior conjecture, Sybils in OSNs do not form tight-knit communities. Next, we examine the fine-grained behaviors of Sybils on Renren using clickstream data. Third, we investigate behind-the-scenes collusion between large groups of Sybils. Our results reveal that Sybils with no explicit social ties still act in concert to launch attacks. Finally, we investigate enhanced techniques to identify stealthy Sybils. In summary, our study advances the understanding of Sybil behavior on OSNs and shows that Sybils can effectively avoid existing community-based Sybil detectors. We hope that our results will foster new research on Sybil detection that is based on novel types of Sybil features.
Zhi Yang 0001, Christo Wilson, Xiao Wang 0018, Tingting Gao, Ben Y. Zhao, Yafei Dai
ACM Trans. Knowl. Discov. Data2
2013 Measuring personalization of web search
abstract
Web search is an integral part of our daily lives. Recently, there has been a trend of personalization in Web search, where different users receive different results for the same search query. The increasing personalization is leading to concerns about Filter Bubble effects, where certain users are simply unable to access information that the search engines' algorithm decides is irrelevant. Despite these concerns, there has been little quantification of the extent of personalization in Web search today, or the user attributes that cause it.
Aniko Hannak, Piotr Sapiezynski, Arash Molavi Kakhki, Balachander Krishnamurthy, David Lazer, Alan Mislove, Christo Wilson
WWW7
2013 Understanding latent interactions in online social networks
abstract
Popular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions , that is, passive actions, such as profile browsing, that cannot be observed by traditional measurement techniques. In this article, we seek a deeper understanding of both active and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 220 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed, publicly viewable visitor logs for each user profile. We capture detailed histories of profile visits over a period of 90 days for users in the Peking University Renren network and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than active events, are nonreciprocal in nature, and that profile popularity is correlated with page views of content rather than with quantity of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior and compare their structural properties, evolution, community structure, and mixing times against those of both active interaction graphs and social graphs.
Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Wenpeng Sha, Peng Huang 0005, Yafei Dai, Ben Y. Zhao
ACM Trans. Web2
2012 Serf and turf: crowdturfing for fun and profit
abstract
Popular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses using crowd-sourcing systems. However, crowd-sourcing systems can also pose a real challenge to existing security mechanisms deployed to protect Internet services. Many of these security techniques rely on the assumption that malicious activity is generated automatically by automated programs. Thus they would perform poorly or be easily bypassed when attacks are generated by real users working in a crowd-sourcing system. Through measurements, we have found surprising evidence showing that not only do malicious crowd-sourcing systems exist, but they are rapidly growing in both user base and total revenue. We describe in this paper a significant effort to study and understand these "crowdturfing" systems in today's Internet. We use detailed crawls to extract data about the size and operational structure of these crowdturfing systems. We analyze details of campaigns offered and performed in these sites, and evaluate their end-to-end effectiveness by running active, benign campaigns of our own. Finally, we study and compare the source of workers on crowdturfing sites in different countries. Our results suggest that campaigns on these systems are highly effective at reaching users, and their continuing growth poses a concrete threat to online communities both in the US and elsewhere.
Gang Wang 0011, Christo Wilson, Xiaohan Zhao, Yibo Zhu 0001, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao
WWW2
2012 Beyond Social Graphs: User Interactions in Online Social Networks and their Implications
abstract
Social networks are popular platforms for interaction, communication, and collaboration between friends. Researchers have recently proposed an emerging class of applications that leverage relationships from social networks to improve security and performance in applications such as email, Web browsing, and overlay routing. While these applications often cite social network connectivity statistics to support their designs, researchers in psychology and sociology have repeatedly cast doubt on the practice of inferring meaningful relationships from social network connections alone. This leads to the question: “Are social links valid indicators of real user interaction? If not, then how can we quantify these factors to form a more accurate model for evaluating socially enhanced applications?” In this article, we address this question through a detailed study of user interactions in the Facebook social network. We propose the use of “interaction graphs” to impart meaning to online social links by quantifying user interactions. We analyze interaction graphs derived from Facebook user traces and show that they exhibit significantly lower levels of the “small-world” properties present in their social graph counterparts. This means that these graphs have fewer “supernodes” with extremely high degree, and overall graph diameter increases significantly as a result. To quantify the impact of our observations, we use both types of graphs to validate several well-known social-based applications that rely on graph properties to infuse new functionality into Internet applications, including Reliable Email (RE), SybilGuard, and the weighted cascade influence maximization algorithm. The results reveal new insights into each of these systems, and confirm our hypothesis that to obtain realistic and accurate results, ongoing research on social network applications studies of social applications should use real indicators of user interactions in lieu of social graphs.
Christo Wilson, Alessandra Sala, Krishna P. N. Puttaswamy, Ben Y. Zhao
ACM Trans. Web1
2010 Measurement-calibrated graph models for social network experiments
abstract
Access to realistic, complex graph datasets is critical to research on social networking systems and applications. Simulations on graph data provide critical evaluation of new systems and applications ranging from community detection to spam filtering and social web search. Due to the high time and resource costs of gathering real graph datasets through direct measurements, researchers are anonymizing and sharing a small number of valuable datasets with the community. However, performing experiments using shared real datasets faces three key disadvantages: concerns that graphs can be de-anonymized to reveal private information, increasing costs of distributing large datasets, and that a small number of available social graphs limits the statistical confidence in the results.
Alessandra Sala, Lili Cao, Christo Wilson, Robert Zablit, Haitao Zheng 0001, Ben Y. Zhao
WWW3