EDBT 2026 Demo / reviewers in the wild / expert
Ke Zhou 0003
dblp:78/2949-3
· DBLP profile ↗
42ranked-venue papers in the field
12as first author
9since 2021 · last 2025
0000-0001-7177-9152ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 39 (11 first)Data Mining & Knowledge Discovery · 2 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | C3AI: Crafting and Evaluating Constitutions for Constitutional AIabstractConstitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge.We introduce the C3AI framework (Crafting Constitutions for CAI models), which serves two key functions: (1) selecting and structuring principles to form effective constitutions before fine-tuning; and (2) evaluating whether finetuned CAI models follow these principles in practice.By analyzing principles from AI and psychology, we found that positively framed, behavior-based principles align more closely with human preferences than negatively framed or trait-based principles.In a safety alignment use case, we applied a graph-based principle selection method to refine an existing CAI constitution, improving safety measures while maintaining strong general reasoning capabilities.Interestingly, fine-tuned CAI models performed well on negatively framed principles but struggled with positively framed ones, in contrast to our human alignment results.This highlights a potential gap between principle design and model adherence.Overall, C3AI provides a structured and scalable approach to both crafting and evaluating CAI constitutions. CCS Concepts• Yara Kyrychenko, Ke Zhou 0003, Edyta Paulina Bogucka, Daniele Quercia |
WWW | 2 |
| 2024 | Characterizing Fake News Targeting CorporationsabstractMisinformation proliferates in the online sphere, with evident impacts on the political and social realms, influencing democratic discourse and posing risks to public health and safety. The corporate world is also a prime target for fake news dissemination. While recent studies have attempted to characterize corporate misinformation and its effects on companies, their findings often suffer from limitations due to qualitative or narrative approaches and a narrow focus on specific industries. To address this gap, we conducted an analysis utilizing social media quantitative methods and crowd-sourcing studies to investigate corporate misinformation across a diverse array of industries within the S&P 500 companies. Our study reveals that corporate misinformation encompasses topics such as products, politics, and societal issues. We discovered companies affected by fake news also get reputable news coverage but less social media attention, leading to heightened negativity in social media comments, diminished stock growth, and increased stress mentions among employee reviews. Additionally, we observe that a company is not targeted by fake news all the time, but there are particular times when a critical mass of fake news emerges. These findings hold significant implications for regulators, business leaders, and investors, emphasizing the necessity to vigilantly monitor the escalating phenomenon of corporate misinformation. Ke Zhou 0003, Sanja Scepanovic, Daniele Quercia |
ICWSM | 1 |
| 2024 | Exploratory Analysis of Recommending Urban Parks for Health-Promoting ActivitiesabstractParks are essential spaces for promoting urban health, and recommender systems could assist individuals in discovering parks for leisure and health-promoting activities. This is particularly important in large cities like London, which has over 1,500 named parks, making it challenging to understand what each park offers. Due to the lack of datasets and the diverse health-promoting activities parks can support (e.g., physical, social, nature-appreciation), it is unclear which recommendation algorithms are best suited for this task. To explore the dynamics of recommending parks for specific activities, we created two datasets: one from a survey of over 250 London residents, and another by inferring visits from over 1 million geotagged Flickr images taken in London parks. Analyzing the geographic patterns of these visits revealed that recommending nearby parks is ineffective, suggesting that this recommendation task is distinct from Point of Interest recommendation. We then tested various recommendation models, identifying a significant popularity bias in the results. Additionally, we found that personalized models have advantages in recommending parks beyond the most popular ones. The data and findings from this study provide a foundation for future research on park recommendations. Linus W. Dietz, Sanja Scepanovic, Ke Zhou 0003, Daniele Quercia |
RecSys | 3 |
| 2023 | How Circadian Rhythms Extracted from Social Media Relate to Physical Activity and SleepabstractCircadian rhythm has been linked to both physical and mental health at an individual level in prior research. Such a link at population level has been long hypothesized but has never been tested, largely because of lack of data. To partly fix this literature gap, we need: a dataset on population-level circadian rhythms, a dataset on population-level health conditions, and strong associations between these two partly independent sets. Recent work has shown that affect on social media data relates to population-level circadian rhythms. Building upon that work, we extracted five circadian rhythm metrics from 6M Reddit posts across 18 major cities (for which the number of residents is highly correlated with the number of users), and paired them with three ground-truth health metrics (daily number of steps, sleep quantity, and sleep quality) extracted from 233K wearable users in these cities. We found that rhythms of online activity approximated sleeping patterns rather than, what the literature previously hypothesized, alertness levels. Despite that, we found that these rhythms, when computed in two specific times of the day (i.e., late at night and early morning), were still predictive of the three ground-truth health metrics: in general, healthier cities had morning spikes on social media, night dips, and expressions of positive affect. These results suggest that circadian rhythms on social media, if taken at two specific times of the day and operationalized with literature-driven metrics, can approximate the temporal evolution of people's shared underlying biological rhythm as it relates to physical activity (R2=0.492), sleep quantity (R2=0.765), and sleep quality (R2=0.624). Ke Zhou 0003, Marios Constantinides, Daniele Quercia, Sanja Scepanovic |
ICWSM | 1 |
| 2023 | Characterization and Prediction of Mobile TasksabstractMobile devices have become an increasingly ubiquitous part of our everyday life. We use mobile services to perform a broad range of tasks (e.g., booking travel or conducting remote office work), leading to often lengthy interactions with several distinct apps and services. Existing mobile systems handle mostly simple user needs, where a single app is taken as the unit of interaction. To understand users’ expectations and to provide context-aware services, it is important to model users’ interactions with their performed task in mind. To provide a comprehensive picture of common mobile tasks, we first conduct a small-scale user study to understand annotated mobile tasks in-depth, while we demonstrate that by using a set of features (temporal, similarity, and log sequence), we can identify if a pair of app usage belong to the same task effectively. Secondly, the proposed best task detection model is applied to a large-scale data set of commercial mobile app usage logs to infer characteristics of complex (multi-app) mobile tasks in the wild. By applying an unsupervised learning framework, we discover common mobile task types that span multiple apps based on various extracted characteristics. We observe that users generally perform 17 common tasks with 47 sub-tasks, ranging from “social media browsing” to “dining out” and “family entertainments”. Finally, we demonstrate that we can predict the next complex mobile task that users are likely to perform by leveraging features from the historically inferred mobile tasks and user contexts. Our work facilitates an in-depth understanding of mobile tasks at scale, enabling applications for promoting task-aware services. Yuan Tian 0022, Ke Zhou 0003, Dan Pelleg |
ACM Trans. Inf. Syst. | 2 |
| 2022 | What and How long: Prediction of Mobile App EngagementabstractUser engagement is crucial to the long-term success of a mobile app. Several metrics, such as dwell time, have been used for measuring user engagement. However, how to effectively predict user engagement in the context of mobile apps is still an open research question. For example, do the mobile usage contexts (e.g., time of day) in which users access mobile apps impact their dwell time? Answers to such questions could help mobile operating system and publishers to optimize advertising and service placement. In this article, we first conduct an empirical study for assessing how user characteristics, temporal features, and the short/long-term contexts contribute to gains in predicting users’ app dwell time on the population level. The comprehensive analysis is conducted on large app usage logs collected through a mobile advertising company. The dataset covers more than 12K anonymous users and 1.3 million log events. Based on the analysis, we further investigate a novel mobile app engagement prediction problem—can we predict simultaneously what app the user will use next and how long he/she will stay on that app? We propose several strategies for this joint prediction problem and demonstrate that our model can improve the performance significantly when compared with the state-of-the-art baselines. Our work can help mobile system developers in designing a better and more engagement-aware mobile app user experience. Yuan Tian 0022, Ke Zhou 0003, Dan Pelleg |
ACM Trans. Inf. Syst. | 2 |
| 2021 | POSSCORE: A Simple Yet Effective Evaluation of Conversational Search with Part of Speech LabellingabstractConversational search systems, such as Google Assistant and Microsoft Cortana, provide a new search paradigm where users are allowed, via natural language dialogues, to communicate with search systems. Evaluating such systems is very challenging since search results are presented in the format of natural language sentences. Given the unlimited number of possible responses, collecting relevance assessments for all the possible responses is infeasible. In this paper, we propose POSSCORE, a simple yet effective automatic evaluation method for conversational search. The proposed embedding-based metric takes the influence of part of speech (POS) of the terms in the response into account. To the best knowledge, our work is the first to systematically demonstrate the importance of incorporating syntactic information, such as POS labels, for conversational search evaluation. Experimental results demonstrate that our metrics can correlate with human preference, achieving significant improvements over state-of-the-art baseline metrics. Zeyang Liu 0004, Ke Zhou 0003, Jiaxin Mao, Max L. Wilson 0001 |
CIKM | 2 |
| 2021 | The Healthy States of America: Creating a Health Taxonomy with Social Media
Sanja Scepanovic, Luca Maria Aiello, Ke Zhou 0003, Sagar Joglekar 0001, Daniele Quercia |
ICWSM | 3 |
| 2021 | Meta-evaluation of Conversational Search Evaluation MetricsabstractConversational search systems, such as Google assistant and Microsoft Cortana, enable users to interact with search systems in multiple rounds through natural language dialogues. Evaluating such systems is very challenging, given that any natural language responses could be generated, and users commonly interact for multiple semantically coherent rounds to accomplish a search task. Although prior studies proposed many evaluation metrics, the extent of how those measures effectively capture user preference remain to be investigated. In this article, we systematically meta-evaluate a variety of conversational search metrics. We specifically study three perspectives on those metrics: (1) reliability : the ability to detect “actual” performance differences as opposed to those observed by chance; (2) fidelity : the ability to agree with ultimate user preference; and (3) intuitiveness : the ability to capture any property deemed important: adequacy, informativeness, and fluency in the context of conversational search. By conducting experiments on two test collections, we find that the performance of different metrics vary significantly across different scenarios, whereas consistent with prior studies, existing metrics only achieve weak correlation with ultimate user preference and satisfaction. METEOR is, comparatively speaking, the best existing single-turn metric considering all three perspectives. We also demonstrate that adapted session-based evaluation metrics can be used to measure multi-turn conversational search, achieving moderate concordance with user satisfaction. To our knowledge, our work establishes the most comprehensive meta-evaluation for conversational search to date. Zeyang Liu 0004, Ke Zhou 0003, Max L. Wilson 0001 |
ACM Trans. Inf. Syst. | 2 |
| 2020 | Identifying Tasks from Mobile App Usage PatternsabstractMobile devices have become an increasingly ubiquitous part of our everyday life. We use mobile services to perform a broad range of tasks (e.g. booking travel or office work), leading to often lengthy interactions within distinct apps and services. Existing mobile systems handle mostly simple user needs, where a single app is taken as the unit of interaction. To understand users' expectations and to provide context-aware services, it is important to model users' interactions in the task space. In this work, we first propose and evaluate a method for the automated segmentation of users' app usage logs into task units. We focus on two problems: (i) given a sequential pair of app usage logs, identify if there exists a task boundary, and (ii) given any pair of two app usage logs, identify if they belong to the same task. We model these as classification problems that use features from three aspects of app usage patterns: temporal, similarity, and log sequence. Our classifiers improve on traditional timeout segmentation, achieving over 89% performance for both problems. Secondly, we use our best task classifier on a large-scale data set of commercial mobile app usage logs to identify common tasks. We observe that users' performed common tasks ranging from regular information checking to entertainment and booking dinner. Our proposed task identification approach provides the means to evaluate mobile services and applications with respect to task completion. Yuan Tian 0022, Ke Zhou 0003, Mounia Lalmas-Roelleke, Dan Pelleg |
SIGIR | 2 |
| 2019 | A Rank-biased Neural Network Model for Click ModelingabstractQuery logs contain rich feedback information from a large number of users interacting with search engines. Various click models have been developed to decode users' search behavior and to extract useful knowledge from query logs. Although the state-of-the-art neural click models have been shown to be very effective in click modeling, the input representations of queries and documents rely on either manually crafted features or on automatic methods suffering from the high-dimensionality issue. Moreover, these neural click models are still rather restrictive when coping with commonly biased user clicks. In this paper, we investigate how to effectively deploy a neural network model for decoding users' click behavior. First, we present two novel rank-biased neural network models ($RBNN$ and $RBNN^* $) for click modeling. The key idea is to deploy different weight matrices across different rank positions. Second, we introduce a new method ($QD\mymathhyphen DCCA$) for automatically learning the vector representations for both queries and documents within the same low-dimensional space, which provides high-quality inputs for $RBNN$ and $RBNN^* $. Finally, a series of experiments are conducted on two different real query logs to validate the effectiveness and efficiency of the proposed neural click models. The experiments demonstrate that: (1) The proposed models can achieve substantially improved performance over the state-of-the-art baseline on two datasets across multiple metrics. By incorporating rank-specific weight matrices, $RBNN$ and $RBNN^* $ are more capable of dealing with the position-bias problem. (2) The input representations of queries, documents and context information significantly affect the performance of neural click models. Thanks to the application of $QD\mymathhyphen DCCA$, not only $RBNN$ and $RBNN^* $ but also the baseline method exhibit enhanced performance. Furthermore, the training cost under the proposed models is greatly reduced. Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Joemon M. Jose, Ke Zhou 0003 |
CHIIR | 5 |
| 2019 | Enhancing Conversational Dialogue Models with Grounded KnowledgeabstractLeveraging external knowledge to enhance conversational models has become a popular research area in recent years. Compared to vanilla generative models, the knowledge-grounded models may produce more informative and engaging responses. Although various approaches have been proposed in the past, how to effectively incorporate knowledge remains an open research question. It is unclear how much external knowledge should be retrieved and what is the optimal way to enhance the conversational model, trading off between relevant information and noise. Therefore, in this paper, we aim to bridge the gap by first extensively evaluating various types of state-of-the-art knowledge-grounded conversational models, including recurrent neural network based, memory networks based, and Transformer based models. We demonstrate empirically that those conversational models can only be enhanced with the right amount of external knowledge. To effectively leverage information originated from external knowledge, we propose a novel Transformer with Expanded Decoder (Transformer-ED or TED for short), which can automatically tune the weights for different sources of evidence when generating responses. Our experiments show that our proposed model outperforms state-of-the-art models in terms of both quality and diversity. Ke Zhou 0003 |
CIKM | 2 |
| 2019 | Predicting Outcomes of Active Sessions Using Multi-action MotifsabstractWeb sites and online services increasingly engage with users through live chats to provide support, advice, and offers. Such approaches require reliable methods to predict the user’s intent and make an informed decision when and how to intervene during an active session. Prior work on predicting purchase intent involved clickstream data mining and feature construction in an ad-hoc manner with a moderate success (AUC 0.70 range). We demonstrate the use of the consumer Purchase Decision Model (PDM) and a principled way of constructing features predictive of the purchase intent. We show that the Logistic Regression (LR) classifiers, trained with multi-action motifs, perform on par with the state-of-the-art LSTM sequence model achieving comparable AUC (0.95 vs 0.96) and performing better for the sparse purchase sessions, with higher recall (0.85 vs 0.61) and higher F1 score (0.73 vs 0.66). While LSTM performs better than LR in terms of weighted averages of F1, precision, and recall, it requires 7 times longer to train and offers no insights about the predictive model in terms of the user actions and the purchase decision stages. The LR predictors are robust and effective in simulating real-time interventions, achieving F1 of 0.84 and AUC of 0.85 after observing only 50% of an active session. For non-purchase sessions that leaves room for live intervention, on average within 8 actions before the session ends. Weiqiang Lin, Natasa Milic-Frayling, Ke Zhou 0003, Eugene Ch'ng |
WI | 3 |
| 2019 | Does Diversity Affect User Satisfaction in Image SearchabstractDiversity has been taken into consideration by existing Web image search engines in ranking search results. However, there is no thorough investigation of how diversity affects user satisfaction in image search. In this article, we address the following questions: (1) How do different factors, such as content and visual presentations, affect users’ perception of diversity? (2) How does search result diversity affect user satisfaction with different search intents? To answer those questions, we conduct a set of laboratory user studies to collect users’ perceived diversity annotations and search satisfaction. We find that the existence of nearly duplicated image results has the largest impact on users’ perceived diversity, followed by the similarity in content and visual presentations. Besides these findings, we also investigate the relationship between diversity and satisfaction in image search. Specifically, we find that users’ preference for diversity varies across different search intents. When users want to collect information or save images for further usage (the Locate search tasks), more diversified result lists lead to higher satisfaction levels. The insights may help commercial image search engines to design better result ranking strategies and evaluation metrics. Zhijing Wu 0001, Ke Zhou 0003, Yiqun Liu 0001, Min Zhang 0006, Shaoping Ma |
ACM Trans. Inf. Syst. | 2 |
| 2018 | How Well do Offline and Online Evaluation Metrics Measure User Satisfaction in Web Image Search?abstractComparing to general Web search engines, image search engines present search results differently, with two-dimensional visual image panel for users to scroll and browse quickly. These differences in result presentation can significantly impact the way that users interact with search engines, and therefore affect existing methods of search evaluation. Although different evaluation metrics have been thoroughly studied in the general Web search environment, how those offline and online metrics reflect user satisfaction in the context of image search is an open question. To shed light on this, we conduct a laboratory user study that collects both explicit user satisfaction feedbacks as well as user behavior signals such as clicks. Based on the combination of both externally assessed topical relevance and image quality judgments, offline image search metrics can be better correlated with user satisfaction than merely using topical relevance. We also demonstrate that existing offline Web search metrics can be adapted to evaluate on a two-dimensional presentation for image search. With respect to online metrics, we find that those based on image click information significantly outperform offline metrics. To our knowledge, our work is the first to thoroughly establish the relationship between different measures and user satisfaction in image search. Fan Zhang 0053, Ke Zhou 0003, Yunqiu Shao, Cheng Luo 0001, Min Zhang 0006, Shaoping Ma |
SIGIR | 2 |
| 2017 | Meta-evaluation of Online and Offline Web Search Evaluation MetricsabstractAs in most information retrieval (IR) studies, evaluation plays an essential part in Web search research. Both offline and online evaluation metrics are adopted in measuring the performance of search engines. Offline metrics are usually based on relevance judgments of query-document pairs from assessors while online metrics exploit the user behavior data, such as clicks, collected from search engines to compare search algorithms. Although both types of IR evaluation metrics have achieved success, to what extent can they predict user satisfaction still remains under-investigated. To shed light on this research question, we meta-evaluate a series of existing online and offline metrics to study how well they infer actual search user satisfaction in different search scenarios. We find that both types of evaluation metrics significantly correlate with user satisfaction while they reflect satisfaction from different perspectives for different search tasks. Offline metrics better align with user satisfaction in homogeneous search (i.e. ten blue links) whereas online metrics outperform when vertical results are federated. Finally, we also propose to incorporate mouse hover information into existing online evaluation metrics, and empirically show that they better align with search user satisfaction than click-based online metrics. Ke Zhou 0003, Yiqun Liu 0001, Min Zhang 0006, Shaoping Ma |
SIGIR | 2 |
| 2017 | Does Document Relevance Affect the Searcher's Perception of Time?abstractTime plays an essential role in multiple areas of Information Retrieval (IR) studies such as search evaluation, user behavior analysis, temporal search result ranking and query understanding. Especially, in search evaluation studies, time is usually adopted as a measure to quantify users' efforts in search processes. Psychological studies have reported that the time perception of human beings can be affected by many stimuli, such as attention and motivation, which are closely related to many cognitive factors in search. Considering the fact that users' search experiences are affected by their subjective feelings of time, rather than the objective time measured by timing devices, it is necessary to look into the different factors that have impacts on search users' perception of time. In this work, we make a first step towards revealing the time perception mechanism of search users with the following contributions: (1) We establish an experimental research framework to measure the subjective perception of time while reading documents in search scenario, which originates from but is also different from traditional time perception measurements in psychological studies. (2) With the framework, we show that while users are reading result documents, document relevance has small yet visible effect on search users' perception of time. By further examining the impact of other factors, we demonstrate that the effect on relevant documents can also be influenced by individuals and tasks. (3) We conduct a preliminary experiment in which the difference between perceived time and dwell time is taken into consideration in a search evaluation task. We found that the revised framework achieved a better correlation with users' satisfaction feedbacks. This work may help us better understand the time perception mechanism of search users and provide insights in how to better incorporate time factor in search evaluation studies. Cheng Luo 0001, Yiqun Liu 0001, Tetsuya Sakai, Ke Zhou 0003, Fan Zhang 0053, Shaoping Ma |
WSDM | 4 |
| 2017 | Detecting Collusive Spamming Activities in Community Question AnsweringabstractCommunity Question Answering (CQA) portals provide rich sources of information on a variety of topics. However, the authenticity and quality of questions and answers (Q&As) has proven hard to control. In a troubling direction, the widespread growth of crowdsourcing websites has created a large-scale, potentially difficult-to-detect workforce to manipulate malicious contents in CQA. The crowd workers who join the same crowdsourcing task about promotion campaigns in CQA collusively manipulate deceptive Q&As for promoting a target (product or service). The collusive spamming group can fully control the sentiment of the target. How to utilize the structure and the attributes for detecting manipulated Q&As? How to detect the collusive group and leverage the group information for the detection task? Yuli Liu, Yiqun Liu 0001, Ke Zhou 0003, Min Zhang 0006, Shaoping Ma |
WWW | 3 |
| 2016 | Playing Your Cards Right: The Effect of Entity Cards on Search Behaviour and WorkloadabstractIn addition to merging results of different types (e.g.~images, videos, news items) into a ranked list of Web documents, modern search engines have also started displaying entity cards (ECs) on the results page. Entity cards are intended to enhance search experience in several ways: (i) they help searchers navigate diversified results, (ii) provide a summary of relevant content directly on the results page and (iii) support exploratory search by highlighting relevant entities associated with a given user query. We conducted a large-scale crowd-sourced user study, with more than $700$ unique searchers, to investigate the effects of entity cards on search behaviour and perceived workload. We find that the presence of ECs has a strong effect on both the way users interact with search results and their perceived task workload. Furthermore, by manipulating EC properties content, coherence and vertical diversity), we uncover different effects and interactions between card properties on measures of search behaviour and workload. Our study contributes an in-depth analysis of the effects of entity cards on user interaction with modern Web search interfaces. Horatiu S. Bota, Ke Zhou 0003, Joemon M. Jose |
CHIIR | 2 |
| 2016 | Detecting Promotion Campaigns in Query Auto CompletionabstractQuery Auto Completion (QAC) aims to provide possible suggestions to Web search users from the moment they start entering a query, which is thought to reduce their physical and cognitive efforts in query formulation. However, the QAC has been misused by malicious users, being transformed into a new form of promotion campaign. These malicious users attack the search engines to replace legitimate auto-completion candidate suggestions with manipulated contents. Through this way, they provide a new malicious advertising service to promote their customers' products or services in QAC. To our best knowledge, we are among the first to investigate this new type of Promotion Campaign in QAC (PCQ). Firstly, we look into the causes of PCQ based on practical commercial search query logs. We found that various queries containing certain promotion intents are submitted multiple times to search engines to promote their rankings in QAC. Secondly, an effective promotion query detection framework is proposed by promotion intent propagation on query-user bipartite graph, which takes into account the behavioral characteristics of promotion campaigns. Finally, we extend the query detection framework to promotion target detection to identify the consistent promotion target which is the inherent goal of the promotion campaign. Large-scale manual annotations on practical data set convey both the effectiveness of our proposed algorithm, and an in-depth understanding of PCQ. Yuli Liu, Yiqun Liu 0001, Ke Zhou 0003, Min Zhang 0006, Shaoping Ma, Hengliang Luo |
CIKM | 3 |
| 2016 | Predicting Search User Examination with Visual SaliencyabstractPredicting users' examination of search results is one of the key concerns in Web search related studies. With more and more heterogeneous components federated into search engine result pages (SERPs), it becomes difficult for traditional position-based models to accurately predict users' actual examination patterns. Therefore, a number of prior works investigate the connection between examination and users' explicit interaction behaviors (e.g.~click-through, mouse movement). Although these works gain much success in predicting users' examination behavior on SERPs, they require the collection of large scale user behavior data, which makes it impossible to predict examination behavior on newly-generated SERPs. To predict user examination on SERPs containing heterogenous components without user interaction information, we propose a new prediction model based on visual saliency map and page content features. Visual saliency, which is designed to measure the likelihood of a given area to attract human visual attention, is used to predict users' attention distribution on heterogenous search components. With an experimental search engine, we carefully design a user study in which users' examination behavior (eye movement) is recorded. Examination prediction results based on this collected data set demonstrate that visual saliency features significantly improve the performance of examination model in heterogeneous search environments. We also found that saliency features help predict internal examination behavior within vertical results. Yiqun Liu 0001, Zeyang Liu 0004, Ke Zhou 0003, Meng Wang 0001, Huan-Bo Luan, Chao Wang 0049, Min Zhang 0006, Shaoping Ma |
SIGIR | 3 |
| 2016 | When does Relevance Mean Usefulness and User Satisfaction in Web Search?abstractRelevance is a fundamental concept in information retrieval (IR) studies. It is however often observed that relevance as annotated by secondary assessors may not necessarily mean usefulness and satisfaction perceived by users. In this study, we confirm the difference by a laboratory study in which we collect relevance annotations by external assessors, usefulness and user satisfaction information by users, for a set of search tasks. We also find that a measure based on usefulness rather than relevance annotated has a better correlation with user satisfaction. However, we show that external assessors are capable of annotating usefulness when provided with more search context information. In addition, we also show that it is possible to generate automatically usefulness labels when some training data is available. Our findings explain why traditional system-centric evaluation metrics are not well aligned with user satisfaction and suggest that a usefulness-based evaluation method can be defined to better reflect the quality of search systems perceived by the users. Jiaxin Mao, Yiqun Liu 0001, Ke Zhou 0003, Jian-Yun Nie, Jingtao Song, Min Zhang 0006, Shaoping Ma, Jiashen Sun, Hengliang Luo |
SIGIR | 3 |
| 2016 | HIA 2016: The 2nd International Workshop on Heterogeneous Information Access at SIGIR 2016abstractInformation access is becoming increasingly heterogeneous. Especially when the user's information need is for exploratory purpose, returning a set of diverse results from different resources could benefit the user. For example, when a user is planning a trip to China, retrieving and showing results from vertical search engines like travel, flight information, map and Q2A sites can satisfy the user's rich and diverse information need. This heterogeneous search paradigm is useful in many contexts and brings many new challenges. Ke Zhou 0003, Yiqun Liu 0001, Roger Jie Luo, Joemon M. Jose |
SIGIR | 1 |
| 2016 | Predicting Pre-click Quality for Native AdvertisementsabstractNative advertising is a specific form of online advertising where ads replicate the look-and-feel of their serving platform. In such context, providing a good user experience with the served ads is crucial to ensure long-term user engagement. In this work, we explore the notion of ad quality, namely the effectiveness of advertising from a user experience perspective. We design a learning framework to predict the pre-click quality of native ads. More specifically, we look at detecting offensive native ads, showing that, to quantify ad quality, ad offensive user feedback rates are more reliable than the commonly used click-through rate metrics. We then conduct a crowd-sourcing study to identify which criteria drive user preferences in native advertising. We translate these criteria into a set of ad quality features that we extract from the ad text, image and advertiser, and then use them to train a model able to identify offensive ads. We show that our model is very effective in detecting offensive ads, and provide in-depth insights on how different features affect ad quality. Finally, we deploy a preliminary version of such model and show its effectiveness in the reduction of the offensive ad feedback rate. Ke Zhou 0003, Miriam Redi, Andrew Haines, Mounia Lalmas-Roelleke |
WWW | 1 |
| 2015 | Does Vertical Bring more Satisfaction?: Predicting Search Satisfaction in a Heterogeneous EnvironmentabstractThe study of search satisfaction is one of the prime concerns in search performance evaluation research. Most existing works on search satisfaction primarily rely on the hypothesis that all results on search engine result pages (SERPs) are homogeneous. However, a variety of heterogeneous vertical results such as videos, images and instant answers are aggregated into SERPs by search engines to improve the diversity and quality of search results. In this paper, we carry out a lab-based user study with specifically designed SERPs to determine how verticals with different qualities and presentation styles affect search satisfaction. Users' satisfaction feedback and external assessors' satisfaction annotations are both collected to make a comparison regarding the perception of search satisfaction. Mouse click-through / movement data and eye movement information are also collected such that we can investigate the influence of vertical results from the perspectives of both benefit and cost. Finally, a vertical-aware learning-based prediction method is proposed to predict search satisfaction on aggregated SERPs. To the best of our knowledge, this paper is the first to analyze the effect of verticals on search satisfaction. The results show that verticals with different qualities, presentation styles and positions have different effects on search satisfaction, among which Encyclopedia verticals, as well as Download verticals, will bring the largest improvement. Furthermore, our proposed vertical-aware prediction method outperforms state-of-the-art methods that are designed for search satisfaction prediction in homogeneous environment. Yiqun Liu 0001, Ke Zhou 0003, Meng Wang 0001, Min Zhang 0006, Shaoping Ma |
CIKM | 3 |
| 2015 | Exploring Composite Retrieval from the Users' Perspective
Horatiu S. Bota, Ke Zhou 0003, Joemon M. Jose |
ECIR | 2 |
| 2015 | Influence of Vertical Result in Web Search ExaminationabstractResearch in how users examine results on search engine result pages (SERPs) helps improve result ranking, advertisement placement, performance evaluation and search UI design. Although examination behavior on organic search results (also known as "ten blue links") has been well studied in existing works, there lacks a thorough investigation on how users examine SERPs with verticals. Considering the fact that a large fraction of SERPs are served with one or more verticals in the practical Web search scenario, it is of vital importance to understand the influence of vertical results on search examination behaviors. In this paper, we focus on five popular vertical types and try to study their influences on users' examination processes in both cases when they are relevant or irrelevant to the search queries. With examination behavior data collected with an eye-tracking device, we show the existence of vertical-aware user behavior effects including vertical attraction effect, examination cut-off effect in the presence of a relevant vertical, and examination spill-over effect in the presence of an irrelevant vertical. Furthermore, we are also among the first to systematically investigate the internal examination behavior within the vertical results. We believe that this work will promote our understanding of user interactions with federated search engines and bring benefit to the construction of search performance evaluations. Zeyang Liu 0004, Yiqun Liu 0001, Ke Zhou 0003, Min Zhang 0006, Shaoping Ma |
SIGIR | 3 |
| 2015 | Incorporating Non-sequential Behavior into Click ModelsabstractClick-through information is considered as a valuable source of users' implicit relevance feedback. As user behavior is usually influenced by a number of factors such as position, presentation style and site reputation, researchers have proposed a variety of assumptions (i.e.~click models) to generate a reasonable estimation of result relevance. The construction of click models usually follow some hypotheses. For example, most existing click models follow the sequential examination hypothesis in which users examine results from top to bottom in a linear fashion. While these click models have been successful, many recent studies showed that there is a large proportion of non-sequential browsing (both examination and click) behaviors in Web search, which the previous models fail to cope with. In this paper, we investigate the problem of properly incorporating non-sequential behavior into click models. We firstly carry out a laboratory eye-tracking study to analyze user's non-sequential examination behavior and then propose a novel click model named Partially Sequential Click Model (PSCM) that captures the practical behavior of users. We compare PSCM with a number of existing click models using two real-world search engine logs. Experimental results show that PSCM outperforms other click models in terms of both predicting click behavior (perplexity) and estimating result relevance (NDCG and user preference test). We also publicize the implementations of PSCM and related datasets for possible future comparison studies. Chao Wang 0049, Yiqun Liu 0001, Meng Wang 0001, Ke Zhou 0003, Jian-Yun Nie, Shaoping Ma |
SIGIR | 4 |
| 2015 | HIA'15: Heterogeneous Information Access Workshop at WSDM 2015abstractThe HIA'15 workshop aims to bring together information retrieval practitioners from industry and academic researchers concerned with heterogeneous information access and search federation. We would like to create a forum to encourage discussion and exchange of ideas on heterogeneous information access in different contexts. To facilitate the discussion, we encourage submissions on ideas and results from different aspects of heterogeneous information access including aggregated search, composite retrieval, personal search, structured search, etc. Another objective of the workshop is to encourage submissions with novel ideas (e.g. new applications) on heterogeneous information access and potential future directions of this area. Ke Zhou 0003, Roger Jie Luo, Djoerd Hiemstra, Joemon M. Jose |
WSDM | 1 |
| 2015 | A Comparative Analysis of Interleaving Methods for Aggregated SearchabstractA result page of a modern search engine often goes beyond a simple list of “10 blue links.” Many specific user needs (e.g., News, Image, Video) are addressed by so-called aggregated or vertical search solutions: specially presented documents, often retrieved from specific sources, that stand out from the regular organic Web search results. When it comes to evaluating ranking systems, such complex result layouts raise their own challenges. This is especially true for so-called interleaving methods that have arisen as an important type of online evaluation: by mixing results from two different result pages, interleaving can easily break the desired Web layout in which vertical documents are grouped together, and hence hurt the user experience. We conduct an analysis of different interleaving methods as applied to aggregated search engine result pages. Apart from conventional interleaving methods, we propose two vertical-aware methods: one derived from the widely used Team-Draft Interleaving method by adjusting it in such a way that it respects vertical document groupings, and another based on the recently introduced Optimized Interleaving framework. We show that our proposed methods are better at preserving the user experience than existing interleaving methods while still performing well as a tool for comparing ranking systems. For evaluating our proposed vertical-aware interleaving methods, we use real-world click data as well as simulated clicks and simulated ranking systems. Aleksandr Chuklin, Anne Schuth, Ke Zhou 0003, Maarten de Rijke |
ACM Trans. Inf. Syst. | 3 |
| 2014 | From Skimming to Reading: A Two-stage Examination Model for Web SearchabstractUser's examination of search results is a key concept involved in all the click models. However, most studies assumed that eye fixation means examination and no further study has been carried out to better understand user's examination behavior. In this study, we design an experimental search engine to collect both the user's feedback on their examinations and the eye-tracking/click-through data. To our surprise, a large proportion (45.8%) of the results fixated by users are not recognized as being "read". Looking into the tracking data, we found that before the user actually "reads" the result, there is often a "skimming" step in which the user quickly looks at the result without reading it. We thus propose a two-stage examination model which composes of a first "from skimming to reading" stage (Stage 1) and a second "from reading to clicking" stage (Stage 2). We found that the biases (e.g. position bias, domain bias, attractiveness bias) considered in many studies impact in different ways in Stage 1 and Stage 2, which suggests that users make judgments according to different signals in different stages. We also show that the two-stage examination behaviors can be predicted with mouse movement behavior, which can be collected at large scale. Relevance estimation with the two-stage examination model also outperforms that with a single-stage examination model. This study shows that the user's examination of search results is a complex cognitive process that needs to be investigated in greater depth and this may have a significant impact on Web search. Yiqun Liu 0001, Chao Wang 0049, Ke Zhou 0003, Jian-Yun Nie, Min Zhang 0006, Shaoping Ma |
CIKM | 3 |
| 2014 | Aligning Vertical Collection Relevance with User IntentabstractSelecting and aggregating different types of content from multiple vertical search engines is becoming popular in web search. The user vertical intent, the verticals the user expects to be relevant for a particular information need, might not correspond to the vertical collection relevance, the verticals containing the most relevant content. In this work we propose different approaches to define the set of relevant verticals based on document judgments. We correlate the collection-based relevant verticals obtained from these approaches to the real user vertical intent, and show that they can be aligned relatively well. The set of relevant verticals defined by those approaches could therefore serve as an approximate but reliable ground-truth for evaluating vertical selection, avoiding the need for collecting explicit user vertical intent, and vice versa. Ke Zhou 0003, Thomas Demeester, Dong Nguyen 0002, Djoerd Hiemstra, Dolf Trieschnigg |
CIKM | 1 |
| 2014 | Evaluating intuitiveness of vertical-aware click modelsabstractModeling user behavior on a search engine result page is important for understanding the users and supporting simulation experiments. As result pages become more complex, click models evolve as well in order to capture additional aspects of user behavior in response to new forms of result presentation. Aleksandr Chuklin, Ke Zhou 0003, Anne Schuth, Floor Sietsma, Maarten de Rijke |
SIGIR | 2 |
| 2014 | Composite retrieval of heterogeneous web searchabstractTraditional search systems generally present a ranked list of documents as answers to user queries. In aggregated search systems, results from different and increasingly diverse verticals (image, video, news, etc.) are returned to users. For instance, many such search engines return to users both images and web documents as answers to the query "flower". Aggregated search has become a very popular paradigm. In this paper, we go one step further and study a different search paradigm: composite retrieval. Rather than returning and merging results from different verticals, as is the case with aggregated search, we propose to return to users a set of "bundles", where a bundle is composed of "cohesive" results from several verticals. For example, for the query "London Olympic", one bundle per sport could be returned, each containing results extracted from news, videos, images, or Wikipedia. Composite retrieval can promote exploratory search in a way that helps users understand the diversity of results available for a specific query and decide what to explore in more detail. In this paper, we propose and evaluate a variety of approaches to construct bundles that are relevant, cohesive and diverse. Compared with three baselines (traditional "general web only" ranking, federated search ranking and aggregated search), our evaluation results demonstrate significant performance improvement for a highly heterogeneous web collection. Horatiu S. Bota, Ke Zhou 0003, Joemon M. Jose, Mounia Lalmas-Roelleke |
WWW | 2 |
| 2013 | On the reliability and intuitiveness of aggregated search metricsabstractAggregating search results from a variety of diverse verticals such as news, images, videos and Wikipedia into a single interface is a popular web search presentation paradigm. Although several aggregated search (AS) metrics have been proposed to evaluate AS result pages, their properties remain poorly understood. In this paper, we compare the properties of existing AS metrics under the assumptions that (1) queries may have multiple preferred verticals; (2) the likelihood of each vertical preference is available; and (3) the topical relevance assessments of results returned from each vertical is available. We compare a wide range of AS metrics on two test collections. Our main criteria of comparison are (1) discriminative power, which represents the reliability of a metric in comparing the performance of systems, and (2) intuitiveness, which represents how well a metric captures the various key aspects to be measured (i.e. various aspects of a user's perception of AS result pages). Our study shows that the AS metrics that capture key AS components (e.g., vertical selection) have several advantages over other metrics. This work sheds new lights on the further developments and applications of AS metrics. Ke Zhou 0003, Mounia Lalmas-Roelleke, Tetsuya Sakai, Ronan Cummins, Joemon M. Jose |
CIKM | 1 |
| 2013 | The Impact of Temporal Intent Variability on Diversity Evaluation
Ke Zhou 0003, Stewart Whiting, Joemon M. Jose, Mounia Lalmas-Roelleke |
ECIR | 1 |
| 2013 | Temporal variance of intents in multi-faceted event-driven information needsabstractTime is often important for understanding user intent during search activity, especially for information needs related to event-driven topics. Diversity for multi-faceted information needs ensures that ranked documents optimally cover multiple facets when a user's intent is uncertain. Effective diversity is reliant on methods to (i) discover and represent facets, and (ii) determine how likely each facet is the user's intent (i.e., its popularity). Past work has developed several techniques addressing these issues, however, they have concentrated on static approaches which do not consider the temporal nature of new and evolving intents and their popularity. In many cases, what a user expects may change dramatically over time as events develop. In this work we study the temporal variance of search intents for event-driven information needs using Wikipedia. First, we model intents based upon the structure represented by the section hierarchy of Wikipedia articles closely related to the information need. Using this technique, we investigate whether temporal changes in the content structure, i.e. in a section's text, reflect the temporal popularity of the intent. We map intents taken from a query-log (as ground-truth) to Wikipedia article sections and found that a large proportion are indeed reflected in topic-related article structure. By correlating the change activity of each section with the use of the intent query over time, we found that section change activity does reflect temporal popularity of many intents. Furthermore, we show that popularity between intents changes over time for event-driven topics. Stewart Whiting, Ke Zhou 0003, Joemon M. Jose, Mounia Lalmas-Roelleke |
SIGIR | 2 |
| 2013 | Which vertical search engines are relevant?abstractAggregating search results from a variety of heterogeneous sources, so-called verticals, such as news, image and video, into a single interface is a popular paradigm in web search. Current approaches that evaluate the effectiveness of aggregated search systems are based on rewarding systems that return highly relevant verticals for a given query, where this relevance is assessed under different assumptions. It is difficult to evaluate or compare those systems without fully understanding the relationship between those underlying assumptions. To address this, we present a formal analysis and a set of extensive user studies to investigate the effects of various assumptions made for assessing query vertical relevance. A total of more than 20,000 assessments on 44 search tasks across 11 verticals are collected through Amazon Mechanical Turk and subsequently analysed. Our results provide insights into various aspects of query vertical relevance and allow us to explain in more depth as well as questioning the evaluation results published in the literature. Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose |
WWW | 1 |
| 2012 | CrowdTiles: presenting crowd-based information for event-driven information needsabstractTime plays a central role in many web search information needs relating to recent events. For recency queries where fresh information is most desirable, there is likely to be a great deal of highly-relevant information created very recently by crowds of people across the world, particularly on platforms such as Wikipedia and Twitter. With so many users, mainstream events are often very quickly reflected in these sources. The English Wikipedia encyclopedia consists of a vast collection of user-edited articles covering a range of topics. During events, users collaboratively create and edit existing articles in near real-time. Simultaneously, users on Twitter disseminate and discuss event details, with a small number of users becoming influential for the topic. Stewart Whiting, Ke Zhou 0003, Joemon M. Jose, Omar Alonso, Teerapong Leelanupab |
CIKM | 2 |
| 2012 | Evaluating reward and risk for vertical selectionabstractThe aggregation of search results from heterogeneous verticals (news, videos, blogs, etc) has become an important consideration in search. When aiming to select suitable verticals, from which items are selected to be shown along with the standard "ten blue links", there exists the potential to both help (selecting relevant verticals) and harm (selecting irrelevant verticals) the existing result set. Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose |
CIKM | 1 |
| 2012 | Assessing and Predicting Vertical Intent for Web Queries
Ke Zhou 0003, Ronan Cummins, Martin Halvey, Mounia Lalmas-Roelleke, Joemon M. Jose |
ECIR | 1 |
| 2012 | Evaluating aggregated search pagesabstractAggregating search results from a variety of heterogeneous sources or verticals such as news, image and video into a single interface is a popular paradigm in web search. Although various approaches exist for selecting relevant verticals or optimising the aggregated search result page, evaluating the quality of an aggregated page is an open question. Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose |
SIGIR | 1 |