EDBT 2026 Demo / reviewers in the wild / expert
Anlei Dong
dblp:28/6385
· DBLP profile ↗
38ranked-venue papers
8as first author
4since 2021 · last 2024
0000-0002-8241-4746ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LEAD: Liberal Feature-based Distillation for Dense RetrievalabstractKnowledge distillation is often used to transfer knowledge from a strong teacher model to a relatively weak student model. Traditional methods include response-based methods and feature-based methods. Response-based methods are widely used but suffer from lower upper limits of performance due to their ignorance of intermediate signals, while feature-based methods have constraints on vocabularies, tokenizers and model architectures. In this paper, we propose a liberal feature-based distillation method (LEAD). LEAD aligns the distribution between the intermediate layers of teacher model and student model, which is effective, extendable, portable and has no requirements on vocabularies, tokenizers, or model architectures. Extensive experiments show the effectiveness of LEAD on widely-used benchmarks, including MS MARCO Passage Ranking, TREC 2019 DL Track, MS MARCO Document Ranking and TREC 2020 DL Track. Our code is available in https://github.com/microsoft/SimXNS/tree/main/LEAD. Hao Sun 0015, Xiao Liu 0029, Yeyun Gong, Anlei Dong, Jingwen Lu, Yan Zhang 0117, Linjun Yang, Rangan Majumder, Nan Duan 0001 |
WSDM | 4 |
| 2023 | CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionabstractThe dual-encoder has become the de facto architecture for dense retrieval.Typically, it computes the latent representations of the query and document independently, thus failing to fully capture the interactions between the query and document.To alleviate this, recent research has focused on obtaining query-informed document representations.During training, it expands the document with a real query, but during inference, it replaces the real query with a generated one.This inconsistency between training and inference causes the dense retrieval model to prioritize query information while disregarding the document when computing the document representation.Consequently, it performs even worse than the vanilla dense retrieval model because its performance heavily relies on the relevance between the generated queries and the real query.In this paper, we propose a curriculum sampling strategy that utilizes pseudo queries during training and progressively enhances the relevance between the generated query and the real query.By doing so, the retrieval model learns to extend its attention from the document alone to both the document and query, resulting in high-quality queryinformed document representations.Experimental results on both in-domain and out-ofdomain datasets demonstrate that our approach outperforms previous dense retrieval models. Xingwei He 0003, Yeyun Gong, A-Long Jin, Hang Zhang 0029, Anlei Dong, Jian Jiao 0007, Siu-Ming Yiu, Nan Duan 0001 |
EMNLP | 5 |
| 2023 | PROD: Progressive Distillation for Dense RetrievalabstractKnowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model. Ideally, we expect the better the teacher is, the better the student performs. However, this expectation does not always come true. It is common that a strong teacher model results in a bad student via distillation due to the nonnegligible gap between teacher and student. To bridge the gap, we propose PROD, a PROgressive Distillation method, for dense retrieval. PROD consists of a teacher progressive distillation and a data progressive distillation to gradually improve the student. To alleviate catastrophic forgetting, we introduce a regularization term in each distillation process. We conduct extensive experiments on seven datasets including five widely-used publicly available benchmarks: MS MARCO Passage, TREC Passage 19, TREC Document 19, MS MARCO Document, and Natural Questions, as well as two industry datasets: Bing-Rel and Bing-Ads. PROD achieves the state-of-the-art in the distillation methods for dense retrieval. Our 6-layer student model even surpasses most of the existing 12-layer models on all five public benchmarks. The code and models are released in https://github.com/microsoft/SimXNS. Zhenghao Lin, Yeyun Gong, Xiao Liu 0029, Hang Zhang 0029, Chen Lin 0001, Anlei Dong, Jian Jiao 0007, Jingwen Lu, Daxin Jiang, Rangan Majumder, Nan Duan 0001 |
WWW | 6 |
| 2022 | Less is Less: When are Snippets Insufficient for Human vs Machine Relevance Estimation?
Gabriella Kazai, Bhaskar Mitra 0001, Anlei Dong, Nick Craswell, Linjun Yang |
ECIR (2) | 3 |
| 2017 | Exploring Query Auto-Completion and Click Logs for Contextual-Aware Web Search and Query SuggestionabstractContextual data plays an important role in modeling search engine users' behaviors on both query auto-completion (QAC) log and normal query (click) log. User's recent search history on each log has been widely studied individually as the context to benefit the modeling of users' behaviors on that log. However, there is no existing work that explores or incorporates both logs together for contextual data. As QAC and click logs actually record users' sequential behaviors while interacting with a search engine, the available context of a user's current behavior based on the same type of log can be strengthened from the user's recent search history shown on the other type of log. Our paper proposes to model users' behaviors on both QAC and click logs simultaneously by utilizing both logs as the contextual data of each other. The key idea is to capture the correlation between users' behavior patterns on both logs. We model such correlation through a novel probabilistic model based on the Latent Dirichlet allocation (LDA) model. The learned users' behavior patterns on both logs are utilized to address not only the application of query auto-completion on QAC logs, but also the click prediction and relevance ranking of web documents on click logs. Experiments on real-world logs demonstrate the effectiveness of the proposed model on both applications. Liangda Li, Hongbo Deng, Anlei Dong, Yi Chang 0001, Ricardo Baeza-Yates, Hongyuan Zha |
WWW | 3 |
| 2016 | Behavior Driven Topic Transition for Search Task IdentificationabstractSearch tasks in users' query sequences are dynamic and interconnected. The formulation of search tasks can be influenced by multiple latent factors such as user characteristics, product features and search interactions, which makes search task identification a challenging problem. In this paper, we propose an unsupervised approach to identify search tasks via topic membership along with topic transition probabilities, thus it becomes possible to interpret how user's search intent emerges and evolves over time. Moreover, a novel hidden semi-Markov model is introduced to model topic transitions by considering not only the semantic information of queries but also the latent search factors originated from user search behaviors. A variational inference algorithm is developed to identify remarkable search behavior patterns, typical topic transition tracks, and the topic membership of each query from query logs. The learned topic transition tracks and the inferred topic memberships enable us to identify both small search tasks, where a user searches the same topic, and big search tasks, where a user searches a series of related topics. We extensively evaluate the proposed approach and compare with several state-of-the-art search task identification methods on both synthetic and real-world query log data, and experimental results illustrate the effectiveness of our proposed model. Liangda Li, Hongbo Deng, Anlei Dong, Yi Chang 0001, Hongyuan Zha |
WWW | 4 |
| 2015 | Propagation-based Sentiment Analysis for Microblogging DataabstractThe explosive popularity of microblogging services encourages more and more online users to share their opinions, and sentiment analysis on such opinion-rich resources has been proven to be an effective way to understand public opinions. On the one hand, the brevity and informality of microblogging data plus its wide variety and rapid evolution of language in microblogging pose new challenges to the vast majority of existing methods. On the other hand, microblogging texts contain various types of emotional signals strongly associated with their sentiment polarity, which brings about new opportunities for sentiment analysis. In this paper, we investigate propagation-based sentiment analysis for microblogging data. In particular, we provide a propagating process to incorporate various types of emotional signals in microblogging data into a coherent model, and propose a novel sentiment analysis framework PSA which learns from both labeled and unlabeled data by iteratively alternating a propagating process and a fitting process. We conduct experiments on real-world microblogging datasets, and the results demonstrate the effectiveness of the proposed framework. Further experiments are conducted to probe the working of the key components of the proposed framework. Jiliang Tang, Chikashi Nobata, Anlei Dong, Yi Chang 0001, Huan Liu 0001 |
SDM | 3 |
| 2015 | Analyzing User's Sequential Behavior in Query Auto-Completion via Markov ProcessesabstractQuery auto-completion (QAC) plays an important role in assisting users typing less while submitting a query. The QAC engine generally offers a list of suggested queries that start with a user's input as a prefix, and the list of suggestions is changed to match the updated input after the user types each keystroke. Therefore rich user interactions can be observed along with each keystroke until a user clicks a suggestion or types the entire query manually. It becomes increasingly important to analyze and understand users' interactions with the QAC engine, to improve its performance. Existing works on QAC either ignored users' interaction data, or assumed that their interactions at each keystroke are independent from others. Our paper pays high attention to users' sequential interactions with a QAC engine in and across QAC sessions, rather than users' interactions at each keystroke of each QAC session separately. Analyzing the dependencies in users' sequential interactions improves our understanding of the following three questions: 1) how is a user's skipping/viewing move at the current keystroke influenced by that at the previous keystroke? 2) how to improve search engines' query suggestions at short keystrokes based on those at latter long keystrokes? and 3) facing a targeted query shown in the suggestion list, why does a user decide to continue typing rather than click the intended suggestion? We propose a probabilistic model that addresses those three questions in a unified way, and illustrate how the model determines users' final click decisions. By comparing with state-of-the-art methods, our proposed model does suggest queries that better satisfy users' intents. Liangda Li, Hongbo Deng, Anlei Dong, Yi Chang 0001, Hongyuan Zha, Ricardo Baeza-Yates |
SIGIR | 3 |
| 2015 | adaQAC: Adaptive Query Auto-Completion via Implicit Negative FeedbackabstractQuery auto-completion (QAC) facilitates user query composition by suggesting queries given query prefix inputs. In 2014, global users of Yahoo! Search saved more than 50% keystrokes when submitting English queries by selecting suggestions of QAC. Users' preference of queries can be inferred during user-QAC interactions, such as dwelling on suggestion lists for a long time without selecting query suggestions ranked at the top. However, the wealth of such implicit negative feedback has not been exploited for designing QAC models. Most existing QAC models rank suggested queries for given prefixes based on certain relevance scores. Aston Zhang, Amit Goyal 0001, Weize Kong, Hongbo Deng, Anlei Dong, Yi Chang 0001, Carl A. Gunter, Jiawei Han 0001 |
SIGIR | 5 |
| 2014 | Identifying and labeling search tasks via query-based hawkes processesabstractWe consider a search task as a set of queries that serve the same user information need. Analyzing search tasks from user query streams plays an important role in building a set of modern tools to improve search engine performance. In this paper, we propose a probabilistic method for identifying and labeling search tasks based on the following intuitive observations: queries that are issued temporally close by users in many sequences of queries are likely to belong to the same search task, meanwhile, different users having the same information needs tend to submit topically coherent search queries. To capture the above intuitions, we directly model query temporal patterns using a special class of point processes called Hawkes processes, and combine topic models with Hawkes processes for simultaneously identifying and labeling search tasks. Essentially, Hawkes processes utilize their self-exciting properties to identify search tasks if influence exists among a sequence of queries for individual users, while the topic model exploits query co-occurrence across different users to discover the latent information needed for labeling search tasks. More importantly, there is mutual reinforcement between Hawkes processes and the topic model in the unified model that enhances the performance of both. We evaluate our method based on both synthetic data and real-world query log data. In addition, we also apply our model to query clustering and search task identification. By comparing with state-of-the-art methods, the results demonstrate that the improvement in our proposed approach is consistent and promising. Liangda Li, Hongbo Deng, Anlei Dong, Yi Chang 0001, Hongyuan Zha |
KDD | 3 |
| 2014 | A two-dimensional click model for query auto-completionabstractQuery auto-completion (QAC) facilitates faster user query input by predicting users' intended queries. Most QAC algorithms take a learning-based approach to incorporate various signals for query relevance prediction. However, such models are trained on simulat- ed user inputs from query log data. The lack of real user interaction data in the QAC process prevents them from further improving the QAC performance. In this work, for the first time we collect a high-resolution QAC query log that records every keystroke in a QAC session. Based on this data, we discover two user behaviors, namely the horizontal skipping bias and vertical position bias which are crucial for rele- vance prediction in QAC. In order to better explain them, we pro- pose a novel two-dimensional click model for modeling the QAC process with emphasis on these behaviors. Extensive experiments on our QAC data set from both PC and mobile devices demonstrate that our proposed model can accurate- ly explain the users' behaviors in interacting with a QAC system, and the resulting relevance model significant improves the QAC performance over existing click models. Furthermore, the learned knowledge about the skipping behavior can be effectively incorpo- rated into existing learning-based models to further improve their performance. Yanen Li, Anlei Dong, Hongning Wang, Hongbo Deng, Yi Chang 0001, ChengXiang Zhai |
SIGIR | 2 |
| 2014 | User modeling in search logs via a nonparametric bayesian approachabstractSearchers' information needs are diverse and cover a broad range of topics; hence, it is important for search engines to accurately understand each individual user's search intents in order to provide optimal search results. Search log data, which records users' search behaviors when interacting with search engines, provides a valuable source of information about users' search intents. Therefore, properly characterizing the heterogeneity among the users' observed search behaviors is the key to accurately understanding their search intents and to further predicting their behaviors. Hongning Wang, ChengXiang Zhai, Anlei Dong, Yi Chang 0001 |
WSDM | 4 |
| 2014 | Exploiting User Preference for Online Learning in Web Content Optimization SystemsabstractWeb portal services have become an important medium to deliver digital content (e.g. news, advertisements, etc.) to Web users in a timely fashion. To attract more users to various content modules on the Web portal, it is necessary to design a recommender system that can effectively achieve Web portal content optimization by automatically estimating content item attractiveness and relevance to user interests. The state-of-the-art online learning methodology adapts dedicated pointwise models to independently estimate the attractiveness score for each candidate content item. Although such pointwise models can be easily adapted for online recommendation, there still remain a few critical problems. First, this pointwise methodology fails to use invaluable user preferences between content items. Moreover, the performance of pointwise models decreases drastically when facing the problem of sparse learning samples. To address these problems, we propose exploring a new dynamic pairwise learning methodology for Web portal content optimization in which we exploit dynamic user preferences extracted based on users' actions on portal services to compute the attractiveness scores of content items. In this article, we introduce two specific pairwise learning algorithms, a straightforward graph-based algorithm and a formalized Bayesian modeling one. Experiments on large-scale data from a commercial Web portal demonstrate the significant improvement of pairwise methodologies over the baseline pointwise models. Further analysis illustrates that our new pairwise learning approaches can benefit personalized recommendation more than pointwise models, since the data sparsity is more critical for personalized content optimization. Jiang Bian 0002, Bo Long, Lihong Li 0001, Taesup Moon, Anlei Dong, Yi Chang 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2013 | Temporal web dynamics and its application to information retrievalabstractThe World Wide Web is highly dynamic and is constantly evolving to cover the latest information about the physical and social updates in the world. At the same time, the changes in web contents are entangled with new information needs and time-sensitive user interactions with information sources. To address these temporal information needs effectively, it is essential for the search engines to model web dynamics and understand the changes in user behavior over time that are caused by them. Kira Radinsky, Fernando Diaz 0001, Susan T. Dumais, Milad Shokouhi, Anlei Dong, Yi Chang 0001 |
WSDM | 5 |
| 2013 | Content-aware click modelingabstractClick models aim at extracting intrinsic relevance of documents to queries from biased user clicks. One basic modeling assumption made in existing work is to treat such intrinsic relevance as an atomic query-document-specific parameter, which is solely estimated from historical clicks without using any content information about a document or relationship among the clicked/skipped documents under the same query. Due to this overly simplified assumption, existing click models can neither fully explore the information about a document's relevance quality nor make predictions of relevance for any unseen documents. Hongning Wang, ChengXiang Zhai, Anlei Dong, Yi Chang 0001 |
WWW | 3 |
| 2013 | Improving recency ranking using twitter dataabstractIn Web search and vertical search, recency ranking refers to retrieving and ranking documents by both relevance and freshness. As impoverished in-links and click information is the the biggest challenge for recency ranking, we advocate the use of Twitter data to address the challenge in this article. We propose a method to utilize Twitter TinyURL to detect fresh and high-quality documents, and leverage Twitter data to generate novel and effective features for ranking. The empirical experiments demonstrate that the proposed approach effectively improves a commercial search engine for both Web search ranking and tweet vertical ranking. Yi Chang 0001, Anlei Dong, Pranam Kolari, Ruiqiang Zhang, Yoshiyuki Inagaki, Fernando Diaz 0001, Hongyuan Zha, Yan Liu 0002 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | User Action Interpretation for Online Content OptimizationabstractWeb portal services have become an important medium to deliver digital content and service, such as news, advertisements, and so on, to Web users in a timely fashion. To attract more users to various content modules on the Web portal, it is necessary to design a recommender system that can effectively achieve online content optimization by automatically estimating content items' attractiveness and relevance to users' interests. User interaction plays a vital role in building effective content optimization, as both implicit user feedbacks and explicit user ratings on the recommended items form the basis for designing and learning recommendation models. However, user actions on real-world Web portal services are likely to represent many implicit signals about users' interests and content attractiveness, which need more accurate interpretation to be fully leveraged in the recommendation models. To address this challenge, we investigate a couple of critical aspects of the online learning framework for personalized content optimization on Web portal services, and, in this paper, we propose deeper user action interpretation to enhance those critical aspects. In particular, we first propose an approach to leverage historical user activity to build behavior-driven user segmentation; then, we introduce an approach for interpreting users' actions from the factors of both user engagement and position bias to achieve unbiased estimation of content attractiveness. Our experiments on the large-scale data from a commercial Web recommender system demonstrate that recommendation models with our user action interpretation can reach significant improvement in terms of online content optimization over the baseline method. The effectiveness of our user action interpretation is also proved by the online test results on real user traffic. Jiang Bian 0002, Anlei Dong, Srihari Reddy, Yi Chang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Iterative Viterbi A* Algorithm for K-Best Sequential Decoding
Zhiheng Huang, Yi Chang 0001, Bo Long, Jean-François Crespo, Anlei Dong, S. Sathiya Keerthi, Su-Lin Wu |
ACL (1) | 5 |
| 2012 | Enhancing product search by best-selling prediction in e-commerceabstractWith the rapid growth of E-Commerce on the Internet, online product search service has emerged as a popular and effective paradigm for customers to find desired products and select transactions. Most product search engines today are based on adaptations of relevance models devised for information retrieval. However, there is still a big gap between the mechanism of finding products that customers really desire to purchase and that of retrieving products of high relevance to customers' query. In this paper, we address this problem by proposing a new ranking framework for enhancing product search based on dynamic best-selling prediction in E-Commerce. Specifically, we first develop an effective algorithm to predict the dynamic best-selling, i.e. the volume of sales, for each product item based on its transaction history. By incorporating such best-selling prediction with relevance, we propose a new ranking model for product search, in which we rank higher the product items that are not only relevant to the customer's need but with higher probability to be purchased by the customer. Results of a large scale evaluation, conducted over the dataset from a commercial product search engine, demonstrate that our new ranking method is more effective for locating those product items that customers really desire to buy at higher rank positions without hurting the search relevance. Bo Long, Jiang Bian 0002, Anlei Dong, Yi Chang 0001 |
CIKM | 3 |
| 2012 | Pairwise cross-domain factor model for heterogeneous transfer rankingabstractLearning to rank arises in many information retrieval applications, ranging from Web search engine, online advertising to recommendation systems. Traditional ranking mainly focuses on one type of data source, and effective modeling relies on a sufficiently large number of labeled examples, which require expensive and time-consuming labeling process. However, in many real-world applications, ranking over multiple related heterogeneous domains becomes a common situation, where in some domains we may have a relatively large amount of training data while in some other domains we can only collect very little. Theretofore, how to leverage labeled information from related heterogeneous domain to improve ranking in a target domain has become a problem of great interests. In this paper, we propose a novel probabilistic model, pairwise cross-domain factor model, to address this problem. The proposed model learns latent factors(features) for multi-domain data in partially-overlapped heterogeneous feature spaces. It is capable of learning homogeneous feature correlation, heterogeneous feature correlation, and pairwise preference correlation for cross-domain knowledge transfer. We also derive two PCDF variations to address two important special cases. Under the PCDF model, we derive a stochastic gradient based algorithm, which facilitates distributed optimization and is flexible to adopt different loss functions and regularization functions to accommodate different data distributions. The extensive experiments on real world data sets demonstrate the effectiveness of the proposed model and algorithm. Bo Long, Yi Chang 0001, Anlei Dong, Jianzhang He |
WSDM | 3 |
| 2012 | Joint relevance and freshness learning from clickthroughs for news searchabstractIn contrast to traditional Web search, where topical relevance is often the main selection criterion, news search is characterized by the increased importance of freshness. However, the estimation of relevance and freshness, and especially the relative importance of these two aspects, are highly specific to the query and the time when the query was issued. In this work, we propose a unified framework for modeling the topical relevance and freshness, as well as their relative importance, based on click logs. We use click statistics and content analysis techniques to define a set of temporal features, which predict the right mix of freshness and relevance for a given query. Experimental results on both historical click data and editorial judgments demonstrate the effectiveness of the proposed approach. Hongning Wang, Anlei Dong, Lihong Li 0001, Yi Chang 0001, Evgeniy Gabrilovich |
WWW | 2 |
| 2011 | User action interpretation for personalized content optimization in recommender systemsabstractUser interaction plays a vital role in recommender systems. Previous studies on algorithmic recommender systems have mainly focused on modeling techniques and feature development. Traditionally, implicit user feedback or explicit user ratings on the recommended items form the basis for designing and training of recommendation algorithms. But user interactions in real-world Web applications (e.g., a portal website with different recommendation modules in the interface) are unlikely to be as ideal as those assumed by previously proposed models. To address this problem, we build an online learning framework for personalized recommendation. We argue that appropriate user action interpretation is critical for a recommender system. The main contribution in this paper is an approach of interpreting users' actions for the online learning to achieve better item relevance estimation. Our experiments on the large-scale data from a commercial Web recommender system demonstrate significant improvement in terms of a precision metric over the baseline model that does not incorporate user action interpretation. The efficacy of this new algorithm is also proved by the online test results on real user traffic. Anlei Dong, Jiang Bian 0002, Srihari Reddy, Yi Chang 0001 |
CIKM | 1 |
| 2010 | Session Based Click Features for Recency RankingabstractRecency ranking refers to the ranking of web results by accounting for both relevance and freshness. This is particularly important for "recency sensitive" queries such as breaking news queries. In this study, we propose a set of novel click features to improve machine learned recency ranking. Rather than computing simple aggregate click through rates, we derive these features using the temporal click through data and query reformulation chains. One of the features that we use is click buzz that captures the spiking interest of a url for a query. We also propose time weighted click through rates which treat recent observations as being exponentially more important. The promotion of fresh content is typically determined by the query intent which can change dynamically over time. Quite often users query reformulations convey clues about the query's intent. Hence we enrich our click features by following query reformulations which typically benefit the first query in the chain of reformulations. Our experiments show these novel features can improve the NDCG5 of a major online search engine's ranking for "recency sensitive" queries by up to 1.57%. This is one of the very few studies that exploits temporal click through data and query reformulations for recency ranking. Yoshiyuki Inagaki, Narayanan Sadagopan, Georges Dupret, Anlei Dong, Ciya Liao, Yi Chang 0001, Zhaohui Zheng 0001 |
AAAI | 4 |
| 2010 | Learning Recurrent Event Queries for Web Search
Ruiqiang Zhang, Yuki Konda, Anlei Dong, Pranam Kolari, Yi Chang 0001, Zhaohui Zheng 0001 |
EMNLP | 3 |
| 2010 | Towards recency ranking in web searchabstractIn web search, recency ranking refers to ranking documents by relevance which takes freshness into account. In this paper, we propose a retrieval system which automatically detects and responds to recency sensitive queries. The system detects recency sensitive queries using a high precision classifier. The system responds to recency sensitive queries by using a machine learned ranking model trained for such queries. We use multiple recency features to provide temporal evidence which effectively represents document recency. Furthermore, we propose several training methodologies important for training recency sensitive rankers. Finally, we develop new evaluation metrics for recency sensitive queries. Our experiments demonstrate the efficacy of the proposed approaches. Anlei Dong, Yi Chang 0001, Zhaohui Zheng 0001, Gilad Mishne, Ruiqiang Zhang, Karolina Buchner, Ciya Liao, Fernando Diaz 0001 |
WSDM | 1 |
| 2010 | Time is of the essence: improving recency ranking using Twitter dataabstractRealtime web search refers to the retrieval of very fresh content which is in high demand. An effective portal web search engine must support a variety of search needs, including realtime web search. However, supporting realtime web search introduces two challenges not encountered in non-realtime web search: quickly crawling relevant content and ranking documents with impoverished link and click information. In this paper, we advocate the use of realtime micro-blogging data for addressing both of these problems. We propose a method to use the micro-blogging data stream to detect fresh URLs. We also use micro-blogging data to compute novel and effective features for ranking fresh URLs. We demonstrate these methods improve effective of the portal web search engine for realtime web search. Anlei Dong, Ruiqiang Zhang, Pranam Kolari, Fernando Diaz 0001, Yi Chang 0001, Zhaohui Zheng 0001, Hongyuan Zha |
WWW | 1 |
| 2009 | Incorporating robustness into web ranking evaluationabstractIn many Web search engines, a ranking function is selected for deployment mainly by comparing the relevance measurements over candidates. Due to the dynamical nature of the Web, the ranking features and the query and URL distribution on which the ranking functions are built, may change dramatically over time. The actual relevance of the function may degrade, and thus the previous function selection conclusions become invalid. In this work we suggest to select Web ranking functions according to both their relevance and robustness to the changes that may lead to relevance degradation over time. We argue that the ranking robustness can be effectively measured by taking into account the ranking score distribution across search results. We then propose two alternatives to the NDCG metric that both incorporate ranking robustness into ranking function evaluation and selection. A machine learning approach is developed to learn the parameters that control the metric sensitivity to score turbulence, from human-judged preference data. Shihao Ji 0001, Zhaohui Zheng 0001, Yi Chang 0001, Anlei Dong |
CIKM | 6 |
| 2009 | Empirical Exploitation of Click Data for Task Specific Ranking
Anlei Dong, Yi Chang 0001, Shihao Ji 0001, Ciya Liao, Zhaohui Zheng 0001 |
EMNLP | 1 |
| 2009 | Enhancing topical ranking with preferences from click-through dataabstractTo overcome the training data insufficiency problem for dedicated model in topical ranking, this paper proposes to utilize click-through data to improve learning. The efficacy of click-through data is explored under the framework of preference learning. The empirical experiment on a commercial search engine shows that, the model trained with the dedicated labeled data combined with skip-next preferences could beat the baseline model and the generic model in NDCG5 for 4.9% and 2.4% respectively. Yi Chang 0001, Anlei Dong, Ciya Liao, Zhaohui Zheng 0001 |
SIGIR | 2 |
| 2008 | Long-Term Cross-Session Relevance Feedback Using Virtual FeaturesabstractRelevance feedback (RF) is an iterative process, which refines the retrievals by utilizing the user's feedback on previously retrieved results. Traditional RF techniques solely use the short-term learning experience and do not exploit the knowledge created during cross sessions with multiple users. In this paper, we propose a novel RF framework, which facilitates the combination of short-term and long-term learning processes by integrating the traditional methods with a new technique called the virtual feature. The feedback history with all the users is digested by the system and is represented in a very efficient form as a virtual feature of the images. As such, the dissimilarity measure can dynamically be adapted, depending on the estimate of the semantic relevance derived from the virtual features. In addition, with a dynamic database, the user's subject concepts may transit from one to another. By monitoring the changes in retrieval performance, the proposed system can automatically adapt the concepts according to the new subject concepts. The experiments are conducted on a real image database. The results manifest that the proposed framework outperforms the traditional within-session and log-based long-term RF techniques. Peng-Yeng Yin, Bir Bhanu, Kuang-Cheng Chang, Anlei Dong |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2008 | Feature synthesized EM algorithm for image retrievalabstractAs a commonly used unsupervised learning algorithm in Content-Based Image Retrieval (CBIR), Expectation-Maximization (EM) algorithm has several limitations, including the curse of dimensionality and the convergence at a local maximum. In this article, we propose a novel learning approach, namely Coevolutionary Feature Synthesized Expectation-Maximization (CFS-EM), to address the above problems. The CFS-EM is a hybrid of coevolutionary genetic programming (CGP) and EM algorithm applied on partially labeled data. CFS-EM is especially suitable for image retrieval because the images can be searched in the synthesized low-dimensional feature space, while a kernel-based method has to make classification computation in the original high-dimensional space. Experiments on real image databases show that CFS-EM outperforms Radial Basis Function Support Vector Machine (RBF-SVM), CGP, Discriminant-EM (D-EM) and Transductive-SVM (TSVM) in the sense of classification performance and it is computationally more efficient than RBF-SVM in the query phase. Rui Li 0083, Bir Bhanu, Anlei Dong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2005 | Coevolutionary feature synthesized EM algorithm for image retrievalabstractAs a commonly used unsupervised learning algorithm in Content-Based Image Retrieval (CBIR), Expectation-Maximization (EM) algorithm has several limitations, especially in high dimensional feature spaces where the data are limited and the computational cost varies exponentially with the number of feature dimensions. Moreover, the convergence is guaranteed only at a local maximum. In this paper, we propose a unified framework of a novel learning approach, namely Coevolutionary Feature Synthesized Expectation-Maximization (CFS-EM), to achieve satisfactory learning in spite of these difficulties. The CFS-EM is a hybrid of coevolutionary genetic programming (CGP) and EM algorithm. The advantages of CFS-EM are: 1) it synthesizes low-dimensional features based on CGP algorithm, which yields near optimal nonlinear transformation and classification precision comparable to kernel methods such as the support vector machine (SVM); 2) the explicitness of feature transformation is especially suitable for image retrieval because the images can be searched in the synthesized low-dimensional space, while kernel-based methods have to make classification computation in the original high-dimensional space; 3) the unlabeled data can be boosted with the help of the class distribution learning using CGP feature synthesis approach. Experimental results show that CFS-EM outperforms pure EM and CGP alone, and is comparable to SVM in the sense of classification. It is computationally more efficient than SVM in query phase. Moreover, it has a high likelihood that it will jump out of a local maximum to provide near optimal results and a better estimation of parameters. Rui Li 0083, Bir Bhanu, Anlei Dong |
ACM Multimedia | 3 |
| 2005 | Integrating Relevance Feedback Techniques for Image Retrieval Using Reinforcement LearningabstractRelevance feedback (RF) is an interactive process which refines the retrievals to a particular query by utilizing the user's feedback on previously retrieved results. Most researchers strive to develop new RF techniques and ignore the advantages of existing ones. In this paper, we propose an image relevance reinforcement learning (IRRL) model for integrating existing RF techniques in a content-based image retrieval system. Various integration schemes are presented and a long-term shared memory is used to exploit the retrieval experience from multiple users. Also, a concept digesting method is proposed to reduce the complexity of storage demand. The experimental results manifest that the integration of multiple RF approaches gives better retrieval performance than using one RF technique alone, and that the sharing of relevance knowledge between multiple query sessions significantly improves the performance. Further, the storage demand is significantly reduced by the concept digesting technique. This shows the scalability of the proposed model with the increasing-size of database. Peng-Yeng Yin, Bir Bhanu, Kuang-Cheng Chang, Anlei Dong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2005 | Active concept learning in image databasesabstractConcept learning in content-based image retrieval systems is a challenging task. This paper presents an active concept learning approach based on the mixture model to deal with the two basic aspects of a database system: the changing (image insertion or removal) nature of a database and user queries. To achieve concept learning, we a) propose a new user directed semi-supervised expectation-maximization algorithm for mixture parameter estimation, and b) develop a novel model selection method based on Bayesian analysis that evaluates the consistency of hypothesized models with the available information. The analysis of exploitation versus exploration in the search space helps to find the optimal model efficiently. Our concept knowledge transduction approach is able to deal with the cases of image insertion and query images being outside the database. The system handles the situation where users may mislabel images during relevance feedback. Experimental results on Corel database show the efficacy of our active concept learning approach and the improvement in retrieval performance by concept transduction. Anlei Dong, Bir Bhanu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | A New Semi-Supervised EM Algorithm for Image RetrievalabstractOne of the main tasks in content-based image retrieval (CBIR) is to reduce the gap between low-level visual features and high-level human concepts. This paper presents a new semi-supervised EM algorithm (NSSEM), where the image distribution in feature space is modeled as a mixture of Gaussian densities. Due to the statistical mechanism of accumulating and processing meta knowledge, the NSS-EM algorithm with long term learning of mixture model parameters can deal with the cases where users may mislabel images during relevance feedback. Our approach that integrates mixture model of the data, relevance feedback and long term learning helps to improve retrieval performance. The concept learning is incrementally refined with increased retrieval experiences. Experiment results on Corel database show the efficacy of our proposed concept learning approach. Anlei Dong, Bir Bhanu |
CVPR (2) | 1 |
| 2003 | Active Concept Learning for Image Retrieval in Dynamic DatabasesabstractConcept learning in content-based image retrieval (CBIR) systems is a challenging task. We present an active concept learning approach based on mixture model to deal with the two basic aspects of a database system: changing (image insertion or removal) nature of a database and user queries. To achieve concept learning, we develop a novel model selection method based on Bayesian analysis that evaluates the consistency of hypothesized models with the available information. The analysis of exploitation vs. exploration in the search space helps to find optimal model efficiently. Experimental results on Corel database show the efficacy of our approach. Anlei Dong, Bir Bhanu |
ICCV | 1 |
| 2003 | Reinforcement Learning for Combining Relevance Feedback TechniquesabstractRelevance feedback (RF) is an interactive process which refines the retrievals by utilizing user's feedback history. Most researchers strive to develop new RF techniques and ignore the advantages of existing ones. We propose an image relevance reinforcement learning (IRRL) model for integrating existing RF techniques. Various integration schemes are presented and a long-term shared memory is used to exploit the retrieval experience from multiple users. Also, a concept digesting method is proposed to reduce the complexity of storage demand. The experimental results manifest that the integration of multiple RF approaches gives better retrieval performance than using one RF technique alone, and that the sharing of relevance knowledge between multiple query sessions also provides significant contributions for improvement. Further, the storage demand is significantly reduced by the concept digesting technique. This shows the scalability of the proposed model against a growing-size database. Peng-Yeng Yin, Bir Bhanu, Kuang-Cheng Chang, Anlei Dong |
ICCV | 4 |
| 2003 | Concept learning and transplantation for dynamic image databasesabstractThe task of a content-based image retrieval (CBIR) system is to cater to users who expect to get relevant images with high precision and efficiency in response to query images. This paper presents a concept learning approach that integrates a mixture model of the data, relevance feedback and long-term continuous learning. The concepts are incrementally refined with increased retrieval experiences. The concept knowledge can be immediately transplanted to deal with the dynamic database situations such as insertion of new images, removal of existing images and query images, which are outside the database. Experimental results on Corel database show the efficacy of our approach. Anlei Dong, Bir Bhanu |
ICME | 1 |