VLDB 2026 Research / reviewers in the wild / expert
Gui-Rong Xue
dblp:72/3906
· DBLP profile ↗
71ranked-venue papers
11as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 57 · 11 first-authorArtificial intelligence and machine learning · 28 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 9Graphics, computer vision, multimedia, augmented reality and games · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
27 papers |
Information retrieval · 63% Data mining · 20% Recommender systems · 8% | |
| Artificial intelligence
14 papers |
Transfer learning and domain adaptation · 64% Information extraction and text analysis · 16% Image recognition and object detection · 6% |
Topics — the 30 heaviest of 81, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking
learning to rank |
0.4 | 3 | 2014 | Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014 Global ranking by exploiting user clicks · SIGIR 2009 Learning to rank with ties · SIGIR 2008 |
Machine learning › Transfer learning and domain adaptation
heterogeneous transfer learning |
0.3 | 3 | 2011 | Heterogeneous Transfer Learning for Image Classification · AAAI 2011 Heterogeneous Transfer Learning for Image Clustering via the SocialWeb · ACL/IJCNLP 2009 Translated Learning: Transfer Learning across Different Feature Spaces · NIPS 2008 |
Information retrieval
ranking |
0.3 | 3 | 2009 | Dataplorer: a scalable search engine for the data web · WWW 2009 Global ranking by exploiting user clicks · SIGIR 2009 Optimizing web search using social annotations · WWW 2007 |
Machine learning › Transfer learning and domain adaptation
cross-modal transfer |
0.2 | 2 | 2011 | Video summarization via transferrable structured learning · WWW 2011 Heterogeneous Transfer Learning for Image Classification · AAAI 2011 |
Information retrieval
text summarization |
0.2 | 2 | 2011 | Video summarization via transferrable structured learning · WWW 2011 Enhancing diversity, coverage and balance for summarization through structure learning · WWW 2009 |
Information retrieval
retrieval models |
0.2 | 3 | 2008 | Learning to rank with ties · SIGIR 2008 Optimizing web search using social annotations · WWW 2007 Multi-model similarity propagation and its application for web image retrieval · ACM Multimedia 2004 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 2 | 2010 | Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA · IEEE Trans. Knowl. Data Eng. 2010 Topic-bridged PLSA for cross-domain text classification · SIGIR 2008 |
Recommender systems
click-through rate prediction |
0.2 | 1 | 2014 | Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014 |
Information retrieval › evaluation › effectiveness metrics
discounted cumulative gain |
0.2 | 1 | 2014 | Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014 |
Information retrieval › online advertising
display advertising |
0.2 | 1 | 2014 | Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014 |
Information retrieval
evaluation |
0.2 | 1 | 2014 | Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014 |
Data mining › dimensionality reduction
feature selection |
0.2 | 1 | 2014 | Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014 |
Data mining › dimensionality reduction › feature selection
group lasso |
0.2 | 1 | 2014 | Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014 |
Machine learning and data management
metric learning |
0.2 | 1 | 2014 | Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014 |
Information retrieval › retrieval evaluation
ranking evaluation |
0.2 | 1 | 2014 | Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014 |
Information retrieval
web search |
0.2 | 3 | 2007 | Optimizing web search using social annotations · WWW 2007 Exploiting the hierarchical structure for link analysis · SIGIR 2005 Implicit link analysis for small web search · SIGIR 2003 |
Data mining › predictive modeling
classification |
0.2 | 2 | 2009 | Web-scale classification with naive bayes · WWW 2009 Co-clustering based classification for out-of-domain documents · KDD 2007 |
Data mining › text mining
text classification |
0.2 | 2 | 2008 | Deep classification in large-scale text hierarchies · SIGIR 2008 Topic-bridged PLSA for cross-domain text classification · SIGIR 2008 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2011 | Heterogeneous Transfer Learning for Image Classification · AAAI 2011 |
Information retrieval › text summarization
video summarization |
0.1 | 1 | 2011 | Video summarization via transferrable structured learning · WWW 2011 |
Natural language and speech › Information extraction and text analysis › topic model
probabilistic latent semantic analysis |
0.1 | 1 | 2010 | Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA · IEEE Trans. Knowl. Data Eng. 2010 |
Information retrieval › online advertising
behavioral targeting |
0.1 | 1 | 2010 | Transfer learning for behavioral targeting · WWW 2010 |
Information retrieval › online advertising
contextual advertising |
0.1 | 1 | 2010 | Visual Contextual Advertising: Bringing Textual Advertisements to Images · AAAI 2010 |
Information retrieval
cross-modal retrieval |
0.1 | 1 | 2010 | Visual Contextual Advertising: Bringing Textual Advertisements to Images · AAAI 2010 |
Information retrieval
multimedia analysis and retrieval |
0.1 | 1 | 2010 | Visual Contextual Advertising: Bringing Textual Advertisements to Images · AAAI 2010 |
Machine learning and data management › weak supervision
positive-unlabeled learning |
0.1 | 1 | 2010 | Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA · IEEE Trans. Knowl. Data Eng. 2010 |
Information retrieval › web search
link analysis |
0.1 | 2 | 2005 | Exploiting the hierarchical structure for link analysis · SIGIR 2005 Implicit link analysis for small web search · SIGIR 2003 |
Information retrieval › ranking
ranking algorithms |
0.1 | 2 | 2005 | Exploiting the hierarchical structure for link analysis · SIGIR 2005 Implicit link analysis for small web search · SIGIR 2003 |
Machine learning › Transfer learning and domain adaptation
cross-domain learning |
0.1 | 1 | 2009 | EigenTransfer: a unified framework for transfer learning · ICML 2009 |
Machine learning › Transfer learning and domain adaptation › knowledge transfer
self-taught learning |
0.1 | 1 | 2009 | EigenTransfer: a unified framework for transfer learning · ICML 2009 |
Methods — techniques the papers use, named apart from their topics
transfer learning · 0.5language model · 0.4linear utility learning · 0.4active learning · 0.4structured SVM · 0.3generative model · 0.2logistic regression · 0.2feature hashing · 0.2distributed implementation · 0.2co-clustering · 0.2matrix factorization · 0.1latent semantic features · 0.1constrained optimization · 0.1EM algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Topic Modeling in Semantic Space with KeywordsabstractA common and convenient approach for user to describe his information need is to provide a set of keywords. Therefore, the technique to understand the need becomes crucial. In this paper, for the information need about a topic or category, we propose a novel method called TDCS(Topic Distilling with Compressive Sensing) for explicit and accurate modeling the topic implied by several keywords. The task is transformed as a topic reconstruction problem in the semantic space with a reasonable intuition that the topic is sparse in the semantic space. The latent semantic space could be mined from documents via unsupervised methods, e.g. LSI. Compressive sensing is leveraged to obtain a sparse representation from only a few keywords. In order to make the distilled topic more robust, an iterative learning approach is adopted. The experiment results show the effectiveness of our method. Moreover, with only a few semantic concepts remained for the topic, our method is efficient for subsequent text mining tasks. Xiaojia Pu, Rong Jin 0001, Gangshan Wu, Dingyi Han, Gui-Rong Xue |
CIKM | 5 |
| 2014 | Coupled Group Lasso for Web-Scale CTR Prediction in Display AdvertisingabstractIn display advertising, click through rate(CTR) prediction is the problem of estimating the probability that an advertisement (ad) is clicked when displayed to a user in a specific context. Due to its easy implementation and promising performance, logistic regression(LR) model has been widely used for CTR prediction, especially in industrial systems. However, it is not easy for LR to capture the nonlinear information, such as the conjunction information, from user features and ad features. In this paper, we propose a novel model, called coupled group lasso(CGL), for CTR prediction in display advertising. CGL can seamlessly integrate the conjunction information from user features and ad features for modeling. Furthermore, CGL can automatically eliminate useless features for both users and ads, which may facilitate fast online prediction. Scalability of CGL is ensured through feature hashing and distributed implementation. Experimental results on real-world data sets show that our CGL model can achieve state-of-the-art performance on web-scale CTR prediction tasks. Wu-Jun Li, Gui-Rong Xue, Dingyi Han |
ICML | 3 |
| 2014 | Learning the Gain Values and Discount Factors of Discounted Cumulative GainsabstractEvaluation metric is an essential and integral part of a ranking system. In the past, several evaluation metrics have been proposed in information retrieval and web search, among them Discounted Cumulative Gain (DCG) has emerged as one that is widely adopted for evaluating the performance of ranking functions used in web search. However, the two sets of parameters, the gain values and discount factors, used in DCG are usually determined in a rather ad-hoc way, and their impacts have not been carefully analyzed. In this paper, we first show that DCG is generally not coherent, i.e., comparing the performance of ranking functions using DCG very much depends on the particular gain values and discount factors used. We then propose a novel methodology that can learn the gain values and discount factors from user preferences over rankings, modeled as a special case of learning linear utility functions. We also discuss how to extend our methods to handle tied preference pairs and how to explore active learning to reduce preference labeling. Numerical simulations illustrate the effectiveness of our proposed methods. Moreover, experiments are also conducted over a side-by-side comparison data set from a commercial search engine to validate the proposed methods on real-world data. Ke Zhou 0002, Hongyuan Zha, Yi Chang 0001, Gui-Rong Xue |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | Multi-task learning to rank for web search
Yi Chang 0001, Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Zhaohui Zheng 0001 |
Pattern Recognit. Lett. | 4 |
| 2012 | Advertising Keywords Recommendation for Short-Text Web Pages Using WikipediaabstractAdvertising keywords recommendation is an indispensable component for online advertising with the keywords selected from the target Web pages used for contextual advertising or sponsored search. Several ranking-based algorithms have been proposed for recommending advertising keywords. However, for most of them performance is still lacking, especially when dealing with short-text target Web pages, that is, those containing insufficient textual information for ranking. In some cases, short-text Web pages may not even contain enough keywords for selection. A natural alternative is then to recommend relevant keywords not present in the target Web pages. In this article, we propose a novel algorithm for advertising keywords recommendation for short-text Web pages by leveraging the contents of Wikipedia, a user-contributed online encyclopedia. Wikipedia contains numerous entities with related entities on a topic linked to each other. Given a target Web page, we propose to use a content-biased PageRank on the Wikipedia graph to rank the related entities. Furthermore, in order to recommend high-quality advertising keywords, we also add an advertisement-biased factor into our model. With these two biases, advertising keywords that are both relevant to a target Web page and valuable for advertising are recommended. In our experiments, several state-of-the-art approaches for keyword recommendation are compared. The experimental results demonstrate that our proposed approach produces substantial improvement in the precision of the top 20 recommended keywords on short-text Web pages over existing approaches. Weinan Zhang 0001, Dingquan Wang, Gui-Rong Xue, Hongyuan Zha |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2012 | Leveraging Auxiliary Data for Learning to RankabstractIn learning to rank, both the quality and quantity of the training data have significant impacts on the performance of the learned ranking functions. However, in many applications, there are usually not sufficient labeled training data for the construction of an accurate ranking model. It is therefore desirable to leverage existing training data from other tasks when learning the ranking function for a particular task, an important problem which we tackle in this article utilizing a boosting framework withtransfer learning. In particular, we propose to adaptively learn transferable representations called super-features from the training data of both the target task and the auxiliary task. Those super-features and the coefficients for combining them are learned in an iterative stage-wise fashion. Unlike previous transfer learning methods, the super-features can be adaptively learned by weak learners from the data. Therefore, the proposed framework is sufficiently flexible to deal with complicated common structures among different learning tasks. We evaluate the performance of the proposed transfer learning method for two datasets from the Letor collection and one dataset collected from a commercial search engine, and we also compare our methods with several existing transfer learning methods. Our results demonstrate that the proposed method can enhance the ranking functions of the target tasks utilizing the training data from the auxiliary tasks. Ke Zhou 0002, Hongyuan Zha, Gui-Rong Xue |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2011 | Heterogeneous Transfer Learning for Image ClassificationabstractTransfer learning as a new machine learning paradigm has gained increasing attention lately. In situations where the training data in a target domain are not sufficient to learn predictive models effectively, transfer learning leverages auxiliary source data from other related source domains for learning. While most of the existing works in this area only focused on using the source data with the same structure as the target data, in this paper, we push this boundary further by proposing a heterogeneous transfer learning framework for knowledge transfer between text and images. We observe that for a target-domain classification problem, some annotated images can be found on many social Web sites, which can serve as a bridge to transfer knowledge from the abundant text documents available over the Web. A key question is how to effectively transfer the knowledge in the source data even though the text can be arbitrarily found. Our solution is to enrich the representation of the target images with semantic concepts extracted from the auxiliary source data through a novel matrix factorization method. By using the latent semantic features generated by the auxiliary data, we are able to build a better integrated image classifier. We empirically demonstrate the effectiveness of our algorithm on the Caltech-256 image dataset. Yuqiang Chen, Zhongqi Lu, Sinno Jialin Pan, Gui-Rong Xue, Yong Yu 0001, Qiang Yang 0001 |
AAAI | 5 |
| 2011 | Cross-Lingual Sentiment Classification via Bi-view Non-negative Matrix Tri-Factorization
Junfeng Pan, Gui-Rong Xue, Yong Yu 0001, Yang Wang 0019 |
PAKDD (1) | 2 |
| 2011 | Video summarization via transferrable structured learningabstractIt is well-known that textual information such as video transcripts and video reviews can significantly enhance the performance of video summarization algorithms. Unfortunately, many videos on the Web such as those from the popular video sharing site YouTube do not have useful textual information. The goal of this paper is to propose a transfer learning framework for video summarization: in the training process both the video features and textual features are exploited to train a summarization algorithm while for summarizing a new video only its video features are utilized. The basic idea is to explore the transferability between videos and their corresponding textual information. Based on the assumption that video features and textual features are highly correlated with each other, we can transfer textual information into knowledge on summarization using video information only. In particular, we formulate the video summarization problem as that of learning a mapping from a set of shots of a video to a subset of the shots using the general framework of SVM-based structured learning. Textual information is transferred by encoding them into a set of constraints used in the structured learning process which tend to provide a more detailed and accurate characterization of the different subsets of shots. Experimental results show significant performance improvement of our approach and demonstrate the utility of textual information for enhancing video summarization. Liangda Li, Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Yong Yu 0001 |
WWW | 3 |
| 2010 | Visual Contextual Advertising: Bringing Textual Advertisements to ImagesabstractAdvertising in the case of textual Web pages has been studied extensively by many researchers. However, with the increasing amount of multimedia data such as image, audio and video on the Web, the need for recommending advertisement for the multimedia data is becoming a reality. In this paper, we address the novel problem of visual contextual advertising, which is to directly advertise when users are viewing images which do not have any surrounding text. A key challenging issue of visual contextual advertising is that images and advertisements are usually represented in image space and word space respectively, which are quite different with each other inherently. As a result, existing methods for Web page advertising are inapplicable since they represent both Web pages and advertisement in the same word space. In order to solve the problem, we propose to exploit the social Web to link these two feature spaces together. In particular, we present a unified generative model to integrate advertisements, words and images. Specifically, our solution combines two parts in a principled approach: First, we transform images from a image feature space to a word space utilizing the knowledge from images with annotations from social Web. Then, a language model based approach is applied to estimate the relevance between transformed images and advertisements. Moreover, in this model, the probability of recommending an advertisement can be inferred efficiently given an image, which enables potential applications to online advertising. Yuqiang Chen, Ou Jin, Gui-Rong Xue, Qiang Yang 0001 |
AAAI | 3 |
| 2010 | Predicting Product Duration for Adaptive Advertisement
Zhongqi Guo, Gui-Rong Xue, Yong Yu 0001 |
ADMA (2) | 3 |
| 2010 | Click Prediction for Product Search on C2C Web Sites
Xiangzhi Wang, Gui-Rong Xue, Yong Yu 0001 |
ADMA (2) | 3 |
| 2010 | Text-Aided Image Classification: Using Labeled Text from Web to Help Image ClassificationabstractAs more and more multimedia data become available on the Web, mining on those data is playing an increasingly important role in Web applications. In this paper, we investigate the interplay between multimedia data mining and text data mining. Specifically, in an approach we called text-aided image classification (TAIC), we address the problem of image classification with very limited amount of labeled images and a large amount of auxiliary labeled text data. This problem is important in practice, since currently on the Web, labeled text data are usually much more than image data. To solve the problem, based on the “bag-of-words” view and the Naive Bayes classification model, we focus our attention on the estimation of the image feature distribution under given concept. We extend the Naive Bayes algorithm by considering a mapping that maps the most discriminative text features into the image feature space. This feature mapping is estimated based on the text-image cooccurrence data on the Web, acting like a bridge that connects text and image knowledge. With this process, we estimate target image feature distribution from a text model based on sufficient labeled data. Our empirical results on real world data sets show that our method makes a good approximation of the image feature distribution when trained with abundant labeled images. In the case amount of labeled images is very limited, the classification performance is improved by using auxiliary labeled text data, which shows that our method can indeed integrate text and image knowledge in a simple yet effective way. Yuqiang Chen, Gui-Rong Xue, Yong Yu 0001 |
APWeb | 3 |
| 2010 | Transfer learning for behavioral targetingabstractRecently, Behavioral Targeting (BT) is attracting much attention from both industry and academia due to its rapid growth in online advertising market. Though a basic assumption of BT, which is, the users who share similar Web browsing behaviors will have similar preference over ads, has been empirically verified, we argue that the users' ad click preference and Web browsing behavior are not reflecting the same user intent though they are correlated. In this paper, we propose to formulate BT as a transfer learning problem. We treat the users' preference over ads and Web browsing behaviors as two different user behavioral domains and propose to utilize transfer learning strategy across these two user behavioral domains to segment users for BT ads delivery. We show that some classical BT solutions could be formulated in transfer learning view. As an example, we propose to leverage translated learning, which is a recent proposed transfer learning algorithm, to benefit the BT ads delivery. Experimental results on real ad click data show that, BT user segmentation by the approach of transfer learning can outperform the classical user segmentation strategies for larger than 20% in terms of smoothed ad Click Through Rate(CTR). Jun Yan 0001, Gui-Rong Xue, Zheng Chen 0001 |
WWW | 3 |
| 2010 | Knowledge transfer for cross domain learning to rank
Depin Chen, Yan Xiong 0001, Jun Yan 0001, Gui-Rong Xue, Gang Wang 0010, Zheng Chen 0001 |
Inf. Retr. | 4 |
| 2010 | Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSAabstractIt is often difficult and time-consuming to provide a large amount of positive and negative examples for training a classification system in many applications such as information retrieval. Instead, users often find it easier to indicate just a few positive examples of what he or she likes, and thus, these are the only labeled examples available for the learning system. A large amount of unlabeled data are easier to obtain. How to make use of the positive and unlabeled data for learning is a critical problem in machine learning and information retrieval. Several approaches for solving this problem have been proposed in the past, but most of these methods do not work well when only a small amount of labeled positive data are available. In this paper, we propose a novel algorithm called Topic-Sensitive pLSA to solve this problem. This algorithm extends the original probabilistic latent semantic analysis (pLSA), which is a purely unsupervised framework, by injecting a small amount of supervision information from the user. The supervision from users is in the form of indicating which documents fit the users' interests. The supervision is encoded into a set of constraints. By introducing the penalty terms for these constraints, we propose an objective function that trades off the likelihood of the observed data and the enforcement of the constraints. We develop an iterative algorithm that can obtain the local optimum of the objective function. Experimental evaluation on three data corpora shows that the proposed method can improve the performance especially only with a small amount of labeled positive data. Ke Zhou 0002, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | Heterogeneous Transfer Learning for Image Clustering via the SocialWeb
Qiang Yang 0001, Yuqiang Chen, Gui-Rong Xue, Wenyuan Dai, Yong Yu 0001 |
ACL/IJCNLP | 3 |
| 2009 | Multi-task learning for learning to rank in web searchabstractBoth the quality and quantity of training data have significant impact on the performance of ranking functions in the context of learning to rank for web search. Due to resource constraints, training data for smaller search engine markets are scarce and we need to leverage existing training data from large markets to enhance the learning of ranking function for smaller markets. In this paper, we present a boosting framework for learning to rank in the multi-task learning context for this purpose. In particular, we propose to learn non-parametric common structures adaptively from multiple tasks in a stage-wise way. An algorithm is developed to iteratively discover super-features that are effective for all the tasks. The estimation of the functions for each task is then learned as a linear combination of those super-features. We evaluate the performance of this multi-task learning method for web search ranking using data from a search engine. Our results demonstrate that multi-task learning methods bring significant relevance improvements over existing baseline methods. Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Gordon Sun, Belle L. Tseng, Zhaohui Zheng 0001, Yi Chang 0001 |
CIKM | 3 |
| 2009 | EigenTransfer: a unified framework for transfer learningabstractThis paper proposes a general framework, called EigenTransfer, to tackle a variety of transfer learning problems, e.g. cross-domain learning, self-taught learning, etc. Our basic idea is to construct a graph to represent the target transfer learning task. By learning the spectra of a graph which represents a learning task, we obtain a set of eigenvectors that reflect the intrinsic structure of the task graph. These eigenvectors can be used as the new features which transfer the knowledge from auxiliary data to help classify target data. Given an arbitrary non-transfer learner (e.g. SVM) and a particular transfer learning task, EigenTransfer can produce a transfer learner accordingly for the target transfer learning task. We apply EigenTransfer on three different transfer learning tasks, cross-domain learning, cross-category learning and self-taught learning, to demonstrate its unifying ability, and show through experiments that EigenTransfer can greatly outperform several representative non-transfer learners. Wenyuan Dai, Ou Jin, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
ICML | 3 |
| 2009 | Global ranking by exploiting user clicksabstractIt is now widely recognized that user interactions with search results can provide substantial relevance information on the documents displayed in the search results. In this paper, we focus on extracting relevance information from one source of user interactions, i.e., user click data, which records the sequence of documents being clicked and not clicked in the result set during a user search session. We formulate the problem as a global ranking problem, emphasizing the importance of the sequential nature of user clicks, with the goal to predict the relevance labels of all the documents in a search session. This is distinct from conventional learning to rank methods that usually design a ranking model defined on a single document; in contrast, in our model the relational information among the documents as manifested by an aggregation of user clicks is exploited to rank all the documents jointly. In particular, we adapt several sequential supervised learning algorithms, including the conditional random field (CRF), the sliding window method and the recurrent sliding window method, to the global ranking problem. Experiments on the click data collected from a commercial search engine demonstrate that our methods can outperform the baseline models for search results re-ranking. Shihao Ji 0001, Ke Zhou 0002, Ciya Liao, Zhaohui Zheng 0001, Gui-Rong Xue, Olivier Chapelle, Gordon Sun, Hongyuan Zha |
SIGIR | 5 |
| 2009 | Efficient query expansion for advertisement searchabstractOnline advertising represents a growing part of the revenues of ma-jor Internet service providers such as Google and Yahoo. A com-monly used strategy is to place advertisements (ads) on the search result pages according to the users ’ submitted queries. Relevant ads are likely to be clicked by a user and to increase the revenues of both advertisers and publishers. However, bid phrases defined by ad-owners are usually contained in limited number of ads. Directly matching user queries with bid phrases often results in finding few appropriate ads. To address this shortcoming, query expansion is often used to increase the chances to match the ads. Nevertheless, query expansion on top of the traditional inverted index faces ef-ficiency issues such as high time complexity and heavy I/O costs. Moreover, precision cannot always be improved, sometimes even hurt due to the involvement of additional noise. Haofen Wang, Linyun Fu, Gui-Rong Xue, Yong Yu 0001 |
SIGIR | 4 |
| 2009 | Enhancing diversity, coverage and balance for summarization through structure learningabstractDocument summarization plays an increasingly important role with the exponential growth of documents on the Web. Many supervised and unsupervised approaches have been proposed to generate summaries from documents. However, these approaches seldom simultaneously consider summary diversity, coverage, and balance issues which to a large extent determine the quality of summaries. In this paper, we consider extract-based summarization emphasizing the following three requirements: 1) diversity in summarization, which seeks to reduce redundancy among sentences in the summary; 2) sufficient coverage, which focuses on avoiding the loss of the document's main information when generating the summary; and 3) balance, which demands that different aspects of the document need to have about the same relative importance in the summary. We formulate the extract-based summarization problem as learning a mapping from a set of sentences of a given document to a subset of the sentences that satisfies the above three requirements. The mapping is learned by incorporating several constraints in a structure learning framework, and we explore the graph structure of the output variables and employ structural SVM for solving the resulted optimization problem. Experiments on the DUC2001 data sets demonstrate significant performance improvements in terms of F1 and ROUGE metrics. Liangda Li, Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Yong Yu 0001 |
WWW | 3 |
| 2009 | Dataplorer: a scalable search engine for the data webabstractMore and more structured information in the form of semantic data is nowadays available. It offers a wide range of new possibilities especially for semantic search and Web data integration. However, their effective exploitation still brings about a number of challenges, e.g. usability, scalability and uncertainty. In this paper, we present Dataplorer, a solution designed to address these challenges. We consider the usability through the use of hybrid queries and faceted search, while still preserving the scalability thanks to an extension of inverted index to support this type of query. Moreover, Dataplorer deals with uncertainty by means of a powerful ranking scheme to find relevant results. Our experimental results show that our proposed approach is promising and it makes us believe that it is possible to extend the current IR infrastructure to query and search the Web of data. Haofen Wang, Qiaoling Liu, Gui-Rong Xue, Yong Yu 0001, Lei Zhang 0007 |
WWW | 3 |
| 2009 | Web-scale classification with naive bayesabstractTraditional Naive Bayes Classifier performs miserably on web-scale taxonomies. In this paper, we investigate the reasons behind such bad performance. We discover that the low performance are not completely caused by the intrinsic limitations of Naive Bayes, but mainly comes from two largely ignored problems: contradiction pair problem and discriminative evidence cancelation problem. We propose modifications that can alleviate the two problems while preserving the advantages of Naive Bayes. The experimental results show our modified Naive Bayes can significantly improve the performance on real web-scale taxonomies. Congle Zhang, Gui-Rong Xue, Yong Yu 0001, Hongyuan Zha |
WWW | 2 |
| 2009 | User language model for collaborative personalized searchabstractTraditional personalized search approaches rely solely on individual profiles to construct a user model. They are often confronted by two major problems: data sparseness and cold-start for new individuals. Data sparseness refers to the fact that most users only visit a small portion of Web pages and hence a very sparse user-term relationship matrix is generated, while cold-start for new individuals means that the system cannot conduct any personalization without previous browsing history. Recently, community-based approaches were proposed to use the group's social behaviors as a supplement to personalization. However, these approaches only consider the commonality of a group of users and still cannot satisfy the diverse information needs of different users. In this article, we present a new approach, called collaborative personalized search. It considers not only the commonality factor among users for defining group user profiles and global user profiles, but also the specialties of individuals. Then, a statistical user language model is proposed to integrate the individual model, group user model and global user model together. In this way, the probability that a user will like a Web page is calculated through a two-step smoothing mechanism. First, a global user model is used to smooth the probability of unseen terms in the individual profiles and provide aggregated behavior of global users. Then, in order to precisely describe individual interests by looking at the behaviors of similar users, users are clustered into groups and group-user models are constructed. The group-user models are integrated into an overall model through a cluster-based language model. The behaviors of the group users can be utilized to enhance the performance of personalized search. This model can alleviate the two aforementioned problems and provide a more effective personalized search than previous approaches. Large-scale experimental evaluations are conducted to show that the proposed approach substantially improves the relevance of a search over several competitive methods. Gui-Rong Xue, Yong Yu 0001, Qiang Yang 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2008 | Iterative Reinforcement Cross-Domain Text Classification
Gui-Rong Xue, Yong Yu 0001 |
ADMA | 2 |
| 2008 | Knowledge Transferring Via Implicit Link Analysis
Wenyuan Dai, Gui-Rong Xue, Yong Yu 0001 |
DASFAA | 3 |
| 2008 | Self-taught clusteringabstractThis paper focuses on a new clustering task, called self-taught clustering. Self-taught clustering is an instance of unsupervised transfer learning, which aims at clustering a small collection of target unlabeled data with the help of a large amount of auxiliary unlabeled data. The target and auxiliary data can be different in topic distribution. We show that even when the target data are not sufficient to allow effective learning of a high quality feature representation, it is possible to learn the useful features with the help of the auxiliary data on which the target data can be clustered effectively. We propose a co-clustering based self-taught clustering algorithm to tackle this problem, by clustering the target and auxiliary data simultaneously to allow the feature representation from the auxiliary data to influence the target data through a common set of features. Under the new data representation, clustering on the target data can be improved. Our experiments on image clustering show that our algorithm can greatly outperform several state-of-the-art clustering methods when utilizing irrelevant unlabeled auxiliary data. Wenyuan Dai, Qiang Yang 0001, Gui-Rong Xue, Yong Yu 0001 |
ICML | 3 |
| 2008 | Spectral domain-transfer learningabstractTraditional spectral classification has been proved to be effective in dealing with both labeled and unlabeled data when these data are from the same domain. In many real world applications, however, we wish to make use of the labeled data from one domain (called in-domain) to classify the unlabeled data in a different domain (out-of-domain). This problem often happens when obtaining labeled data in one domain is difficult while there are plenty of labeled data from a related but different domain. In general, this is a transfer learning problem where we wish to classify the unlabeled data through the labeled data even though these data are not from the same domain. In this paper, we formulate this domain-transfer learning problem under a novel spectral classification framework, where the objective function is introduced to seek consistency between the in-domain supervision and the out-of-domain intrinsic structure. Through optimization of the cost function, the label information from the in-domain data is effectively transferred to help classify the unlabeled data from the out-of-domain. We conduct extensive experiments to evaluate our method and show that our algorithm achieves significant improvements on classification performance over many state-of-the-art algorithms. Wenyuan Dai, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
KDD | 3 |
| 2008 | Translated Learning: Transfer Learning across Different Feature SpacesabstractThis paper investigates a new machine learning strategy called translated learning. Unlike many previous learning tasks, we focus on how to use labeled data from one feature space to enhance the classification of other entirely different learning spaces. For example, we might wish to use labeled text data to help learn a model for classifying image data, when the labeled images are difficult to obtain. An important aspect of translated learning is to build a "bridge" to link one feature space (known as the "source space") to another space (known as the "target space") through a translator in order to migrate the knowledge from source to target. The translated learning solution uses a language model to link the class labels to the features in the source spaces, which in turn is translated to the features in the target spaces. Finally, this chain of linkages is completed by tracing back to the instances in the target spaces. We show that this path of linkage can be modeled using a Markov chain and risk minimization. Through experiments on the text-aided image classification and cross-language classification tasks, we demonstrate that our translated learning framework can greatly outperform many state-of-the-art baseline methods. Wenyuan Dai, Yuqiang Chen, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
NIPS | 3 |
| 2008 | Knowledge Supervised Text Classification with No Labeled Documents
Congle Zhang, Gui-Rong Xue, Yong Yu 0001 |
PRICAI | 2 |
| 2008 | Topic-bridged PLSA for cross-domain text classificationabstractIn many Web applications, such as blog classification and new-sgroup classification, labeled data are in short supply. It often happens that obtaining labeled data in a new domain is expensive and time consuming, while there may be plenty of labeled data in a related but different domain. Traditional text classification ap-proaches are not able to cope well with learning across different domains. In this paper, we propose a novel cross-domain text classification algorithm which extends the traditional probabilistic latent semantic analysis (PLSA) algorithm to integrate labeled and unlabeled data, which come from different but related domains, into a unified probabilistic model. We call this new model Topic-bridged PLSA, or TPLSA. By exploiting the common topics between two domains, we transfer knowledge across different domains through a topic-bridge to help the text classification in the target domain. A unique advantage of our method is its ability to maximally mine knowledge that can be transferred between domains, resulting in superior performance when compared to other state-of-the-art text classification approaches. Experimental eval-uation on different kinds of datasets shows that our proposed algorithm can improve the performance of cross-domain text classification significantly. Gui-Rong Xue, Wenyuan Dai, Qiang Yang 0001, Yong Yu 0001 |
SIGIR | 1 |
| 2008 | Deep classification in large-scale text hierarchiesabstractMost classification algorithms are best at categorizing the Web documents into a few categories, such as the top two levels in the Open Directory Project. Such a classification method does not give very detailed topic-related class information for the user because the first two levels are often too coarse. However, classification on a large-scale hierarchy is known to be intractable for many target categories with cross-link relationships among them. In this paper, we propose a novel deep-classification approach to categorize Web documents into categories in a large-scale taxonomy. The approach consists of two stages: a search stage and a classification stage. In the first stage, a category-search algorithm is used to acquire the category candidates for a given document. Based on the category candidates, we prune the large-scale hierarchy to focus our classification effort on a small subset of the original hierarchy. As a result, the classification model is trained on the small subset before being applied to assign the category for a new document. Since the category candidates are sufficiently close to each other in the hierarchy, a statistical-language-model based classifier using n-gram features is exploited. Furthermore, the structure of the taxonomy can be utilized in this stage to improve the performance of classification. We demonstrate the performance of our proposed algorithms on the Open Directory Project with over 130,000 categories. Experimental results show that our proposed approach can reach 51.8% on the measure of Mi-F1 at the 5th level, which is 77.7% improvement over top-down based SVM classification algorithms. Gui-Rong Xue, Dikan Xing, Qiang Yang 0001, Yong Yu 0001 |
SIGIR | 1 |
| 2008 | Learning to rank with tiesabstractDesigning effective ranking functions is a core problem for information retrieval and Web search since the ranking functions directly impact the relevance of the search results. The problem has been the focus of much of the research at the intersection of Web search and machine learning, and learning ranking functions from preference data in particular has recently attracted much interest. The objective of this paper is to empirically examine several objective functions that can be used for learning ranking functions from preference data. Specifically, we investigate the roles of ties in the learning process. By ties, we mean preference judgments that two documents have equal degree of relevance with respect to a query. This type of data has largely been ignored or not properly modeled in the past. In this paper, we analyze the properties of ties and develop novel learning frameworks which combine ties and preference data using statistical paired comparison models to improve the performance of learned ranking functions. The resulting optimization problems explicitly incorporating ties and preference data are solved using gradient boosting methods. Experimental studies are conducted using three publicly available data sets which demonstrate the effectiveness of the proposed new methods. Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Yong Yu 0001 |
SIGIR | 2 |
| 2008 | Advertising keyword suggestion based on concept hierarchyabstractThe increasing growth of the World Wide Web constantly enlarges the revenue generated by search engine advertising. Advertisers bid on keywords associated with their products to display their ads on the search result pages. Keyword suggestion methods are proposed to fill the gap between the keywords chosen by advertisers and the popular queries, through finding new relevant keywords according to some statistical information (for example, the keyword co-occurrence). However, there is little effort taking semantic information, such as concept hierarchy, into account. In this paper, we propose a novel keyword suggestion method that fully exploits the semantic knowledge among concept hierarchy. Given a keyword, we first match it with some relevant concepts. Then the relevant concepts are used with their hierarchy to fertilize the meanings of the keywords. Finally new keywords are suggested according to the concept information rather than the statistical co-occurrence of the keyword itself. Experimental results show that our proposed method can successfully provide suggestion that meets the accuracy and coverage requirements Gui-Rong Xue, Yong Yu 0001 |
WSDM | 2 |
| 2008 | Deep classifier: automatically categorizing search results into large-scale hierarchiesabstractOrganizing Web search results into hierarchical categories facilitates users' browsing through Web search results, especially for ambiguous queries where the potential results are mixed together. Previous methods on search result classification are usually based on pre-training a classification model on some fixed and shallow hierarchical categories, where only the top-two-level categories of a Web taxonomy is used. Such classification methods may be too coarse for users to browse, since most search results would be classified into only two or three shallow categories. Instead, a deep hierarchical classifier must provide many more categories. However, the performance of such classifiers is usually limited because their classification effectiveness can deteriorate rapidly at the third or fourth level of a hierarchy. In this paper, we propose a novel algorithm known as Deep Classifier to classify the search results into detailed hierarchical categories with higher effectiveness than previous approaches. Given the search results in response to a query, the algorithm first prunes a wide-ranged hierarchy into a narrow one with the help of some Web directories. Different strategies are proposed to select the training data by utilizing the hierarchical structures. Finally, a discriminative naíve Bayesian classifier is developed to perform efficient and effective classification. As a result, the algorithm can provide more meaningful and specific class labels for search result browsing than shallow style of classification. We conduct experiments to show that the Deep Classifier can achieve significant improvement over state-of-the-art algorithms. In addition, with sufficient off-line preparation, the efficiency of the proposed algorithm is suitable for online application Dikan Xing, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
WSDM | 2 |
| 2008 | WWW 2008 workshop on social web search and mining: SWSM2008abstractNo abstract available. Juan-Zi Li, Gui-Rong Xue, Jie Tang 0001, Ying Ding 0001 |
WWW | 2 |
| 2008 | Can chinese web pages be classified with english data source?abstractAs the World Wide Web in China grows rapidly, mining knowledge in Chinese Web pages becomes more and more important. Mining Web information usually relies on the machine learning techniques which require a large amount of labeled data to train credible models. Although the number of Chinese Web pages increases quite fast, it still lacks Chinese labeled data. However, there are relatively sufficient English labeled Web pages. These labeled data, though in different linguistic representations, share a substantial amount of semantic information with Chinese ones, and can be utilized to help classify Chinese Web pages. In this paper, we propose an information bottleneck based approach to address this cross-language classification problem. Our algorithm first translates all the Chinese Web pages to English. Then, all the Web pages, including Chinese and English ones, are encoded through an information bottleneck which can allow only limited information to pass. Therefore, in order to retain as much useful information as possible, the common part between Chinese and English Web pages is inclined to be encoded to the same code (i.e. class label), which makes the cross-language classification accurate. We evaluated our approach using the Web pages collected from Open Directory Project (ODP). The experimental results show that our method significantly improves several existing supervised and semi-supervised classifiers. Gui-Rong Xue, Wenyuan Dai, Qiang Yang 0001, Yong Yu 0001 |
WWW | 2 |
| 2007 | Transferring Naive Bayes Classifiers for Text Classification
Wenyuan Dai, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
AAAI | 2 |
| 2007 | Boosting for transfer learningabstractTraditional machine learning makes a basic assumption: the training and test data should be under the same distribution. However, in many cases, this identical-distribution assumption does not hold. The assumption might be violated when a task from one new domain comes, while there are only labeled data from a similar old domain. Labeling the new data can be costly and it would also be a waste to throw away all the old data. In this paper, we present a novel transfer learning framework called TrAdaBoost, which extends boosting-based learning algorithms (Freund & Schapire, 1997). TrAdaBoost allows users to utilize a small amount of newly labeled data to leverage the old data to construct a high-quality classification model for the new data. We show that this method can allow us to learn an accurate model using only a tiny amount of new data and a large amount of old data, even when the new data are not sufficient to train a model alone. We show that TrAdaBoost allows knowledge to be effectively transferred from the old data to the new. The effectiveness of our algorithm is analyzed theoretically and empirically to show that our iterative algorithm can converge well to an accurate model. Wenyuan Dai, Qiang Yang 0001, Gui-Rong Xue, Yong Yu 0001 |
ICML | 3 |
| 2007 | Co-clustering based classification for out-of-domain documentsabstractIn many real world applications, labeled data are in short supply. It often happens that obtaining labeled data in a new domain is expensive and time consuming, while there may be plenty of labeled data from a related but different domain. Traditional machine learning is not able to cope well with learning across different domains. In this paper, we address this problem for a text-mining task, where the labeled data are under one distribution in one domain known as in-domain data, while the unlabeled data are under a related but different domain known as out-of-domain data. Our general goal is to learn from the in-domain and apply the learned knowledge to out-of-domain. We propose a co-clustering based classification (CoCC) algorithm to tackle this problem. Co-clustering is used as a bridge to propagate the class structure and knowledge from the in-domain to the out-of-domain. We present theoretical and empirical analysis to show that our algorithm is able to produce high quality classification results, even when the distributions between the two data are different. The experimental results show that our algorithm greatly improves the classification performance over the traditional learning algorithms. Wenyuan Dai, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001 |
KDD | 2 |
| 2007 | Bridged Refinement for Transfer Learning
Dikan Xing, Wenyuan Dai, Gui-Rong Xue, Yong Yu 0001 |
PKDD | 3 |
| 2007 | Adaptive Email Spam Filtering Based on Information Theory
Wenyuan Dai, Gui-Rong Xue, Yong Yu 0001 |
WISE | 3 |
| 2007 | Optimizing web search using social annotationsabstractThis paper explores the use of social annotations to improve web search. Nowadays, many services, e.g. del.icio.us, have been developed for web users to organize and share their favorite web pages on line by using social annotations. We observe that the social annotations can benefit web search in two aspects: 1) the annotations are usually good summaries of corresponding web pages; 2) the count of annotations indicates the popularity of web pages. Two novel algorithms are proposed to incorporate the above information into page ranking: 1) SocialSimRank (SSR) calculates the similarity between social annotations and web queries; 2) SocialPageRank (SPR) captures the popularity of web pages. Preliminary experimental results show that SSR can find the latent semantic association between queries and annotations, while SPR successfully measures the quality (popularity) of a web page from the web users ’ perspective. We further evaluate the proposed methods empirically with 50 manually constructed queries and 3000 auto-generated queries on a dataset crawled from del.icio.us. Experiments show that both SSR and SPR benefit web search significantly. Shenghua Bao, Gui-Rong Xue, Xiaoyuan Wu, Yong Yu 0001, Ben Fei, Zhong Su |
WWW | 2 |
| 2007 | Exploring in the weblog space by detecting informative and affective articlesabstractWeblogs have become a prevalent source of information for people to express themselves. In general, there are two genres of contents in weblogs. The first kind is about the webloggers' personal feelings, thoughts or emotions. We call this kind of weblogs affective articles. The second kind of weblogs is about technologies and different kinds of informative news. In this paper, we present a machine learning method for classifying informative and affective articles among weblogs. We consider this problem as a binary classification problem. By using machine learning approaches, we achieve about 92% on information retrieval performance measures including precision, recall and F1. We set up three studies on the applications of above classification approach in both research and industrial fields. The above classification approach is used to improve the performance of classification of emotions from weblog articles. We also develop an intent-driven weblog-search engine based on the classification techniques to improve the satisfaction of Web users. Finally, our approach is applied to search for weblogs with a great deal of informative articles. Xiaochuan Ni 0001, Gui-Rong Xue, Yong Yu 0001, Qiang Yang 0001 |
WWW | 2 |
| 2006 | A Hierarchical Model of Web Graph
Yong Yu 0001, Dingyi Han, Gui-Rong Xue |
ADMA | 5 |
| 2006 | IRFCF: Iterative Rating Filling Collaborative Filtering Algorithm
Gui-Rong Xue, Fan-De Zhu, Ai-Guo Yao |
APWeb | 3 |
| 2006 | Image Description Mining and Hierarchical Clustering on Data Records Using HR-Tree
Congle Zhang, Gui-Rong Xue, Yong Yu 0001 |
APWeb | 3 |
| 2006 | An Empirical Study of Data Smoothing Methods for Memory-Based and Hybrid Collaborative Filtering
Dingyi Han, Gui-Rong Xue, Yong Yu 0001 |
PRICAI | 2 |
| 2006 | A Novel Web Page Categorization Algorithm Based on Block Propagation Using Query-Log Information
Wenyuan Dai, Yong Yu 0001, Congle Zhang, Gui-Rong Xue |
WAIM | 5 |
| 2006 | Exploiting Rating Behaviors for Effective Collaborative Filtering
Dingyi Han, Yong Yu 0001, Gui-Rong Xue |
WISE | 3 |
| 2006 | Reinforcing Web-object Categorization Through Interrelationships
Gui-Rong Xue, Yong Yu 0001, Dou Shen, Qiang Yang 0001, Hua-Jun Zeng, Zheng Chen 0001 |
Data Min. Knowl. Discov. | 1 |
| 2006 | TSSP: Multi-features based reinforcement algorithm to find related papers
Shen Huang, Yong Yu 0001, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Wei-Ying Ma |
Web Intell. Agent Syst. | 3 |
| 2005 | An Algorithm for Enumerating SCCs in Web Graph
Yong Yu 0001, Gui-Rong Xue |
APWeb | 4 |
| 2005 | Block-Based Language Modeling Approach Towards Web Search
Shengping Li, Shen Huang, Gui-Rong Xue, Yong Yu 0001 |
APWeb | 3 |
| 2005 | Using Probabilistic Latent Semantic Analysis for Personalized Web Search
Gui-Rong Xue, Hua-Jun Zeng, Yong Yu 0001 |
APWeb | 2 |
| 2005 | China Web Graph Measurements and Evolution
Yong Yu 0001, Gui-Rong Xue |
APWeb | 4 |
| 2005 | Scalable collaborative filtering using cluster-based smoothingabstractMemory-based approaches for collaborative filtering identify the similarity between two users by comparing their ratings on a set of items. In the past, the memory-based approach has been shown to suffer from two fundamental problems: data sparsity and difficulty in scalability. Alternatively, the model-based approach has been proposed to alleviate these problems, but this approach tends to limit the range of users. In this paper, we present a novel approach that combines the advantages of these two approaches by introducing a smoothing-based method. In our approach, clusters generated from the training data provide the basis for data smoothing and neighborhood selection. As a result, we provide higher accuracy as well as increased efficiency in recommendations. Empirical studies on two datasets (EachMovie and MovieLens) show that our new proposed approach consistently outperforms other state-of-art collaborative filtering algorithms. Gui-Rong Xue, Qiang Yang 0001, Wensi Xi, Hua-Jun Zeng, Yong Yu 0001, Zheng Chen 0001 |
SIGIR | 1 |
| 2005 | Exploiting the hierarchical structure for link analysisabstractLink analysis algorithms have been extensively used in Web information retrieval. However, current link analysis algorithms generally work on a flat link graph, ignoring the hierarchal structure of the Web graph. They often suffer from two problems: the sparsity of link graph and biased ranking of newly-emerging pages. In this paper, we propose a novel ranking algorithm called Hierarchical Rank as a solution to these two problems, which considers both the hierarchical structure and the link structure of the Web. In this algorithm, Web pages are first aggregated based on their hierarchical structure at directory, host or domain level and link analysis is performed on the aggregated graph. Then, the importance of each node on the aggregated graph is distributed to individual pages belong to the node based on the hierarchical structure. This algorithm allows the importance of linked Web pages to be distributed in the Web page space even when the space is sparse and contains new pages. Experimental results on the .GOV collection of TREC 2003 and 2004 show that hierarchical ranking algorithm consistently outperforms other well-known ranking algorithms, including the PageRank, BlockRank and LayerRank. In addition, experimental results show that link aggregation at the host level is much better than link aggregation at either the domain or directory levels. Gui-Rong Xue, Qiang Yang 0001, Hua-Jun Zeng, Yong Yu 0001, Zheng Chen 0001 |
SIGIR | 1 |
| 2005 | Interactive Chinese Search Results Clustering for Personalization
Gui-Rong Xue, Shen Huang, Yong Yu 0001 |
WAIM | 2 |
| 2005 | Importance-Based Web Page Classification Using Cost-Sensitive SVM
Gui-Rong Xue, Yong Yu 0001, Hua-Jun Zeng |
WAIM | 2 |
| 2004 | Optimizing web search using web click-through dataabstractThe performance of web search engines may often deteriorate due to the diversity and noisy information contained within web pages. User click-through data can be used to introduce more accurate description (metadata) for web pages, and to improve the search performance. However, noise and incompleteness, sparseness, and the volatility of web pages and are three major challenges for research work on user click-through log mining. In this paper, we propose a novel iterative reinforced algorithm to utilize the user click-through data to improve search performance. The algorithm fully explores the interrelations between and web pages, and effectively finds virtual queries for web pages and overcomes the challenges discussed above. Experiment results on a large set of MSN click-through log data show a significant improvement on search performance over the naive query log mining algorithm as well as the baseline search engine. Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Weiguo Fan |
CIKM | 1 |
| 2004 | MRSSA: an iterative algorithm for similarity spreading over interrelated objectsabstractWe introduce the Multiple Relationship Similarity Spreading Algorithm (MRSSA) to enhance IR effectiveness. This method has similarity computed in an iterative "spreading" fashion for multiple object types, combining both inter- and intra-object relationships. We demonstrate the value of this approach in the context of the WWW, where the key objects are web pages and queries, Relationships considered are derived from hyperlinks (in- and out-links) and click-through logs. Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Edward A. Fox |
CIKM | 1 |
| 2004 | IRC: An Iterative Reinforcement Categorization Algorithm for Interrelated Web ObjectsabstractMost existing categorization algorithms deal with homogeneous Web data objects, and consider interrelated objects as additional features when taking the interrelationships with other types of objects into account. However, focusing on any single aspects of these interrelationships and objects does not fully reveal their true categories. In this paper, we propose a categorization algorithm, the iterative reinforcement categorization algorithm (IRC), to exploit the full interrelationships between the heterogeneous objects on the Web. IRC attempts to classify the interrelated Web objects by iterative reinforcement between individual classification results of different types via the interrelationships. Experiments on a clickthrough log dataset from MSN search engine show that, with the Fl measures, IRC achieves a 26.4% improvement over a pure content-based classification method, a 21% improvement over a query metadata-based method, and a 16.4% improvement over a virtual document-based method. Furthermore, our experiments show that IRC converges rapidly. Gui-Rong Xue, Dou Shen, Qiang Yang 0001, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wensi Xi, Wei-Ying Ma |
ICDM | 1 |
| 2004 | Multi-model similarity propagation and its application for web image retrievalabstractIn this paper, we propose an iterative similarity propagation approach to explore the inter-relationships between Web images and their textual annotations for image retrieval. By considering Web images as one type of objects, their surrounding texts as another type, and constructing the links structure between them via webpage analysis, we can iteratively reinforce the similarities between images. The basic idea is that if two objects of the same type are both related to one object of another type, these two objects are similar; likewise, if two objects of the same type are related to two different, but similar objects of another type, then to some extent, these two objects are also similar. The goal of our method is to fully exploit the mutual reinforcement between images and their textual annotations. Our experiments based on 10,628 images crawled from the Web show that our proposed approach can significantly improve Web image retrieval performance. Xin-Jing Wang, Wei-Ying Ma, Gui-Rong Xue, Xing Li 0001 |
ACM Multimedia | 3 |
| 2004 | DHT Based Searching Improved by Sliding Window
Shen Huang, Gui-Rong Xue, Yan-Feng Ge, Yong Yu 0001 |
WAIM | 2 |
| 2004 | TSSP: A Reinforcement Algorithm to Find Related PapersabstractContent analysis and citation analysis are two common methods in recommending system. Compared with content analysis, citation analysis can discover more implicitly related papers. However, the citation-based methods may introduce more noise in citation graph and cause topic drift. Some work combine content with citation to improve similarity measurement. The problem is that the two features are not used to reinforce each other to get better result. To solve the problem, we propose a new algorithm, Topic Sensitive Similarity Propagation (TSSP), to effectively integrate content similarity into similarity propagation. TSSP has two parts: citation context based propagation and iterative reinforcement. First, citation contexts provide clues for which papers are topic related to and filter out less irrelevant citations. Second, iteratively integrating content and citation similarity enable them to reinforce each other during the propagation. The experimental results of a user study show TSSP outperforms other algorithms in almost all cases. Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma |
Web Intelligence | 2 |
| 2004 | Multi-type Features Based Web Document Clustering
Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma |
WISE | 2 |
| 2004 | Exploiting PageRank at Different Block Level
Xue-Mei Jiang, Gui-Rong Xue, Wen-Guan Song, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma |
WISE | 2 |
| 2004 | Optimizing Web Search Using Spreading Activation on the Clickthrough Data
Gui-Rong Xue, Shen Huang, Yong Yu 0001, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma |
WISE | 1 |
| 2003 | Implicit link analysis for small web searchabstractCurrent Web search engines generally impose link analysis-based re-ranking on web-page retrieval. However, the same techniques, when applied directly to small web search such as intranet and site search, cannot achieve the same performance because their link structures are different from the global Web. In this paper, we propose an approach to constructing implicit links by mining users' access patterns, and then apply a modified PageRank algorithm to re-rank web-pages for small web search. Our experimental results indicate that the proposed method outperforms content-based method by 16%, explicit link-based PageRank by 20% and DirectHit by 14%, respectively. Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma, HongJiang Zhang, Chao-Jun Lu |
SIGIR | 1 |