Gui-Rong Xue

dblp:72/3906 · DBLP profile ↗
← Back
71ranked-venue papers
11as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 57 · 11 first-authorArtificial intelligence and machine learning · 28 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 9Graphics, computer vision, multimedia, augmented reality and games · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
27 papers
Information retrieval · 63% Data mining · 20% Recommender systems · 8%
Artificial intelligence
14 papers
Transfer learning and domain adaptation · 64% Information extraction and text analysis · 16% Image recognition and object detection · 6%

Topics — the 30 heaviest of 81, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking
learning to rank
0.432014
Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014
Global ranking by exploiting user clicks · SIGIR 2009
Learning to rank with ties · SIGIR 2008
Machine learning › Transfer learning and domain adaptation
heterogeneous transfer learning
0.332011
Heterogeneous Transfer Learning for Image Classification · AAAI 2011
Heterogeneous Transfer Learning for Image Clustering via the SocialWeb · ACL/IJCNLP 2009
Translated Learning: Transfer Learning across Different Feature Spaces · NIPS 2008
Information retrieval
ranking
0.332009
Dataplorer: a scalable search engine for the data web · WWW 2009
Global ranking by exploiting user clicks · SIGIR 2009
Optimizing web search using social annotations · WWW 2007
Machine learning › Transfer learning and domain adaptation
cross-modal transfer
0.222011
Video summarization via transferrable structured learning · WWW 2011
Heterogeneous Transfer Learning for Image Classification · AAAI 2011
Information retrieval
text summarization
0.222011
Video summarization via transferrable structured learning · WWW 2011
Enhancing diversity, coverage and balance for summarization through structure learning · WWW 2009
Information retrieval
retrieval models
0.232008
Learning to rank with ties · SIGIR 2008
Optimizing web search using social annotations · WWW 2007
Multi-model similarity propagation and its application for web image retrieval · ACM Multimedia 2004
Natural language and speech › Information extraction and text analysis
topic model
0.222010
Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA · IEEE Trans. Knowl. Data Eng. 2010
Topic-bridged PLSA for cross-domain text classification · SIGIR 2008
Recommender systems
click-through rate prediction
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Information retrieval › evaluation › effectiveness metrics
discounted cumulative gain
0.212014
Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014
Information retrieval › online advertising
display advertising
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Information retrieval
evaluation
0.212014
Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014
Data mining › dimensionality reduction
feature selection
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Data mining › dimensionality reduction › feature selection
group lasso
0.212014
Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising · ICML 2014
Machine learning and data management
metric learning
0.212014
Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014
Information retrieval › retrieval evaluation
ranking evaluation
0.212014
Learning the Gain Values and Discount Factors of Discounted Cumulative Gains · IEEE Trans. Knowl. Data Eng. 2014
Information retrieval
web search
0.232007
Optimizing web search using social annotations · WWW 2007
Exploiting the hierarchical structure for link analysis · SIGIR 2005
Implicit link analysis for small web search · SIGIR 2003
Data mining › predictive modeling
classification
0.222009
Web-scale classification with naive bayes · WWW 2009
Co-clustering based classification for out-of-domain documents · KDD 2007
Data mining › text mining
text classification
0.222008
Deep classification in large-scale text hierarchies · SIGIR 2008
Topic-bridged PLSA for cross-domain text classification · SIGIR 2008
Computer vision › Image recognition and object detection
image classification
0.112011
Heterogeneous Transfer Learning for Image Classification · AAAI 2011
Information retrieval › text summarization
video summarization
0.112011
Video summarization via transferrable structured learning · WWW 2011
Natural language and speech › Information extraction and text analysis › topic model
probabilistic latent semantic analysis
0.112010
Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA · IEEE Trans. Knowl. Data Eng. 2010
Information retrieval › online advertising
behavioral targeting
0.112010
Transfer learning for behavioral targeting · WWW 2010
Information retrieval › online advertising
contextual advertising
0.112010
Visual Contextual Advertising: Bringing Textual Advertisements to Images · AAAI 2010
Information retrieval
cross-modal retrieval
0.112010
Visual Contextual Advertising: Bringing Textual Advertisements to Images · AAAI 2010
Information retrieval
multimedia analysis and retrieval
0.112010
Visual Contextual Advertising: Bringing Textual Advertisements to Images · AAAI 2010
Machine learning and data management › weak supervision
positive-unlabeled learning
0.112010
Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA · IEEE Trans. Knowl. Data Eng. 2010
Information retrieval › web search
link analysis
0.122005
Exploiting the hierarchical structure for link analysis · SIGIR 2005
Implicit link analysis for small web search · SIGIR 2003
Information retrieval › ranking
ranking algorithms
0.122005
Exploiting the hierarchical structure for link analysis · SIGIR 2005
Implicit link analysis for small web search · SIGIR 2003
Machine learning › Transfer learning and domain adaptation
cross-domain learning
0.112009
EigenTransfer: a unified framework for transfer learning · ICML 2009
Machine learning › Transfer learning and domain adaptation › knowledge transfer
self-taught learning
0.112009
EigenTransfer: a unified framework for transfer learning · ICML 2009

Methods — techniques the papers use, named apart from their topics

transfer learning · 0.5language model · 0.4linear utility learning · 0.4active learning · 0.4structured SVM · 0.3generative model · 0.2logistic regression · 0.2feature hashing · 0.2distributed implementation · 0.2co-clustering · 0.2matrix factorization · 0.1latent semantic features · 0.1constrained optimization · 0.1EM algorithm · 0.1
YearPublicationVenuePosition
2015 Topic Modeling in Semantic Space with Keywords
abstract
A common and convenient approach for user to describe his information need is to provide a set of keywords. Therefore, the technique to understand the need becomes crucial. In this paper, for the information need about a topic or category, we propose a novel method called TDCS(Topic Distilling with Compressive Sensing) for explicit and accurate modeling the topic implied by several keywords. The task is transformed as a topic reconstruction problem in the semantic space with a reasonable intuition that the topic is sparse in the semantic space. The latent semantic space could be mined from documents via unsupervised methods, e.g. LSI. Compressive sensing is leveraged to obtain a sparse representation from only a few keywords. In order to make the distilled topic more robust, an iterative learning approach is adopted. The experiment results show the effectiveness of our method. Moreover, with only a few semantic concepts remained for the topic, our method is efficient for subsequent text mining tasks.
Xiaojia Pu, Rong Jin 0001, Gangshan Wu, Dingyi Han, Gui-Rong Xue
CIKM5
2014 Coupled Group Lasso for Web-Scale CTR Prediction in Display Advertising
abstract
In display advertising, click through rate(CTR) prediction is the problem of estimating the probability that an advertisement (ad) is clicked when displayed to a user in a specific context. Due to its easy implementation and promising performance, logistic regression(LR) model has been widely used for CTR prediction, especially in industrial systems. However, it is not easy for LR to capture the nonlinear information, such as the conjunction information, from user features and ad features. In this paper, we propose a novel model, called coupled group lasso(CGL), for CTR prediction in display advertising. CGL can seamlessly integrate the conjunction information from user features and ad features for modeling. Furthermore, CGL can automatically eliminate useless features for both users and ads, which may facilitate fast online prediction. Scalability of CGL is ensured through feature hashing and distributed implementation. Experimental results on real-world data sets show that our CGL model can achieve state-of-the-art performance on web-scale CTR prediction tasks.
Wu-Jun Li, Gui-Rong Xue, Dingyi Han
ICML3
2014 Learning the Gain Values and Discount Factors of Discounted Cumulative Gains
abstract
Evaluation metric is an essential and integral part of a ranking system. In the past, several evaluation metrics have been proposed in information retrieval and web search, among them Discounted Cumulative Gain (DCG) has emerged as one that is widely adopted for evaluating the performance of ranking functions used in web search. However, the two sets of parameters, the gain values and discount factors, used in DCG are usually determined in a rather ad-hoc way, and their impacts have not been carefully analyzed. In this paper, we first show that DCG is generally not coherent, i.e., comparing the performance of ranking functions using DCG very much depends on the particular gain values and discount factors used. We then propose a novel methodology that can learn the gain values and discount factors from user preferences over rankings, modeled as a special case of learning linear utility functions. We also discuss how to extend our methods to handle tied preference pairs and how to explore active learning to reduce preference labeling. Numerical simulations illustrate the effectiveness of our proposed methods. Moreover, experiments are also conducted over a side-by-side comparison data set from a commercial search engine to validate the proposed methods on real-world data.
Ke Zhou 0002, Hongyuan Zha, Yi Chang 0001, Gui-Rong Xue
IEEE Trans. Knowl. Data Eng.4
2012 Multi-task learning to rank for web search
Yi Chang 0001, Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Zhaohui Zheng 0001
Pattern Recognit. Lett.4
2012 Advertising Keywords Recommendation for Short-Text Web Pages Using Wikipedia
abstract
Advertising keywords recommendation is an indispensable component for online advertising with the keywords selected from the target Web pages used for contextual advertising or sponsored search. Several ranking-based algorithms have been proposed for recommending advertising keywords. However, for most of them performance is still lacking, especially when dealing with short-text target Web pages, that is, those containing insufficient textual information for ranking. In some cases, short-text Web pages may not even contain enough keywords for selection. A natural alternative is then to recommend relevant keywords not present in the target Web pages. In this article, we propose a novel algorithm for advertising keywords recommendation for short-text Web pages by leveraging the contents of Wikipedia, a user-contributed online encyclopedia. Wikipedia contains numerous entities with related entities on a topic linked to each other. Given a target Web page, we propose to use a content-biased PageRank on the Wikipedia graph to rank the related entities. Furthermore, in order to recommend high-quality advertising keywords, we also add an advertisement-biased factor into our model. With these two biases, advertising keywords that are both relevant to a target Web page and valuable for advertising are recommended. In our experiments, several state-of-the-art approaches for keyword recommendation are compared. The experimental results demonstrate that our proposed approach produces substantial improvement in the precision of the top 20 recommended keywords on short-text Web pages over existing approaches.
Weinan Zhang 0001, Dingquan Wang, Gui-Rong Xue, Hongyuan Zha
ACM Trans. Intell. Syst. Technol.3
2012 Leveraging Auxiliary Data for Learning to Rank
abstract
In learning to rank, both the quality and quantity of the training data have significant impacts on the performance of the learned ranking functions. However, in many applications, there are usually not sufficient labeled training data for the construction of an accurate ranking model. It is therefore desirable to leverage existing training data from other tasks when learning the ranking function for a particular task, an important problem which we tackle in this article utilizing a boosting framework withtransfer learning. In particular, we propose to adaptively learn transferable representations called super-features from the training data of both the target task and the auxiliary task. Those super-features and the coefficients for combining them are learned in an iterative stage-wise fashion. Unlike previous transfer learning methods, the super-features can be adaptively learned by weak learners from the data. Therefore, the proposed framework is sufficiently flexible to deal with complicated common structures among different learning tasks. We evaluate the performance of the proposed transfer learning method for two datasets from the Letor collection and one dataset collected from a commercial search engine, and we also compare our methods with several existing transfer learning methods. Our results demonstrate that the proposed method can enhance the ranking functions of the target tasks utilizing the training data from the auxiliary tasks.
Ke Zhou 0002, Hongyuan Zha, Gui-Rong Xue
ACM Trans. Intell. Syst. Technol.4
2011 Heterogeneous Transfer Learning for Image Classification
abstract
Transfer learning as a new machine learning paradigm has gained increasing attention lately. In situations where the training data in a target domain are not sufficient to learn predictive models effectively, transfer learning leverages auxiliary source data from other related source domains for learning. While most of the existing works in this area only focused on using the source data with the same structure as the target data, in this paper, we push this boundary further by proposing a heterogeneous transfer learning framework for knowledge transfer between text and images. We observe that for a target-domain classification problem, some annotated images can be found on many social Web sites, which can serve as a bridge to transfer knowledge from the abundant text documents available over the Web. A key question is how to effectively transfer the knowledge in the source data even though the text can be arbitrarily found. Our solution is to enrich the representation of the target images with semantic concepts extracted from the auxiliary source data through a novel matrix factorization method. By using the latent semantic features generated by the auxiliary data, we are able to build a better integrated image classifier. We empirically demonstrate the effectiveness of our algorithm on the Caltech-256 image dataset.
Yuqiang Chen, Zhongqi Lu, Sinno Jialin Pan, Gui-Rong Xue, Yong Yu 0001, Qiang Yang 0001
AAAI5
2011 Cross-Lingual Sentiment Classification via Bi-view Non-negative Matrix Tri-Factorization
Junfeng Pan, Gui-Rong Xue, Yong Yu 0001, Yang Wang 0019
PAKDD (1)2
2011 Video summarization via transferrable structured learning
abstract
It is well-known that textual information such as video transcripts and video reviews can significantly enhance the performance of video summarization algorithms. Unfortunately, many videos on the Web such as those from the popular video sharing site YouTube do not have useful textual information. The goal of this paper is to propose a transfer learning framework for video summarization: in the training process both the video features and textual features are exploited to train a summarization algorithm while for summarizing a new video only its video features are utilized. The basic idea is to explore the transferability between videos and their corresponding textual information. Based on the assumption that video features and textual features are highly correlated with each other, we can transfer textual information into knowledge on summarization using video information only. In particular, we formulate the video summarization problem as that of learning a mapping from a set of shots of a video to a subset of the shots using the general framework of SVM-based structured learning. Textual information is transferred by encoding them into a set of constraints used in the structured learning process which tend to provide a more detailed and accurate characterization of the different subsets of shots. Experimental results show significant performance improvement of our approach and demonstrate the utility of textual information for enhancing video summarization.
Liangda Li, Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Yong Yu 0001
WWW3
2010 Visual Contextual Advertising: Bringing Textual Advertisements to Images
abstract
Advertising in the case of textual Web pages has been studied extensively by many researchers. However, with the increasing amount of multimedia data such as image, audio and video on the Web, the need for recommending advertisement for the multimedia data is becoming a reality. In this paper, we address the novel problem of visual contextual advertising, which is to directly advertise when users are viewing images which do not have any surrounding text. A key challenging issue of visual contextual advertising is that images and advertisements are usually represented in image space and word space respectively, which are quite different with each other inherently. As a result, existing methods for Web page advertising are inapplicable since they represent both Web pages and advertisement in the same word space. In order to solve the problem, we propose to exploit the social Web to link these two feature spaces together. In particular, we present a unified generative model to integrate advertisements, words and images. Specifically, our solution combines two parts in a principled approach: First, we transform images from a image feature space to a word space utilizing the knowledge from images with annotations from social Web. Then, a language model based approach is applied to estimate the relevance between transformed images and advertisements. Moreover, in this model, the probability of recommending an advertisement can be inferred efficiently given an image, which enables potential applications to online advertising.
Yuqiang Chen, Ou Jin, Gui-Rong Xue, Qiang Yang 0001
AAAI3
2010 Predicting Product Duration for Adaptive Advertisement
Zhongqi Guo, Gui-Rong Xue, Yong Yu 0001
ADMA (2)3
2010 Click Prediction for Product Search on C2C Web Sites
Xiangzhi Wang, Gui-Rong Xue, Yong Yu 0001
ADMA (2)3
2010 Text-Aided Image Classification: Using Labeled Text from Web to Help Image Classification
abstract
As more and more multimedia data become available on the Web, mining on those data is playing an increasingly important role in Web applications. In this paper, we investigate the interplay between multimedia data mining and text data mining. Specifically, in an approach we called text-aided image classification (TAIC), we address the problem of image classification with very limited amount of labeled images and a large amount of auxiliary labeled text data. This problem is important in practice, since currently on the Web, labeled text data are usually much more than image data. To solve the problem, based on the “bag-of-words” view and the Naive Bayes classification model, we focus our attention on the estimation of the image feature distribution under given concept. We extend the Naive Bayes algorithm by considering a mapping that maps the most discriminative text features into the image feature space. This feature mapping is estimated based on the text-image cooccurrence data on the Web, acting like a bridge that connects text and image knowledge. With this process, we estimate target image feature distribution from a text model based on sufficient labeled data. Our empirical results on real world data sets show that our method makes a good approximation of the image feature distribution when trained with abundant labeled images. In the case amount of labeled images is very limited, the classification performance is improved by using auxiliary labeled text data, which shows that our method can indeed integrate text and image knowledge in a simple yet effective way.
Yuqiang Chen, Gui-Rong Xue, Yong Yu 0001
APWeb3
2010 Transfer learning for behavioral targeting
abstract
Recently, Behavioral Targeting (BT) is attracting much attention from both industry and academia due to its rapid growth in online advertising market. Though a basic assumption of BT, which is, the users who share similar Web browsing behaviors will have similar preference over ads, has been empirically verified, we argue that the users' ad click preference and Web browsing behavior are not reflecting the same user intent though they are correlated. In this paper, we propose to formulate BT as a transfer learning problem. We treat the users' preference over ads and Web browsing behaviors as two different user behavioral domains and propose to utilize transfer learning strategy across these two user behavioral domains to segment users for BT ads delivery. We show that some classical BT solutions could be formulated in transfer learning view. As an example, we propose to leverage translated learning, which is a recent proposed transfer learning algorithm, to benefit the BT ads delivery. Experimental results on real ad click data show that, BT user segmentation by the approach of transfer learning can outperform the classical user segmentation strategies for larger than 20% in terms of smoothed ad Click Through Rate(CTR).
Jun Yan 0001, Gui-Rong Xue, Zheng Chen 0001
WWW3
2010 Knowledge transfer for cross domain learning to rank
Depin Chen, Yan Xiong 0001, Jun Yan 0001, Gui-Rong Xue, Gang Wang 0010, Zheng Chen 0001
Inf. Retr.4
2010 Learning with Positive and Unlabeled Examples Using Topic-Sensitive PLSA
abstract
It is often difficult and time-consuming to provide a large amount of positive and negative examples for training a classification system in many applications such as information retrieval. Instead, users often find it easier to indicate just a few positive examples of what he or she likes, and thus, these are the only labeled examples available for the learning system. A large amount of unlabeled data are easier to obtain. How to make use of the positive and unlabeled data for learning is a critical problem in machine learning and information retrieval. Several approaches for solving this problem have been proposed in the past, but most of these methods do not work well when only a small amount of labeled positive data are available. In this paper, we propose a novel algorithm called Topic-Sensitive pLSA to solve this problem. This algorithm extends the original probabilistic latent semantic analysis (pLSA), which is a purely unsupervised framework, by injecting a small amount of supervision information from the user. The supervision from users is in the form of indicating which documents fit the users' interests. The supervision is encoded into a set of constraints. By introducing the penalty terms for these constraints, we propose an objective function that trades off the likelihood of the observed data and the enforcement of the constraints. We develop an iterative algorithm that can obtain the local optimum of the objective function. Experimental evaluation on three data corpora shows that the proposed method can improve the performance especially only with a small amount of labeled positive data.
Ke Zhou 0002, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
IEEE Trans. Knowl. Data Eng.2
2009 Heterogeneous Transfer Learning for Image Clustering via the SocialWeb
Qiang Yang 0001, Yuqiang Chen, Gui-Rong Xue, Wenyuan Dai, Yong Yu 0001
ACL/IJCNLP3
2009 Multi-task learning for learning to rank in web search
abstract
Both the quality and quantity of training data have significant impact on the performance of ranking functions in the context of learning to rank for web search. Due to resource constraints, training data for smaller search engine markets are scarce and we need to leverage existing training data from large markets to enhance the learning of ranking function for smaller markets. In this paper, we present a boosting framework for learning to rank in the multi-task learning context for this purpose. In particular, we propose to learn non-parametric common structures adaptively from multiple tasks in a stage-wise way. An algorithm is developed to iteratively discover super-features that are effective for all the tasks. The estimation of the functions for each task is then learned as a linear combination of those super-features. We evaluate the performance of this multi-task learning method for web search ranking using data from a search engine. Our results demonstrate that multi-task learning methods bring significant relevance improvements over existing baseline methods.
Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Gordon Sun, Belle L. Tseng, Zhaohui Zheng 0001, Yi Chang 0001
CIKM3
2009 EigenTransfer: a unified framework for transfer learning
abstract
This paper proposes a general framework, called EigenTransfer, to tackle a variety of transfer learning problems, e.g. cross-domain learning, self-taught learning, etc. Our basic idea is to construct a graph to represent the target transfer learning task. By learning the spectra of a graph which represents a learning task, we obtain a set of eigenvectors that reflect the intrinsic structure of the task graph. These eigenvectors can be used as the new features which transfer the knowledge from auxiliary data to help classify target data. Given an arbitrary non-transfer learner (e.g. SVM) and a particular transfer learning task, EigenTransfer can produce a transfer learner accordingly for the target transfer learning task. We apply EigenTransfer on three different transfer learning tasks, cross-domain learning, cross-category learning and self-taught learning, to demonstrate its unifying ability, and show through experiments that EigenTransfer can greatly outperform several representative non-transfer learners.
Wenyuan Dai, Ou Jin, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
ICML3
2009 Global ranking by exploiting user clicks
abstract
It is now widely recognized that user interactions with search results can provide substantial relevance information on the documents displayed in the search results. In this paper, we focus on extracting relevance information from one source of user interactions, i.e., user click data, which records the sequence of documents being clicked and not clicked in the result set during a user search session. We formulate the problem as a global ranking problem, emphasizing the importance of the sequential nature of user clicks, with the goal to predict the relevance labels of all the documents in a search session. This is distinct from conventional learning to rank methods that usually design a ranking model defined on a single document; in contrast, in our model the relational information among the documents as manifested by an aggregation of user clicks is exploited to rank all the documents jointly. In particular, we adapt several sequential supervised learning algorithms, including the conditional random field (CRF), the sliding window method and the recurrent sliding window method, to the global ranking problem. Experiments on the click data collected from a commercial search engine demonstrate that our methods can outperform the baseline models for search results re-ranking.
Shihao Ji 0001, Ke Zhou 0002, Ciya Liao, Zhaohui Zheng 0001, Gui-Rong Xue, Olivier Chapelle, Gordon Sun, Hongyuan Zha
SIGIR5
2009 Efficient query expansion for advertisement search
abstract
Online advertising represents a growing part of the revenues of ma-jor Internet service providers such as Google and Yahoo. A com-monly used strategy is to place advertisements (ads) on the search result pages according to the users ’ submitted queries. Relevant ads are likely to be clicked by a user and to increase the revenues of both advertisers and publishers. However, bid phrases defined by ad-owners are usually contained in limited number of ads. Directly matching user queries with bid phrases often results in finding few appropriate ads. To address this shortcoming, query expansion is often used to increase the chances to match the ads. Nevertheless, query expansion on top of the traditional inverted index faces ef-ficiency issues such as high time complexity and heavy I/O costs. Moreover, precision cannot always be improved, sometimes even hurt due to the involvement of additional noise.
Haofen Wang, Linyun Fu, Gui-Rong Xue, Yong Yu 0001
SIGIR4
2009 Enhancing diversity, coverage and balance for summarization through structure learning
abstract
Document summarization plays an increasingly important role with the exponential growth of documents on the Web. Many supervised and unsupervised approaches have been proposed to generate summaries from documents. However, these approaches seldom simultaneously consider summary diversity, coverage, and balance issues which to a large extent determine the quality of summaries. In this paper, we consider extract-based summarization emphasizing the following three requirements: 1) diversity in summarization, which seeks to reduce redundancy among sentences in the summary; 2) sufficient coverage, which focuses on avoiding the loss of the document's main information when generating the summary; and 3) balance, which demands that different aspects of the document need to have about the same relative importance in the summary. We formulate the extract-based summarization problem as learning a mapping from a set of sentences of a given document to a subset of the sentences that satisfies the above three requirements. The mapping is learned by incorporating several constraints in a structure learning framework, and we explore the graph structure of the output variables and employ structural SVM for solving the resulted optimization problem. Experiments on the DUC2001 data sets demonstrate significant performance improvements in terms of F1 and ROUGE metrics.
Liangda Li, Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Yong Yu 0001
WWW3
2009 Dataplorer: a scalable search engine for the data web
abstract
More and more structured information in the form of semantic data is nowadays available. It offers a wide range of new possibilities especially for semantic search and Web data integration. However, their effective exploitation still brings about a number of challenges, e.g. usability, scalability and uncertainty. In this paper, we present Dataplorer, a solution designed to address these challenges. We consider the usability through the use of hybrid queries and faceted search, while still preserving the scalability thanks to an extension of inverted index to support this type of query. Moreover, Dataplorer deals with uncertainty by means of a powerful ranking scheme to find relevant results. Our experimental results show that our proposed approach is promising and it makes us believe that it is possible to extend the current IR infrastructure to query and search the Web of data.
Haofen Wang, Qiaoling Liu, Gui-Rong Xue, Yong Yu 0001, Lei Zhang 0007
WWW3
2009 Web-scale classification with naive bayes
abstract
Traditional Naive Bayes Classifier performs miserably on web-scale taxonomies. In this paper, we investigate the reasons behind such bad performance. We discover that the low performance are not completely caused by the intrinsic limitations of Naive Bayes, but mainly comes from two largely ignored problems: contradiction pair problem and discriminative evidence cancelation problem. We propose modifications that can alleviate the two problems while preserving the advantages of Naive Bayes. The experimental results show our modified Naive Bayes can significantly improve the performance on real web-scale taxonomies.
Congle Zhang, Gui-Rong Xue, Yong Yu 0001, Hongyuan Zha
WWW2
2009 User language model for collaborative personalized search
abstract
Traditional personalized search approaches rely solely on individual profiles to construct a user model. They are often confronted by two major problems: data sparseness and cold-start for new individuals. Data sparseness refers to the fact that most users only visit a small portion of Web pages and hence a very sparse user-term relationship matrix is generated, while cold-start for new individuals means that the system cannot conduct any personalization without previous browsing history. Recently, community-based approaches were proposed to use the group's social behaviors as a supplement to personalization. However, these approaches only consider the commonality of a group of users and still cannot satisfy the diverse information needs of different users. In this article, we present a new approach, called collaborative personalized search. It considers not only the commonality factor among users for defining group user profiles and global user profiles, but also the specialties of individuals. Then, a statistical user language model is proposed to integrate the individual model, group user model and global user model together. In this way, the probability that a user will like a Web page is calculated through a two-step smoothing mechanism. First, a global user model is used to smooth the probability of unseen terms in the individual profiles and provide aggregated behavior of global users. Then, in order to precisely describe individual interests by looking at the behaviors of similar users, users are clustered into groups and group-user models are constructed. The group-user models are integrated into an overall model through a cluster-based language model. The behaviors of the group users can be utilized to enhance the performance of personalized search. This model can alleviate the two aforementioned problems and provide a more effective personalized search than previous approaches. Large-scale experimental evaluations are conducted to show that the proposed approach substantially improves the relevance of a search over several competitive methods.
Gui-Rong Xue, Yong Yu 0001, Qiang Yang 0001
ACM Trans. Inf. Syst.1
2008 Iterative Reinforcement Cross-Domain Text Classification
Gui-Rong Xue, Yong Yu 0001
ADMA2
2008 Knowledge Transferring Via Implicit Link Analysis
Wenyuan Dai, Gui-Rong Xue, Yong Yu 0001
DASFAA3
2008 Self-taught clustering
abstract
This paper focuses on a new clustering task, called self-taught clustering. Self-taught clustering is an instance of unsupervised transfer learning, which aims at clustering a small collection of target unlabeled data with the help of a large amount of auxiliary unlabeled data. The target and auxiliary data can be different in topic distribution. We show that even when the target data are not sufficient to allow effective learning of a high quality feature representation, it is possible to learn the useful features with the help of the auxiliary data on which the target data can be clustered effectively. We propose a co-clustering based self-taught clustering algorithm to tackle this problem, by clustering the target and auxiliary data simultaneously to allow the feature representation from the auxiliary data to influence the target data through a common set of features. Under the new data representation, clustering on the target data can be improved. Our experiments on image clustering show that our algorithm can greatly outperform several state-of-the-art clustering methods when utilizing irrelevant unlabeled auxiliary data.
Wenyuan Dai, Qiang Yang 0001, Gui-Rong Xue, Yong Yu 0001
ICML3
2008 Spectral domain-transfer learning
abstract
Traditional spectral classification has been proved to be effective in dealing with both labeled and unlabeled data when these data are from the same domain. In many real world applications, however, we wish to make use of the labeled data from one domain (called in-domain) to classify the unlabeled data in a different domain (out-of-domain). This problem often happens when obtaining labeled data in one domain is difficult while there are plenty of labeled data from a related but different domain. In general, this is a transfer learning problem where we wish to classify the unlabeled data through the labeled data even though these data are not from the same domain. In this paper, we formulate this domain-transfer learning problem under a novel spectral classification framework, where the objective function is introduced to seek consistency between the in-domain supervision and the out-of-domain intrinsic structure. Through optimization of the cost function, the label information from the in-domain data is effectively transferred to help classify the unlabeled data from the out-of-domain. We conduct extensive experiments to evaluate our method and show that our algorithm achieves significant improvements on classification performance over many state-of-the-art algorithms.
Wenyuan Dai, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
KDD3
2008 Translated Learning: Transfer Learning across Different Feature Spaces
abstract
This paper investigates a new machine learning strategy called translated learning. Unlike many previous learning tasks, we focus on how to use labeled data from one feature space to enhance the classification of other entirely different learning spaces. For example, we might wish to use labeled text data to help learn a model for classifying image data, when the labeled images are difficult to obtain. An important aspect of translated learning is to build a "bridge" to link one feature space (known as the "source space") to another space (known as the "target space") through a translator in order to migrate the knowledge from source to target. The translated learning solution uses a language model to link the class labels to the features in the source spaces, which in turn is translated to the features in the target spaces. Finally, this chain of linkages is completed by tracing back to the instances in the target spaces. We show that this path of linkage can be modeled using a Markov chain and risk minimization. Through experiments on the text-aided image classification and cross-language classification tasks, we demonstrate that our translated learning framework can greatly outperform many state-of-the-art baseline methods.
Wenyuan Dai, Yuqiang Chen, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
NIPS3
2008 Knowledge Supervised Text Classification with No Labeled Documents
Congle Zhang, Gui-Rong Xue, Yong Yu 0001
PRICAI2
2008 Topic-bridged PLSA for cross-domain text classification
abstract
In many Web applications, such as blog classification and new-sgroup classification, labeled data are in short supply. It often happens that obtaining labeled data in a new domain is expensive and time consuming, while there may be plenty of labeled data in a related but different domain. Traditional text classification ap-proaches are not able to cope well with learning across different domains. In this paper, we propose a novel cross-domain text classification algorithm which extends the traditional probabilistic latent semantic analysis (PLSA) algorithm to integrate labeled and unlabeled data, which come from different but related domains, into a unified probabilistic model. We call this new model Topic-bridged PLSA, or TPLSA. By exploiting the common topics between two domains, we transfer knowledge across different domains through a topic-bridge to help the text classification in the target domain. A unique advantage of our method is its ability to maximally mine knowledge that can be transferred between domains, resulting in superior performance when compared to other state-of-the-art text classification approaches. Experimental eval-uation on different kinds of datasets shows that our proposed algorithm can improve the performance of cross-domain text classification significantly.
Gui-Rong Xue, Wenyuan Dai, Qiang Yang 0001, Yong Yu 0001
SIGIR1
2008 Deep classification in large-scale text hierarchies
abstract
Most classification algorithms are best at categorizing the Web documents into a few categories, such as the top two levels in the Open Directory Project. Such a classification method does not give very detailed topic-related class information for the user because the first two levels are often too coarse. However, classification on a large-scale hierarchy is known to be intractable for many target categories with cross-link relationships among them. In this paper, we propose a novel deep-classification approach to categorize Web documents into categories in a large-scale taxonomy. The approach consists of two stages: a search stage and a classification stage. In the first stage, a category-search algorithm is used to acquire the category candidates for a given document. Based on the category candidates, we prune the large-scale hierarchy to focus our classification effort on a small subset of the original hierarchy. As a result, the classification model is trained on the small subset before being applied to assign the category for a new document. Since the category candidates are sufficiently close to each other in the hierarchy, a statistical-language-model based classifier using n-gram features is exploited. Furthermore, the structure of the taxonomy can be utilized in this stage to improve the performance of classification. We demonstrate the performance of our proposed algorithms on the Open Directory Project with over 130,000 categories. Experimental results show that our proposed approach can reach 51.8% on the measure of Mi-F1 at the 5th level, which is 77.7% improvement over top-down based SVM classification algorithms.
Gui-Rong Xue, Dikan Xing, Qiang Yang 0001, Yong Yu 0001
SIGIR1
2008 Learning to rank with ties
abstract
Designing effective ranking functions is a core problem for information retrieval and Web search since the ranking functions directly impact the relevance of the search results. The problem has been the focus of much of the research at the intersection of Web search and machine learning, and learning ranking functions from preference data in particular has recently attracted much interest. The objective of this paper is to empirically examine several objective functions that can be used for learning ranking functions from preference data. Specifically, we investigate the roles of ties in the learning process. By ties, we mean preference judgments that two documents have equal degree of relevance with respect to a query. This type of data has largely been ignored or not properly modeled in the past. In this paper, we analyze the properties of ties and develop novel learning frameworks which combine ties and preference data using statistical paired comparison models to improve the performance of learned ranking functions. The resulting optimization problems explicitly incorporating ties and preference data are solved using gradient boosting methods. Experimental studies are conducted using three publicly available data sets which demonstrate the effectiveness of the proposed new methods.
Ke Zhou 0002, Gui-Rong Xue, Hongyuan Zha, Yong Yu 0001
SIGIR2
2008 Advertising keyword suggestion based on concept hierarchy
abstract
The increasing growth of the World Wide Web constantly enlarges the revenue generated by search engine advertising. Advertisers bid on keywords associated with their products to display their ads on the search result pages. Keyword suggestion methods are proposed to fill the gap between the keywords chosen by advertisers and the popular queries, through finding new relevant keywords according to some statistical information (for example, the keyword co-occurrence). However, there is little effort taking semantic information, such as concept hierarchy, into account. In this paper, we propose a novel keyword suggestion method that fully exploits the semantic knowledge among concept hierarchy. Given a keyword, we first match it with some relevant concepts. Then the relevant concepts are used with their hierarchy to fertilize the meanings of the keywords. Finally new keywords are suggested according to the concept information rather than the statistical co-occurrence of the keyword itself. Experimental results show that our proposed method can successfully provide suggestion that meets the accuracy and coverage requirements
Gui-Rong Xue, Yong Yu 0001
WSDM2
2008 Deep classifier: automatically categorizing search results into large-scale hierarchies
abstract
Organizing Web search results into hierarchical categories facilitates users' browsing through Web search results, especially for ambiguous queries where the potential results are mixed together. Previous methods on search result classification are usually based on pre-training a classification model on some fixed and shallow hierarchical categories, where only the top-two-level categories of a Web taxonomy is used. Such classification methods may be too coarse for users to browse, since most search results would be classified into only two or three shallow categories. Instead, a deep hierarchical classifier must provide many more categories. However, the performance of such classifiers is usually limited because their classification effectiveness can deteriorate rapidly at the third or fourth level of a hierarchy. In this paper, we propose a novel algorithm known as Deep Classifier to classify the search results into detailed hierarchical categories with higher effectiveness than previous approaches. Given the search results in response to a query, the algorithm first prunes a wide-ranged hierarchy into a narrow one with the help of some Web directories. Different strategies are proposed to select the training data by utilizing the hierarchical structures. Finally, a discriminative naíve Bayesian classifier is developed to perform efficient and effective classification. As a result, the algorithm can provide more meaningful and specific class labels for search result browsing than shallow style of classification. We conduct experiments to show that the Deep Classifier can achieve significant improvement over state-of-the-art algorithms. In addition, with sufficient off-line preparation, the efficiency of the proposed algorithm is suitable for online application
Dikan Xing, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
WSDM2
2008 WWW 2008 workshop on social web search and mining: SWSM2008
abstract
No abstract available.
Juan-Zi Li, Gui-Rong Xue, Jie Tang 0001, Ying Ding 0001
WWW2
2008 Can chinese web pages be classified with english data source?
abstract
As the World Wide Web in China grows rapidly, mining knowledge in Chinese Web pages becomes more and more important. Mining Web information usually relies on the machine learning techniques which require a large amount of labeled data to train credible models. Although the number of Chinese Web pages increases quite fast, it still lacks Chinese labeled data. However, there are relatively sufficient English labeled Web pages. These labeled data, though in different linguistic representations, share a substantial amount of semantic information with Chinese ones, and can be utilized to help classify Chinese Web pages. In this paper, we propose an information bottleneck based approach to address this cross-language classification problem. Our algorithm first translates all the Chinese Web pages to English. Then, all the Web pages, including Chinese and English ones, are encoded through an information bottleneck which can allow only limited information to pass. Therefore, in order to retain as much useful information as possible, the common part between Chinese and English Web pages is inclined to be encoded to the same code (i.e. class label), which makes the cross-language classification accurate. We evaluated our approach using the Web pages collected from Open Directory Project (ODP). The experimental results show that our method significantly improves several existing supervised and semi-supervised classifiers.
Gui-Rong Xue, Wenyuan Dai, Qiang Yang 0001, Yong Yu 0001
WWW2
2007 Transferring Naive Bayes Classifiers for Text Classification
Wenyuan Dai, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
AAAI2
2007 Boosting for transfer learning
abstract
Traditional machine learning makes a basic assumption: the training and test data should be under the same distribution. However, in many cases, this identical-distribution assumption does not hold. The assumption might be violated when a task from one new domain comes, while there are only labeled data from a similar old domain. Labeling the new data can be costly and it would also be a waste to throw away all the old data. In this paper, we present a novel transfer learning framework called TrAdaBoost, which extends boosting-based learning algorithms (Freund & Schapire, 1997). TrAdaBoost allows users to utilize a small amount of newly labeled data to leverage the old data to construct a high-quality classification model for the new data. We show that this method can allow us to learn an accurate model using only a tiny amount of new data and a large amount of old data, even when the new data are not sufficient to train a model alone. We show that TrAdaBoost allows knowledge to be effectively transferred from the old data to the new. The effectiveness of our algorithm is analyzed theoretically and empirically to show that our iterative algorithm can converge well to an accurate model.
Wenyuan Dai, Qiang Yang 0001, Gui-Rong Xue, Yong Yu 0001
ICML3
2007 Co-clustering based classification for out-of-domain documents
abstract
In many real world applications, labeled data are in short supply. It often happens that obtaining labeled data in a new domain is expensive and time consuming, while there may be plenty of labeled data from a related but different domain. Traditional machine learning is not able to cope well with learning across different domains. In this paper, we address this problem for a text-mining task, where the labeled data are under one distribution in one domain known as in-domain data, while the unlabeled data are under a related but different domain known as out-of-domain data. Our general goal is to learn from the in-domain and apply the learned knowledge to out-of-domain. We propose a co-clustering based classification (CoCC) algorithm to tackle this problem. Co-clustering is used as a bridge to propagate the class structure and knowledge from the in-domain to the out-of-domain. We present theoretical and empirical analysis to show that our algorithm is able to produce high quality classification results, even when the distributions between the two data are different. The experimental results show that our algorithm greatly improves the classification performance over the traditional learning algorithms.
Wenyuan Dai, Gui-Rong Xue, Qiang Yang 0001, Yong Yu 0001
KDD2
2007 Bridged Refinement for Transfer Learning
Dikan Xing, Wenyuan Dai, Gui-Rong Xue, Yong Yu 0001
PKDD3
2007 Adaptive Email Spam Filtering Based on Information Theory
Wenyuan Dai, Gui-Rong Xue, Yong Yu 0001
WISE3
2007 Optimizing web search using social annotations
abstract
This paper explores the use of social annotations to improve web search. Nowadays, many services, e.g. del.icio.us, have been developed for web users to organize and share their favorite web pages on line by using social annotations. We observe that the social annotations can benefit web search in two aspects: 1) the annotations are usually good summaries of corresponding web pages; 2) the count of annotations indicates the popularity of web pages. Two novel algorithms are proposed to incorporate the above information into page ranking: 1) SocialSimRank (SSR) calculates the similarity between social annotations and web queries; 2) SocialPageRank (SPR) captures the popularity of web pages. Preliminary experimental results show that SSR can find the latent semantic association between queries and annotations, while SPR successfully measures the quality (popularity) of a web page from the web users ’ perspective. We further evaluate the proposed methods empirically with 50 manually constructed queries and 3000 auto-generated queries on a dataset crawled from del.icio.us. Experiments show that both SSR and SPR benefit web search significantly.
Shenghua Bao, Gui-Rong Xue, Xiaoyuan Wu, Yong Yu 0001, Ben Fei, Zhong Su
WWW2
2007 Exploring in the weblog space by detecting informative and affective articles
abstract
Weblogs have become a prevalent source of information for people to express themselves. In general, there are two genres of contents in weblogs. The first kind is about the webloggers' personal feelings, thoughts or emotions. We call this kind of weblogs affective articles. The second kind of weblogs is about technologies and different kinds of informative news. In this paper, we present a machine learning method for classifying informative and affective articles among weblogs. We consider this problem as a binary classification problem. By using machine learning approaches, we achieve about 92% on information retrieval performance measures including precision, recall and F1. We set up three studies on the applications of above classification approach in both research and industrial fields. The above classification approach is used to improve the performance of classification of emotions from weblog articles. We also develop an intent-driven weblog-search engine based on the classification techniques to improve the satisfaction of Web users. Finally, our approach is applied to search for weblogs with a great deal of informative articles.
Xiaochuan Ni 0001, Gui-Rong Xue, Yong Yu 0001, Qiang Yang 0001
WWW2
2006 A Hierarchical Model of Web Graph
Yong Yu 0001, Dingyi Han, Gui-Rong Xue
ADMA5
2006 IRFCF: Iterative Rating Filling Collaborative Filtering Algorithm
Gui-Rong Xue, Fan-De Zhu, Ai-Guo Yao
APWeb3
2006 Image Description Mining and Hierarchical Clustering on Data Records Using HR-Tree
Congle Zhang, Gui-Rong Xue, Yong Yu 0001
APWeb3
2006 An Empirical Study of Data Smoothing Methods for Memory-Based and Hybrid Collaborative Filtering
Dingyi Han, Gui-Rong Xue, Yong Yu 0001
PRICAI2
2006 A Novel Web Page Categorization Algorithm Based on Block Propagation Using Query-Log Information
Wenyuan Dai, Yong Yu 0001, Congle Zhang, Gui-Rong Xue
WAIM5
2006 Exploiting Rating Behaviors for Effective Collaborative Filtering
Dingyi Han, Yong Yu 0001, Gui-Rong Xue
WISE3
2006 Reinforcing Web-object Categorization Through Interrelationships
Gui-Rong Xue, Yong Yu 0001, Dou Shen, Qiang Yang 0001, Hua-Jun Zeng, Zheng Chen 0001
Data Min. Knowl. Discov.1
2006 TSSP: Multi-features based reinforcement algorithm to find related papers
Shen Huang, Yong Yu 0001, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Wei-Ying Ma
Web Intell. Agent Syst.3
2005 An Algorithm for Enumerating SCCs in Web Graph
Yong Yu 0001, Gui-Rong Xue
APWeb4
2005 Block-Based Language Modeling Approach Towards Web Search
Shengping Li, Shen Huang, Gui-Rong Xue, Yong Yu 0001
APWeb3
2005 Using Probabilistic Latent Semantic Analysis for Personalized Web Search
Gui-Rong Xue, Hua-Jun Zeng, Yong Yu 0001
APWeb2
2005 China Web Graph Measurements and Evolution
Yong Yu 0001, Gui-Rong Xue
APWeb4
2005 Scalable collaborative filtering using cluster-based smoothing
abstract
Memory-based approaches for collaborative filtering identify the similarity between two users by comparing their ratings on a set of items. In the past, the memory-based approach has been shown to suffer from two fundamental problems: data sparsity and difficulty in scalability. Alternatively, the model-based approach has been proposed to alleviate these problems, but this approach tends to limit the range of users. In this paper, we present a novel approach that combines the advantages of these two approaches by introducing a smoothing-based method. In our approach, clusters generated from the training data provide the basis for data smoothing and neighborhood selection. As a result, we provide higher accuracy as well as increased efficiency in recommendations. Empirical studies on two datasets (EachMovie and MovieLens) show that our new proposed approach consistently outperforms other state-of-art collaborative filtering algorithms.
Gui-Rong Xue, Qiang Yang 0001, Wensi Xi, Hua-Jun Zeng, Yong Yu 0001, Zheng Chen 0001
SIGIR1
2005 Exploiting the hierarchical structure for link analysis
abstract
Link analysis algorithms have been extensively used in Web information retrieval. However, current link analysis algorithms generally work on a flat link graph, ignoring the hierarchal structure of the Web graph. They often suffer from two problems: the sparsity of link graph and biased ranking of newly-emerging pages. In this paper, we propose a novel ranking algorithm called Hierarchical Rank as a solution to these two problems, which considers both the hierarchical structure and the link structure of the Web. In this algorithm, Web pages are first aggregated based on their hierarchical structure at directory, host or domain level and link analysis is performed on the aggregated graph. Then, the importance of each node on the aggregated graph is distributed to individual pages belong to the node based on the hierarchical structure. This algorithm allows the importance of linked Web pages to be distributed in the Web page space even when the space is sparse and contains new pages. Experimental results on the .GOV collection of TREC 2003 and 2004 show that hierarchical ranking algorithm consistently outperforms other well-known ranking algorithms, including the PageRank, BlockRank and LayerRank. In addition, experimental results show that link aggregation at the host level is much better than link aggregation at either the domain or directory levels.
Gui-Rong Xue, Qiang Yang 0001, Hua-Jun Zeng, Yong Yu 0001, Zheng Chen 0001
SIGIR1
2005 Interactive Chinese Search Results Clustering for Personalization
Gui-Rong Xue, Shen Huang, Yong Yu 0001
WAIM2
2005 Importance-Based Web Page Classification Using Cost-Sensitive SVM
Gui-Rong Xue, Yong Yu 0001, Hua-Jun Zeng
WAIM2
2004 Optimizing web search using web click-through data
abstract
The performance of web search engines may often deteriorate due to the diversity and noisy information contained within web pages. User click-through data can be used to introduce more accurate description (metadata) for web pages, and to improve the search performance. However, noise and incompleteness, sparseness, and the volatility of web pages and are three major challenges for research work on user click-through log mining. In this paper, we propose a novel iterative reinforced algorithm to utilize the user click-through data to improve search performance. The algorithm fully explores the interrelations between and web pages, and effectively finds virtual queries for web pages and overcomes the challenges discussed above. Experiment results on a large set of MSN click-through log data show a significant improvement on search performance over the naive query log mining algorithm as well as the baseline search engine.
Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Weiguo Fan
CIKM1
2004 MRSSA: an iterative algorithm for similarity spreading over interrelated objects
abstract
We introduce the Multiple Relationship Similarity Spreading Algorithm (MRSSA) to enhance IR effectiveness. This method has similarity computed in an iterative "spreading" fashion for multiple object types, combining both inter- and intra-object relationships. We demonstrate the value of this approach in the context of the WWW, where the key objects are web pages and queries, Relationships considered are derived from hyperlinks (in- and out-links) and click-through logs.
Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Edward A. Fox
CIKM1
2004 IRC: An Iterative Reinforcement Categorization Algorithm for Interrelated Web Objects
abstract
Most existing categorization algorithms deal with homogeneous Web data objects, and consider interrelated objects as additional features when taking the interrelationships with other types of objects into account. However, focusing on any single aspects of these interrelationships and objects does not fully reveal their true categories. In this paper, we propose a categorization algorithm, the iterative reinforcement categorization algorithm (IRC), to exploit the full interrelationships between the heterogeneous objects on the Web. IRC attempts to classify the interrelated Web objects by iterative reinforcement between individual classification results of different types via the interrelationships. Experiments on a clickthrough log dataset from MSN search engine show that, with the Fl measures, IRC achieves a 26.4% improvement over a pure content-based classification method, a 21% improvement over a query metadata-based method, and a 16.4% improvement over a virtual document-based method. Furthermore, our experiments show that IRC converges rapidly.
Gui-Rong Xue, Dou Shen, Qiang Yang 0001, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wensi Xi, Wei-Ying Ma
ICDM1
2004 Multi-model similarity propagation and its application for web image retrieval
abstract
In this paper, we propose an iterative similarity propagation approach to explore the inter-relationships between Web images and their textual annotations for image retrieval. By considering Web images as one type of objects, their surrounding texts as another type, and constructing the links structure between them via webpage analysis, we can iteratively reinforce the similarities between images. The basic idea is that if two objects of the same type are both related to one object of another type, these two objects are similar; likewise, if two objects of the same type are related to two different, but similar objects of another type, then to some extent, these two objects are also similar. The goal of our method is to fully exploit the mutual reinforcement between images and their textual annotations. Our experiments based on 10,628 images crawled from the Web show that our proposed approach can significantly improve Web image retrieval performance.
Xin-Jing Wang, Wei-Ying Ma, Gui-Rong Xue, Xing Li 0001
ACM Multimedia3
2004 DHT Based Searching Improved by Sliding Window
Shen Huang, Gui-Rong Xue, Yan-Feng Ge, Yong Yu 0001
WAIM2
2004 TSSP: A Reinforcement Algorithm to Find Related Papers
abstract
Content analysis and citation analysis are two common methods in recommending system. Compared with content analysis, citation analysis can discover more implicitly related papers. However, the citation-based methods may introduce more noise in citation graph and cause topic drift. Some work combine content with citation to improve similarity measurement. The problem is that the two features are not used to reinforce each other to get better result. To solve the problem, we propose a new algorithm, Topic Sensitive Similarity Propagation (TSSP), to effectively integrate content similarity into similarity propagation. TSSP has two parts: citation context based propagation and iterative reinforcement. First, citation contexts provide clues for which papers are topic related to and filter out less irrelevant citations. Second, iteratively integrating content and citation similarity enable them to reinforce each other during the propagation. The experimental results of a user study show TSSP outperforms other algorithms in almost all cases.
Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma
Web Intelligence2
2004 Multi-type Features Based Web Document Clustering
Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma
WISE2
2004 Exploiting PageRank at Different Block Level
Xue-Mei Jiang, Gui-Rong Xue, Wen-Guan Song, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma
WISE2
2004 Optimizing Web Search Using Spreading Activation on the Clickthrough Data
Gui-Rong Xue, Shen Huang, Yong Yu 0001, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma
WISE1
2003 Implicit link analysis for small web search
abstract
Current Web search engines generally impose link analysis-based re-ranking on web-page retrieval. However, the same techniques, when applied directly to small web search such as intranet and site search, cannot achieve the same performance because their link structures are different from the global Web. In this paper, we propose an approach to constructing implicit links by mining users' access patterns, and then apply a modified PageRank algorithm to re-rank web-pages for small web search. Our experimental results indicate that the proposed method outperforms content-based method by 16%, explicit link-based PageRank by 20% and DirectHit by 14%, respectively.
Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma, HongJiang Zhang, Chao-Jun Lu
SIGIR1