Jian-Tao Sun

dblp:75/3428 · DBLP profile ↗
← Back
46ranked-venue papers
6as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 39 · 6 first-authorArtificial intelligence and machine learning · 25 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
23 papers
Information retrieval · 51% Data mining · 16% Recommender systems · 10%
Artificial intelligence
9 papers
Information extraction and text analysis · 36% Transfer learning and domain adaptation · 23% Learning theory · 11%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 100%

Topics — the 30 heaviest of 75, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
click-through rate prediction
0.422014
Exploiting contextual factors for click modeling in sponsored search · WSDM 2014
Estimating ad group performance in sponsored search · WSDM 2014
Information retrieval › online advertising
sponsored search
0.422014
Exploiting contextual factors for click modeling in sponsored search · WSDM 2014
Estimating ad group performance in sponsored search · WSDM 2014
Data integration and cleaning
missing data
0.322014
Searching Dimension Incomplete Databases · IEEE Trans. Knowl. Data Eng. 2014
Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009
Query processing and optimization
similarity query processing
0.322014
Searching Dimension Incomplete Databases · IEEE Trans. Knowl. Data Eng. 2014
Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009
Data mining › text mining › topic model
multilingual topic models
0.222011
Cross lingual text classification by mining multilingual topics from wikipedia · WSDM 2011
Mining multilingual topics from wikipedia · WWW 2009
Information retrieval › user behavior › search behavior
click model
0.212014
Exploiting contextual factors for click modeling in sponsored search · WSDM 2014
Information retrieval
online advertising
0.212014
Estimating ad group performance in sponsored search · WSDM 2014
Information retrieval
query understanding
0.222009
Understanding user's query intent with wikipedia · WWW 2009
Context-aware query classification · SIGIR 2009
Information retrieval › query understanding
query classification
0.222009
Context-aware query classification · SIGIR 2009
Building bridges for web query classification · SIGIR 2006
Machine learning › Transfer learning and domain adaptation › cross-domain learning
multi-domain learning
0.112012
Multi-domain active learning for text classification · KDD 2012
Information retrieval
image retrieval
0.112012
Browse-to-search · ACM Multimedia 2012
Information retrieval › image retrieval
product image retrieval
0.112012
Browse-to-search · ACM Multimedia 2012
Interaction techniques and input
gesture input
0.112012
Browse-to-search · ACM Multimedia 2012
Interaction techniques and input › touch interaction
touch gesture
0.112012
Browse-to-search · ACM Multimedia 2012
Information retrieval
retrieval models
0.122008
DirichletRank: Solving the zero-one gap problem of PageRank · ACM Trans. Inf. Syst. 2008
Supervised Latent Semantic Indexing for Document Categorization · ICDM 2004
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.112011
Distance Metric Learning under Covariate Shift · IJCAI 2011
Natural language and speech › Information extraction and text analysis › text classification › transfer learning for text classification
cross-lingual text classification
0.112011
Cross lingual text classification by mining multilingual topics from wikipedia · WSDM 2011
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.112011
Distance Metric Learning under Covariate Shift · IJCAI 2011
Data mining › text mining
topic modeling
0.112011
Cross lingual text classification by mining multilingual topics from wikipedia · WSDM 2011
Natural language and speech › Information extraction and text analysis
text classification
0.132012
Text Classification Improved through Automatically Extracted Sequences · ICDE 2006
Multi-domain active learning for text classification · KDD 2012
Supervised Latent Semantic Indexing for Document Categorization · ICDM 2004
Information retrieval › retrieval models › latent semantic models
latent semantic indexing
0.122006
Latent semantic analysis for multiple-type interrelated data objects · SIGIR 2006
Supervised Latent Semantic Indexing for Document Categorization · ICDM 2004
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
cross-domain sentiment classification
0.112010
Cross-domain sentiment classification via spectral feature alignment · WWW 2010
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment classification
0.112010
Cross-domain sentiment classification via spectral feature alignment · WWW 2010
Machine learning and data management
active learning
0.112009
Effective multi-label active learning for text classification · KDD 2009
Machine learning and data management › active learning
multi-label active learning
0.112009
Effective multi-label active learning for text classification · KDD 2009
Data mining › predictive modeling › classification
multi-label classification
0.112009
Effective multi-label active learning for text classification · KDD 2009
Knowledge graphs
multilingual knowledge base
0.112009
Mining multilingual topics from wikipedia · WWW 2009
Data mining
pattern mining
0.112009
Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009
Query processing and optimization › probabilistic query processing
probabilistic similarity query
0.112009
Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009
Information retrieval › query understanding › query classification
query intent classification
0.112009
Understanding user's query intent with wikipedia · WWW 2009

Methods — techniques the papers use, named apart from their topics

topic modeling · 0.3visual-based search · 0.3regression · 0.2probabilistic framework · 0.2filtering bounds · 0.2probabilistic modeling · 0.2shared subspace learning · 0.1kernel methods · 0.1basis set construction · 0.1active learning · 0.1importance weighting · 0.1importance sampling · 0.1convex optimization · 0.1spectral feature alignment · 0.1co-clustering · 0.1version space · 0.1support vector machine · 0.1probability triangle inequality · 0.1
YearPublicationVenuePosition
2014 Estimating ad group performance in sponsored search
abstract
In modern commercial search engines, the pay-per-click (PPC) advertising model is widely used in sponsored search. The search engines try to deliver ads which can produce greater click yields (the total number of clicks for the list of ads per impression). Therefore, predicting user clicks plays a critical role in sponsored search. The current ad-delivery strategy is a two-step approach which first predicts individual ad CTR for the given query and then selects the ads with higher predicted CTR. However, this strategy is naturally suboptimal and correlation between ads is often ignored under this strategy. The learning problem is focused on predicting individual performance rather than group performance which is the more important measurement.
Dawei Yin 0001, Bin Cao 0001, Jian-Tao Sun, Brian D. Davison 0001
WSDM3
2014 Exploiting contextual factors for click modeling in sponsored search
abstract
Sponsored search is the primary business for today's commercial search engines. Accurate prediction of the Click-Through Rate (CTR) for ads is key to displaying relevant ads to users. In this paper, we systematically study the two kinds of contextual factors influencing the CTR: 1) In micro factors, we focus on the factors for mainline ads, including ad depth, query diversity, ad interaction. 2) In macro factors, we try to understand the correlations of clicks between organic search and sponsored search. Based on this data analysis, we propose novel click models which harvest these new explored factors. To the best of our knowledge, this is the first paper to examine and model the effects of the above contextual factors in sponsored search. Extensive experiments on large-scale real-world datasets show that by incorporating these contextual factors, our novel click models can outperform state-of-the-art methods.
Dawei Yin 0001, Shike Mei, Bin Cao 0001, Jian-Tao Sun, Brian D. Davison 0001
WSDM4
2014 Searching Dimension Incomplete Databases
abstract
Similarity query is a fundamental problem in database, data mining and information retrieval research. Recently, querying incomplete data has attracted extensive attention as it poses new challenges to traditional querying techniques. The existing work on querying incomplete data addresses the problem where the data values on certain dimensions are unknown. However, in many real-life applications, such as data collected by a sensor network in a noisy environment, not only the data values but also the dimension information may be missing. In this work, we propose to investigate the problem of similarity search on dimension incomplete data. A probabilistic framework is developed to model this problem so that the users can find objects in the database that are similar to the query with probability guarantee. Missing dimension information poses great computational challenge, since all possible combinations of missing dimensions need to be examined when evaluating the similarity between the query and the data objects. We develop the lower and upper bounds of the probability that a data object is similar to the query. These bounds enable efficient filtering of irrelevant data objects without explicitly examining all missing dimension combinations. A probability triangle inequality is also employed to further prune the search space and speed up the query process. The proposed probabilistic framework and techniques can be applied to both whole and subsequence queries. Extensive experimental results on real-life data sets demonstrate the effectiveness and efficiency of our approach.
Wei Cheng 0002, Xiaoming Jin, Jian-Tao Sun, Xuemin Lin 0001, Xiang Zhang 0001, Wei Wang 0010
IEEE Trans. Knowl. Data Eng.3
2013 Sentiment Topic Model with Decomposed Prior
abstract
This paper deals with the problem of jointly mining topics, sentiments, and the association between them from online reviews in an unsupervised way. Previous methods often treat a sentiment as a special topic and assume a word is generated from a flat mixture of topics, where the discriminative performance of sentiment analysis is not satisfied. A key reason is that providing rich priors on the polarity of a word for the flat mixture is difficult as the polarity often depends on the topic. To solve the problem we propose a novel model. We decompose the generative process of a word's sentiment polarity to a two-level hierarchy: the first level determines whether a word is used as a sentiment word or just an ordinary topic word, and the second level (if the word is used as a sentiment word) determines the polarity of it. With the decomposition, we provide separate prior for the the first level to encourage the discrimination between sentiment words and ordinary topic words. This prior is relatively easy to obtain compared to the concrete prior of the word polarities. We construct the prior based on part-of-speech tags of words and embed the prior into the model. Experiments on four real online review data sets show that our model consistently outperforms previous methods in the task of sentiment analysis, and simultaneously performs well in the sub-tasks of discovering ordinary topics, sentiment-specific topics, and extracting topic-specific sentiment words.
Zheng Chen 0001, Chengtao Li, Jian-Tao Sun
SDM3
2012 Probabilistic sequential POIs recommendation via check-in data
abstract
While on the go, people are using their phones as a personal concierge discovering what is around and deciding what to do. Mobile phone has become a recommendation terminal customized for individuals. While existing research predominantly focuses on one-step recommendation---recommending the next single activity according to current context, this work moves one step beyond by recommending a series of activities, which is a package of sequential Points of Interest (POIs). The recommended POIs are not only relevant to user context (i.e., current location, time, and check-in), but also personalized to his/her check-in history. We presents a probabilistic approach, which is highly motivated from a large-scale commercial mobile check-in data analysis, to ranking a list of sequential POI categories (e.g., "Japanese food" and "bar") and POIs (e.g., "I love sushi"). The approach enables users to plan consecutive activities on the move. Specifically, the probabilistic recommendation approach estimates the transition probability from one POI to another, conditioned on current context and check-in history in a Markov chain. To alleviate the discritization error and sparsity problem, we further introduce context collaboration and integrate prior information. Experiments on over 100k real-world check-in records and 20k POIs validate the effectiveness of the proposed approach.
Jitao Sang 0001, Tao Mei 0001, Jian-Tao Sun, Changsheng Xu, Shipeng Li 0001
SIGSPATIAL/GIS3
2012 Multi-domain active learning for text classification
abstract
Active learning has been proven to be effective in reducing labeling efforts for supervised learning. However, existing active learning work has mainly focused on training models for a single domain. In practical applications, it is common to simultaneously train classifiers for multiple domains. For example, some merchant web sites (like Amazon.com) may need a set of classifiers to predict the sentiment polarity of product reviews collected from various domains (e.g., electronics, books, shoes). Though different domains have their own unique features, they may share some common latent features. If we apply active learning on each domain separately, some data instances selected from different domains may contain duplicate knowledge due to the common features. Therefore, how to choose the data from multiple domains to label is crucial to further reducing the human labeling efforts in multi-domain learning. In this paper, we propose a novel multi-domain active learning framework to jointly select data instances from all domains with duplicate information considered. In our solution, a shared subspace is first learned to represent common latent features of different domains. By considering the common and the domain-specific features together, the model loss reduction induced by each data instance can be decomposed into a common part and a domain-specific part. In this way, the duplicate information across domains can be encoded into the common part of model loss reduction and taken into account when querying. We compare our method with the state-of-the-art active learning approaches on several text classification tasks: sentiment classification, newsgroup classification and email spam filtering. The experiment results show that our method reduces the human labeling efforts by 33.2%, 42.9% and 68.7% on the three tasks, respectively.
Lianghao Li, Xiaoming Jin, Sinno Jialin Pan, Jian-Tao Sun
KDD4
2012 Browse-to-search
abstract
This demonstration presents a novel interactive online shopping application based on visual search technologies. When users want to buy something on a shopping site, they usually have the requirement of looking for related information from other web sites. Therefore users need to switch between the web page being browsed and other websites that provide search results. The proposed application enables users to naturally search products of interest when they browse a web page, and make their even causal purchase intent easily satisfied. The interactive shopping experience is characterized by: 1) in session---it allows users to specify the purchase intent in the browsing session, instead of leaving the current page and navigating to other websites; 2) in context---the browsed web page provides implicit context information which helps infer user purchase preferences; 3) in focus---users easily specify their search interest using gesture on touch devices and do not need to formulate queries in search box; 4) natural-gesture inputs and visual-based search provides users a natural shopping experience. The system is evaluated against a data set consisting of several millions commercial product images.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng, Jian-Tao Sun, Shipeng Li 0001
ACM Multimedia7
2012 Mining Mobile Users' Activities Based on Search Query Text and Context
Bingyue Peng, Jian-Tao Sun
PAKDD (2)3
2012 A classification approach to coreference in discharge summaries: 2011 i2b2 challenge
abstract
OBJECTIVE: To create a highly accurate coreference system in discharge summaries for the 2011 i2b2 challenge. The coreference categories include Person, Problem, Treatment, and Test. DESIGN: An integrated coreference resolution system was developed by exploiting Person attributes, contextual semantic clues, and world knowledge. It includes three subsystems: Person coreference system based on three Person attributes, Problem/Treatment/Test system based on numerous contextual semantic extractors and world knowledge, and Pronoun system based on a multi-class support vector machine classifier. The three Person attributes are patient, relative and hospital personnel. Contextual semantic extractors include anatomy, position, medication, indicator, temporal, spatial, section, modifier, equipment, operation, and assertion. The world knowledge is extracted from external resources such as Wikipedia. MEASUREMENTS: Micro-averaged precision, recall and F-measure in MUC, BCubed and CEAF were used to evaluate results. RESULTS: The system achieved an overall micro-averaged precision, recall and F-measure of 0.906, 0.925, and 0.915, respectively, on test data (from four hospitals) released by the challenge organizers. It achieved a precision, recall and F-measure of 0.905, 0.920 and 0.913, respectively, on test data without Pittsburgh data. We ranked the first out of 20 competing teams. Among the four sub-tasks on Person, Problem, Treatment, and Test, the highest F-measure was seen for Person coreference. CONCLUSIONS: This system achieved encouraging results. The Person system can determine whether personal pronouns and proper names are coreferent or not. The Problem/Treatment/Test system benefits from both world knowledge in evaluating the similarity of two mentions and contextual semantic extractors in identifying semantic clues. The Pronoun system can automatically detect whether a Pronoun mention is coreferent to that of the other four types. This study demonstrates that it is feasible to accomplish the coreference task in discharge summaries.
Yan Xu 0001, Jiahua Liu, Jiajun Wu 0001, Yue Wang 0035, Zhuowen Tu, Jian-Tao Sun, Jun'ichi Tsujii, Eric I-Chao Chang
J. Am. Medical Informatics Assoc.6
2011 Improving context-aware query classification via adaptive self-training
abstract
Topical classification of user queries is critical for general-purpose web search systems. It is also a challenging task, due to the sparsity of query terms and the lack of labeled queries. On the other hand, search contexts embedded in query sessions and unlabeled queries free on the web have not been fully utilized in most query classification systems. In this work, we leverage these information to improve query classification accuracy.
Minmin Chen, Jian-Tao Sun, Xiaochuan Ni 0001, Yixin Chen 0001
CIKM2
2011 Unsupervised transactional query classification based on webpage form understanding
abstract
Query type classification aims to classify search queries into categories like navigational, informational and transactional, etc., according to the type of information need behind the queries. Although this problem has drawn many research attentions, previous methods usually require editors to label queries as training data or need domain knowledge to edit rules for predicting query type. Also, the existing work has been mainly focusing on the classification of informational and navigational query types. Transactional query classification has not been well addressed. In this work, we propose an unsupervised approach for transactional query classification. This method is based on the observation that, after the transactional queries are issued to a search engine, many users will click the search result pages and then have interactions with Web forms on these pages. The interactions, e.g., typing in text box, making selections from dropdown list, clicking on a button to execute actions, are used to specify detailed information of the transaction. By mining toolbar search log data, which records the associations between queries and Web forms clicked by users, we can get a set of good quality transactional queries without using manual labeling efforts. By matching these automatically acquired transactional queries and their associated Web form contents, we can generalize these queries into patterns. These patterns can be used to classify queries which are not covered by search log. Our experiments indicate that transactional queries produced by this method have good quality. The pattern based classifier achieves 83% F1 classification result. This is very effective considering the fact that we do not adopt any labeling efforts to train the classifier.
Xiaochuan Ni 0001, Jian-Tao Sun, Zheng Chen 0001
CIKM3
2011 Representing document as dependency graph for document clustering
abstract
In traditional clustering methods, a document is often represented as "bag of words" (in BOW model) or n-grams (in suffix tree document model) without considering the natural language relationships between the words. In this paper, we propose a novel approach DGDC (Dependency Graph-based Document Clustering algorithm) to address this issue. In our algorithm, each document is represented as a dependency graph where the nodes correspond to words which can be seen as meta-descriptions of the document; whereas the edges stand for the relations between pairs of words. A new similarity measure is proposed to compute the pairwise similarity of documents based on their corresponding dependency graphs. By applying the new similarity measure in the Group-average Agglomerative Hierarchial Clustering (GAHC) algorithm, the final clusters of documents can be obtained. The experiments were carried out on five public document datasets. The empirical results have indicated that the DGDC algorithm can achieve better performance in document clustering tasks compared with other approaches based on the BOW model and suffix tree document model.
Xiaochuan Ni 0001, Jian-Tao Sun, Yunhai Tong, Zheng Chen 0001
CIKM3
2011 Distance Metric Learning under Covariate Shift
abstract
Learning distance metrics is a fundamental problem in machine learning. Previous distance-metric learning research assumes that the training and test data are drawn from the same distribution, which may be violated in practical applications. When the distributions differ, a situation referred to as covariate shift, the metric learned from training data may not work well on the test data. In this case the metric is said to be inconsistent. In this paper, we address this problem by proposing a novel metric learning framework known as consistent distance metric learning (CDML), which solves the problem under covariate shift situations. We theoretically analyze the conditions when the metrics learned under covariate shift are consistent. Based on the analysis, a convex optimization problem is proposed to deal with the CDML problem. An importance sampling method is proposed for metric learning and two importance weighting strategies are proposed and compared in this work. Experiments are carried out on synthetic and real world datasets to show the effectiveness of the proposed method.
Bin Cao 0001, Xiaochuan Ni 0001, Jian-Tao Sun, Gang Wang 0010, Qiang Yang 0001
IJCAI3
2011 Cross lingual text classification by mining multilingual topics from wikipedia
abstract
This paper investigates how to effectively do cross lingual text classification by leveraging a large scale and multilingual knowledge base, Wikipedia. Based on the observation that each Wikipedia concept is described by documents of different languages, we adapt existing topic modeling algorithms for mining multilingual topics from this knowledge base. The extracted topics have multiple types of representations, with each type corresponding to one language. In this work, we regard such topics extracted from Wikipedia documents as universal-topics, since each topic corresponds with same semantic information of different languages. Thus new documents of different languages can be represented in a space using a group of universal-topics. We use these universal-topics to do cross lingual text classification. Given the training data labeled for one language, we can train a text classifier to classify the documents of another language by mapping all documents of both languages into the universal-topic space. This approach does not require any additional linguistic resources, like bilingual dictionaries, machine translation tools, or labeling data for the target language. The evaluation results indicate that our topic modeling approach is effective for building cross lingual text classifier.
Xiaochuan Ni 0001, Jian-Tao Sun, Jian Hu 0001, Zheng Chen 0001
WSDM2
2011 Learning bidirectional asymmetric similarity for collaborative filtering via matrix factorization
Bin Cao 0001, Qiang Yang 0001, Jian-Tao Sun, Zheng Chen 0001
Data Min. Knowl. Discov.3
2010 Cross-domain sentiment classification via spectral feature alignment
abstract
Sentiment classification aims to automatically predict sentiment polarity (e.g., positive or negative) of users publishing sentiment data (e.g., reviews, blogs). Although traditional classification algorithms can be used to train sentiment classifiers from manually labeled text data, the labeling work can be time-consuming and expensive. Meanwhile, users often use some different words when they express sentiment in different domains. If we directly apply a classifier trained in one domain to other domains, the performance will be very low due to the differences between these domains. In this work, we develop a general solution to sentiment classification when we do not have any labels in a target domain but have some labeled data in a different domain, regarded as source domain. In this cross-domain sentiment classification setting, to bridge the gap between the domains, we propose a spectral feature alignment (SFA) algorithm to align domain-specific words from different domains into unified clusters, with the help of domain-independent words as a bridge. In this way, the clusters can be used to reduce the gap between domain-specific words of the two domains, which can be used to train sentiment classifiers in the target domain accurately. Compared to previous approaches, SFA can discover a robust representation for cross-domain data by fully exploiting the relationship between the domain-specific and domain-independent words via simultaneously co-clustering them in a common latent space. We perform extensive experiments on two real world datasets, and demonstrate that SFA significantly outperforms previous approaches to cross-domain sentiment classification.
Sinno Jialin Pan, Xiaochuan Ni 0001, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
WWW3
2009 Context-Aware Online Commercial Intention Detection
Derek Hao Hu, Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
ACML3
2009 PQC: personalized query classification
abstract
Query classification (QC) is a task that aims to classify Web queries into topical categories. Since queries are usually short in length and ambiguous, the same query may need to be classified to different categories according to different people's perspectives. In this paper, we propose the Personalized Query Classification (PQC) task and develop an algorithm based on user preference learning as a solution. Users' preferences that are hidden in clickthrough logs are quite helpful for search engines to improve their understandings of users' queries. We propose to connect query classification with users' preference learning from clickthrough logs for PQC. To tackle the sparseness problem in clickthrough logs, we propose a collaborative ranking model to leverage similar users' information. Experiments on a real world clickthrough log data show that our proposed PQC algorithm can gain significant improvement compared with general QC as well as natural baselines. Our method can be applied to a wide range of applications including personalized search and online advertising.
Bin Cao 0001, Jian-Tao Sun, Evan Wei Xiang, Derek Hao Hu, Qiang Yang 0001, Zheng Chen 0001
CIKM2
2009 Exploiting term relationship to boost text classification
abstract
Document classification provides an effective way to handle the explosive online textual data. However, in practical classification settings, we face the so-called feature sparsity problem caused by a lack of training documents or the shortness of text to be classified. In this paper, we solve the sparsity problem by exploiting term relationships along with Naive Bayes classifiers. The first method is to estimate term relationships based on the co-occurrence information of two terms in a certain context. The second method estimates the term relationships based on the distribution of terms over different hierarchical categories in a publicly available document taxonomy. Thereafter, term relationship is used to augment Naive Bayes classifiers. We test our methods on two open-domain data sets to demonstrate its advantages. The experimental results show that our method can significantly improve the classification performance, especially when we do not have enough training data or the texts are Web search queries.
Dou Shen, Jianmin Wu, Bin Cao 0001, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001, Ying Li 0040
CIKM4
2009 Probabilistic Similarity Query on Dimension Incomplete Data
abstract
Retrieving similar data has drawn many research efforts in the literature due to its importance in data mining, database and information retrieval. This problem is challenging when the data is incomplete. In previous research, data incompleteness refers to the fact that data values for some dimensions are unknown. However, in many practical applications (e.g., data collection by sensor network under bad environment), not only data values but even data dimension information may also be missing, which will make most similarity query algorithms infeasible. In this work, we propose the novel similarity query problem on dimension incomplete data and adopt a probabilistic framework to model this problem. For this problem, users can give a distance threshold and a probability threshold to specify their retrieval requirements. The distance threshold is used to specify the allowed distance between query and data objects and the probability threshold is used to require that the retrieval results satisfy the distance condition at least with the given probability. Instead of enumerating all possible cases to recover the missed dimensions, we propose an efficient approach to speed up the retrieval process by leveraging the inherent relations between query and dimension incomplete data objects. During the query process, we estimate the lower/upper bounds of the probability that the query is satisfied by a given data object, and utilize these bounds to filter irrelevant data objects efficiently. Furthermore, a probability triangle inequality is proposed to further speed up query processing. According to our experiments on real data sets, the proposed similarity query method is verified to be effective and efficient on dimension incomplete data.
Wei Cheng 0002, Xiaoming Jin, Jian-Tao Sun
ICDM3
2009 Effective multi-label active learning for text classification
abstract
Labeling text data is quite time-consuming but essential for automatic text classification. Especially, manually creating multiple labels for each document may become impractical when a very large amount of data is needed for training multi-label text classifiers. To minimize the human-labeling efforts, we propose a novel multi-label active learning approach which can reduce the required labeled data without sacrificing the classification accuracy. Traditional active learning algorithms can only handle single-label problems, that is, each data is restricted to have one label. Our approach takes into account the multi-label information, and select the unlabeled data which can lead to the largest reduction of the expected model loss. Specifically, the model loss is approximated by the size of version space, and the reduction rate of the size of version space is optimized with Support Vector Machines (SVM). An effective label prediction method is designed to predict possible labels for each unlabeled data point, and the expected loss for multi-label data is approximated by summing up losses on all labels according to the most confident result of label prediction. Experiments on several real-world data sets (all are publicly available) demonstrate that our approach can obtain promising classification result with much fewer labeled data than state-of-the-art methods.
Bishan Yang, Jian-Tao Sun, Tengjiao Wang 0003, Zheng Chen 0001
KDD2
2009 Context-aware query classification
abstract
Understanding users'search intent expressed through their search queries is crucial to Web search and online advertisement. Web query classification (QC) has been widely studied for this purpose. Most previous QC algorithms classify individual queries without considering their context information. However, as exemplified by the well-known example on query "jaguar", many Web queries are short and ambiguous, whose real meanings are uncertain without the context information. In this paper, we incorporate context information into the problem of query classification by using conditional random field (CRF) models. In our approach, we use neighboring queries and their corresponding clicked URLs (Web pages) in search sessions as the context information. We perform extensive experiments on real world search logs and validate the effectiveness and effciency of our approach. We show that we can improve the F1 score by 52% as compared to other state-of-the-art baselines.
Huanhuan Cao, Derek Hao Hu, Dou Shen, Daxin Jiang, Jian-Tao Sun, Enhong Chen, Qiang Yang 0001
SIGIR5
2009 Understanding user's query intent with wikipedia
abstract
Understanding the intent behind a user's query can help search engine to automatically route the query to some corresponding vertical search engines to obtain particularly relevant contents, thus, greatly improving user satisfaction. There are three major challenges to the query intent classification problem: (1) Intent representation; (2) Domain coverage and (3) Semantic interpretation. Current approaches to predict the user's intent mainly utilize machine learning techniques. However, it is difficult and often requires many human efforts to meet all these challenges by the statistical machine learning approaches. In this paper, we propose a general methodology to the problem of query intent classification. With very little human effort, our method can discover large quantities of intent concepts by leveraging Wikipedia, one of the best human knowledge base. The Wikipedia concepts are used as the intent representation space, thus, each intent domain is represented as a set of Wikipedia articles and categories. The intent of any input query is identified through mapping the query into the Wikipedia representation space. Compared with previous approaches, our proposed method can achieve much better coverage to classify queries in an intent domain even through the number of seed intent examples is very small. Moreover, the method is very general and can be easily applied to various intent domains. We demonstrate the effectiveness of this method in three different applications, i.e., travel, job, and person name. In each of the three cases, only a couple of seed intent queries are provided. We perform the quantitative evaluations in comparison with two baseline methods, and the experimental results shows that our method significantly outperforms other methods in each intent domain.
Jian Hu 0001, Gang Wang 0004, Frederick H. Lochovsky, Jian-Tao Sun, Zheng Chen 0001
WWW4
2009 Mining multilingual topics from wikipedia
abstract
In this paper, we try to leverage a large-scale and multilingual knowledge base, Wikipedia, to help effectively analyze and organize Web information written in different languages. Based on the observation that one Wikipedia concept may be described by articles in different languages, we adapt existing topic modeling algorithm for mining multilingual topics from this knowledge base. The extracted 'universal' topics have multiple types of representations, with each type corresponding to one language. Accordingly, new documents of different languages can be represented in a space using a group of universal topics, which makes various multilingual Web applications feasible.
Xiaochuan Ni 0001, Jian-Tao Sun, Jian Hu 0001, Zheng Chen 0001
WWW2
2008 Learning Bidirectional Similarity for Collaborative Filtering
Bin Cao 0001, Jian-Tao Sun, Jianmin Wu, Qiang Yang 0001, Zheng Chen 0001
ECML/PKDD (1)2
2008 DirichletRank: Solving the zero-one gap problem of PageRank
abstract
Link-based ranking algorithms are among the most important techniques to improve web search. In particular, the PageRank algorithm has been successfully used in the Google search engine and has been attracting much attention recently. However, we find that PageRank has a “zero-one gap” problem which, to the best of our knowledge, has not been addressed in any previous work. This problem can be potentially exploited to spam PageRank results and make the state-of-the-art link-based antispamming techniques ineffective. The zero-one gap problem arises as a result of the current ad hoc way of computing transition probabilities in the random surfing model. We therefore propose a novel DirichletRank algorithm which calculates these probabilities using Bayesian estimation with a Dirichlet prior. DirichletRank is a variant of PageRank, but does not have the problem of zero-one gap and can be analytically shown substantially more resistant to some link spams than PageRank. Experiment results on TREC data show that DirichletRank can achieve better retrieval accuracy than PageRank due to its more reasonable allocation of transition probabilities. More importantly, experiments on the TREC dataset and another real web dataset from the Webgraph project show that, compared with the original PageRank, DirichletRank is more stable under link perturbation and is significantly more robust against both manually identified web spams and several simulated link spams. DirichletRank can be computed as efficiently as PageRank, and thus is scalable to large-scale web applications.
Xuanhui Wang, Tao Tao 0003, Jian-Tao Sun, Azadeh Shakery, ChengXiang Zhai
ACM Trans. Inf. Syst.3
2007 Shine: search heterogeneous interrelated entities
abstract
Heterogeneous entities or objects are very common and are usually interrelated with each other in many scenarios. For example, typical Web search activities involve multiple types of interrelated entities such as end users, Web pages, and search queries. In this paper, we define and study a novel problem: S earch H eterogeneous IN terrelated E ntities (SHINE). Given a SHINE-query which can be any type(s) of entities, the task of SHINE is to retrieve multiple types of related entities to answer this query. This is in contrast to the traditional search,which only deals with a single type of entities (e.g., Web pages). The advantages of SHINE include: (1) It is feasible for end users to specify their information need along different dimensions by accepting queries with different types. (2) Answering a query by multiple types of entities provides informative context for users to better understand the search results and facilitate their information exploration. (3) Multiple relations among heterogeneous entities can be utilized to improve the ranking of any particular type of entities. To attain the goal of SHINE, we propose to represent all entities in a unified space through utilizing their interaction relationships. Two approaches, M-LSA and E-VSM, are discussed and compared in this paper. The experiments on 3 data sets (i.e., a literature data set, a search engine log data set, and a recommendation data set) show the effectiveness and flexibility of our proposed methods.
Xuanhui Wang, Jian-Tao Sun, Zheng Chen 0001
CIKM2
2007 Feature selection in a kernel space
abstract
We address the problem of feature selection in a kernel space to select the most discriminative and informative features for classification and data analysis. This is a difficult problem because the dimension of a kernel space may be infinite. In the past, little work has been done on feature selection in a kernel space. To solve this problem, we derive a basis set in the kernel space as a first step for feature selection. Using the basis set, we then extend the margin-based feature selection algorithms that are proven effective even when many features are dependent. The selected features form a subspace of the kernel space, in which different state-of-the-art classification algorithms can be applied for classification. We conduct extensive experiments over real and simulated data to compare our proposed method with four baseline algorithms. Both theoretical analysis and experimental results validate the effectiveness of our proposed method.
Bin Cao 0001, Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
ICML3
2007 Detect and Track Latent Factors with Online Nonnegative Matrix Factorization
Bin Cao 0001, Dou Shen, Jian-Tao Sun, Xuanhui Wang, Qiang Yang 0001, Zheng Chen 0001
IJCAI3
2007 Document Summarization Using Conditional Random Fields
Dou Shen, Jian-Tao Sun, Hua Li 0001, Qiang Yang 0001, Zheng Chen 0001
IJCAI2
2006 Text classification improved through multigram models
abstract
Classification algorithms and document representation approaches are two key elements for a successful document classification system. In the past, much work has been conducted to find better ways to represent documents. However, most of the attempts rely on certain extra resources such as WordNet, or they face the problem of extremely high dimension. In this paper, we propose a new document representation approach based on n-multigram language models. This approach can automatically discover the hidden semantic sequences in the documents under each category. Based on n-multigram language models and n-gram language models, we put forward two text classification algorithms. The experiments on RCV1 show that our proposed algorithm based on n-multigram models alone can achieve the similar or even better classification performance compared with the classifier based on n-gram models but the model size of our algorithm is much smaller than that of the latter. Another proposed algorithm based on the combination of n-multigram models and n-gram models improves the micro-F1 and macro-F1 values from 89.5% to 92.6% and 87.2% to 91.1% respectively. All these observations support the validity of our proposed document representation approach.
Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
CIKM2
2006 Text Classification Improved through Automatically Extracted Sequences
abstract
We propose to use the n-multigram model to help the automatic text classification task. This model could automatically discover the latent semantic sequences contained in the document set of each category. Based on the n-multigram model and the n-gram language model, we put forward two text classification algorithms. The experiments on RCV1 show that our proposed algorithm based on n-multigram model can achieve the similar classification performance compared with the one based on n-gram model. However, the model size of our algorithm is only 4.21% of the latter one. Another proposed algorithm based on the combination of nmultigram model and n-gram model improves the micro- F1 and macro-F1 values by 3.5% and 4.5% respectively which support the validity of our approach.
Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
ICDE2
2006 Subjectivity Categorization of Weblog with Part-of-Speech Based Smoothing
abstract
Experts from different domains try to mine users' comments on Weblogs for different reasons such as politics or commerce. All these needs necessitate automatically distinguishing subjective Weblog contents from objective ones, namely subjectivity categorization. Since Weblogs contain various topics from different domains, limited training data can hardly cover all the topics and "unseen words" becomes a serious problem for categorization tasks. In this paper, part-of-speech (POS) based smoothing is proposed to alleviate the "unseen words" problem. In conjunction with a naive Bayes model constructed from limited training data, the probability of an unseen word in a new domain can be well smoothed by the probability of its POS result. Empirical studies on five datasets show that our approach consistently outperforms the basic naive Bayes with Laplace smoothing. In a cross-domain experiment, our approach achieves 22.0% improvement in Macro Fl and 24.4% in Micro Fl over basic naive Bayes. These verify that POS based smoothing can indeed benefit subjectivity categorization, especially in the cases with a large number of unseen words.
Shen Huang, Jian-Tao Sun, Xuanhui Wang, Hua-Jun Zeng, Zheng Chen 0001
ICDM2
2006 Latent Friend Mining from Blog Data
abstract
The rapid growth of blog (also known as "weblog") data provides a rich resource for social community mining. In this paper, we put forward a novel research problem of mining the latent friends of bloggers based on the contents of their blog entries. Latent friends are defined in this paper as people who share the similar topic distribution in their blogs. These people may not actually know each other, but they have the interest and potential to find each other out. Three approaches are designed for latent friend detection. The first one, called cosine similarity-based method, determines the similarity between bloggers by calculating the cosine similarity between the contents of the blogs. The second approach, known as topic-based method, is based on the discovery of latent topics using a latent topic model and then calculating the similarity at the topic level. The third one is two-level similarity-based, which is conducted in two stages. In the first stage, an existing topic hierarchy is exploited to build a topic distribution for a blogger. Then, in the second stage, a detailed similarity comparison is conducted for bloggers that are close in interest to each other which are discovered in the first stage. Our experimental results show that both the topic-based and two-level similarity-based methods work well, and the last approach performs much better than the first two. In this paper, we give a detailed analysis of the advantages and disadvantages of different approaches.
Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
ICDM2
2006 Building bridges for web query classification
abstract
Web query classification (QC) aims to classify Web users' queries, which are often short and ambiguous, into a set of target categories. QC has many applications including page ranking in Web search, targeted advertisement in response to queries, and personalization. In this paper, we present a novel approach for QC that outperforms the winning solution of the ACM KDDCUP 2005 competition, whose objective is to classify 800,000 real user queries. In our approach, we first build a bridging classifier on an intermediate taxonomy in an offline mode. This classifier is then used in an online mode to map user queries to the target categories via the above intermediate taxonomy. A major innovation is that by leveraging the similarity distribution over the intermediate taxonomy, we do not need to retrain a new classifier for each new set of target categories, and therefore the bridging classifier needs to be trained only once. In addition, we introduce category selection as a new method for narrowing down the scope of the intermediate taxonomy based on which we classify the queries. Category selection can improve both efficiency and effectiveness of the online classification. By combining our algorithm with the winning solution of KDDCUP 2005, we made an improvement by 9.7% and 3.8% in terms of precision and F1 respectively compared with the best results of KDDCUP 2005.
Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
SIGIR2
2006 Thread detection in dynamic text message streams
abstract
Text message stream is a newly emerging type of Web data which is produced in enormous quantities with the popularity of Instant Messaging and Internet Relay Chat. It is beneficial for detecting the threads contained in the text stream for various applications, including information retrieval, expert recognition and even crime prevention. Despite its importance, not much research has been conducted so far on this problem due to the characteristics of the data in which the messages are usually very short and incomplete. In this paper, we present a stringent definition of the thread detection task and our preliminary solution to it. We propose three variations of a single-pass clustering algorithm for exploiting the temporal information in the streams. An algorithm based on linguistic features is also put forward to exploit the discourse structure information. We conducted several experiments to compare our approaches with some existing algorithms on a real dataset. The results show that all three variations of the single-pass algorithm outperform the basic single-pass algorithm. Our proposed algorithm based on linguistic features improves the performance relatively by 69.5% and 9.7% when compared with the basic single-pass algorithm and the best variation algorithm in terms of F1 respectively.
Dou Shen, Qiang Yang 0001, Jian-Tao Sun, Zheng Chen 0001
SIGIR3
2006 Latent semantic analysis for multiple-type interrelated data objects
abstract
Co-occurrence data is quite common in many real applications. Latent Semantic Analysis (LSA) has been successfully used to identify semantic relations in such data. However, LSA can only handle a single co-occurrence relationship between two types of objects. In practical applications, there are many cases where multiple types of objects exist and any pair of these objects could have a pairwise co-occurrence relation. All these co-occurrence relations can be exploited to alleviate data sparseness or to represent objects more meaningfully. In this paper, we propose a novel algorithm, M-LSA, which conducts latent semantic analysis by incorporating all pairwise co-occurrences among multiple types of objects. Based on the mutual reinforcement principle, M-LSA identifies the most salient concepts among the co-occurrence data and represents all the objects in a unified semantic space. M-LSA is general and we show that several variants of LSA are special cases of our algorithm. Experiment results show that M-LSA outperforms LSA on multiple applications, including collaborative filtering, text clustering, and text categorization.
Xuanhui Wang, Jian-Tao Sun, Zheng Chen 0001, ChengXiang Zhai
SIGIR2
2006 A comparison of implicit and explicit links for web page classification
abstract
It is well known that Web-page classification can be enhanced by using hyperlinks that provide linkages between Web pages. However, in the Web space, hyperlinks are usually sparse, noisy and thus in many situations can only provide limited help in classification. In this paper, we extend the concept of linkages from explicit hyperlinks to implicit links built between Web pages. By observing that people who search the Web with the same queries often click on different, but related documents together, we draw implicit links between Web pages that are clicked after the same queries. Those pages are implicitly linked. We provide an approach for automatically building the implicit links between Web pages using Web query logs, together with a thorough comparison between the uses of implicit and explicit links in Web page classification. Our experimental results on a large dataset confirm that the use of the implicit links is better than using explicit links in classification performance, with an increase of more than 10.5% in terms of the Macro-F1 measurement.
Dou Shen, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001
WWW2
2006 CWS: a comparative web search system
abstract
In this paper, we define and study a novel search problem: Comparative Web Search (CWS). The task of CWS is to seek relevant and comparative information from the Web to help users conduct comparisons among a set of topics. A system called CWS is developed to effectively facilitate Web users' comparison needs. Given a set of queries, which represent the topics that a user wants to compare, the system is characterized by: (1) automatic retrieval and ranking of Web pages by incorporating both their relevance to the queries and the comparative contents they contain; (2) automatic clustering of the comparative contents into semantically meaningful themes; (3) extraction of representative keyphrases to summarize the commonness and differences of the comparative contents in each theme. We developed a novel interface which supports two types of view modes: a pair-view which displays the result in the page level, and a cluster-view which organizes the comparative pages into the themes and displays the extracted phrases to facilitate users' comparison. Experiment results show the CWS system is effective and efficient.
Jian-Tao Sun, Xuanhui Wang, Dou Shen, Hua-Jun Zeng, Zheng Chen 0001
WWW1
2006 Mining clickthrough data for collaborative web search
abstract
This paper is to investigate the group behavior patterns of search activities based on Web search history data, i.e., clickthrough data, to boost search performance. We propose a Collaborative Web Search (CWS) framework based on the probabilistic modeling of the co-occurrence relationship among the heterogeneous web objects: users, queries, and Web pages. The CWS framework consists of two steps: (1) a cube-clustering approach is put forward to estimate the semantic cluster structures of the Web objects; (2) Web search activities are conducted by leveraging the probabilistic relations among the estimated cluster structures. Experiments on a real-world clickthrough data set validate the effectiveness of our CWS approach.
Jian-Tao Sun, Xuanhui Wang, Dou Shen, Hua-Jun Zeng, Zheng Chen 0001
WWW1
2006 Query enrichment for web-query classification
abstract
Web-search queries are typically short and ambiguous. To classify these queries into certain target categories is a difficult but important problem. In this article, we present a new technique called query enrichment, which takes a short query and maps it to intermediate objects. Based on the collected intermediate objects, the query is then mapped to target categories. To build the necessary mapping functions, we use an ensemble of search engines to produce an enrichment of the queries. Our technique was applied to the ACM Knowledge Discovery and Data Mining competition (ACM KDDCUP) in 2005, where we won the championship on all three evaluation metrics (precision, F1 measure, which combines precision and recall, and creativity, which is judged by the organizers) among a total of 33 teams worldwide. In this article, we show that, despite the difficulty of an abundance of ambiguous queries and lack of training data, our query-enrichment technique can solve the problem satisfactorily through a two-phase classification framework. We present a detailed description of our algorithm and experimental evaluation. Our best result for F1 and precision is 42.4% and 44.4%, respectively, which is 9.6% and 24.3% higher than those from the runner-ups, respectively.
Dou Shen, Jian-Tao Sun, Jeffrey Junfeng Pan, Kangheng Wu, Jie Yin 0001, Qiang Yang 0001
ACM Trans. Inf. Syst.3
2005 A practical system of keyphrase extraction for web pages
abstract
Keyphrases can be used to facilitate Web users grasping the main topic(s) of a Web page. We present a practical system of automatic keyphrase extraction for Web pages. In this system, a regression model was first trained based on a set of human-labeled documents. Then it was used to extract keyphrases from new pages automatically. This paper makes three contributions. First, the structure information in a Web page was investigated for keyphrase extraction task. Second, the query log data associated with a Web page collected by a search engine server were used to help keyphrase extraction. Third, a method was put forward in this paper in order to evaluate the similarity of phrases.
Jian-Tao Sun, Hua-Jun Zeng, Kwok-Yan Lam
CIKM2
2005 Web-page summarization using clickthrough data
abstract
Most previous Web-page summarization methods treat a Web page as plain text. However, such methods fail to uncover the full knowledge associated with a Web page needed in building a high-quality summary, because many of these methods do not consider the hidden relationships in the Web. Uncovering the hidden knowledge is important in building good Web-page summarizers. In this paper, we extract the extra knowledge from the clickthrough data of a Web search engine to improve Web-page summarization. Wefirst analyze the feasibility in utilizing the clickthrough data to enhance Web-page summarization and then propose two adapted summarization methods that take advantage of the relationships discovered from the clickthrough data. For those pages that are not covered by the clickthrough data, we design a thematic lexicon approach to generate implicit knowledge for them. Our methods are evaluated on a dataset consisting of manually annotated pages as well as a large dataset that is crawled from the Open Directory Project website. The experimental results indicate that significant improvements can be achieved through our proposed summarizer as compared to the summarizers that do not use the clickthrough data.
Jian-Tao Sun, Dou Shen, Hua-Jun Zeng, Qiang Yang 0001, Yuchang Lu, Zheng Chen 0001
SIGIR1
2005 CubeSVD: a novel approach to personalized Web search
abstract
As the competition of Web search market increases, there is a high demand for personalized Web search to conduct retrieval incorporating Web users' information needs. This paper focuses on utilizing clickthrough data to improve Web search. Since millions of searches are conducted everyday, a search engine accumulates a large volume of clickthrough data, which records who submits queries and which pages he/she clicks on. The clickthrough data is highly sparse and contains different types of objects (user, query and Web page), and the relationships among these objects are also very complicated. By performing analysis on these data, we attempt to discover Web users' interests and the patterns that users locate information. In this paper, a novel approach CubeSVD is proposed to improve Web search. The clickthrough data is represented by a 3-order tensor, on which we perform 3-mode analysis using the higher-order singular value decomposition technique to automatically capture the latent factors that govern the relations among these multi-type objects: users, queries and Web pages. A tensor reconstructed based on the CubeSVD analysis reflects both the observed interactions among these objects and the implicit associations among them. Therefore, Web search activities can be carried out based on CubeSVD analysis. Experimental evaluations using a real-world data set collected from an MSN search engine show that CubeSVD achieves encouraging search results in comparison with some standard methods.
Jian-Tao Sun, Hua-Jun Zeng, Huan Liu 0001, Yuchang Lu, Zheng Chen 0001
WWW1
2004 Supervised Latent Semantic Indexing for Document Categorization
abstract
Latent semantic indexing (LSI) is a successful technology in information retrieval (IR) which attempts to explore the latent semantics implied by a query or a document through representing them in a dimension-reduced space. However, LSI is not optimal for document categorization tasks because it aims to find the most representative features for document representation rather than the most discriminative ones. In this paper, we propose supervised LSI (SLSI) which selects the most discriminative basis vectors using the training data iteratively. The extracted vectors are then used to project the documents into a reduced dimensional space for better classification. Experimental evaluations show that the SLSI approach leads to dramatic dimension reduction while achieving good classification results.
Jian-Tao Sun, Zheng Chen 0001, Hua-Jun Zeng, Yuchang Lu, Chun-Yi Shi, Wei-Ying Ma
ICDM1
2004 GE-CKO: A Method to Optimize Composite Kernels for Web Page Classification
abstract
Most of current researches on Web page classification focus on leveraging heterogeneous features such as plain text, hyperlinks and anchor texts in an effective and efficient way. Composite kernel method is one topic of interest among them. It first selects a bunch of initial kernels, each of which is determined separately by a certain type of features. Then a classifier is trained based on a linear combination of these kernels. In this paper, we propose an effective way to optimize the linear combination of kernels. We proved that this problem is equivalent to solving a generalized eigenvalue problem. And the weight vector of the kernels is the eigenvector associated with the largest eigen-value. A support vector machine (SVM) classifier is then trained based on this optimized combination of kernels. Our experiment on the WebKB dataset has shown the effectiveness of our proposed method.
Jian-Tao Sun, Benyu Zhang, Zheng Chen 0001, Yuchang Lu, Chunyi Shi, Wei-Ying Ma
Web Intelligence1