Boqin Feng

dblp:80/4882 · DBLP profile ↗
← Back
12ranked-venue papers in the field
0as first author
1since 2021 · last 2021
0000-0002-4879-6090ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 4Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 2Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2021 A Multi-label Propagation Community Detection Algorithm for Dynamic Complex Networks
Hanning Zhang, Bo Dong 0001, Haiyu Wu, Boqin Feng
CAiSE4
2014 Motif-Based Hyponym Relation Extraction from Wikipedia Hyperlinks
abstract
Discovering hyponym relations among domain-specific terms is a fundamental task in taxonomy learning and knowledge acquisition. However, the great diversity of various domain corpora and the lack of labeled training sets make this task very challenging for conventional methods that are based on text content. The hyperlink structure of Wikipedia article pages was found to contain recurring network motifs in this study, indicating the probability of a hyperlink being a hyponym hyperlink. Hence, a novel hyponym relation extraction approach based on the network motifs of Wikipedia hyperlinks was proposed. This approach automatically constructs motif-based features from the hyperlink structure of a domain; every hyperlink is mapped to a 13-dimensional feature vector based on the 13 types of three-node motifs. The approach extracts structural information from Wikipedia and heuristically creates a labeled training set. Classification models were determined from the training sets for hyponym relation extraction. Two experiments were conducted to validate our approach based on seven domain-specific datasets obtained from Wikipedia. The first experiment, which utilized manually labeled data, verified the effectiveness of the motif-based features. The second experiment, which utilized an automatically labeled training set of different domains, showed that the proposed approach performs better than the approach based on lexico-syntactic patterns and achieves comparable result to the approach based on textual features. Experimental results show the practicability and fairly good domain scalability of the proposed approach.
Bifan Wei, Jun Liu 0002, Wei Zhang 0053, Boqin Feng
IEEE Trans. Knowl. Data Eng.6
2011 QoRank: A query-dependent ranking model using LSE-based weighted multiple hyperplanes aggregation for information retrieval
abstract
Ranking is a core problem for information retrieval since the performance of the search system is directly impacted by the accuracy of ranking results. Ranking model construction has been the focus of both the fields of information retrieval and machine learning, and learning to rank in particular has attracted much interest. Many ranking models have been proposed, for example, RankSVM is a state-of-the-art method for learning to rank and has been empirically demonstrated to be effective. However, most of the proposed methods do not consider about the significant differences between queries, only resort to a single function in ranking. In this paper, we present a novel ranking model named QoRank, which performs the learning task dependent on queries. We also propose a LSE (least-squares estimation) -based weighted method to aggregate the ranking lists produced by base decision functions as the final ranking. Comparison of QoRank with other ranking techniques is conducted, and several evaluation criteria are employed to evaluate its performance. Experimental results on the LETOR OHSUMED data set show that QoRank strikes a good balance of accuracy and complexity, and outperforms the baseline methods. © 2010 Wiley Periodicals, Inc.
Heli Sun, Boqin Feng
Int. J. Intell. Syst.3
2010 gSkeletonClu: Density-Based Network Clustering via Structure-Connected Tree Division or Agglomeration
abstract
Community detection is an important task for mining the structure and function of complex networks. Many pervious approaches are difficult to detect communities with arbitrary size and shape, and are unable to identify hubs and outliers. A recently proposed network clustering algorithm, SCAN, is effective and can overcome this difficulty. However, it depends on a sensitive parameter: minimum similarity threshold ε, but provides no automated way to find it. In this paper, we propose a novel density-based network clustering algorithm, called gSkeletonClu (graph-skeleton based clustering). By projecting a network to its Core-Connected Maximal Spanning Tree (CCMST), the network clustering problem is converted to finding core-connected components in the CCMST. We discover that all possible values of the parameter ε lie in the edge weights of the corresponding CCMST. By means of tree divisive or agglomerative clustering, our algorithm can find the optimal parameter ε and detect communities, hubs and outliers in large-scale undirected networks automatically without any user interaction. Extensive experiments on both real-world and synthetic networks demonstrate the superior performance of gSkeletonClu over the baseline methods.
Heli Sun, Jiawei Han 0001, Hongbo Deng, Peixiang Zhao 0001, Boqin Feng
ICDM6
2009 OrdRank: Learning to Rank with Ordered Multiple Hyperplanes
abstract
Ranking is a central problem for information retrieval systems, because the performance of an information retrieval system is mainly evaluated by the effectiveness of its ranking results. Learning to rank has received much attention in recent years due to its importance in information retrieval. This paper focuses on learning to rank in document retrieval and presents a ranking model named OrdRank that ranks documents with ordered multiple hyperplanes. Comparison of OrdRank with other state-of-the-art ranking techniques is conducted and several evaluation criteria are employed to evaluate its performance. Experimental results on the OHSUMED dataset show that OrdRank outperforms other methods, both in terms of quality of ranking results and efficiency.
Heli Sun, Boqin Feng, Yingliang Zhao, Jun Liu 0002
Web Intelligence3
2008 A Co-occurrence Based Hierarchical Method for Clustering Web Search Results
abstract
This study proposes a novel method to group and organize search results. We apply statistical techniques to term co-occurrence information in a corpus to retrieve bi-grams firstly, and then combine bi-grams into n-grams. After eliminating redundant n-grams, the remaining ones are ranked and selected as cluster labels. Base clusters are constructed according to these cluster labels and then agglomerated into higher-level clusters. We refer to the proposed algorithm as CoHC (co-occurrence based hierarchical clustering). we compare CoHC with three other search results clustering (SRC) algorithms: suffix tree clustering (STC), Lingo, and Vivisimo. We also analyze the properties of cluster labels produced by different SRC algorithms. The experimental results show that our method outperforms the other three SRC algorithms, and is helpful to the user for browsing and locating the results of interest.
Boqin Feng
Web Intelligence2
2005 Retrieval Based on Language Model with Relative Entropy and Feedback
Hua Huo, Boqin Feng
PAKDD2
2004 Support Vector Machines Learning for Web-Based Adaptive and Active Information Retrieval
Zhaofeng Ma, Boqin Feng
APWeb2
2004 Collaborative Filtering Algorithm Based on Mutual Information
Boqin Feng
APWeb2
2004 Fail-Stop Authentication Protocol for Digital Rights Management in Adaptive Information Retrieval System
Zhaofeng Ma, Boqin Feng
WAIM2
2004 Adaptive Load Balancing and Fault Tolerance in Push-Based Information Delivery Middleware Service
Zhaofeng Ma, Boqin Feng
WAIM2
2004 XML-Based Meta-Retrieval of Networked Data Sources
abstract
The XML-based meta-retrieval described in this paper is an integrated retrieval over multiple academic data sources on the premise that each data source has one web interface visible in the Internet environment. The goal of this model is to present an integrated retrieval interface in the Internet, to offer transparent accessing by formulating new queries suitable for heterogeneous data sources based on three adaptation rules, and to return the consolidated data results retrieved from multiple data sources. Technologies like eXtensible Markup Language(XML) and eXtensible Stylesheet Language Family(XSL) are used to facilitate the fulfillment of the goal. Experiment results prove that the model solve the heterogeneity of the data sources and provide transparent accessing for the end users.
Boqin Feng
Web Intelligence3