Weiguo Fan

dblp:f/WeiguoFan · also Weiguo (Patrick) Fan · DBLP profile ↗
← Back
57ranked-venue papers in the field
6as first author
18since 2021 · last 2026
0000-0003-1272-5538ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 38 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 11 (1 first)Database Systems & Data Management · 6 (2 first)Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2026 A reliability-enhanced deep ensemble learning framework for recommendation
Junpeng Guo, Weiguo Fan
Inf. Manag.4
2026 Asymmetric cross-side network effects on online knowledge-sharing platforms and the role of platform recommendations
Wei Liu 0172, Xueying Sun, Weiguo Fan, Zhengfa Yang
Inf. Manag.3
2026 User's online status and knowledge contribution behavior in the Q&A community-Based on core and non-core contributors
Mi Zhou 0001, Weiguo Fan
Inf. Manag.3
2026 MLAFormer: Multi-scale transformer with local convolutional auto-correlation and pre-training for time series forecasting
Ming Gao 0008, Jiafu Tang, Weiguo Fan, Jingmin An
Inf. Process. Manag.4
2026 Integrating managerial and investor textual data for financial distress prediction: A framework combining multi-source financial information fusion network with LLM
Shaoze Cui, Weiguo Fan
Inf. Process. Manag.3
2025 Collaborative local-global context modeling for session-based recommendation
Weiyue Li, Bowei Chen 0001, Ming Gao 0008, Jingmin An, Cheng Chen 0040, Weiguo Fan, Zhiguo Zhu
Inf. Process. Manag.7
2024 Lightweight privacy-preserving authentication mechanism in 5G-enabled industrial cyber physical systems
Xinyin Xiang, Jin Cao 0001, Weiguo Fan
Inf. Sci.3
2023 What makes user-generated content more helpful on social media platforms? Insights from creator interactivity perspective
Qingfeng Zeng, Yu Zhang 0238, Weiguo Fan
Inf. Process. Manag.5
2023 More is better? Understanding the effects of online interactions on patients health anxiety
abstract
Abstract Online health platforms play an important role in chronic disease management. Patients participate in online health platforms to receive and provide health‐related support from each other. However, there remains a debate about whether the influence of social interaction on patient health anxiety is linearly positive. Based on uncertainty, information overload, and the theory of motivational information management, we develop and test a model considering a potential curvilinear relationship between social interaction and health anxiety, as well as a moderating effect of health literacy. We collect patient interaction data from an online health platform based on chronic disease management in China and use text mining and econometrics to test our hypotheses. Specifically, we find an inverted U‐shaped relationship between informational provision and health anxiety. Our results also show that information receipt and emotion provision have U‐shaped relationships with health anxiety. Interestingly, health literacy can effectively alleviate the U‐shaped relationship between information receipt and health anxiety. These findings not only provide new insights into the literature on online patient interactions but also provide decision support for patients and platform managers.
Zhaohua Deng, Guorui Fan, Bin Wang 0001, Weiguo Fan, Shan Liu 0004
J. Assoc. Inf. Sci. Technol.5
2023 Aspect sentiment mining of short bullet screen comments from online TV series
abstract
Abstract Bullet screen comments (BSCs) are user‐generated short comments that appear as real‐time overlays on many video platforms, expressing the audience opinions and emotions about different aspects of the ongoing video. Unlike traditional long comments after a show, BSCs are often incomplete, ambiguous in context, and correlated over time. Current studies in sentiment analysis of BSCs rarely address these challenges, motivating us to develop an aspect‐level sentiment analysis framework. Our framework, BSCNET, is a pre‐trained language encoder‐based deep neural classifier designed to enhance semantic understanding. A novel neighbor context construction method is proposed to uncover latent contextual correlation among BSCs over time, and we also incorporate semi‐supervised learning to reduce labeling costs. The framework increases F1 (Macro) and accuracy by up to 10% and 10.2%, respectively. Additionally, we have developed two novel downstream tasks. The first is noisy BSCs identification, which reached F1 (Macro) and accuracy of 90.1% and 98.3%, respectively, through fine‐tuning the BSCNET. The second is the prediction of future episode popularity, where the MAPE is reduced by 11%–19.0% when incorporating sentiment features. Overall, this study provides a methodology reference for aspect‐level sentiment analysis of BSCs and highlights its potential for viewing experience or forthcoming content optimization.
Jiayue Liu, Ziyao Zhou, Ming Gao 0008, Jiafu Tang, Weiguo Fan
J. Assoc. Inf. Sci. Technol.5
2023 Leveraging User Ideas for Product Innovation in Open Innovation Communities: A Study of Two Stages of the Idea Adoption
abstract
The abundance of user ideas in the open innovation community poses a challenge of information overload for enterprises, making it difficult to effectively adopt ideas for product innovation. Existing studies primarily focus on evaluating the usefulness of user ideas, while overlooking the actual adoption outcomes of these ideas. Based on the information adoption model, this paper proposes a user idea adoption model, which integrates the two stages of idea collection and idea application. The authors creatively discuss the factors influencing the process from “perceived usefulness” to “adoption” of a user idea, and explore the complete mechanism of user idea adoption for product innovation. The results show that the information quality and source credibility of user ideas affect the perceived usefulness. The degree of user demand, innovation complexity, and competition intensity have moderating effects on the relationship between the perceived usefulness of user ideas and the results of user ideas adoption.
Ning Zhang 0016, Wenfei Zhao, Zhiliang Pang, Lifeng He, Weiguo Fan
J. Glob. Inf. Manag.6
2022 Integrating the sentiments of multiple news providers for stock market index movement prediction: A deep learning approach based on evidential reasoning rule
Shaoze Cui, Hongshan Xiao, Weiguo Fan, Hongwu Zhang, Yu Wang 0135
Inf. Sci.4
2022 The more, the better? The effect of feedback and user's past successes on idea implementation in open innovation communities
abstract
Abstract Establishing open innovation communities has evolved as an important product innovation and development strategy for companies. Yet, the success of such communities relies on the successful implementation of many user‐submitted ideas. Although extant literature has examined the impact of user experience and idea characteristics on idea implementation, little is known from the information input perspective, for example, feedback. Based on the information overload theory and knowledge content framework, we propose that the amount and types of feedback content have different effects on the likelihood of subsequent idea implementation, and such effects depend on the level of users' success experience. We tested the research model using a panel logistic model with the data of MIUI Forum. The study results revealed that the amount of feedback has an inverted U‐shaped effect on idea implementation, and such effect is moderated by a user's past success. Moreover, the type of feedback content (cost and benefit‐related feedback and functionality‐related feedback) positively affects idea implementation, and a user's past success positively moderated the above effects. Finally, we discuss the theoretical and practical implications, limitations of our research, and suggestions for future research.
Qian Liu 0013, Zhengfa Yang, Xiaofang Cai, Qianzhou Du, Weiguo Fan
J. Assoc. Inf. Sci. Technol.5
2022 Understanding the spread of COVID-19 misinformation on social media: The effects of topics and a political leader's nudge
abstract
The spread of misinformation on social media has become a major societal issue during recent years. In this work, we used the ongoing COVID-19 pandemic as a case study to systematically investigate factors associated with the spread of multi-topic misinformation related to one event on social media based on the heuristic-systematic model. Among factors related to systematic processing of information, we discovered that the topics of a misinformation story matter, with conspiracy theories being the most likely to be retweeted. As for factors related to heuristic processing of information, such as when citizens look up to their leaders during such a crisis, our results demonstrated that behaviors of a political leader, former US President Donald J. Trump, may have nudged people's sharing of COVID-19 misinformation. Outcomes of this study help social media platform and users better understand and prevent the spread of misinformation on social media.
Xiangyu Wang 0009, Min Zhang 0014, Weiguo Fan, Kang Zhao 0001
J. Assoc. Inf. Sci. Technol.3
2021 Mining product innovation ideas from online reviews
Min Zhang 0014, Brandon Fan, Ning Zhang 0016, Weiguo Fan
Inf. Process. Manag.5
2021 Predicting crowdfunding project success based on backers' language preferences
abstract
Abstract Project success is critical in the crowdfunding domain. Rather than the existing project‐centric prediction methods, we propose a novel backer‐centric prediction method. We identify each backer's preferences based on their pledge history and calculate the cosine similarity between backer's preferences and the project as each backer's persuasibility. Finally, we aggregate all the backers' persuasibility to predict project success. To validate our method, we crawled data on 183,886 projects launched during or before December 2014 on Kickstarter, a crowdfunding website. We selected 4,922 backers with a total of 442,793 pledges to identify backers' preferences. The results show that a backer is more likely to be persuaded by a project that is more similar to the backer's preferences. Our findings not only demonstrate the efficacy of backers' pledge history for predicting crowdfunding project success but also verify that a backer‐centric method can supplement the existing project‐centric approaches. Our model and findings enable crowdfunding platform agencies, fund‐seeking entrepreneurs, and investors to predict the success of a crowdfunding project.
Qianzhou Du, Jing Li 0096, Yanqing Du, G. Alan Wang, Weiguo Fan
J. Assoc. Inf. Sci. Technol.5
2021 Cross-modal retrieval with dual multi-angle self-attention
abstract
Abstract In recent years, cross‐modal retrieval has been a popular research topic in both fields of computer vision and natural language processing. There is a huge semantic gap between different modalities on account of heterogeneous properties. How to establish the correlation among different modality data faces enormous challenges. In this work, we propose a novel end‐to‐end framework named Dual Multi‐Angle Self‐Attention (DMASA) for cross‐modal retrieval. Multiple self‐attention mechanisms are applied to extract fine‐grained features for both images and texts from different angles. We then integrate coarse‐grained and fine‐grained features into a multimodal embedding space, in which the similarity degrees between images and texts can be directly compared. Moreover, we propose a special multistage training strategy, in which the preceding stage can provide a good initial value for the succeeding stage and make our framework work better. Very promising experimental results over the state‐of‐the‐art methods can be achieved on three benchmark datasets of Flickr8k, Flickr30k, and MSCOCO.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
J. Assoc. Inf. Sci. Technol.6
2021 Deep Cross-Modal Face Naming for People News Retrieval
abstract
How to integrate multimodal information sources for face naming in multimodal news is a hot and yet challenging problem. A novel deep cross-modal face naming scheme is developed in this paper to facilitate more effective people news retrieval for large-scale multimodal news. This scheme integrates deep multimodal analysis, cross-modal correlation learning, and multimodal information mining, in which the efficient naming mechanism aims to cluster the deep features of different modalities into a common space to explore their inter-related correlations, and a special Web mining pattern is designed to optimize the name-face matching for rare non-celebrity. Such a cross-modal face naming model can be treated as a problem of bi-media semantic mapping and modeled as an inter-related correlation distribution over deep representations of multimodal news, in which the most important is to create more effective cross-modal name-face correlation and measure to what degree they are correlated. The experiments on a large number of public data from Yahoo! News have obtained very positive results and demonstrated the effectiveness of the proposed model.
Lian Zhou, Yuejie Zhang, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Knowl. Data Eng.5
2020 Adverse drug event detection and extraction from open data: A deep learning approach
Brandon Fan, Weiguo Fan, Carly Smith, Harold Garner
Inf. Process. Manag.2
2020 Crowd characteristics and crowd wisdom: Evidence from an online investment community
abstract
Fueled by the explosive growth of Web 2.0 and social media, online investment communities have become a popular venue for individual investors to interact with each other. Investor opinions extracted from online investment communities capture “crowd wisdom” and have begun to play an important role in financial markets. Existing research confirms the importance of crowd wisdom in stock predictions, but fails to investigate factors influencing crowd performance (that is, crowd prediction accuracy). In order to help improve crowd performance, our research strives to investigate the impact of crowd characteristics on crowd performance. We conduct an empirical study using a large data set collected from a popular online investment community, StockTwits. Our findings show that experience diversity, participant independence, and network decentralization are all positively related to crowd performance. Furthermore, crowd size moderates the influence of crowd characteristics on crowd performance. From a theoretical perspective, our work enriches extant literature by empirically testing the relationship between crowd characteristics and crowd performance. From a practical perspective, our findings help investors better evaluate social sensors embedded in user‐generated stock predictions, based upon which they can make better investment decisions.
Hong Hong 0002, Qiang Ye 0004, Qianzhou Du, G. Alan Wang, Weiguo Fan
J. Assoc. Inf. Sci. Technol.5
2020 User adoption of physician's replies in an online health community: An empirical study
abstract
Abstract Online health question‐and‐answer consultation with physicians is becoming a common phenomenon. However, it is unclear how users identify the most satisfying reply. Based on the dual‐process theory of knowledge adoption, we developed a conceptual model and empirical method to study which factors influence adoption of a reply. We extracted 6 variables for argument quality (Ease of understanding, Relevance, Completeness, Objectivity, Timeliness, Structure) and 4 for source credibility (Physician's online experience, Physician's offline expertise, Hospital location, Hospital level). The empirical results indicate that both central and peripheral routes affect user's adoption of a response. Physician's offline expertise negatively affects user's adoption decision, while physician's online experience positively affects it; this effect is positively moderated by user involvement.
Xinmiao Li, Weiguo Fan
J. Assoc. Inf. Sci. Technol.3
2019 Discovering Product Defects and Solutions from Online User Generated Contents
abstract
The recent increase in online user generated content (UGC) has led to the availability of a large number of posts about products and services. Often, these posts contain complaints that the consumers purchasing the products and services have. However, discovering and summarizing product defects and the related knowledge from large quantities of user posts is a difficult task. Traditional aspect opinion mining models, that aim to discover the product aspects and their corresponding opinions, are not sufficient to discover the product defect information from the user posts. In this paper, we propose the Product Defect Latent Dirichlet Allocation model (PDLDA), a probabilistic model that identifies domain-specific knowledge about product issues using interdependent three-dimensional topics: Component, Symptom, and Resolution. A Gibbs sampling based inference method for PDLDA is also introduced. To evaluate our model, we introduce three novel product review datasets. Both qualitative and quantitative evaluations show that the proposed model results in apparent improvement in the quality of discovered product defect information. Our model has the potential to benefit customers, manufacturers, and policy makers, by automatically discovering product defects from online data.
Xuan Zhang 0005, Zhilei Qiao, Aman Ahuja, Weiguo Fan, Edward A. Fox, Chandan K. Reddy
WWW4
2015 Novel applications of social media analytics
Weiguo Fan, Xiangbin Yan
Inf. Manag.1
2014 Determinants of users' continuance of social networking sites: A self-regulation perspective
Hui Lin 0003, Weiguo Fan, Patrick Y. K. Chau
Inf. Manag.2
2014 The relativity of privacy preservation based on social tagging
Baozhen Lee, Weiguo Fan, Anna Cinzia Squicciarini
Inf. Sci.2
2012 Examining Commitment in Electronic Knowledge Repository Usage
Hui Lin 0003, Weiguo Fan
J. Comput. Inf. Syst.2
2011 Digital Library 2.0 for Educational Resources
Monika Akbar, Weiguo Fan, Clifford A. Shaffer, Yinlin Chen, Lillian N. Cassel, Lois M. L. Delcambre, Dan Garcia 0001, Gregory W. Hislop, Frank M. Shipman III, Richard Furuta, B. Stephen Carpenter II, Hao-wei Hsieh, Bob Siegfried, Edward A. Fox
TPDL2
2011 Trace-Oriented Feature Analysis for Large-Scale Text Data Dimension Reduction
abstract
Dimension reduction for large-scale text data is attracting much attention nowadays due to the rapid growth of the World Wide Web. We can categorize those popular dimension reduction algorithms into two groups: feature extraction and feature selection algorithms. In the former, new features are combined from their original features through algebraic transformation. Though many of them have been validated to be effective, these algorithms are typically associated with high computational overhead, making them difficult to be applied on real-world text data. In the latter, subsets of features are selected directly. These algorithms are widely used in real-world tasks owing to their efficiency, but are often based on greedy strategies rather than optimal solutions. An important problem remains: it has been troublesome to integrate these two types of algorithms into a single framework, making it difficult to reap the benefits from both. In this paper, we formulate the two algorithm categories through a unified optimization framework, under which we develop a novel feature selection algorithm called Trace-Oriented Feature Analysis (TOFA). In detail, we integrate the objective functions of several state-of-the-art feature extraction algorithms into a unified one under the optimization framework, and then we propose to optimize this objective function in the solution space of feature selection algorithms for dimensionality reduction. Since the proposed objective function of TOFA integrates many prominent feature extraction algorithms' objective functions, such as unsupervised Principal Component Analysis (PCA) and supervised Maximum Margin Criterion (MMC), TOFA can handle both supervised and unsupervised problems. In addition, by tuning a weight value, TOFA is also suitable to solve semisupervised learning problems. Experimental results on several real-world data sets validate the effectiveness and efficiency of TOFA in text data for dimensionality reduction purpose.
Jun Yan 0001, Ning Liu 0001, Shuicheng Yan, Qiang Yang 0001, Weiguo Fan, Wei Wei 0002, Zheng Chen 0001
IEEE Trans. Knowl. Data Eng.5
2009 ExSearch: a novel vertical search engine for online barter business
abstract
E-Commerce has shown its exponentially-growing business value in the past decade. However, in contrast to the successful examples in online sales, such as Amazon1 and eBay2, the online barter business is still underexplored due to the lack of corresponding information aggregation service. In this paper, we design and implement a novel vertical search engine, called ExSearch, to aggregate online barter information for developing the barter market. Different from classical general purpose Web search engines, ExSearch adopts a focused crawler to gather related information from various websites. We propose to automatically extract the barter information from free-text Web pages such that the unstructured information is represented in structured databases. In addition, we utilize the data mining techniques such as regression to fulfill the missing information, which cannot be extracted from the Web pages. Finally, we validate and rank the search results according to user queries. Experimental results show that each component module in our proposed ExSearch system is efficient and effective. The volunteer users are satisfied by and interested in this novel vertical search engine.
Lei Ji 0001, Jun Yan 0001, Ning Liu 0001, Weiguo Fan, Zheng Chen 0001
CIKM5
2009 Identifying vertical search intention of query through social tagging propagation
abstract
A pressing task during the unification process is to identify a user's vertical search intention based on the user's query. In this paper, we propose a novel method to propagate social annotation, which includes user-supplied tag data, to both queries and VSEs for semantically bridging them. Our proposed algorithm consists of three key steps: query annotation, vertical annotation and query intention identification. Our algorithm, referred to as TagQV, verifies that the social tagging can be propagated to represent Web objects such as queries and VSEs besides Web pages. Experiments on real Web search queries demonstrate the effectiveness of TagQV in query intention identification.
Ning Liu 0001, Jun Yan 0001, Weiguo Fan, Qiang Yang 0001, Zheng Chen 0001
WWW3
2009 Examining the Success of Websites beyond E-Commerce: An Extension of the is Success Model
Ludwig Christian Schaupp, France Bélanger, Weiguo Fan
J. Comput. Inf. Syst.3
2008 TOFA: Trace Oriented Feature Analysis in Text Categorization
abstract
Dimension reduction for large-scale text data is attracting much attention lately due to the rapid growth of World Wide Web. We can consider dimension reduction algorithms in two categories: feature extraction and feature selection. An important problem remains: it has been difficult to integrate these two algorithm categories into a single framework, making it difficult to reap the benefit of both. In this paper, we formulate the two algorithm categories through a unified optimization framework. Under this framework, we develop a novel feature selection algorithm called Trace Oriented Feature Analysis (TOFA). The novel objective function of TOFA is a unified framework that integrates many prominent feature extraction algorithms such as unsupervised Principal Component Analysis and supervised Maximum Margin Criterion are special cases of it. Thus TOFA can process not only supervised problem but also unsupervised and semi-supervised problems. Experimental results on real text datasets demonstrate the effectiveness and efficiency of TOFA.
Jun Yan 0001, Ning Liu 0001, Qiang Yang 0001, Weiguo Fan, Zheng Chen 0001
ICDM4
2008 Integration of complex archeology digital libraries: An ETANA-DL experience
Rao Shen, Naga Srinivas Vemuri, Weiguo Fan, Edward A. Fox
Inf. Syst.3
2007 A novel clustering-based RSS aggregator
abstract
In recent years, different commercial Weblog subscribing systems have been proposed to return stories from users. subscribed feeds. In this paper, we propose a novel clustering-based RSS aggregator called as RSS Clusgator System (RCS) for Weblog reading. Note that an RSS feed may have several different topics. A user may only be interested in a subset of these topics. In addition there could be many different stories from multiple RSS feeds, which discuss similar topic from different perspectives. A user may be interested in this topic but do not know how to collect all feeds related to this topic. In contrast to many previous works, we cluster all stories in RSS feeds into hierarchical structure to better serve the readers. Through this way, users can easily find all their interested stories. To make the system current, we propose a flexible time window for incremental clustering. RCS utilizes both link information and content information for efficient clustering. Experiments show the effectiveness of RCS.
Jun Yan 0001, Zhi-Hong Deng 0001, Lei Ji 0001, Weiguo Fan, Benyu Zhang, Zheng Chen 0001
WWW5
2006 Learning to advertise
abstract
Content-targeted advertising, the task of automatically associating ads to a Web page, constitutes a key Web monetization strategy nowadays. Further, it introduces new challenging technical problems and raises interesting questions. For instance, how to design ranking functions able to satisfy conflicting goals such as selecting advertisements (ads) that are relevant to the users and suitable and profitable to the publishers and advertisers? In this paper we propose a new framework for associating ads with web pages based on Genetic Programming (GP). Our GP method aims at learning functions that select the most appropriate ads, given the contents of a Web page. These ranking functions are designed to optimize overall precision and minimize the number of misplacements. By using a real ad collection and web pages from a newspaper, we obtained a gain over a state-of-the-art baseline method of 61.7% in average precision. Further, by evolving individuals to provide good ranking estimations, GP was able to discover ranking functions that are very effective in placing ads in web pages while avoiding irrelevant ones.
Anísio Lacerda, Marco Cristo, Marcos André Gonçalves, Weiguo Fan, Nivio Ziviani, Berthier A. Ribeiro-Neto
SIGIR4
2006 Effective and Efficient Dimensionality Reduction for Large-Scale and Streaming Data Preprocessing
abstract
Dimensionality reduction is an essential data preprocessing technique for large-scale and streaming data classification tasks. It can be used to improve both the efficiency and the effectiveness of classifiers. Traditional dimensionality reduction approaches fall into two categories: feature extraction and feature selection. Techniques in the feature extraction category are typically more effective than those in feature selection category. However, they may break down when processing large-scale data sets or data streams due to their high computational complexities. Similarly, the solutions provided by the feature selection approaches are mostly solved by greedy strategies and, hence, are not ensured to be optimal according to optimized criteria. In this paper, we give an overview of the popularly used feature extraction and selection algorithms under a unified framework. Moreover, we propose two novel dimensionality reduction algorithms based on the orthogonal centroid algorithm (OC). The first is an incremental OC (IOC) algorithm for feature extraction. The second algorithm is an orthogonal centroid feature selection (OCFS) method which can provide optimal solutions according to the OC criterion. Both are designed under the same optimization criterion. Experiments on Reuters Corpus Volume-1 data set and some public large-scale text data sets indicate that the two algorithms are favorable in terms of their effectiveness and efficiency when compared with other state-of-the-art algorithms.
Jun Yan 0001, Benyu Zhang, Ning Liu 0001, Shuicheng Yan, Weiguo Fan, Qiang Yang 0001, Wensi Xi, Zheng Chen 0001
IEEE Trans. Knowl. Data Eng.6
2005 A Similarity Reinforcement Algorithm for Heterogeneous Web Pages
Ning Liu 0001, Jun Yan 0001, Fengshan Bai, Benyu Zhang, Wensi Xi, Weiguo Fan, Zheng Chen 0001, Lei Ji 0001, Chenyong Hu, Wei-Ying Ma
APWeb6
2005 Semantic verification for fact seeking engines
abstract
We present the architecture of our web question answering (fact seeking) system and introduce a novel algorithm to validate semantic categories of the expected answers. When tested on the questions used by the prior research, our system demonstrated the performance comparable to the current state of the art systems. Our semantic verification algorithm has improved the accuracy of answers of the affected questions by 30%.
Dmitri Roussinov, Weiguo Fan, Fernando Adrian Das Neves
CIKM2
2005 Discretization based learning approach to information retrieval
abstract
We have designed a representation scheme, which is based on the discrete representation of a document ranking function, which is capable of reproducing and enhancing the properties of such popular ranking functions as tf.idf, BM25 or those based on language models. Our tests have demonstrated the capability of our approach to achieve the performance of the best known scoring functions solely through training, without using any known heuristic or analytic formulas.
Dmitri Roussinov, Weiguo Fan, Fernando Adrian Das Neves
CIKM2
2005 A new framework to combine descriptors for content-based image retrieval
abstract
In this paper, we propose a novel framework using Genetic Programming to combine image database descriptors for content-based image retrieval (CBIR). Our framework is validated through several experiments involving two image databases and specific domains, where the images are retrieved based on the shape of their objects.
Ricardo da Silva Torres, Alexandre X. Falcão, Baoping Zhang, Weiguo Fan, Edward A. Fox, Marcos André Gonçalves, Pável Calado
CIKM4
2005 Intelligent GP fusion from multiple sources for text classification
abstract
This paper shows how citation-based information and structural content (e.g., title, abstract) can be combined to improve classification of text documents into predefined categories. We evaluate different measures of similarity -- five derived from the citation information of the collection, and three derived from the structural content -- and determine how they can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our experiments with the ACM Computing Classification Scheme, using documents from the ACM Digital Library, indicate that GP can discover similarity functions superior to those based solely on a single type of evidence. Effectiveness of the similarity functions discovered through simple majority voting is better than that of content-based as well as combination-based Support Vector Machine classifiers. Experiments also were conducted to compare the performance between GP techniques and other fusion techniques such as Genetic Algorithms (GA) and linear fusion. Empirical results show that GP was able to discover better similarity functions than GA or other fusion techniques.
Baoping Zhang, Yuxin Chen 0003, Weiguo Fan, Edward A. Fox, Marcos André Gonçalves, Marco Cristo, Pável Calado
CIKM3
2005 SimFusion: measuring similarity using unified relationship matrix
abstract
In this paper we use a Unified Relationship Matrix (URM) to represent a set of heterogeneous data objects (e.g., web pages, queries) and their interrelationships (e.g., hyperlinks, user click-through sequences). We claim that iterative computations over the URM can help overcome the data sparseness problem and detect latent relationships among heterogeneous data objects, thus, can improve the quality of information applications that require com- bination of information from heterogeneous sources. To support our claim, we present a unified similarity-calculating algorithm, SimFusion. By iteratively computing over the URM, SimFusion can effectively integrate relationships from heterogeneous sources when measuring the similarity of two data objects. Experiments based on a web search engine query log and a web page collection demonstrate that SimFusion can improve similarity measurement of web objects over both traditional content based algorithms and the cutting edge SimRank algorithm.
Wensi Xi, Edward A. Fox, Weiguo Fan, Benyu Zhang, Zheng Chen 0001, Jun Yan 0001, Dong Zhuang
SIGIR3
2005 OCFS: optimal orthogonal centroid feature selection for text categorization
abstract
Text categorization is an important research area in many Information Retrieval (IR) applications. To save the storage space and computation time in text categorization, efficient and effective algorithms for reducing the data before analysis are highly desired. Traditional techniques for this purpose can generally be classified into feature extraction and feature selection. Because of efficiency, the latter is more suitable for text data such as web documents. However, many popular feature selection techniques such as Information Gain (IG) andχ2-test (CHI) are all greedy in nature and thus may not be optimal according to some criterion. Moreover, the performance of these greedy methods may be deteriorated when the reserved data dimension is extremely low. In this paper, we propose an efficient optimal feature selection algorithm by optimizing the objective function of Orthogonal Centroid (OC) subspace learning algorithm in a discrete solution space, called Orthogonal Centroid Feature Selection (OCFS). Experiments on 20 Newsgroups (20NG), Reuters Corpus Volume 1 (RCV1) and Open Directory Project (ODP) data show that OCFS is consistently better than IG and CHI with smaller computation time especially when the reduced dimension is extremely small.
Jun Yan 0001, Ning Liu 0001, Benyu Zhang, Shuicheng Yan, Zheng Chen 0001, Weiguo Fan, Wei-Ying Ma
SIGIR7
2005 Intelligent fusion of structural and citation-based evidence for text classification
abstract
This paper shows how different measures of similarity derived from the citation information and the structural content (e.g., title, abstract) of the collection can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our experiments with the ACM Computing Classification Scheme, using documents from the ACM Digital Library, indicate that GP can discover similarity functions superior to those based solely on a single type of evidence. Effectiveness of the similarity functions discovered through simple majority voting is better than that of content-based as well as combination-based Support Vector Machine classifiers. Experiments also were conducted to compare the performance between GP techniques and other fusion techniques such as Genetic Algorithms (GA) and linear fusion. Empirical results show that GP was able to discover better similarity functions than other fusion techniques.
Baoping Zhang, Yuxin Chen 0003, Weiguo Fan, Edward A. Fox, Marcos André Gonçalves, Marco Cristo, Pável Calado
SIGIR3
2005 Improving web search results using affinity graph
abstract
In this paper, we propose a novel ranking scheme named Affinity Ranking (AR) to re-rank search results by optimizing two metrics: (1) diversity -- which indicates the variance of topics in a group of documents; (2) information richness -- which measures the coverage of a single document to its topic. Both of the two metrics are calculated from a directed link graph named Affinity Graph (AG). AG models the structure of a group of documents based on the asymmetric content similarities between each pair of documents. Experimental results in Yahoo! Directory, ODP Data, and Newsgroup data demonstrate that our proposed ranking algorithm significantly improves the search performance. Specifically, the algorithm achieves 31% improvement in diversity and 12% improvement in information richness relatively within the top 10 search results.
Benyu Zhang, Hua Li 0001, Yi Liu 0054, Lei Ji 0001, Wensi Xi, Weiguo Fan, Zheng Chen 0001, Wei-Ying Ma
SIGIR6
2005 Probabilistic question answering on the Web
abstract
Abstract Web‐based search engines such as Google and NorthernLight return documents that are relevant to a user query, not answers to user questions. We have developed an architecture that augments existing search engines so that they support natural language question answering. The process entails five steps: query modulation, document retrieval, passage extraction, phrase extraction, and answer ranking. In this article, we describe some probabilistic approaches to the last three of these stages. We show how our techniques apply to a number of existing search engines, and we also present results contrasting three different methods for question answering. Our algorithm, probabilistic phrase reranking (PPR), uses proximity and question type features and achieves a total reciprocal document rank of .20 on the TREC8 corpus. Our techniques have been implemented as a Web‐accessible system, called NSIR.
Dragomir R. Radev, Weiguo Fan, Harris Wu, Amardeep Grewal
J. Assoc. Inf. Sci. Technol.2
2004 Optimizing web search using web click-through data
abstract
The performance of web search engines may often deteriorate due to the diversity and noisy information contained within web pages. User click-through data can be used to introduce more accurate description (metadata) for web pages, and to improve the search performance. However, noise and incompleteness, sparseness, and the volatility of web pages and are three major challenges for research work on user click-through log mining. In this paper, we propose a novel iterative reinforced algorithm to utilize the user click-through data to improve search performance. The algorithm fully explores the interrelations between and web pages, and effectively finds virtual queries for web pages and overcomes the challenges discussed above. Experiment results on a large set of MSN click-through log data show a significant improvement on search performance over the naive query log mining algorithm as well as the baseline search engine.
Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Weiguo Fan
CIKM7
2004 Combining structural and citation-based evidence for text classification
abstract
This paper discusses how citation-based information and structural content (e.g., title, abstract) can be combined to improve classification of text documents into predefined categories. We evaluate different measures of similarity derived from the citation structure and the structural content of the collection, and determine how they can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our empirical experiments using documents from the ACM Digital Library and the ACM Computing Classification System show that we can discover similarity functions that work better than using evidence in isolation and whose combined performance through a simple majority voting is comparable to that of Support Vector Machine classifiers.
Baoping Zhang, Marcos André Gonçalves, Weiguo Fan, Yuxin Chen 0003, Edward A. Fox, Pável Calado, Marco Cristo
CIKM3
2004 IMMC: incremental maximum margin criterion
abstract
Subspace learning approaches have attracted much attention in academia recently. However, the classical batch algorithms no longer satisfy the applications on streaming data or large-scale data. To meet this desirability, Incremental Principal Component Analysis (IPCA) algorithm has been well established, but it is an unsupervised subspace learning approach and is not optimal for general classification tasks, such as face recognition and Web document categorization. In this paper, we propose an incremental supervised subspace learning algorithm, called Incremental Maximum Margin Criterion (IMMC), to infer an adaptive subspace by optimizing the Maximum Margin Criterion. We also present the proof for convergence of the proposed algorithm. Experimental results on both synthetic dataset and real world datasets show that IMMC converges to the similar subspace as that of batch approach.
Jun Yan 0001, Benyu Zhang, Shuicheng Yan, Qiang Yang 0001, Hua Li 0001, Zheng Chen 0001, Wensi Xi, Weiguo Fan, Wei-Ying Ma
KDD8
2004 Tuning before feedback: combining ranking discovery and blind feedback for robust retrieval
abstract
Both ranking functions and user queries are very important factors affecting a search engine's performance. Prior research has looked at how to improve ad-hoc retrieval performance for existing queries while tuning the ranking function, or modify and expand user queries using a fixed ranking scheme using blind feedback. However, almost no research has looked at how to combine ranking function tuning and blind feedback together to improve ad-hoc retrieval performance. In this paper, we look at the performance improvement for ad-hoc retrieval from a more integrated point of view by combining the merits of both techniques. In particular, we argue that the ranking function should be tuned first, using user-provided queries, before applying the blind feedback technique. The intuition is that highly-tuned ranking offers more high quality documents at the top of the hit list, thus offers a stronger baseline for blind feedback. We verify this integrated model in a large scale heterogeneous collection and the experimental results show that combining ranking function tuning and blind feedback can improve search performance by almost 30% over the baseline Okapi system.
Weiguo Fan, Ming Luo 0001, Li Wang 0007, Wensi Xi, Edward A. Fox
SIGIR1
2004 A generic ranking function discovery framework by genetic programming for information retrieval
Weiguo Fan, Michael D. Gordon, Praveen Pathak
Inf. Process. Manag.1
2004 The effects of fitness functions on genetic programming-based ranking discovery forWeb search
abstract
Abstract Genetic‐based evolutionary learning algorithms, such as genetic algorithms (GAs) and genetic programming (GP), have been applied to information retrieval (IR) since the 1980s. Recently, GP has been applied to a new IR task—discovery of ranking functions for Web search—and has achieved very promising results. However, in our prior research, only one fitness function has been used for GP‐based learning. It is unclear how other fitness functions may affect ranking function discovery for Web search, especially since it is well known that choosing a proper fitness function is very important for the effectiveness and efficiency of evolutionary algorithms. In this article, we report our experience in contrasting different fitness function designs on GP‐based learning using a very large Web corpus. Our results indicate that the design of fitness functions is instrumental in performance improvement. We also give recommendations on the design of fitness functions for genetic‐based information retrieval experiments.
Weiguo Fan, Edward A. Fox, Praveen Pathak, Harris Wu
J. Assoc. Inf. Sci. Technol.1
2004 Discovery of Context-Specific Ranking Functions for Effective Information Retrieval Using Genetic Programming
abstract
The Internet and corporate intranets have brought a lot of information. People usually resort to search engines to find required information. However, these systems tend to use only one fixed ranking strategy regardless of the contexts. This poses serious performance problems when characteristics of different users, queries, and text collections are taken into account. We argue that the ranking strategy should be context specific and we propose a , new systematic method that can automatically generate ranking strategies for different contexts based on genetic programming (GP). The new method was tested on TREC data and the results are very promising.
Weiguo Fan, Michael D. Gordon, Praveen Pathak
IEEE Trans. Knowl. Data Eng.1
2002 Probabilistic question answering on the web
abstract
Web-based search engines such as Google and NorthernLight return documents that are relevant to a user query, not answers to user questions. We have developed an architecture that augments existing search engines so that they support natural language question answering. The process entails five steps: query modulation, document retrieval, passage extraction, phrase extraction, and answer ranking. In this paper we describe some probabilistic approaches to the last three of these stages. We show how our techniques apply to a number of existing search engines and we also present results contrasting three different methods for question answering. Our algorithm, probabilistic phrase reranking (PPR) using proximity and question type features achieves a total reciprocal document rank of .20 on the TREC 8 corpus. Our techniques have been implemented as a Web-accessible system, called NSIR.
Dragomir R. Radev, Weiguo Fan, Harris Wu, Amardeep Grewal
WWW2
2002 Getting answers to natural language questions on the Web
abstract
Abstract Most popular search engines are not designed for answering natural language questions. However, when we asked hundreds of natural language questions of nine leading search engines, all retrieved at least one correct answer on more than three‐quarters of the questions. We identified the best‐performing search engines overall for factual natural language questions. We found performance differences depending on the domain of factual question asked. Other aspects of questions also predicted significantly different performance: the number of words in the question, the presence of a proper noun, and whether the question is time dependent. An additional analysis tested for differential performance by specific search engines on these four question factors. The analysis found no evidence for such interactions.
Dragomir R. Radev, Kelsey Libner, Weiguo Fan
J. Assoc. Inf. Sci. Technol.3
2001 Mining the Web for Answers to Natural Language Questions
abstract
The web is now becoming one of the largest information and knowledge repositories. Many large scale search engines (Google, Fast, Northern Light, etc.) have emerged to help users find information. In this paper, we study how we can effectively use these existing search engines to mine the Web and discover the "correct" answers to factual natural language questions.We propose a probabilistic algorithm called QASM (Question Answering using Statistical Models) that learns the best query paraphrase of a natural language question. We validate our approach for both local and web search engines using questions from the TREC evaluation. We also show how this algorithm can be combined with another algorithm (AnSel) to produce precise answers to natural language questions.
Dragomir R. Radev, Zhiping Zheng, Sasha Blair-Goldensohn, Weiguo Fan, John M. Prager
CIKM6
2001 Discovering and reconciling value conflicts for numerical data integration
Weiguo Fan, Hongjun Lu, Stuart E. Madnick, David Wai-Lok Cheung
Inf. Syst.1