VLDB 2026 Research / reviewers in the wild / expert
Wanli Zuo
dblp:64/1676
· DBLP profile ↗
49ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 9 since 2021Databases, data management, data science and information retrieval · 20 · 6 since 2021Systems, architecture and hardware · 5Graphics, computer vision, multimedia, augmented reality and games · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OPHEP-Miner: One-phase high-efficiency pattern mining utilizing tree structures
Genlang Chen, Fangyu Wu 0001, Wanli Zuo, Youxi Wu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Effective approaches for mining correlated and low-average-cost patterns
Genlang Chen, Shiting Wen, Wanli Zuo |
Knowl. Based Syst. | 4 |
| 2023 | TC-GAT: Graph Attention Network for Temporal Causality DiscoveryabstractThe present study explores the intricacies of causal relationship extraction, a vital component in the pursuit of causality knowledge. Causality is frequently intertwined with temporal elements, as the progression from cause to effect is not instantaneous but rather ensconced in a temporal dimension. Thus, the extraction of temporal causality holds paramount significance in the field. In light of this, we propose a method for extracting causality from the text that integrates both temporal and causal relations, with a particular focus on the time aspect. To this end, we first compile a dataset that encompasses temporal relationships. Subsequently, we present a novel model, TC-GAT, which employs a graph attention mechanism to assign weights to the temporal relationships and leverages a causal knowledge graph to determine the adjacency matrix. Additionally, we implement an equilibrium mechanism to regulate the interplay between temporal and causal relations. Our experiments demonstrate that our proposed method significantly surpasses baseline models in the task of causality extraction. Xiaosong Yuan, Wanli Zuo, Yijia Zhang 0003 |
IJCNN | 3 |
| 2023 | The Causal Strength Bank: A New Benchmark for Causal Strength Classification
Xiaosong Yuan, Renchu Guan, Wanli Zuo, Yijia Zhang 0003 |
PAKDD (1) | 3 |
| 2023 | Mining top-k high average-utility itemsets based on breadth-first search
Genlang Chen, Fangyu Wu 0001, Shiting Wen, Wanli Zuo |
Appl. Intell. | 5 |
| 2023 | Dual-core mutual learning between scoring systems and clinical features for ICU mortality prediction
Zhenkun Shi, Sen Wang 0001, Lin Yue, Yijia Zhang 0003, Binod Kumar Adhikari, Wanli Zuo, Xue Li 0001 |
Inf. Sci. | 7 |
| 2022 | Label-aware Multi-level Contrastive Learning for Cross-lingual Spoken Language UnderstandingabstractDespite the great success of spoken language understanding (SLU) in high-resource languages, it remains challenging in low-resource languages mainly due to the lack of labeled training data.The recent multilingual codeswitching approach achieves better alignments of model representations across languages by constructing a mixed-language context in zeroshot cross-lingual SLU.However, current codeswitching methods are limited to implicit alignment and disregard the inherent semantic structure in SLU, i.e., the hierarchical inclusion of utterances, slots, and words.In this paper, we propose to model the utterance-slot-word structure by a multi-level contrastive learning framework at the utterance, slot, and word levels to facilitate explicit alignment.Novel codeswitching schemes are introduced to generate hard negative examples for our contrastive learning framework.Furthermore, we develop a label-aware joint model leveraging label semantics to enhance the implicit alignment and feed to contrastive learning.Our experimental results show that our proposed methods significantly improve the performance compared with the strong baselines on two zero-shot crosslingual SLU benchmark datasets. Shining Liang, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Wanli Zuo, Xianglin Zuo, Daxin Jiang |
EMNLP | 5 |
| 2022 | Effective algorithms to mine skyline frequent-utility itemsets
Genlang Chen, Wanli Zuo |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | A multi-level neural network for implicit causality detection in web texts
Shining Liang, Wanli Zuo, Zhenkun Shi, Sen Wang 0001, Junhu Wang, Xianglin Zuo |
Neurocomputing | 2 |
| 2021 | Reinforced Iterative Knowledge Distillation for Cross-Lingual Named Entity RecognitionabstractNamed entity recognition (NER) is a fundamental component in many applications, such as Web Search and Voice Assistants. Although deep neural networks greatly improve the performance of NER, due to the requirement of large amounts of training data, deep neural networks can hardly scale out to many languages in an industry setting. To tackle this challenge, cross-lingual NER transfers knowledge from a rich-resource language to languages with low resources through pre-trained multilingual language models. Instead of using training data in target languages, cross-lingual NER has to rely on only training data in source languages, and optionally adds the translated training data derived from source languages. However, the existing cross-lingual NER methods do not make good use of rich unlabeled data in target languages, which is relatively easy to collect in industry applications. To address the opportunities and challenges, in this paper we describe our novel practice in Microsoft to leverage such large amounts of unlabeled data in target languages in real production settings. To effectively extract weak supervision signals from the unlabeled data, we develop a novel approach based on the ideas of semi-supervised learning and reinforcement learning. The empirical study on three benchmark data sets verifies that our approach establishes the new state-of-the-art performance with clear edges. Now, the NER techniques reported in this paper are on their way to become a fundamental component for Web ranking, Entity Pane, Answers Triggering, and Question Answering in the Microsoft Bing search engine. Moreover, our techniques will also serve as part of the Spoken Language Understanding module for a commercial voice assistant. We plan to open source the code of the prototype framework after deployment. Shining Liang, Ming Gong 0001, Jian Pei 0001, Linjun Shou, Wanli Zuo, Xianglin Zuo, Daxin Jiang |
KDD | 5 |
| 2021 | CalibreNet: Calibration Networks for Multilingual Sequence LabelingabstractLack of training data in low-resource languages presents huge challenges to sequence labeling tasks such as named entity recognition (NER) and machine reading comprehension (MRC). One major obstacle is the errors on the boundary of predicted answers. To tackle this problem, we propose CalibreNet, which predicts answers in two steps. In the first step, any existing sequence labeling method can be adopted as a base model to generate an initial answer. In the second step, CalibreNet refines the boundary of the initial answer. To tackle the challenge of lack of training data in low-resource languages, we dedicatedly develop a novel unsupervised phrase boundary recovery pre-training task to enhance the multilingual boundary detection capability of CalibreNet. Experiments on two cross-lingual benchmark datasets show that the proposed approach achieves SOTA results on zero-shot cross-lingual NER and MRC tasks. Shining Liang, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Wanli Zuo, Daxin Jiang |
WSDM | 5 |
| 2021 | Deep dynamic imputation of clinical time series for mortality prediction
Zhenkun Shi, Sen Wang 0001, Lin Yue, Lixin Pang, Xianglin Zuo, Wanli Zuo, Xue Li 0001 |
Inf. Sci. | 6 |
| 2021 | A Scalable Redefined Stochastic BlockmodelabstractStochastic blockmodel (SBM) is a widely used statistical network representation model, with good interpretability, expressiveness, generalization, and flexibility, which has become prevalent and important in the field of network science over the last years. However, learning an optimal SBM for a given network is an NP-hard problem. This results in significant limitations when it comes to applications of SBMs in large-scale networks, because of the significant computational overhead of existing SBM models, as well as their learning methods. Reducing the cost of SBM learning and making it scalable for handling large-scale networks, while maintaining the good theoretical properties of SBM, remains an unresolved problem. In this work, we address this challenging task from a novel perspective of model redefinition. We propose a novel redefined SBM with Poisson distribution and its block-wise learning algorithm that can efficiently analyse large-scale networks. Extensive validation conducted on both artificial and real-world data shows that our proposed method significantly outperforms the state-of-the-art methods in terms of a reasonable trade-off between accuracy and scalability. 1 Xueyan Liu 0001, Bo Yang 0002, Hechang Chen, Katarzyna Musial, Hongxu Chen 0002, Yang Li 0030, Wanli Zuo |
ACM Trans. Knowl. Discov. Data | 7 |
| 2021 | A block-based generative model for attributed network embedding
Xueyan Liu 0001, Bo Yang 0002, Wenzhuo Song, Katarzyna Musial, Wanli Zuo, Hongxu Chen 0002, Hongzhi Yin |
World Wide Web | 5 |
| 2020 | A Review of Dataset and Labeling Methods for Causality ExtractionabstractCausality represents the most important kind of correlation between events.Extracting causality from text has become a promising hot topic in NLP.However, there is no mature research systems, evaluation rules and datasets for public evaluation.Moreover, there is a lack of unified causal sequence labeling methods, which constitute the key factors that hinder the progress of causality extraction research.We survey the limitations and shortcomings of existing causality research field comprehensively from the aspects of basic concepts, extraction methods, experimental data, and labeling methods, so as to provide reference for future research on causality extraction.We summarize the existing causality datasets, explore their practicability and extensibility from multiple perspectives.Aiming at the problem of causal sequence labeling, we analyze the existing methods of causal sequence labeling, with a summarizations of its regulation.Multiple candidate causal labeling sequences are put forward according to labeling controversy to explore the optimal labeling method through experiments, and suggestions are provided for selecting labeling method. Jinghang Xu, Wanli Zuo, Shining Liang, Xianglin Zuo |
COLING | 2 |
| 2020 | Effective sanitization approaches to protect sensitive knowledge in high-utility itemset mining
Shiting Wen, Wanli Zuo |
Appl. Intell. | 3 |
| 2020 | Joint Personalized Markov Chains with social network embedding for cold-start recommendation
Yijia Zhang 0003, Zhenkun Shi, Wanli Zuo, Lin Yue, Shining Liang, Xue Li 0001 |
Neurocomputing | 3 |
| 2020 | Semi-supervised stochastic blockmodel for structure analysis of signed networksabstractFinding hidden structural patterns is a critical problem for all types of networks, including signed networks. Among all of the methods for structural analysis of complex network, stochastic blockmodel (SBM) is an important research tool because it is flexible and can generate networks with many different types of structures. However, most existing SBM learning methods for signed networks are unsupervised, leading to poor performance in terms of finding hidden structural patterns, especially when handling noisy and sparse networks. Learning SBM in a semi-supervised way is a promising avenue for overcoming the above difficulty. In this type of model, a small number of labelled nodes and a large number of unlabelled nodes, coupled with their network structures, are simultaneously used to train SBM. We propose a novel semi-supervised signed stochastic blockmodel and its learning algorithm based on variational Bayesian inference, with the goal of discovering both assortative (the nodes connect more densely in same clusters than that in different clusters) and disassortative (the nodes link more sparsely in same clusters than that in different clusters) structures from signed networks. The proposed model is validated through a number of experiments wherein it compared with the state-of-the-art methods using both synthetic and real-world data. The carefully designed tests, allowing to account for different scenarios, show our method outperforms other approaches existing in this space. It is especially relevant in the case of noisy and sparse networks as they constitute the majority of the real-world networks. Xueyan Liu 0001, Wenzhuo Song, Katarzyna Musial, Xuehua Zhao, Wanli Zuo, Bo Yang 0002 |
Knowl. Based Syst. | 5 |
| 2020 | Adaptive Deep Modeling of Users and Items Using Side Information for RecommendationabstractIn the existing recommender systems, matrix factorization (MF) is widely applied to model user preferences and item features by mapping the user-item ratings into a low-dimension latent vector space. However, MF has ignored the individual diversity where the user's preference for different unrated items is usually different. A fixed representation of user preference factor extracted by MF cannot model the individual diversity well, which leads to a repeated and inaccurate recommendation. To this end, we propose a novel latent factor model called adaptive deep latent factor model (ADLFM), which learns the preference factor of users adaptively in accordance with the specific items under consideration. We propose a novel user representation method that is derived from their rated item descriptions instead of original user-item ratings. Based on this, we further propose a deep neural networks framework with an attention factor to learn the adaptive representations of users. Extensive experiments on Amazon data sets demonstrate that ADLFM outperforms the state-of-the-art baselines greatly. Also, further experiments show that the attention factor indeed makes a great contribution to our method. Lei Zheng 0001, Yuanbo Xu, Bangzuo Zhang, Fuzhen Zhuang, Philip S. Yu, Wanli Zuo |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2019 | Deep Interpretable Mortality Model for Intensive Care Unit Risk Prediction
Zhenkun Shi, Weitong Chen 0001, Shining Liang, Wanli Zuo, Lin Yue, Sen Wang 0001 |
ADMA | 4 |
| 2019 | DMMAM: Deep Multi-source Multi-task Attention Model for Intensive Care Unit Diagnosis
Zhenkun Shi, Wanli Zuo, Weitong Chen 0001, Lin Yue, Yuwei Hao, Shining Liang |
DASFAA (2) | 2 |
| 2019 | Deep Latent Factor Model with Hierarchical Similarity Measure for recommender systems
Lei Zheng 0001, He Huang 0008, Yuanbo Xu, Philip S. Yu, Wanli Zuo |
Inf. Sci. | 6 |
| 2019 | A survey of sentiment analysis in social media
Lin Yue, Weitong Chen 0001, Xue Li 0001, Wanli Zuo, Minghao Yin |
Knowl. Inf. Syst. | 4 |
| 2019 | Low Cost Edge Sensing for High Quality DemosaickingabstractDigital cameras that use Color Filter Arrays (CFA) entail a demosaicking procedure to form full RGB images. To digital camera industry, demosaicking speed is as important as demosaicking accuracy, because camera users have been accustomed to viewing captured photos instantly. Moreover, the cost associated with demosaicking should not go beyond the cost saved by using CFA. For this purpose, we revisit the classical Hamilton-Adams (HA) algorithm, which outperforms many sophisticated techniques in both speed and accuracy. Our analysis shows that the HA pipeline is highly efficient to exploit the originally captured data, but its oversimplified inter- and intra-channel smoothness formulation hinders its accuracy. We therefore propose a very low cost edge sensing scheme, which guides demosaicking by a logistic functional of the difference between directional variations. We extensively compare our algorithm with 27 demosaicking algorithms by running their open source codes on benchmark datasets. Compared to methods of similar computational cost, our method achieves substantially higher accuracy; Whereas compared to methods of similar accuracy, our method has significantly lower cost. On test images of currently popular resolution, the quality of our algorithm is comparable to top performers, yet our speed is tens of times faster. Source code for this work will be released with paper publication. Yan Niu, Jihong Ouyang, Wanli Zuo, Fuxin Wang |
IEEE Trans. Image Process. | 3 |
| 2018 | Sensitive Data Detection Using NN and KNN from Big Data
Binod Kumar Adhikari, Wanli Zuo, Ramesh Maharjan, Lin Guo 0005 |
ICA3PP (4) | 2 |
| 2018 | Prognosis of Thyroid Disease Using MS-Apriori Improved Decision Tree
Yuwei Hao, Wanli Zuo, Zhenkun Shi, Lin Yue, Fengling He |
KSEM (1) | 2 |
| 2018 | Social Bayesian Personal Ranking for Missing Data in Implicit Feedback Recommendation
Yijia Zhang 0003, Wanli Zuo, Zhenkun Shi, Lin Yue, Shining Liang |
KSEM (1) | 2 |
| 2018 | DEDSC: A Domain Expert Discovery Method Based on Structure and ContentabstractResearchers usually extract domain experts only through analyzing network structure or partitioning users into several communities according to their label information. Combining structure and content to discovery domain experts is a new attempt. Motivated by that, this paper proposes a domain expert discovery method based on network structure and content semantics, called DEDSC, which can extract authority nodes in overlapping communities. To analyze the overall authority for each user in the social network, two definitions, structure authority value and content authority value, are proposed to evaluate the authority of users in different perspectives. Partitioning users into communities can make the results more accurate. Experimental results show that our proposed method can discover domain experts effectively. In addition, when we need to extract domain experts in a new test dataset, we do not need to re-train the data in the training dataset. Lu Liu 0013, Wanli Zuo, Tao Peng 0003 |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2018 | Multi-modal multi-layered topic classification model for social event analysis
Yongheng Chen, Chunyan Yin, Yaojin Lin, Wanli Zuo |
Multim. Tools Appl. | 4 |
| 2017 | Detecting outlier pairs in complex network based on link structure and semantic relationship
Lu Liu 0013, Wanli Zuo, Tao Peng 0003 |
Expert Syst. Appl. | 2 |
| 2017 | Multi-factors based sentence ordering for cross-document fusion from multimodal content
Lin Yue, Zhenkun Shi, Sen Wang 0001, Weitong Chen 0001, Wanli Zuo |
Neurocomputing | 6 |
| 2017 | Structure2Content: An Incremental Method for Detecting Outlier Correlation in Heterogeneous NetworkabstractHeterogeneous networks are ubiquitous. People like to discover rare but meaningful objects and patterns from such networks. Regardless of high structure similarity or high content similarity, the corresponding objects can be used in data analysis. However, the vast differences between structure and contents should be paid more attention. In this paper, we propose an outlier correlation detection method, called Structure2Content, which discovers outlier correlation incrementally in structure-level and content-level. Structure2Content addresses three important challenges: (1) how can we measure the target object’s structure and content similarity? (2) how can we find the representative features of target objects? (3) how can we insert new data or delete the obsoleted data incrementally. To tackle these challenges, Structure2Content applies four main techniques: (1) two matrices are used to store structure and content similarity, respectively, (2) 3-tuples are used to represent the closeness degree between objects, (3) a mirror step and an iterative process are combined to obtain the top-K outlier correlations, and (4) only updating 3-tuples can help insert or delete data incrementally instead of training all data from the beginning. Substantial experiments show that our proposed method is very effective for outlier correlations detection. Lu Liu 0013, Wanli Zuo, Tao Peng 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2016 | Building text classifiers using positive, unlabeled and 'outdated' examplesabstractSummary Learning from positive and unlabeled examples (PU learning) is a partially supervised classification that is frequently used in Web and text retrieval system. The merit of PU learning is that it can get good performance with less manual work. Motivated by transfer learning, this paper presents a novel method that transfers the ‘outdated data’ into the process of PU learning. We first propose a way to measure the strength of the features and select the strong features and the weak features according to the strength of the features. Then, we extract the reliable negative examples and the candidate negative examples using the strong and the weak features (Transfer‐1DNF). Finally, we construct a classifier called weighted voting iterative support vector machine (SVM) that is made up of several subclassifiers by applying SVM iteratively, and each subclassifier is assigned a weight in each iteration. We conduct the experiments on two datasets: 20 Newsgroups and Reuters‐21578, and compare our method with three baseline algorithms: positive example‐based learning, weighted voting classifier and SVM. The results show that our proposed method Transfer‐1DNF can extract more reliable negative examples with lower error rates, and our classifier outperforms the baseline algorithms. Copyright © 2016 John Wiley & Sons, Ltd. Wanli Zuo, Lu Liu 0013, Yuanbo Xu, Tao Peng 0003 |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | An improved density peaks-based clustering method for social circle discovery in social networks
Wanli Zuo, Ying Wang 0009 |
Neurocomputing | 2 |
| 2015 | Modeling Status Theory in Trust PredictionabstractWith the pervasion of social media, trust has been playing more of an important role in helping online users collect reliable information. In reality, user-specified trust relations are often very sparse; hence, inferring unknown trust relations has attracted increasing attention in recent years. Social status is one of the most important concepts in trust, and status theory is developed to help us understand the important role of social status in the formation of trust relations. In this paper, we investigate how to exploit social status in trust prediction by modeling status theory. We first vertify status theory in trust relations, then provide a principled way to model it mathematically, and propose a novel framework sTrust which incorporates status theory for trust prediction. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework. Futher experiments are conducted to understand the importance of status theory in trust prediction. Ying Wang 0009, Xin Wang 0035, Jiliang Tang, Wanli Zuo, Guoyong Cai |
AAAI | 4 |
| 2015 | Exploring Social Context for Topic Identification in Short and Noisy TextsabstractWith the pervasion of social media, topic identification in short texts attracts increasing attention in recent years. However, in nature the texts of social media are short and noisy, and the structures are sparse and dynamic, resulting in difficulty to identify topic categories exactly from online social media. Inspired by social science findings that preference consistency and social contagion are observed in social media, we investigate topic identification in short and noisy texts by exploring social context from the perspective of social sciences. In particular, we present a mathematical optimization formulation that incorporates the preference consistency and social contagion theories into a supervised learning method, and conduct feature selection to tackle short and noisy texts in social media, which result in a Sociological framework for Topic Identification (STI). Experimental results on real-world datasets from Twitter and Citation Network demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the importance of social context in topic identification. Xin Wang 0035, Ying Wang 0009, Wanli Zuo, Guoyong Cai |
AAAI | 3 |
| 2015 | A fuzzy document clustering approach based on domain-specified ontology
Lin Yue, Wanli Zuo, Tao Peng 0003, Ying Wang 0009, Xuming Han |
Data Knowl. Eng. | 2 |
| 2015 | Research on Trust Prediction from a Sociological Perspective
Ying Wang 0009, Xin Wang 0035, Wanli Zuo |
J. Comput. Sci. Technol. | 3 |
| 2014 | PU text classification enhanced by term frequency-inverse document frequency-improved weightingabstractSUMMARY Term frequency–inverse document frequency (TF–IDF), one of the most popular feature (also called term or word) weighting methods used to describe documents in the vector space model and the applications related to text mining and information retrieval, can effectively reflect the importance of the term in the collection of documents, in which all documents play the same roles. But, TF–IDF does not take into account the difference of term IDF weighting if the documents play different roles in the collection of documents, such as positive and negative training set in text classification. In view of the aforementioned text, this paper presents a novel TF–IDF‐improved feature weighting approach, which reflects the importance of the term in the positive and the negative training examples, respectively. We also build a weighted voting classifier by iteratively applying the support vector machine algorithm and implement one‐class support vector machine and Positive Example Based Learning methods used for comparison. During classifying, an improved 1‐DNF algorithm, called 1‐DNFC, is also adopted, aiming at identifying more reliable negative documents from the unlabeled examples set. The experimental results show that the performance of term frequency inverse positive–negative document frequency‐based classifier outperforms that of TF–IDF‐based one, and the performance of weighted voting classifier also exceeds that of one‐class support vector machine‐based classifier and Positive Example Based Learning‐based classifier. Copyright © 2013 John Wiley & Sons, Ltd. Tao Peng 0003, Lu Liu 0013, Wanli Zuo |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | Tunneling enhanced by web page content block partition for focused crawlingabstractConcurrency and Computation: Practice and Experience 20(1):61–74 (January 2008) (DOI: 10.1002/cpe.1211) A bug was identified after publication in the testing program of this article. Figures 6 and 7 were subsequently affected. The correct results are shown as follows: Dynamic plot of average harvest rate versus number of crawled pages. Performance is averaged across topics and standard errors are also shown. The error bars correspond to ± standard error. Dynamic plot of average target recall versus number of crawled pages. Performance is averaged across topics and standard errors are also shown. The error bars correspond to ± standard error. Tao Peng 0003, Changli Zhang, Wanli Zuo |
Concurr. Comput. Pract. Exp. | 3 |
| 2009 | Sentiment analysis of Chinese documents: From sentence to document levelabstractAbstract User‐generated content on the Web has become an extremely valuable source for mining and analyzing user opinions on any topic. Recent years have seen an increasing body of work investigating methods to recognize favorable and unfavorable sentiments toward specific subjects from online text. However, most of these efforts focus on English and there have been very few studies on sentiment analysis of Chinese content. This paper aims to address the unique challenges posed by Chinese sentiment analysis. We propose a rule‐based approach including two phases: (1) determining each sentence's sentiment based on word dependency, and (2) aggregating sentences to predict the document sentiment. We report the results of an experimental study comparing our approach with three machine learning‐based approaches using two sets of Chinese articles. These results illustrate the effectiveness of our proposed method and its advantages against learning‐based approaches. Changli Zhang, Daniel Dajun Zeng, Jiexun Li, Fei-Yue Wang 0001, Wanli Zuo |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2008 | Research of PU Text Semi-supervised Classification Based on Ontology Feature ExtractionabstractFor the shortcomings in the method of traditional statistics-based feature extraction on PU issues, we put forward feature extraction based on ontology to improve the performance of PU classification. We improved PEBL algorithm, and get the document vector of positive set using ontology-based feature extraction, then find the strong positive features, which include the crossing semantics in the positive documents and have higher frequency in positive set. The improved algorithm scans the documents twice. First, we get the semantic of the documents by ontology. Second, we filtrate the terms which include none of these semantic to reduce the dimension and obtain the document vector. Experiments had shown that the improved PEBL classifier increases the F1 score by 0.7389%. Fuyu Yuan, Wanli Zuo, Fengling He |
ICMLA | 3 |
| 2008 | Tunneling enhanced by web page content block partition for focused crawlingabstractAbstract The complexity of web information environments and multiple‐topic web pages are negative factors significantly affecting the performance of focused crawling. A highly relevant region in a web page may be obscured because of low overall relevance of that page. Segmenting the web pages into smaller units will significantly improve the performance. Conquering and traversing irrelevant page to reach a relevant one (tunneling) can improve the effectiveness of focused crawling by expanding its reach. This paper presents a heuristic‐based method to enhance focused crawling performance. The method uses a Document Object Model (DOM)‐based page partition algorithm to segment a web page into content blocks with a hierarchical structure and investigates how to take advantage of block‐level evidence to enhance focused crawling by tunneling. Page segmentation can transform an uninteresting multi‐topic web page into several single topic context blocks and some of which may be interesting. Accordingly, focused crawler can pursue the interesting content blocks to retrieve the relevant pages. Experimental results indicate that this approach outperforms Breadth‐First, Best‐First and Link‐context algorithm both in harvest rate, target recall and target length. Copyright © 2007 John Wiley & Sons, Ltd. Tao Peng 0003, Changli Zhang, Wanli Zuo |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | SVM based adaptive learning method for text classification from positive and unlabeled documents
Tao Peng 0003, Wanli Zuo, Fengling He |
Knowl. Inf. Syst. | 2 |
| 2007 | First-order focused crawlingabstractThis paper reports a new general framework of focused web crawling based on "relational subgroup discovery". Predicates are used explicitly to represent the relevance clues of those unvisited pages in the crawl frontier, and then first-order classification rules are induced using subgroup discovery technique. The learned relational rules with sufficient support and confidence will guide the crawling process afterwards. We present the many interesting features of our proposed first-order focused crawler, together with preliminary promising experimental results. Qingyang Xu, Wanli Zuo |
WWW | 2 |
| 2006 | Multi-dimensional Sequential Pattern Mining Based on Concept Lattice
Wanli Zuo |
ADMA | 2 |
| 2005 | An Auto-stopped Hierarchical Clustering Algorithm for Analyzing 3D Model Database
Tian-yang Lv, Yu-hui Xing, Shaobin Huang, Zhengxuan Wang, Wanli Zuo |
PKDD | 5 |
| 2005 | An Auto-stopped Hierarchical Clustering Algorithm Integrating Outlier Detection Algorithm
Tian-yang Lv, Tai-xue Su, Zhengxuan Wang, Wanli Zuo |
WAIM | 4 |
| 2004 | Extracting Precise Link Context Using NLP Parsing TechniqueabstractLink context has been exploited extensively ever since the advent of the World Wide Web, but the approach to extracting precise link context has not been fully explored and many state-of-the-art extraction methods are based on simplistic heuristics and require ad-hoc parameters. In this paper, we propose a novel two-step extraction model, which aims to systematically derive link context of quality as high as anchor text. In the macroscopic analysis step, a systematic web page structure analysis is performed to locate the content cohesive text region and potential relevant header or header like tags. In the microscopic extraction step, an English parser is used to extract the relevant sentence fragments in the text region and the nearest heading text is encompassed if the need arises. Preliminary experimental results proved our approach's effectiveness. Qingyang Xu, Wanli Zuo |
Web Intelligence | 2 |