EDBT 2026 Demo / reviewers in the wild / expert
Yiyu Yao
dblp:y/YiyuYao · also Y. Y. Yao
· DBLP profile ↗
96ranked-venue papers in the field
27as first author
8since 2021 · last 2026
0000-0001-6502-6226ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 45 (16 first)Data Mining & Knowledge Discovery · 16 (3 first)Other / Interdisciplinary · 16 (6 first)Information Retrieval & Web Search · 14 (2 first)Database Systems & Data Management · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A high-performance method for handling dual mixed data based on three-way decision
Wanting Wang 0003, Qingzhao Kong, Yiyu Yao, Eric C. C. Tsang, Conghao Yan |
Inf. Sci. | 3 |
| 2025 | A trilevel framework of rough sets and granular rough sets: Characterizing existing models and formulating new modelsabstractWe propose a trilevel framework for studying rough sets and granular rough sets by applying the principles of three-way decision as thinking in threes. The framework builds and interprets any model of rough sets at three levels: the binary relations level concerning the relationships between objects, the granular space level concerning granules of objects, namely, sets of objects called granular objects, and the approximation level concerning the approximations of sets of objects by granular objects. We identify and characterize eight classes of rough set models, including Pawlak, covering-based, and granular rough sets. By reviewing the existing studies within the framework, we find that there is a lack of investigations on three classes. To fill in these gaps, we investigate two types of granular spaces induced by any binary relations: neighborhood-induced granular spaces and maximal-clique-induced granular spaces. We examine the properties of the two types of granular space and the properties of rough set approximations in the corresponding two classes of models. We also consider a third class of models of granular rough sets based on granular spaces without referencing a binary relation. Junfang Luo, Chengjun Shi, Yiyu Yao |
Inf. Sci. | 3 |
| 2024 | An axiomatic framework for three-way clustering
Yingxiao Chen, Ping Zhu 0001, Yiyu Yao |
Inf. Sci. | 3 |
| 2024 | Shadowed set approximations of L-fuzzy setsabstractPedrycz shadowed sets are three-way approximations of fuzzy sets by transforming the infinite levels of fuzzy set membership grades in the unit interval [ 0 , 1 ] into three levels. The three levels represent qualitatively the sets of the white, grey, and black members of a shadowed set. In this paper, we generalize the notion of shadowed sets to the case of L-fuzzy sets by making three new contributions. First, we consider two representations of a shadowed set. One is a three-valued L-fuzzy set and the other is three pairwise disjoint sets. Second, we introduce two methods for constructing a shadowed set. One divides a finite lattice based on the notion of a pair of a set of designated core membership grades and a set of designated null membership grades. The other uses a pair of threshold sets, which generalizes the method that uses a pair of thresholds. We study formal properties of the two methods and show that they are equivalent. Finally, based on a distance function on a lattice, we present a simple method to build the sets of designated core and null membership grades. Li Zhang 0088, Yiyu Yao, Ping Zhu 0001 |
Inf. Sci. | 2 |
| 2022 | Aggregation operators on shadowed sets
Stefania Boffa, Andrea Campagner, Davide Ciucci, Yiyu Yao |
Inf. Sci. | 4 |
| 2022 | BMW-TOPSIS: A generalized TOPSIS model based on three-way decision
Yumei Wang, Peide Liu, Yiyu Yao |
Inf. Sci. | 3 |
| 2022 | Matrix approach for fuzzy description reduction and group decision-making with fuzzy β-covering
Jingqian Wang 0001, Xiaohong Zhang 0001, Yiyu Yao |
Inf. Sci. | 3 |
| 2021 | A three-way decision based construction of shadowed sets from Atanassov intuitionistic fuzzy sets
Jilin Yang, Yiyu Yao |
Inf. Sci. | 2 |
| 2020 | Synergy of Granular Computing, Shadowed Sets, and Three-way Decisions
Davide Ciucci, Yiyu Yao |
Inf. Sci. | 2 |
| 2020 | Granularity-driven sequential three-way decisions: A cost-sensitive approach to classification
Yu Fang 0009, Cong Gao 0001, Yiyu Yao |
Inf. Sci. | 3 |
| 2020 | Covering-based variable precision fuzzy rough sets with PROMETHEE-EDAS methods
Jianming Zhan 0001, Haibo Jiang, Yiyu Yao |
Inf. Sci. | 3 |
| 2020 | Intuitionistic fuzzy TOPSIS method based on CVPIFRS models: An application to biomedical problems
Li Zhang 0088, Jianming Zhan 0001, Yiyu Yao |
Inf. Sci. | 3 |
| 2019 | On the properties of subsethood measures
Mengjun Hu, Xiaofei Deng, Yiyu Yao |
Inf. Sci. | 3 |
| 2019 | Min-max attribute-object bireducts: On unifying models of reducts in rough set theory
Xi'ao Ma, Yiyu Yao |
Inf. Sci. | 2 |
| 2019 | TOPSIS method based on a fuzzy covering approximation space: An application to biological nano-materials selection
Kai Zhang 0049, Jianming Zhan 0001, Yiyu Yao |
Inf. Sci. | 3 |
| 2019 | Three-way fuzzy partitions defined by shadowed sets
Xue Rong Zhao, Yiyu Yao |
Inf. Sci. | 2 |
| 2018 | Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way Decisions (Extended Abstract)abstractText classification techniques are playing a crucial role in identifying relevant texts from a large data set, e.g., various online crimes such as Cyberbullying, terrorist recruiting, propaganda or attack planning. Until now, supervised deep learning has brought about breakthroughs in processing multimedia data; however, there was no good practical way to harvest this opportunity for text classification because acquiring and maintaining a massive amount of training examples are too expensive for a large number of categories (e.g., Yahoo! taxonomy contains nearly 300,000 categories and the Library of Congress Subject Headings (LCSH) contains 394,070 subjects). Therefore, the question of how to effectively learn from sparse or small set of training examples is crucial for the true success of text classification. Semi-supervised approaches have been proposed for this challenge, which usually use a pair or several existing classifiers to extend a small training set. However, extracted pseudo training samples are uncertain because they are determined by a machine rather than people. Also, the massive volume and high variability of text data are creating a number of challenging issues such as the scalability and complicated relations between words. There are two fundamental issues with regards to the performance of existing classifiers: overlook and overload. Overlook means that some objects relevant to a class have been omitted, whereas overload means that some objects assigned to a class are actually not relevant to that class. The two issues are even more serious in the following two cases: (1) large uncertain boundary - the decision boundary between two classes includes many mixed examples (e.g., relevant and nonrelevant documents together), and (2) unbalanced classes - one class (e.g., information about terrorist attacks) is much smaller than another class (e.g., normal descriptions). We propose a three-way decision model [1] for dealing with the uncertain boundary for improving text classification performance based on rough set techniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary and negative regions) by two main boundary vectors created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors produced according to the structure of the boundary region. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. The experimental results on the standard data sets RCV1 and Reuters-21578 show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of F1 measure and AUC area compared with six other popular baseline models. Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001 |
ICDE | 4 |
| 2018 | A Linear Model for Three-Way Analysis of Facial SimilarityabstractCard sorting was used to gather information about facial similarity judgments. A group of raters put a set of facial photos into an unrestricted number of different piles according to each rater’s judgment of similarity. This paper proposes a linear model for 3-way analysis of similarity. An overall rating function is a weighted linear combination of ratings from individual raters. A pair of photos is considered to be similar, dissimilar, or divided, respectively, if the overall rating function is greater than or equal to a certain threshold, is less than or equal to another threshold, or is between the two thresholds. The proposed framework for 3-way analysis of similarity is complementary to studies of similarity based on features of photos. Daryl H. Hepting, Hadeel Hatim Bin Amer, Yiyu Yao |
IPMU (2) | 3 |
| 2018 | Modes of Sequential Three-Way Classifications
Yiyu Yao, Mengjun Hu, Xiaofei Deng |
IPMU (2) | 1 |
| 2018 | Three-way decision perspectives on class-specific attribute reducts
Xi'ao Ma, Yiyu Yao |
Inf. Sci. | 2 |
| 2017 | Dynamic probabilistic rough sets with incomplete data
Chuan Luo 0001, Tianrui Li 0001, Yiyu Yao |
Inf. Sci. | 3 |
| 2017 | Constructing shadowed sets and three-way approximations of fuzzy sets
Yiyu Yao, Xiaofei Deng |
Inf. Sci. | 1 |
| 2017 | Class-specific attribute reducts in rough set theory
Yiyu Yao, Xianyong Zhang |
Inf. Sci. | 1 |
| 2017 | Measurement of general granules
Liquan Zhao, Yiyu Yao |
Inf. Sci. | 2 |
| 2017 | Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way DecisionsabstractText classification is a process of classifying documents into predefined categories through different classifiers learned from labelled or unlabelled training samples. Many researchers who work on binary text classification attempt to find a more effective way to separate relevant texts from a large data set. However, current text classifiers cannot unambiguously describe the decision boundary between positive and negative objects because of uncertainties caused by text feature selection and the knowledge learning process. This paper proposes a three-way decision model for dealing with the uncertain boundary to improve the binary text classification performance based on therough settechniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary, and negative regions) by two main boundary vectors$\vec{C_{P}}$and$\vec{C_{N}}$, created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors$\vec{B_{P}}$and$\vec{B_{N}}$, produced according to the structure of the boundary region. It involves an indirect strategy which is composed of two successive steps in the whole classification process: ‘two-way to three-way’ and ‘three-way to two-way’. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. A large number of experiments have been conducted based on the standard data sets RCV1 and Reuters-21578. The experimental results show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of$F_{1}$measure and$AUC$area compared with six other popular baseline models. Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | A Semantical Approach to Rough Sets and Dominance-Based Rough Sets
Lynn D'eer, Chris Cornelis, Yiyu Yao |
IPMU (2) | 3 |
| 2016 | Rough-set concept analysis: Interpreting RS-definable concepts based on ideas from formal concept analysis
Yiyu Yao |
Inf. Sci. | 1 |
| 2016 | Rough set models in multigranulation spaces
Yiyu Yao |
Inf. Sci. | 1 |
| 2016 | Detecting and refining overlapping regions in complex networks with three-way decisions
Hong Yu 0007, Peng Jiao, Yiyu Yao, Guoyin Wang 0001 |
Inf. Sci. | 3 |
| 2014 | Decision-theoretic three-way approximations of fuzzy sets
Xiaofei Deng, Yiyu Yao |
Inf. Sci. | 2 |
| 2014 | Quantitative rough sets based on subsethood measures
Yiyu Yao, Xiaofei Deng |
Inf. Sci. | 1 |
| 2014 | Cost-sensitive three-way email spam filtering
Bing Zhou 0002, Yiyu Yao, Jigang Luo |
J. Intell. Inf. Syst. | 2 |
| 2012 | Covering based rough set approximations
Yiyu Yao, Bingxue Yao |
Inf. Sci. | 1 |
| 2012 | A measurement theory view on the granularity of partitions
Yiyu Yao, Liquan Zhao |
Inf. Sci. | 1 |
| 2011 | Probabilistic rule induction with the LERS data mining systemabstractBased on classical rough set approximations, the LERS (Learning from Examples based on Rough Sets) data mining system induces two types of rules, namely, certain rules from lower approximations and possible rules from upper approximations. By relaxing the stringent requirement of the classical rough sets, one can obtain probabilistic approximations. The LERS can be easily applied to induce probabilistic positive and boundary rules from probabilistic positive and boundary regions. This paper discusses several fundamental issues related to probabilistic rule induction with LERS, including rule induction algorithm, quantitative measures associated with rules, and the rule conflict resolution method. © 2011 Wiley Periodicals, Inc. Jerzy W. Grzymala-Busse, Yiyu Yao |
Int. J. Intell. Syst. | 2 |
| 2011 | The superiority of three-way decisions in probabilistic rough set models
Yiyu Yao |
Inf. Sci. | 1 |
| 2011 | Record-level peculiarity-based data analysis and classifications
Jian Yang 0016, Ning Zhong 0001, Yiyu Yao, Jue Wang 0004 |
Knowl. Inf. Syst. | 3 |
| 2011 | User-centric query refinement and processing using granularity-based strategies
Yi Zeng 0001, Ning Zhong 0001, Yulin Qin, Zhisheng Huang, Yiyu Yao, Frank van Harmelen |
Knowl. Inf. Syst. | 7 |
| 2010 | MGRS: A multi-granulation rough set
Jiye Liang, Yiyu Yao, Chuangyin Dang |
Inf. Sci. | 3 |
| 2010 | Three-way decisions with probabilistic rough sets
Yiyu Yao |
Inf. Sci. | 1 |
| 2010 | Rough implication operator based on strong topological rough algebras
Xiaohong Zhang 0001, Yiyu Yao, Hong Yu 0007 |
Inf. Sci. | 2 |
| 2010 | Evaluating information retrieval system performance based on user preference
Bing Zhou 0002, Yiyu Yao |
J. Intell. Inf. Syst. | 2 |
| 2010 | ACM TKDD Special Issue on Knowledge Discovery for Web IntelligenceabstractNo abstract available. Ning Zhong 0001, Gregory Piatetsky-Shapiro, Yiyu Yao, Philip S. Yu |
ACM Trans. Knowl. Discov. Data | 3 |
| 2009 | Peculiarity Analysis for ClassificationsabstractPeculiarity-oriented mining (POM) is a new data mining method consisting of peculiar data identification and peculiar data analysis. Peculiarity factor (PF) and local peculiarity factor (LPF) are important concepts employed to describe the peculiarity of points in the identification step. One can study the notions at both attribute and record levels. In this paper, a new record LPF called distance based record LPF (D-record LPF) is proposed, which is defined as the sum of distances between a point and its nearest neighbors. It is proved mathematically that D-record LPF can characterize accurately the probability density function of a continuous m-dimensional distribution. This provides a theoretical basis for some existing distance based anomaly detection techniques. More important, it also provides an effective method for describing the class conditional probabilities in the Bayesian classifier. The result enables us to apply peculiarity analysis for classification problems. A novel algorithm called LPF-Bayes classifier and its kernelized implementation are presented, which have some connection to the Bayesian classifier. Experimental results on several benchmark data sets demonstrate that the proposed classifiers are effective. Jian Yang 0016, Ning Zhong 0001, Yiyu Yao, Jue Wang 0004 |
ICDM | 3 |
| 2009 | DBLP-SSE: A DBLP Search Support EngineabstractA Search Support Engine (SSE) is implemented based on the basic principles of Information Retrieval Support Systems (IRSS) and Information Seeking Support Systems (ISSS). An SSE aims at meeting the diversity needs from different users, providing various supporting functionalities, tools, etc. for users to perform various tasks beyond the traditional search and browsing provided by current search engines. As an illustrative example, we developed a DBLP search support engine (DBLP-SSE), and we discuss some concrete supporting functionalities, namely, search refinement support, domain analysis support, etc. Each of the functionality focus on a unique perspective supporting users finding useful information and knowledge from the DBLP dataset. The search support engine can be considered as a step towards Knowledge Retrieval (KR) and Web Intelligence (WI). Yi Zeng 0001, Yiyu Yao, Ning Zhong 0001 |
Web Intelligence | 2 |
| 2009 | Relative reducts in consistent and inconsistent decision tables of the Pawlak rough set model
Duoqian Miao 0001, Yan Zhao 0001, Yiyu Yao, Huaxiong Li |
Inf. Sci. | 3 |
| 2009 | Discernibility matrix simplification for constructing attribute reducts
Yiyu Yao, Yan Zhao 0001 |
Inf. Sci. | 1 |
| 2008 | Local peculiarity factor and its application in outlier detectionabstractPeculiarity oriented mining (POM), aiming to discover peculiarity rules hidden in a dataset, is a new data mining method. In the past few years, many results and applications on POM have been reported. However, there is still a lack of theoretical analysis. In this paper, we prove that the peculiarity factor (PF), one of the most important concepts in POM, can accurately characterize the peculiarity of data with respect to the probability density function of a normal distribution, but is unsuitable for more general distributions. Thus, we propose the concept of local peculiarity factor (LPF). It is proved that the LPF has the same ability as the PF for a normal distribution and is the so-called µ-sensitive peculiarity description for general distributions. To demonstrate the effectiveness of the LPF, we apply it to outlier detection problems and give a new outlier detection algorithm called LPF-Outlier. Experimental results show that LPF-Outlier is an effective outlier detection algorithm. Jian Yang 0016, Ning Zhong 0001, Yiyu Yao, Jue Wang 0004 |
KDD | 3 |
| 2008 | A multiview approach for intelligent data analysis based on data operators
Yaohua Chen, Yiyu Yao |
Inf. Sci. | 2 |
| 2008 | Attribute reduction in decision-theoretic rough set models
Yiyu Yao, Yan Zhao 0001 |
Inf. Sci. | 1 |
| 2007 | Granular Computing for Web Intelligence and Brain Informatics
Yiyu Yao |
Web Intelligence | 1 |
| 2007 | Knowledge Retrieval (KR)abstractWith the ever-increasing growth of data and information, finding the right knowledge becomes a real challenge and an urgent task. Traditional data and information retrieval systems that support the current web are no longer adequate for knowledge seeking tasks. Knowledge retrieval systems will be the next generation of retrieval system serving those purposes. Basic issues of knowledge retrieval systems are examined and a conceptual framework of such systems is proposed. Theories and Technologies such as theory of knowledge, machine learning and knowledge discovery, psychology, logic and inference, linguistics, etc. are briefly mentioned for the implementation of knowledge retrieval systems. Two applications of knowledge retrieval in rough sets and biomedical domains are presented. Yiyu Yao, Yi Zeng 0001, Ning Zhong 0001, Jimmy Huang 0001 |
Web Intelligence | 1 |
| 2007 | Relational peculiarity-oriented mining
Muneaki Ohshima, Ning Zhong 0001, Yiyu Yao, Chunnian Liu |
Data Min. Knowl. Discov. | 3 |
| 2007 | Data analysis based on discernibility and indiscernibility
Yan Zhao 0001, Yiyu Yao |
Inf. Sci. | 2 |
| 2006 | The PDD Framework for Detecting Categories of Peculiar DataabstractPeculiar data are objects that are relatively few in number and significantly different from the other objects in a data set. In this paper, we propose the PDD framework for detecting multiple categories of peculiar data. This framework provides an extensible set of perspectives for viewing data, currently including viewing data as a set of records, attributes, frequencies, intervals, sequences, or sequences of changes. By using these six views of the data, multiple categories of peculiar data can be detected to reveal different aspects of the data. For each view, the framework provides an extensible set of peculiarity measures to detect outliers and other kinds of peculiar data. The PDD framework has been implemented for Oracle and Access. Experiments are reported for data sets concerning Regina weather and NHL hockey. Mahesh Shrestha, Howard J. Hamilton, Yiyu Yao, Ken Konkel, Liqiang Geng |
ICDM | 3 |
| 2006 | Perspective of Applying the Global E-mail NetworkabstractRecently, research on social network and Web intelligence (WI) has shown that social intelligence techniques act as the imperative channel for automated email-centric tasks. This paper gives a complete picture of what can we do in the global social e-mail network and how to do. Our main contributions include: (1) we describe two mechanisms for implementing applications in the global social e-mail network; (2) we design operable e-mail for communicating in that network. To our best knowledge, it is the first time to discuss how to implement social intelligence in such a network under the notion of WI. We believe that this work consequentially explores a new and absolutely necessarily needed research field of WI Wenbin Li 0008, Ning Zhong 0001, Jiming Liu 0001, Yiyu Yao, Chunnian Liu |
Web Intelligence | 4 |
| 2006 | Neighborhood systems and approximate retrieval
Yiyu Yao |
Inf. Sci. | 1 |
| 2005 | Visualization Support for Interactive Query RefinementabstractIt has been well documented that Web searchers have difficulties crafting queries to fulfill their information needs. In this work, we use a concept knowledge base generated from the ACM computing classification system to generate a query space that represents the query terms in relation to the concepts they describe and the other terms that are related to these concepts. A visual representation of this query space allows the user to interpret the relationships between their query terms and the query space. Interactive query refinement within this visual representation takes advantage of the user's visual information processing abilities, and allows the user to choose terms that accurately represent their information need. A preview of the search results from Google provides the user with an indication of the current state of their query refinement process. This work allows the user to take an active role in the information retrieval process, supporting the fundamental shift from information retrieval systems to information retrieval support systems. Orland Hoeber, Xue Dong Yang, Yiyu Yao |
Web Intelligence | 3 |
| 2005 | Multilevel Web PersonalizationabstractWeb personalization is one of the major concerns of Web intelligence. It is noticed that the two components of Web, users and services, can be understood from multiple views in forms of hierarchies, and each hierarchy is organized by a multilevel structure. A unified model is proposed for including all personalization styles. The unified model with multilevel hierarchy structures, add new understandings and insights into the Web personalization issues. More concise and precise recommendations can then be studied and pursued based on this multilevel model. Yan Zhao 0001, Yiyu Yao, Ning Zhong 0001 |
Web Intelligence | 2 |
| 2004 | Relational Peculiarity Oriented Data MiningabstractPeculiarity rules are a new type of interesting rules which can be discovered by searching the relevance among peculiar data. A main task of mining peculiarity rules is the identification of peculiarity. Traditional methods of finding peculiar data are attribute-based approaches. This paper extends peculiarity oriented mining to relational peculiarity oriented mining. Peculiar data are identified on record level, and peculiar rules are mined and explained in a relational mining framework. The results from preliminary experiments show that relational peculiarity oriented mining is very effective. Ning Zhong 0001, Chunnian Liu, Yiyu Yao, Muneaki Ohshima, Mingxin Huang, Jiajin Huang |
ICDM | 3 |
| 2004 | Peculiarity Oriented Analysis in Multi-people Tracking Images
Muneaki Ohshima, Ning Zhong 0001, Yiyu Yao, Shinichi Murata |
PAKDD | 3 |
| 2004 | TMS: Targeted Marketing System Based on Market Value FunctionsabstractTMS (Targeted Marketing System) is an integrated system and toolkit for profit-driven and cost-effective marketing. The system consists of three components: a Web-based user interface, a market value inference engine, and a presentation and evaluation module. It supports marketing decision making for a company or an organization by combining results from information retrieval, data mining, information theory, and utility theory. Jiajin Huang, Ning Zhong 0001, Chunnian Liu, Yiyu Yao, Dejun Qiu, Chuangxin Ou |
Web Intelligence | 4 |
| 2004 | Web-based Support Systems (WSS): A Report of the WIC Canada Research CentreabstractWIC Canada promotes the collaboration between Canadian researchers on Web Intelligence, facilitates exchange with other WIC centers. WIC Canada researchers work on a diversity of WI related research areas: foundations of Web Intelligence, Web-based support systems, Bayesian networks, and IntelligentWeb Information Systems (IWIS). Their results have appeared in reputable journals and the annual IEEE/WIC/ACM International Conference onWeb Intelligence. In the next few years, WIC Canada will focus on promotingWeb Intelligence research, attracting new members, and forming new centres. The WIC Canada will play the role of coordinating those activities. WIC Canada needs your input for its growth. Your participation will be greatly appreciated. All questions and suggestions should be directed to our Co-ordinator, Dr. Jingtao Yao, at [email protected]. Yiyu Yao, JingTao Yao 0001, Cory J. Butz, Pawan Lingras, Dawn N. Jutla |
Web Intelligence | 1 |
| 2003 | Attribute Reduction of Rough Sets in Mining Market Value FunctionsabstractThe linear model of market value functions is a new method for direct marketing. Just like other methods in direct marketing, attribute reduction is very important to deal with large databases. We apply the algorithm of attribute reduction, which is based on the combination of rough set theory with the boosting algorithm, to the linear model of market value functions. Experimental results compared with the ELSA/ANN model show that the proposed algorithms can be used effectively in the linear model of market value functions. Jiajin Huang, Chunnian Liu, Chuangxin Ou, Yiyu Yao, Ning Zhong 0001 |
Web Intelligence | 4 |
| 2003 | Web-Based Information Retrieval Support Systems: Building Research Tools for Scientists in the New Information AgeabstractThe concept of Web-based information retrieval support systems (WIRSS) is introduced. The needs for WIRSS are shown by a detailed case study of existing research article indexing and citation analysis systems, such as current content, DBLP, science citation index and CiteSeer. The objective of WIRSS is to build new and effective research tools for scientists to access, explore and use information on the Web, which may lead to improved research productivity and quality. JingTao Yao 0001, Yiyu Yao |
Web Intelligence | 2 |
| 2003 | The Wisdom Web: New Challenges for Web Intelligence (WI)
Jiming Liu 0001, Ning Zhong 0001, Yiyu Yao, Zbigniew W. Ras |
J. Intell. Inf. Syst. | 3 |
| 2003 | Peculiarity Oriented Multidatabase MiningabstractPeculiarity rules are a new class of rules which can be discovered by searching relevance among a relatively small number of peculiar data. Peculiarity oriented mining in multiple data sources is different from, and complementary to, existing approaches for discovering new, surprising, and interesting patterns hidden in data. A theoretical framework for peculiarity oriented mining is presented. Within the proposed framework, we give a formal interpretation and comparison of three classes of rules, namely, association rules, exception rules, and peculiarity rules, as well as describe how to mine interesting peculiarity rules in multiple databases. Ning Zhong 0001, Yiyu Yao, Muneaki Ohshima |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2002 | Value Added Association Rules
Tsau Young Lin, Yiyu Yao, Eric Louie |
PAKDD | 2 |
| 2001 | Information Retrieval by Possibilistic Reasoning
Churn-Jung Liau, Yiyu Yao |
DEXA | 2 |
| 2001 | Data Analysis and Mining in Ordered Information TablesabstractMany real-world problems deal with ordering objects instead of classifying objects, although the majority of the research in machine learning and data mining has been focused on the latter. For the modeling of ordering problems, we generalize the notion of information tables to ordered information tables by adding order relations on attribute values. The problem of mining ordering rules is formulated as finding associations between the orderings of attribute values and the overall ordering of objects. An ordering rule may state, for example, that "if the value of an object x on an attribute a is ordered ahead of the value of another object y on the same attribute, then x is ordered ahead of y". For mining ordering rules, we first transform an ordered information table into binary information, and then apply any standard machine learning and data mining algorithms. As an illustration, we analyze in detail the Maclean's university ranking for the year 2000. Ying Sai, Yiyu Yao, Ning Zhong 0001 |
ICDM | 2 |
| 2001 | Interestingness, Peculiarity, and Multi-Database MiningabstractIn order to discover new, surprising, interesting patterns hidden in data, peculiarity oriented mining and multidatabase mining are required. In the paper, we introduce peculiarity rules as a new class of rules, which can be discovered from a relatively low number of peculiar data by searching the relevance among the peculiar data. We give a formal interpretation and comparison of three classes of rules: association rules, exception rules, and peculiarity rules, as well as describe how to mine more interesting peculiarity rules in multiple databases. Ning Zhong 0001, Yiyu Yao, Muneaki Ohshima, Setsuo Ohsuga |
ICDM | 2 |
| 2001 | Web Intelligence (WI)
Yiyu Yao, Ning Zhong 0001, Jiming Liu 0001, Setsuo Ohsuga |
Web Intelligence | 1 |
| 2001 | Information granulation and rough set approximationabstractInformation granulation and concept approximation are some of the fundamental issues of granular computing. Granulation of a universe involves grouping of similar elements into granules to form coarse-grained views of the universe. Approximation of concepts, represented by subsets of the universe, deals with the descriptions of concepts using granules. In the context of rough set theory, this paper examines the two related issues. The granulation structures used by standard rough set theory and the corresponding approximation structures are reviewed. Hierarchical granulation and approximation structures are studied, which results in stratified rough set approximations. A nested sequence of granulations induced by a set of nested equivalence relations leads to a nested sequence of rough set approximations. A multi-level granulation, characterized by a special class of equivalence relations, leads to a more general approximation structure. The notion of neighborhood systems is also explored. © 2001 John Wiley & Sons, Inc. Yiyu Yao |
Int. J. Intell. Syst. | 1 |
| 2000 | On Association, Similarity and Dependency of Attributes
Yiyu Yao, Ning Zhong 0001 |
PAKDD | 1 |
| 1999 | On Information-Theoretic Measures of Attribute Importance
Yiyu Yao, S. K. Michael Wong, Cory J. Butz |
PAKDD | 1 |
| 1999 | An Analysis of Quantitative Measures Associated with Rules
Yiyu Yao, Ning Zhong 0001 |
PAKDD | 1 |
| 1999 | Peculiarity Oriented Multi-database Mining
Ning Zhong 0001, Yiyu Yao, Setsuo Ohsuga |
PKDD | 2 |
| 1998 | Constructive and Algebraic Methods of the Theory of Rough Sets
Yiyu Yao |
Inf. Sci. | 1 |
| 1998 | A Comparative Study of Fuzzy Sets and Rough Sets
Yiyu Yao |
Inf. Sci. | 1 |
| 1998 | Relational Interpretations of Neigborhood Operators and Rough Set Approximation Operators
Yiyu Yao |
Inf. Sci. | 1 |
| 1998 | Interpretation of Belief Functions in The Theory of Rough Sets
Yiyu Yao, Pawan Lingras |
Inf. Sci. | 1 |
| 1998 | Data Mining Using Extensions of the Rough Set ModelabstractThis article examines basic issues of data mining using the theory of rough sets, which is a recent proposal for generalizing classical set theory. The Pawlak rough set model is based on the concept of an equivalence relation. Recent research has shown that a generalized rough set model need not be based on equivalence relation axioms. The Pawlak rough set model has been used for deriving deterministic as well as probabilistic rules from a complete database. This article demonstrates that a generalized rough set model can be used for generating rules from incomplete databases. These rules are based on plausibility functions proposed by Shafer. The article also discusses the importance of rule extraction from incomplete databases in data mining. © 1998 John Wiley & Sons, Inc. Pawan Lingras, Yiyu Yao |
J. Am. Soc. Inf. Sci. | 2 |
| 1995 | Measuring Retrieval Effectiveness Based on User Preference of DocumentsabstractThe notion of user preference is adopted for the representation, interpretation, and measurement of the relevance or usefulness of documents. User judgments on documents may be formally described by a weak order (i.e., user ranking) and measured using an ordinal scale. Within this framework, a new measure of system performance is suggested based on the distance between user ranking and system ranking. It only uses the relative order of documents and therefore confirms to the valid use of an ordinal scale measuring relevance. It is also applicable to multilevel relevance judgments and ranked system output. The appropriateness of the proposed measure is demonstrated through an axiomatic approach. The inherent relationships between the new measure and many existing measures provide further supporting evidence. © 1995 John Wiley & Sons, Inc. Yiyu Yao |
J. Am. Soc. Inf. Sci. | 1 |
| 1995 | On Modeling Information Retrieval with Probabilistic InferenceabstractThis article examines and extends the logical models of information retrieval in the context of probability theory. The fundamental notions of term weights and relevance are given probabilistic interpretations. A unified framework is developed for modeling the retrieval process with probabilistic inference. This new approach provides a common conceptual and mathematical basis for many retrieval models, such as the Boolean, fuzzy set, vector space, and conventional probabilistic models. Within this framework, the underlying assumptions employed by each model are identified, and the inherent relationships between these models are analyzed. Although this article is mainly a theoretical analysis of probabilistic inference for information retrieval, practical methods for estimating the required probabilities are provided by simple examples. S. K. Michael Wong, Yiyu Yao |
ACM Trans. Inf. Syst. | 2 |
| 1993 | Computation of Term Associations by a Neural NetworkabstractThis paper suggests a method for computing term associations based on an adaptive bilinear retrieval model.Such a model can be implemented by using a three-layer feedforward neural network.Term associations are modeled by weighted links connecting different neurons, and are derived by the perception learning algorithm without the need for introducing any ad hoc parameters.The preliminary results indicate the usefulness of neural networks in the design of adaptive information retrieval systems. S. K. Michael Wong, Y. J. Cai, Yiyu Yao |
SIGIR | 3 |
| 1993 | A Probabilistic Method for Computing Term-by-Term RelationshipsabstractThis article suggests a probabilistic method to compute the term relationships from relevance information, which complements the studies on a non-probabilistic technique called pseudo-classification. A quadratic ranking function (i.e., a bilinear function) on the components of document and query vectors is derived by incorporating the term-by-term relationships. The conventional probabilistic indexing model, the probabilistic retrieval model, and our earlier generalized model are special cases of the proposed model. By exploring the different views of probability, procedures for estimating the required parameters are provided. © 1993 John Wiley & Sons, Inc. S. K. Michael Wong, Yiyu Yao |
J. Am. Soc. Inf. Sci. | 2 |
| 1992 | An Analysis of Vector Space Models Based on Computational GeometryabstractThis paper analyzes the properties, structures and limitations of vector-based models for information retrieval from the computational geometry point of view. It is shown that both the pseudo-cosine and the standard vector space models can be viewed as special cases of a generalized linear model. More importantly, both the necessary and sufficient conditions have been identified, under which ranking functions such as the inner-product, cosine, pseudo-cosine, Dice, covariance and product-moment correlation measures can be used to rank the documents. The structure of the solution region for acceptable ranking is analyzed and an algorithm for finding all the solution vectors is suggested. S. K. Michael Wong, Yiyu Yao |
SIGIR | 3 |
| 1992 | An Information-Theoretic Measure of Term SpecificityabstractThe inverse document frequency (IDF) and signal-noise ratio (S/N) approaches are two well known term weighting schemes based on term specificity. However, the existing justifications for these methods are still somewhat inconclusive and sometimes even based on incompatible assumptions. Although both methods are related to term specificity, their relationship has not been thoroughly investigated. An information-theoretic measure for term specificity is introduced in this study. It is explicitly shown that the IDF weighting scheme can be derived from the proposed approach by assuming that the frequency of occurrence of each index term is uniform within the set of documents containing the term. The information-theoretic interpretation of term specificity also establishes the relationship between the IDF and S/N methods. © 1992 John Wiley & Sons, Inc. S. K. Michael Wong, Yiyu Yao |
J. Am. Soc. Inf. Sci. | 2 |
| 1991 | Preference Structure, Inference and Set-Oriented RetrievalabstractIn this paper, a framework for modeling information retrieval is introduced by coml)ining the salient features of many inference-bad and set-orient,ed retrieval models.The degrees of relevance of different subsets of clocumeuts are inferred from the user preference judgments on subsets of index terms.In order to demonstrate the usefulness of the proposed framework, the Boolean and the binary vector space models are analyzed.This analysis reveals the structures implicitly used in these models. Yiyu Yao, S. K. Michael Wong |
SIGIR | 1 |
| 1991 | A probabilistic inference model for information retrieval
S. K. Michael Wong, Yiyu Yao |
Inf. Syst. | 2 |
| 1991 | Evaluation of an adaptive linear modelabstractThis article reports on the experimental evaluation of an adaptive linear model that constructs improved query vectors from the user preference judgments on a sample set of documents. The performance of this method is compared with that of the standard relevance feedback techniques. The experimental results seem to demonstrate the effectiveness of the adaptive method. © 1991 John Wiley & Sons, Inc. S. K. Michael Wong, Yiyu Yao, Gerard Salton, Chris Buckley |
J. Am. Soc. Inf. Sci. | 2 |
| 1990 | A generalized binary probabilistic independence modelabstractThe main objective of this article is to develop a theoretical model for information retrieval. By utilizing more complete statistical information, a generalized binary probabilistic independence model is proposed. Earlier probabilistic independence models are shown to be special cases of the new model. A quadratic discriminant function is obtained, which can be reduced to the linear discriminant functions of the conventional probabilistic models. © 1990 John Wiley & Sons, Inc. S. K. Michael Wong, Yiyu Yao |
J. Am. Soc. Inf. Sci. | 2 |
| 1990 | Query formulation in linear retrieval modelsabstractThe subject of query formulation is analyzed within the framework of adaptive linear models. Our study is based on the notions of user preference and an acceptable ranking strategy. Such an approach enables us to adopt a gradient descent algorithm to formulate the query vector by an inductive process. We also present a critical analysis of the existing relevance feedback and the probabilistic approaches. It is shown that Rocchio's method is a special case of our linear model and the independence assumption may be stronger than required for a linear system. Our method has the added advantages that it is applicable to both nonbinary document representation and a user preference relation inducing more than two classes. © 1990 John Wiley & Sons, Inc. S. K. Michael Wong, Yiyu Yao |
J. Am. Soc. Inf. Sci. | 2 |
| 1989 | A probability distribution model for information retrieval
S. K. Michael Wong, Yiyu Yao |
Inf. Process. Manag. | 2 |
| 1988 | Linear Structure in Information RetrievalabstractBased on the concept of user preference, we investigate the linear structure in information retrieval. We also discuss a practical procedure to determine the linear decision function and present an analysis of term weighting. Our experimental results seem to demonstrate that our model provides a useful framework for the design of an adaptive system. S. K. Michael Wong, Yiyu Yao, Peter Bollmann-Sdorra |
SIGIR | 2 |
| 1987 | A Statistical Similarity MeasureabstractWithin the framework of the vector space models, a statistical similarity measure between document and query is proposed. In this approach the assumption that term (or atomic) vectors are pairwise orthogonal is not required. In addition, it provides a natural and consistent interpretation of term occurrence frequencies obtained from autoindexing. S. K. Michael Wong, Yiyu Yao |
SIGIR | 2 |