Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Richard Jensen

dblp:34/545 · DBLP profile ↗
← Back
47ranked-venue papers
20as first author
1since 2021 · last 2026
0000-0002-1016-1524ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 16 first-authorDatabases, data management, data science and information retrieval · 9 · 4 first-author · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 94% Data integration and cleaning · 4% Database system architecture and tuning · 1%
Theoretical computer science
1 paper
Logic in computer science · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 56% Trustworthy machine learning · 44%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › dimensionality reduction
feature selection
0.112010
A Distance Measure Approach to Exploring the Rough Set Boundary Region for Attribute Reduction · IEEE Trans. Knowl. Data Eng. 2010
Data mining › dimensionality reduction › feature selection
rough set feature selection
0.112010
A Distance Measure Approach to Exploring the Rough Set Boundary Region for Attribute Reduction · IEEE Trans. Knowl. Data Eng. 2010
Logic in computer science › knowledge representation and reasoning › uncertainty reasoning
attribute reduction
0.112010
A Distance Measure Approach to Exploring the Rough Set Boundary Region for Attribute Reduction · IEEE Trans. Knowl. Data Eng. 2010
Logic in computer science › knowledge representation and reasoning › uncertainty reasoning
rough set theory
0.112010
A Distance Measure Approach to Exploring the Rough Set Boundary Region for Attribute Reduction · IEEE Trans. Knowl. Data Eng. 2010
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.012004
Semantics-Preserving Dimensionality Reduction: Rough and Fuzzy-Rough-Based Approaches · IEEE Trans. Knowl. Data Eng. 2004
Machine learning › Trustworthy machine learning
interpretability
0.012004
Semantics-Preserving Dimensionality Reduction: Rough and Fuzzy-Rough-Based Approaches · IEEE Trans. Knowl. Data Eng. 2004
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › feature selection
feature grouping
0.012004
Semantics-Preserving Dimensionality Reduction: Rough and Fuzzy-Rough-Based Approaches · IEEE Trans. Knowl. Data Eng. 2004
Data integration and cleaning › schema mapping
object-relational mapping
0.011993
Persistence Software: Bridging Object-Oriented Programming and Relational Databases · SIGMOD Conference 1993
Database system architecture and tuning
relational database system
0.011993
Persistence Software: Bridging Object-Oriented Programming and Relational Databases · SIGMOD Conference 1993

Methods — techniques the papers use, named apart from their topics

distance metric · 0.2dependency function · 0.2boundary region analysis · 0.2rough set theory · 0.0fuzzy rough set · 0.0client-side caching · 0.0automatic code generation · 0.0
YearPublicationVenuePosition
2026 Feature selection for high-dimensional imbalanced class datasets using Harmony Search and Kullback-Leibler divergence
Alireza Moayedikia, Richard Jensen, Sara Fin
Data Min. Knowl. Discov.2
2020 Fuzzy-Rough Set Bireducts for Data Reduction
abstract
Data reduction is an important step that helps ease the computational intractability for learning techniques when data are large. This is particularly true for the huge datasets that have become commonplace in recent times. The main problem facing both data preprocessors and learning techniques is that data are expanding both in terms of dimensionality and also in terms of the number of data instances. Approaches based on fuzzy-rough sets offer many advantages for both feature selection and classification, particularly for real-valued and noisy data; however, the majority of recent approaches tend to address the task of data reduction in terms of either dimensionality or training data size in isolation. This paper demonstrates how the notion of fuzzy-rough bireducts can be used for the simultaneous reduction of data size and dimensionality. It also shows how bireducts and, therefore, reduced subtables of data can be used not only as a preprocessing tool but also for the learning of compact and robust classifiers. Furthermore, the ideas can also be extended to the unsupervised domain when dealing with unlabeled data. Experimental evaluation of various techniques demonstrate that high levels of simultaneous reduction of both dimensionality and data size can be achieved whilst maintaining robust performance.
Neil Mac Parthaláin, Richard Jensen, Ren Diao
IEEE Trans. Fuzzy Syst.2
2019 Effective instance selection using the fuzzy-rough lower approximation
abstract
Fuzzy-rough set theory has been applied with much success to the problem of feature selection, where there is a clear link between the constructs (i.e., lower approximation, positive region, etc) and the problem (i.e., finding reducts). However, there has not been much development with regards to instance selection. Previous techniques have focused on preserving the positive region or the dependency, or have concentrated on prototype selection. This paper proposes a general instance selection approach that is efficient and effective in finding reductions that maintain the integrity of the underlying class structure. By utilizing the lower approximation information, instances can be removed that have a high similarity with more representative instances of a class. In addition to a threshold-based approach, a fully automated method that requires no user input is also presented. The experimentation demonstrates that the proposed method is indeed effective at reducing the number of instances with minimal information loss.
Richard Jensen, Mehran Amiri, Neil Mac Parthaláin
FUZZ-IEEE1
2017 Feature selection for high dimensional imbalanced class data using harmony search
Alireza Moayedikia, Kok-Leong Ong, Yee Ling Boo, William Yeoh 0002, Richard Jensen
Eng. Appl. Artif. Intell.5
2016 Dataset condensation using OWA fuzzy-rough set-based nearest neighbor classifier
abstract
The application of fuzzy-rough sets for the task of feature selection and rule induction has been the topic of much interest recently. However, applications for data instance or object selection have attracted much less attention. In this paper a novel approach for dataset condensation based on ordered weighted aggregation (OWA) fuzzy-rough sets is proposed in the context of the KNN classifier. Initially, a rank is assigned to each data instance of the dataset, based upon a novel measure inspired by fuzzy-rough sets. High quality data instances which possess higher ranks are retained based on another new metric whilst others can then be removed. An additional novel innovation is the elimination of any subjective user-specified threshold in order to determine which particular data instances are candidates for removal. This is in keeping with the rough set ideology of data-driven approaches. A series of non-parametric statistical tests demonstrate that the technique is very effective and can produce useful condensations of the data.
Mehran Amiri, Richard Jensen, Mahdi Eftekhari, Neil Mac Parthaláin
FUZZ-IEEE2
2016 Missing data imputation using fuzzy-rough methods
Mehran Amiri, Richard Jensen
Neurocomputing2
2015 Discovering fuzzy-rough reducts through Estimation of Distribution Algorithms
abstract
Due to the explosive growth of stored information worldwide, feature selection (FS) is becoming an increasingly important step, particularly given the abundance of noisy, irrelevant or misleading features. The main aim of FS is to determine a minimal feature subset from a problem domain while retaining a suitably high accuracy in representing the original set of features. However, the problem of finding optimal reductions is challenging as there is always a trade-off between the extent of reduction and the resulting information loss. This topic has been of particular interest in rough and fuzzy-rough set theory, as these provide a mechanism for defining optimality using only the data itself. Evolutionary methods have been used to try to find rough and fuzzy-rough optimal reductions, but these approaches ignore the fact that not all equally-sized reducts have the same utility for classifiers. This paper presents a novel approach for fuzzy-rough feature selection that uses Estimation of Distribution Algorithms to maintain information about the quality of features, to then obtain a better quality reduct that is more useful in general.
Richard Jensen, Neil Mac Parthaláin
FUZZ-IEEE1
2015 Fuzzy-rough feature selection using flock of starlings optimisation
abstract
Much use has been made of particle swarm optimisation as a tool to solve complex optimisation tasks, and many extensions and modifications to the original algorithm have been proposed. One such extension is related to the murmuration or flocking behaviour of starling birds and their flight trajectories in relation to flock cohesion giving rise to the so-called flock of starlings optimisation algorithm. This algorithm uses the topological model of starling bird flocks as a basis for modifying the original particle swarm optimisation approach. In this paper, two novel approaches for feature selection using fuzzy-rough sets and based upon two different interpretations of the flock of starlings algorithm are proposed. The results demonstrate that the approach can converge quickly and can discover subsets of smaller size and which are more stable than traditional PSO.
Neil Mac Parthaláin, Richard Jensen
FUZZ-IEEE2
2015 Weighted bee colony algorithm for discrete optimization problems with application to feature selection
Alireza Moayedikia, Richard Jensen, Uffe Kock Wiil, Rana Forsati
Eng. Appl. Artif. Intell.2
2015 Towards scalable fuzzy-rough feature selection
Richard Jensen, Neil Mac Parthaláin
Inf. Sci.1
2014 Heuristic search for fuzzy-rough bireducts and its use in classifier ensembles
abstract
Rough set theory has proven to be a useful mathematical basis for developing automated computational approaches which are able to deal with and utilise imperfect knowledge. Fuzzy-rough set theory is an extension to rough set theory and enhances the ability to model uncertainty and vagueness more effectively. There have been many developments in this area which offer robust methods for feature selection or instance selection. However, these are often carried out in isolation rather than considering both types of selection simultaneously. For this purpose, the notion of a bireduct has been proposed recently but the task of finding bireducts of high quality remains a significant challenge. This paper presents a heuristic strategy for the identification of fuzzy-rough bireducts, which is based on a music-inspired global optimisation algorithm called harmony search. The concept of e-bireducts is employed in this approach for the evaluation and improvisation of the candidate solutions. The stochastically-selected bireducts are also utilised to construct classifier ensembles. The presented technique is experimentally evaluated using a number of real-valued benchmark data sets.
Ren Diao, Neil Mac Parthaláin, Richard Jensen, Qiang Shen 0001
FUZZ-IEEE3
2014 Feature grouping-based fuzzy-rough feature selection
abstract
Data dimensionality has become a pervasive problem in many areas that require the learning of interpretable models. This has become particularly pronounced in recent years with the seemingly relentless growth in the size of datasets. Indeed, as the number of dimensions increases, the number of data instances required in order to generate accurate models increases exponentially. Feature selection has therefore become not only a useful step in the process of model learning, but rather an increasingly necessary one. Rough set and fuzzy-rough set theory have been used as such dataset pre-processors with much success, however the underlying time/space complexity of the subset evaluation metric is an obstacle to the processing of very large data. This paper proposes a general approach to this problem that employs a novel feature grouping step in order to alleviate the processing overhead for large datasets. The approach is framed within the context of (and applied to) fuzzy-rough sets, although it can be used with other subset evaluation techniques. The experimental evaluation demonstrates that considerable computational effort can be avoided, and as a result efficiency can be improved considerably for larger datasets.
Richard Jensen, Neil Mac Parthaláin, Chris Cornelis
FUZZ-IEEE1
2014 Enriched ant colony optimization and its application in feature selection
Rana Forsati, Alireza Moayedikia, Richard Jensen, Mehrnoush Shamsfard, Mohammad Reza Meybodi
Neurocomputing3
2014 Finding rough and fuzzy-rough set reducts with SAT
Richard Jensen, Andrew Tuson, Qiang Shen 0001
Inf. Sci.1
2013 Simultaneous feature and instance selection using fuzzy-rough bireducts
abstract
Rough set theory has proven to be a useful mathematical basis for developing automated computational approaches which are able to deal with and utilise imperfect knowledge. Ever since its inception, this theory has been successfully employed for developing computationally efficient techniques for addressing problems such as the discovery of hidden patterns in data, decision rule induction, and feature selection. As an extension to this theory, fuzzy-rough sets enhance the ability to model uncertainty and vagueness more effectively. The efficacy of fuzzy-rough set based approaches for the tasks of feature selection and rule induction is now well established in the literature. Although some work has been carried out using fuzzy-rough set theory for the tasks of feature selection and instance selection in isolation, the potential of this theory for its application to tasks for the simultaneous selection of both features and instances has not been investigated thus far. This paper proposes a novel method for simultaneous instance and feature selection based on fuzzy-rough sets. The initial experimentation demonstrates that the method can significantly reduce both the number of instances and features whilst maintaining high classification accuracies.
Neil Mac Parthaláin, Richard Jensen
FUZZ-IEEE2
2013 Quality, frequency and similarity based fuzzy nearest neighbor classification
abstract
This paper proposes an approach based on fuzzy rough set theory to improve nearest neighbor based classification. Six measures are introduced to evaluate the quality of the nearest neighbors. This quality is combined with the frequency at which classes occur among the nearest neighbors and the similarity w.r.t. the nearest neighbor, to decide which class to pick among the neighbor's classes. The importance of each aspect is weighted using optimized weights. An experimental study shows that our method, Quality, Frequency and Similarity based Fuzzy Nearest Neighbor (QFSNN), outperforms state-of-the-art nearest neighbor classifiers.
Nele Verbiest, Chris Cornelis, Richard Jensen
FUZZ-IEEE3
2013 Unsupervised fuzzy-rough set-based dimensionality reduction
Neil Mac Parthaláin, Richard Jensen
Inf. Sci.2
2012 Fuzzy rough positive region based nearest neighbour classification
abstract
This paper proposes a classifier that uses fuzzy rough set theory to improve the Fuzzy Nearest Neighbour (FNN) classifier.We show that previous attempts to use fuzzy rough set theory to improve the FNN algorithm have some shortcomings and we overcome them by using the fuzzy positive region to measure the quality of the nearest neighbours in the FNN classifier.A preliminary experimental evaluation shows that the new approach generally improves upon existing methods.
Nele Verbiest, Chris Cornelis, Richard Jensen
FUZZ-IEEE3
2011 Fuzzy-rough set based semi-supervised learning
abstract
Much work has been carried out in the area of fuzzy-rough sets for supervised learning. However, very little has been accomplished for the unsupervised or semi-supervised tasks. For many real-word applications, it is often expensive, time-consuming and difficult to obtain labels for all data objects. This often results in large quantities of data which may only have very few labelled data objects. This paper proposes a novel fuzzy-rough based semi-supervised self-learning or self-training approach for the assignment of labels to unlabelled data. Unlike other semi-supervised approaches, the proposed technique requires no subjective thresholding or domain information. An experimental evaluation is performed on artificial data and also applied to a real-world mammographic risk assessment problem with encouraging results.
Neil Mac Parthaláin, Richard Jensen
FUZZ-IEEE2
2011 Fuzzy-rough nearest neighbour classification and prediction
Richard Jensen, Chris Cornelis
Theor. Comput. Sci.1
2010 Fuzzy-rough instance selection
abstract
Rough set theory provides a useful mathematical foundation for developing automated computational systems that can help understand and make use of imperfect knowledge. Since its introduction, this theory has been successfully utilised to devise mathematically sound and often, computationally efficient techniques for addressing problems such as hidden pattern discovery from data, feature selection and decision rule generation. Fuzzy-rough set theory improves upon this by enabling uncertainty and vagueness to be modeled more effectively. Recently, the value of fuzzy-rough sets for feature selection and rule induction has been established. However, the potential of this theory for instance selection has not been investigated at all. This paper proposes three novel methods for instance selection based on fuzzy-rough sets. The initial experimentation demonstrates that the methods can significantly reduce the number of instances whilst maintaining high classification accuracies.
Richard Jensen, Chris Cornelis
FUZZ-IEEE1
2010 Extending propositional satisfiability to determine minimal fuzzy-rough reducts
abstract
This paper describes a novel, principled approach to real-valued dataset reduction based on fuzzy and rough set theory. The approach is based on the formulation of fuzzy-rough discernibility matrices, that can be transformed into a satisfiability problem; an extension of rough set approaches that only apply to discrete datasets. The fuzzy-rough hybrid reduction method is then realised algorithmically by a modified version of a traditional satisifability approach. This produces an efficient and provably optimal approach to data reduction that works well on a number of machine learning benchmarks in terms of both time and classification accuracy.
Richard Jensen, Andrew Tuson, Qiang Shen 0001
FUZZ-IEEE1
2010 Fuzzy-rough approaches for mammographic risk analysis
abstract
The accuracy of methods for the assessment of mammographic risk analysis is heavily related to breast tissue characteristics. Previous work has demonstrated considerable success in developing an automatic breast tissue classification methodology which overcomes this difficulty. This paper proposes a unified approach for the application of a number of rough and fuzzy-rough set methods to the analysis of mammographic data. Indeed this is the first time that fuzzy-rough approaches have been applied to this particular problem domain. In the unified approach detailed here feature selection methods are employed for dimensionality reduction developed using rough sets and fuzzy-rough sets. A number of classifiers are then used to examine the data reduced by the feature selection approaches and assess the positive impact of these methods on classification accuracy. Additionally, this paper also employs a new fuzzy-rough classifier based on the nearest neighbour classification algorithm. The novel use of such an approach demonstrates its efficiency in improving classification accuracy for mammographic data, as well as considerably removing redundant, irrelevant, and noisy features. This is supported with experimental application to two well-known datasets. The overall result of employing the proposed unified approach is that feature selection can identify only those features which require extraction. This can have the positive effect of increasing the risk assessment accuracy rate whilst additionally reducing the time required for expert scrutiny, which in-turn means the risk analysis process is potentially quicker and involves less screening.
Neil Mac Parthaláin, Richard Jensen, Qiang Shen 0001, Reyer Zwiggelaar
Intell. Data Anal.2
2010 Attribute selection with fuzzy decision reducts
Chris Cornelis, Richard Jensen, Germán Hurtado Martín, Dominik Slezak
Inf. Sci.2
2010 A Distance Measure Approach to Exploring the Rough Set Boundary Region for Attribute Reduction
abstract
Feature Selection (FS) or Attribute Reduction techniques are employed for dimensionality reduction and aim to select a subset of the original features of a data set which are rich in the most useful information. The benefits of employing FS techniques include improved data visualization and transparency, a reduction in training and utilization times and potentially, improved prediction performance. Many approaches based on rough set theory up to now, have employed the dependency function, which is based on lower approximations as an evaluation step in the FS process. However, by examining only that information which is considered to be certain and ignoring the boundary region, or region of uncertainty, much useful information is lost. This paper examines a rough set FS technique which uses the information gathered from both the lower approximation dependency value and a distance metric which considers the number of objects in the boundary region and the distance of those objects from the lower approximation. The use of this measure in rough set feature selection can result in smaller subset sizes than those obtained using the dependency function alone. This demonstrates that there is much valuable information to be extracted from the boundary region. Experimental results are presented for both crisp and real-valued data and compared with two other FS techniques in terms of subset size, runtimes, and classification accuracy.
Neil Mac Parthaláin, Qiang Shen 0001, Richard Jensen
IEEE Trans. Knowl. Data Eng.3
2009 Hybrid fuzzy-rough rule induction and feature selection
abstract
The automated generation of feature pattern-based if-then rules is essential to the success of many intelligent pattern classifiers, especially when their inference results are expected to be directly human-comprehensible. Fuzzy and rough set theory have been applied with much success to this area as well as to feature selection. Since both applications of rough set theory involve the processing of equivalence classes for their successful operation, it is natural to combine them into a single integrated method that generates concise, meaningful and accurate rules. This paper proposes such an approach, based on fuzzy-rough sets. The algorithm is experimentally evaluated against leading classifiers, including fuzzy and rough rule inducers, and shown to be effective.
Richard Jensen, Chris Cornelis, Qiang Shen 0001
FUZZ-IEEE1
2009 Interval-valued fuzzy-rough feature selection in datasets with missing values
abstract
One of the many successful applications of rough set theory has been to the area of feature selection. The rough set principle of using only the supplied data and no other information has many benefits, where most other methods require supplementary knowledge. Fuzzy-rough set theory has recently been proposed as an extension of this, in order to better handle the uncertainty present in real data. However, following this approach, there has been no investigation (theoretical or otherwise) into how to deal with missing values effectively, another problem encountered when using real world data. This paper proposes an extension of the fuzzy-rough feature selection methodology, based on interval-valued fuzzy sets, as a means to counter this problem via the representation of missing values in an intuitive way.
Richard Jensen, Qiang Shen 0001
FUZZ-IEEE1
2009 Measures for Unsupervised Fuzzy-Rough Feature Selection
abstract
For supervised learning, feature selection algorithms attempt to maximise a given function of predictive accuracy. This function usually considers the ability of feature vectors to reflect decision class labels. It is therefore intuitive to retain only those features that are related to or lead to these decision classes. However, in unsupervised learning, decision class labels are not provided, which poses questions such as; which features should be retained? and, why not use all of the information? The problem is that not all features are important. Some of the features may be redundant, and others may be irrelevant and noisy. In this paper, some new fuzzy-rough set-based approaches to unsupervised feature selection are proposed. These approaches require no thresholding or domain information, and result in a significant reduction in dimensionality whilst retaining the semantics of the data.
Neil Mac Parthaláin, Richard Jensen
ISDA2
2009 Feature selection for aiding glass forensic evidence analysis
abstract
The evaluation of glass evidence in forensic science is an important issue. Traditionally, this has depended on the comparison of the physical and chemical attributes of an unknown fragment with a control fragment. A high degree of discrimination between glass fragments is now achievable due to adv ances in analytical capabilities. A random effects model using two levels of hierarchical nesting is applied to the calculation of a likelihood ratio (LR) as a solution to the problem of comparison between two sets of replicated continuous observations where it is unknown whether the sets of measurements shared a common origin. Replicate measurements from a population of such measurements allow the calculation of both within-group and between-group variances. Univariate normal kernel estimation procedures have been used for this, where the between-group distribution is considered to be non-normal. However, the choice of variable for use in LR estimation is critical to the quality of LR produced. This paper investigates the use of feature selection for the purpose of selecting the variable for estimation without the need for expert knowledge. Results are recorded for several selectors using normal, exponential, adaptive and biweight kernel estimation techniques. Misclassification rates for the LR estimators are used to measure performance. The experiments performed reveal the capability of the proposed approach for this task.
Richard Jensen, Qiang Shen 0001
Intell. Data Anal.1
2009 New Approaches to Fuzzy-Rough Feature Selection
abstract
There has been great interest in developing methodologies that are capable of dealing with imprecision and uncertainty. The large amount of research currently being carried out in fuzzy and rough sets is representative of this. Many deep relationships have been established, and recent studies have concluded as to the complementary nature of the two methodologies. Therefore, it is desirable to extend and hybridize the underlying concepts to deal with additional aspects of data imperfection. Such developments offer a high degree of flexibility and provide robust solutions and advanced tools for data analysis. Fuzzy-rough set-based feature (FS) selection has been shown to be highly useful at reducing data dimensionality but possesses several problems that render it ineffective for large datasets. This paper proposes three new approaches to fuzzy-rough FS-based on fuzzy similarity relations. In particular, a fuzzy extension to crisp discernibility matrices is proposed and utilized. Initial experimentation shows that the methods greatly reduce dimensionality while preserving classification accuracy.
Richard Jensen, Qiang Shen 0001
IEEE Trans. Fuzzy Syst.1
2009 Are More Features Better? A Response to Attributes Reduction Using Fuzzy Rough Sets
abstract
A recent TRANSACTIONS ON FUZZY SYSTEMS paper proposing a new fuzzy-rough feature selector (FRFS) has claimed that the more attributes remain in datasets, the better the approximations and hence resulting models. [Tsang ,IEEE Trans. Fuzzy Syst., vol. 16, no. 5, pp. 1130-1141]. This claim has been used as a primary criticism of the original FRFS method [Jensen and Shen,IEEE Trans. Fuzzy Syst., vol. 15, no. 1, pp. 73-89, Feb. 2007]. Although, in certain applications, it may be necessary to consider as many features as possible, the claim is contrary to the motivation behind feature selection concerning the curse of dimensionality, the presence of redundant and irrelevant features, and the large amount of literature documenting observed improvements in modeling techniques following data reduction. This letter discusses this issue, as well as two other issues raised by Tsang [IEEE Trans. Fuzzy Syst., vol. 16, no. 5, pp. 1130-1141, Oct. 2008] regarding the original algorithm.
Richard Jensen, Qiang Shen 0001
IEEE Trans. Fuzzy Syst.1
2008 A noise-tolerant approach to fuzzy-rough feature selection
abstract
In rough set based feature selection, the goal is to omit attributes (features) from decision systems such that objects in different decision classes can still be discerned. A popular way to evaluate attribute subsets with respect to this criterion is based on the notion of dependency degree. In the standard approach, attributes are expected to be qualitative; in the presence of quantitative attributes, the methodology can be generalized using fuzzy rough sets, to handle gradual (in) discernibility between attribute values more naturally. However, both the extended approach, as well as its crisp counterpart, exhibit a strong sensitivity to noise: a change in a single object may significantly influence the outcome of the reduction procedure. Therefore, in this paper, we consider a more flexible methodology based on the recently introduced vaguely quantified rough set (VQRS) model. The method can handle both crisp (discrete-valued) and fuzzy (real-valued) data, and encapsulates the existing noise-tolerant data reduction approach using variable precision rough sets (VPRS), as well as the traditional rough set model, as special cases.
Chris Cornelis, Richard Jensen
FUZZ-IEEE2
2008 Finding fuzzy-rough reducts with fuzzy entropy
abstract
Dataset dimensionality is undoubtedly the single most significant obstacle which exasperates any attempt to apply effective computational intelligence techniques to problem domains. In order to address this problem a technique which reduces dimensionality is employed prior to the application of any classification learning. Such feature selection (FS) techniques attempt to select a subset of the original features of a dataset which are rich in the most useful information. The benefits can include improved data visualisation and transparency, a reduction in training and utilisation times and potentially, improved prediction performance. Methods based on fuzzy-rough set theory have demonstrated this with much success. Such methods have employed the dependency function which is based on the information contained in the lower approximation as an evaluation step in the FS process. This paper presents three novel feature selection techniques employing fuzzy entropy to locate fuzzy-rough reducts. This approach is compared with two other fuzzy-rough feature selection approaches which utilise other measures for the selection of subsets.
Neil Mac Parthaláin, Richard Jensen, Qiang Shen 0001
FUZZ-IEEE2
2008 Approximation-based feature selection and application for algae population estimation
Qiang Shen 0001, Richard Jensen
Appl. Intell.2
2007 Tolerance-based and Fuzzy-Rough Feature Selection
abstract
One of the main obstacles facing the application of computational intelligence technologies in pattern recognition (and indeed in many other tasks) is that of dataset dimensionality. To enable pattern classifiers to be effective, a dimensionality minimization step is usually carried out beforehand. Rough set theory has been successfully applied for this as it requires only the supplied data and no other information; most other methods require supplementary knowledge. However, the main limitation of traditional rough set-based selection in the literature is the restrictive requirement that all data is discrete; it is not possible to consider real-valued or noisy data. This has been tackled previously via the use of discretization methods, but may result in information loss. This paper investigates two approaches based on rough set extensions, namely fuzzy-rough and tolerance rough sets, that address these problems and retain dataset semantics. The methods are compared experimentally and utilized for the task of forensic glass fragment identification.
Richard Jensen, Qiang Shen 0001
FUZZ-IEEE1
2007 Survey of Rough and Fuzzy Hybridization
abstract
This paper provides a broad overview of logical and black box approaches to fuzzy and rough hybridization. The logical approaches include theoretical, supervised learning, feature selection, and unsupervised learning. The black box approaches consist of neural and evolutionary computing. Since both theories originated in the expert system domain, there are a number of research proposals that combine rough and fuzzy concepts in supervised learning. However, continuing developments of rough and fuzzy extensions to clustering, neurocomputing, and genetic algorithms make hybrid approaches in these areas a potentially rewarding research opportunity as well.
Pawan Lingras, Richard Jensen
FUZZ-IEEE2
2007 Distance Measure Assisted Rough Set Feature Selection
abstract
Feature selection (FS) is a technique for dimensionality reduction. Its aims are to select a subset of the original features of a dataset which are rich in the most useful information. The benefits include improved data visualisation, transparency, a reduction in training and utilisation times and potentially, improved prediction performance. Many approaches based on rough set theory have employed the dependency function which is based on the information contained in the lower approximation as an evaluation step in the FS process with much success. This paper presents a novel rough set FS technique which uses the information of both the lower approximation dependency value and a distance metric for the consideration of objects in the boundary region. The use of this measure in rough set feature selection can result in smaller subset sizes than those obtained using the dependency function alone.
Neil Mac Parthaláin, Qiang Shen 0001, Richard Jensen
FUZZ-IEEE3
2007 Feature selection based on rough sets and particle swarm optimization
Xiangyang Wang 0003, Jie Yang 0002, Xiaolong Teng, Weijun Xia, Richard Jensen
Pattern Recognit. Lett.5
2007 Fuzzy-Rough Sets Assisted Attribute Selection
abstract
Attribute selection (AS) refers to the problem of selecting those input attributes or features that are most predictive of a given outcome; a problem encountered in many areas such as machine learning, pattern recognition and signal processing. Unlike other dimensionality reduction methods, attribute selectors preserve the original meaning of the attributes after reduction. This has found application in tasks that involve datasets containing huge numbers of attributes (in the order of tens of thousands) which, for some learning algorithms, might be impossible to process further. Recent examples include text processing and web content classification. AS techniques have also been applied to small and medium-sized datasets in order to locate the most informative attributes for later use. One of the many successful applications of rough set theory has been to this area. The rough set ideology of using only the supplied data and no other information has many benefits in AS, where most other methods require supplementary knowledge. However, the main limitation of rough set-based attribute selection in the literature is the restrictive requirement that all data is discrete. In classical rough set theory, it is not possible to consider real-valued or noisy data. This paper investigates a novel approach based on fuzzy-rough sets, fuzzy rough feature selection (FRFS), that addresses these problems and retains dataset semantics. FRFS is applied to two challenging domains where a feature reducing step is important; namely, web content classification and complex systems monitoring. The utility of this approach is demonstrated and is compared empirically with several dimensionality reducers. In the experimental studies, FRFS is shown to equal or improve classification accuracy when compared to the results from unreduced data. Classifiers that use a lower dimensional set of attributes which are retained by fuzzy-rough reduction outperform those that employ more attributes returned by the existing crisp rough reduction method. In addition, it is shown that FRFS is more powerful than the other AS techniques in the comparative study
Richard Jensen, Qiang Shen 0001
IEEE Trans. Fuzzy Syst.1
2006 Fuzzy Entropy-assisted Fuzzy-Rough Feature Selection
abstract
Feature selection (FS) is a dimensionality reduction technique that aims to select a subset of the original features of a dataset which offer the most useful information. The benefits of feature selection include improved data visualisation, transparency, reduction in training and utilisation times and improved prediction performance. Methods based on fuzzy-rough set theory (FRFS) have employed the dependency function to guide the process with much success. This paper presents a novel fuzzy-rough FS technique which is guided by fuzzy entropy. The use of this measure in fuzzy-rough feature selection can result in smaller subset sizes than those obtained through FRFS alone, with little loss or even an increase in overall classification accuracy.
Neil Mac Parthaláin, Richard Jensen, Qiang Shen 0001
FUZZ-IEEE2
2005 Fuzzy-rough data reduction with ant colony optimization
Richard Jensen, Qiang Shen 0001
Fuzzy Sets Syst.1
2004 Fuzzy-rough attribute reduction with application to web categorization
Richard Jensen, Qiang Shen 0001
Fuzzy Sets Syst.1
2004 Selecting informative features with fuzzy-rough sets and its application for complex systems monitoring
Qiang Shen 0001, Richard Jensen
Pattern Recognit.2
2004 Semantics-Preserving Dimensionality Reduction: Rough and Fuzzy-Rough-Based Approaches
abstract
Semantics-preserving dimensionality reduction refers to the problem of selecting those input features that are most predictive of a given outcome; a problem encountered in many areas such as machine learning, pattern recognition, and signal processing. This has found successful application in tasks that involve data sets containing huge numbers of features (in the order of tens of thousands), which would be impossible to process further. Recent examples include text processing and Web content classification. One of the many successful applications of rough set theory has been to this feature selection area. This paper reviews those techniques that preserve the underlying semantics of the data, using crisp and fuzzy rough set-based methodologies. Several approaches to feature selection based on rough set theory are experimentally compared. Additionally, a new area in feature selection, feature grouping, is highlighted and a rough set-based feature grouping technique is detailed.
Richard Jensen, Qiang Shen 0001
IEEE Trans. Knowl. Data Eng.1
2002 Fuzzy-rough sets for descriptive dimensionality reduction
abstract
One of the main obstacles facing current fuzzy modelling techniques is that of dataset dimensionality. To enable these techniques to be effective, a redundancy-removing step is usually carried out beforehand. Rough set theory (RST) has been used as such a dataset pre-processor with much success, however it is reliant upon a crisp dataset; important information may be lost as a result of quantization. The paper proposes a dimensionality reduction technique that employs a hybrid variant of rough sets, fuzzy-rough sets, to avoid this information loss.
Richard Jensen, Qiang Shen 0001
FUZZ-IEEE1
2001 A Rough Set-Aided System for Sorting WWW Bookmarks
Richard Jensen, Qiang Shen 0001
Web Intelligence1
1993 Persistence Software: Bridging Object-Oriented Programming and Relational Databases
abstract
Building object-oriented applications which access relational data introduces a number of technical issues for developers who are making the transition to C++. We describe these issues and discuss how we have addressed them in Persistence, an application development tool that uses an automatic code generator to merge C++ applications with relational data. We use client-side caching to provide the application program with efficient access to the data.
Arthur M. Keller, Richard Jensen, Shailesh Agrawal
SIGMOD Conference2