EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Gu 0001
dblp:128/5017-1
· DBLP profile ↗
18ranked-venue papers in the field
13as first author
8since 2021 · last 2024
0000-0001-9116-4761ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 15 (12 first)Other / Interdisciplinary · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A soft prototype-based autonomous fuzzy inference system for network intrusion detectionabstractNowadays, cyber-attacks have become a common and persistent issue affecting various human activities in modern societies. Due to the continuously evolving landscape of cyber-attacks and the growing concerns around “black box” models, there has been a strong demand for novel explainable and interpretable intrusion detection systems with online learning abilities. In this paper, a novel soft prototype-based autonomous fuzzy inference system (SPAFIS) is proposed for network intrusion detection . SPAFIS learns from network traffic data streams online on a chunk-by-chunk basis and autonomously identifies a set of meaningful, human-interpretable soft prototypes to build an IF-THEN fuzzy rule base for classification. Thanks to the utilization of soft prototypes, SPAFIS can precisely capture the underlying data structure and local patterns, and perform internal reasoning and decision-making in a human-interpretable manner based on the ensemble properties and mutual distances of data. To maintain a healthy and compact knowledge base , a pruning scheme is further introduced to SPAFIS, allowing itself to periodically examine the learned solution and remove redundant soft prototypes from its knowledge base. Numerical examples on public network intrusion detection datasets demonstrated the efficacy of the proposed SPAFIS in both offline and online application scenarios, outperforming the state-of-the-art alternatives. Xiaowei Gu 0001, Gareth Howells 0001, Haiyue Yuan |
Inf. Sci. | 1 |
| 2023 | Self-adaptive fuzzy learning ensemble systems with dimensionality compression from data streamsabstractEnsemble learning is a widely used methodology to build powerful predictors from multiple individual weaker ones. However, the vast majority of ensemble learning models are designed for offline application scenarios, the use of evolving fuzzy systems in ensemble learning for online learning from data streams has not been sufficiently explored, yet. In this paper, a novel self-adaptive fuzzy learning ensemble system is introduced for data stream prediction. The proposed ensemble system employs the very sparse random projection technique to compress the consequent parts of the learned fuzzy rules by individual base models to a more compressed form, thereby reducing redundant information and improving computational efficiency. To improve the overall prediction performance, a dynamical base model pruning scheme is introduced to the proposed ensemble system together with a novel inferencing scheme, such that less accurate base models will be removed from the ensemble structure at each learning cycle automatically and only these more accurate ones will be involved in joint decision-making. Numerical examples based on a wide range of benchmark datasets demonstrate the stronger prediction performance of the proposed ensemble system over the state-of-the-art alternatives. Xiaowei Gu 0001 |
Inf. Sci. | 1 |
| 2023 | A multi-objective q-rung orthopair fuzzy programming approach to heterogeneous group decision making
Guolin Tang, Xiaowei Gu 0001, Francisco Chiclana, Peide Liu, Kedong Yin |
Inf. Sci. | 2 |
| 2022 | An explainable semi-supervised self-organizing fuzzy inference system for streaming data classification
Xiaowei Gu 0001 |
Inf. Sci. | 1 |
| 2022 | Self-organizing Divisive Hierarchical Voronoi Tessellation-based classifier
Xiaowei Gu 0001, Qiang Shen 0001 |
Inf. Sci. | 1 |
| 2022 | Interval type-2 fuzzy programming method for risky multicriteria decision-making with heterogeneous relationship
Guolin Tang, Jianpeng Long, Xiaowei Gu 0001, Francisco Chiclana, Peide Liu, Fubin Wang |
Inf. Sci. | 3 |
| 2021 | A multi-granularity locally optimal prototype-based approach for classification
Xiaowei Gu 0001, Miqing Li |
Inf. Sci. | 1 |
| 2021 | A self-adaptive fuzzy learning system for streaming data prediction
Xiaowei Gu 0001, Qiang Shen 0001 |
Inf. Sci. | 1 |
| 2020 | A self-adaptive synthetic over-sampling technique for imbalanced classificationabstractTraditionally, in supervised machine learning, (a significant) part of the available data (usually 50%-80%) is used for training and the rest—for validation. In many problems, however, the data are highly imbalanced in regard to different classes or does not have good coverage of the feasible data space which, in turn, creates problems in validation and usage phase. In this paper, we propose a technique for synthesizing feasible and likely data to help balance the classes as well as to boost the performance in terms of confusion matrix as well as overall. The idea, in a nutshell, is to synthesize data samples in close vicinity to the actual data samples specifically for the less represented (minority) classes. This has also implications to the so-called fairness of machine learning. In this paper, we propose a specific method for synthesizing data in a way to balance the classes and boost the performance, especially of the minority classes. It is generic and can be applied to different base algorithms, for example, support vector machines, k-nearest neighbour classifiers deep neural, rule-based classifiers, decision trees, and so forth. The results demonstrated that (a) a significantly more balanced (and fair) classification results can be achieved and (b) that the overall performance as well as the performance per class measured by confusion matrix can be boosted. In addition, this approach can be very valuable for the cases when the number of actual available labelled data is small which itself is one of the problems of the contemporary machine learning. Xiaowei Gu 0001, Plamen Angelov 0001, Eduardo A. Soares 0001 |
Int. J. Intell. Syst. | 1 |
| 2020 | A self-training hierarchical prototype-based approach for semi-supervised classification
Xiaowei Gu 0001 |
Inf. Sci. | 1 |
| 2019 | Local optimality of self-organising neuro-fuzzy inference systems
Xiaowei Gu 0001, Plamen Angelov 0001, Hai-Jun Rong |
Inf. Sci. | 1 |
| 2019 | A hierarchical prototype-based approach for classification
Xiaowei Gu 0001, Weiping Ding 0001 |
Inf. Sci. | 1 |
| 2018 | Empirical Fuzzy SetsabstractIn this paper, we introduce a new form of describing fuzzy sets (FSs) and a new form of fuzzy rule-based (FRB) systems, namely, empirical fuzzy sets (εFSs) and empirical fuzzy rule-based (εFRB) systems. Traditionally, the membership functions (MFs), which are the key mathematical representation of FSs, are designed subjectively or extracted from the data by clustering projections or defined subjectively. εFSs, on the contrary, are described by the empirically derived membership functions (εMFs). The new proposal made in this paper is based on the recently introduced Empirical Data Analytics (EDA) computational framework and is closely linked with the density of the data. This allows to keep and improve the link between the objective data and the subjective labels, linguistic terms, and classes definition. Furthermore, εFSs can deal with heterogeneous data combining categorical with continuous and/or discrete data in a natural way. εFRB systems can be extracted from data including data streams and can have dynamically evolving structure. However, they can also be used as a tool to represent expert knowledge. The main difference from the traditional FSs and FRB systems is that the expert does not need to define the MF per variable; instead, possibly multimodal, densities will be extracted automatically from the data and used as εMFs in a vector form for all numerical variables. This is done in a seamless way whereby the human involvement is only required to label the classes and linguistic terms. Moreover, even this intervention is optional. Thus, the proposed new approach to define and design the FSs and FRB systems puts the human “in the driving seat.” Instead of asking experts to define features and MFs correspondingly, to parameterize them, to define algorithm parameters, to choose types of MFs, or to label each individual item, it only requires (optionally) to select prototypes from data and (again, optionally) to label them. Numerical examples as well as a naïve empirical fuzzy (εF) classifier are presented with an illustrative purpose. Due to the very fundamental nature of the proposal, it can have a very wide area of applications resulting in a series of new algorithms such as εF classifiers, εF predictors, εF controllers, and so on. This is left for the future research. Plamen Angelov 0001, Xiaowei Gu 0001 |
Int. J. Intell. Syst. | 2 |
| 2018 | Deep rule-based classifier with human-level performance and characteristics
Plamen Angelov 0001, Xiaowei Gu 0001 |
Inf. Sci. | 2 |
| 2018 | Self-organising fuzzy logic classifier
Xiaowei Gu 0001, Plamen Angelov 0001 |
Inf. Sci. | 1 |
| 2018 | Self-Organised direction aware data partitioning algorithm
Xiaowei Gu 0001, Plamen Angelov 0001, Dmitry Kangin, José C. Príncipe |
Inf. Sci. | 1 |
| 2018 | A method for autonomous data partitioning
Xiaowei Gu 0001, Plamen Angelov 0001, José C. Príncipe |
Inf. Sci. | 1 |
| 2017 | Empirical Data AnalyticsabstractIn this paper, we propose an approach to data analysis, which is based entirely on the empirical observations of discrete data samples and the relative proximity of these points in the data space. At the core of the proposed new approach is the typicality—an empirically derived quantity that resembles probability. This nonparametric measure is a normalized form of the square centrality (centrality is a measure of closeness used in graph theory). It is also closely linked to the cumulative proximity and eccentricity (a measure of the tail of the distributions that is very useful for anomaly detection and analysis of extreme values). In this paper, we introduce and study two types of typicality, namely its local and global versions. The local typicality resembles the well-known probability density function (pdf), probability mass function, and fuzzy set membership but differs from all of them. The global typicality, on the other hand, resembles well-known histograms but also differs from them. A distinctive feature of the proposed new approach, empirical data analysis (EDA), is that it is not limited by restrictive impractical prior assumptions about the data generation model as the traditional probability theory and statistical learning approaches are. Moreover, it does not require an explicit and binary assumption of either randomness or determinism of the empirically observed data, their independence, or even their number (it can be as low as a couple of data samples). The typicality is considered as a fundamental quantity in the pattern analysis, which is derived directly from data and is stated in a discrete form in contrast to the traditional approach where a continuous pdf is assumed a priori and estimated from data afterward. The typicality introduced in this paper is free from the paradoxes of the pdf. Typicality is objectivist while the fuzzy sets and the belief-based branch of the probability theory are subjectivist. The local typicality is expressed in a closed analytical form and can be calculated recursively, thus, computationally very efficiently. The other nonparametric ensemble properties of the data introduced and studied in this paper, namely, the square centrality, cumulative proximity, and eccentricity, can also be updated recursively for various types of distance metrics. Finally, a new type of classifier called naïve typicality-based EDA class is introduced, which is based on the newly introduced global typicality. This is only one of the wide range of possible applications of EDA including but not limited for anomaly detection, clustering, classification, control, prediction, control, rare events analysis, etc., which will be the subject of further research. Plamen Angelov 0001, Xiaowei Gu 0001, Dmitry Kangin |
Int. J. Intell. Syst. | 2 |