EDBT 2026 Demo / reviewers in the wild / expert
Suyun Zhao
dblp:15/5578
· DBLP profile ↗
50ranked-venue papers
8as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 13 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSystems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Personalized Clustering via Targeted Representation LearningabstractClustering traditionally aims to reveal a natural grouping structure within unlabeled data. However, this structure may not always align with users' preferences. In this paper, we propose a personalized clustering method that explicitly performs targeted representation learning by interacting with users via modicum task information (e.g., must-link or cannot-link pairs) to guide the clustering direction. We query users with the most informative pairs, i.e., those pairs most hard to cluster and those most easy to miscluster, to facilitate the representation learning in terms of the clustering preference. Moreover, by exploiting attention mechanism, the targeted representation is learned and augmented. By leveraging the targeted representation and constrained contrastive loss as well, personalized clustering is obtained. Theoretically, we verify that the risk of personalized clustering is tightly bounded, guaranteeing that active queries to users do mitigate the clustering risk. Experimentally, extensive results show that our method performs well across different clustering tasks and datasets, even when only a limited number of queries are available. Xiwen Geng, Suyun Zhao, Borui Peng, Pan Du 0002, Hong Chen 0001, Cuiping Li 0001, Mengdie Wang |
AAAI | 2 |
| 2025 | Decoupled Self-knowledge Distillation Makes Differentially Private Deep Learning Stronger
Dengfeng Zhao, DaXuan Xue, Suyun Zhao, Hong Chen 0001 |
DASFAA (5) | 3 |
| 2025 | Unsupervised Learning for Class Distribution MismatchabstractClass distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an "other" category. However, they focus on semi-supervised scenarios and heavily rely on labeled data, limiting their applicability and performance.
To address this, we propose Unsupervised Learning for Class Distribution Mismatch (UCDM), which constructs positive-negative pairs from unlabeled data for classifier training. Our approach randomly samples images and uses a diffusion model to add or erase semantic classes, synthesizing diverse training pairs. Additionally, we introduce a confidence-based labeling mechanism that iteratively assigns pseudo-labels to valuable real-world data and incorporates them into the training process.
Extensive experiments on three datasets demonstrate UCDM’s superiority over previous semi-supervised methods. Specifically, with a 60\% mismatch proportion on Tiny-ImageNet dataset, our approach, without relying on labeled data, surpasses OpenMatch (with 40 labels per class) by 35.1%, 63.7%, and 72.5% in classifying known, unknown, and new classes. Pan Du 0002, Wangbo Zhao, Xinai Lu, Zhikai Li, Chaoyu Gong, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Kai Wang 0036, Yang You 0001 |
ICML | 7 |
| 2025 | Towards a Theoretical Understanding of Semi-Supervised Learning Under Class Distribution MismatchabstractSemi-supervised learning (SSL) confronts a formidable challenge under class distribution mismatch, wherein unlabeled data contain numerous categories absent in the labeled dataset. Traditional SSL methods undergo performance deterioration in such mismatch scenarios due to the invasion of those instances from unknown categories. Despite some technical efforts to enhance SSL by mitigating the invasion, the profound theoretical analysis of SSL under class distribution mismatch is still under study. Accordingly, in this work, we propose Bi-Objective Optimization Mechanism (BOOM) to theoretically analyze the excess risk between the empirical optimal solution and the population-level optimal solution. Specifically, BOOM reveals that the SSL error is the essential contributor behind excess risk, resulting from both the pseudo-labeling error and invasion error. Meanwhile, BOOM unveils that the optimization objectives of SSL under mismatch are binary: high-quality pseudo-labels and adaptive weights on the unlabeled instances, which contribute to alleviating the pseudo-labeling error and the invasion error, respectively. Moreover, BOOM explicitly discovers the fundamental factors crucial for optimizing the bi-objectives, guided by which an approach is then proposed as a strong baseline for SSL under mismatch. Extensive experiments on benchmark and real datasets confirm the effectiveness of our proposed algorithm. Pan Du 0002, Suyun Zhao, Puhui Tan, Zisen Sheng, Zeyu Gan, Hong Chen 0001, Cuiping Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Fuzzy rough unlearning model for feature selection
Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Jun-Hai Zhai, Qiangjun Zhou |
Int. J. Approx. Reason. | 2 |
| 2023 | Echo of Neighbors: Privacy Amplification for Personalized Private Federated Learning with Shuffle ModelabstractFederated Learning, as a popular paradigm for collaborative training, is vulnerable against privacy attacks. Different privacy levels regarding users' attitudes need to be satisfied locally, while a strict privacy guarantee for the global model is also required centrally. Personalized Local Differential Privacy (PLDP) is suitable for preserving users' varying local privacy, yet only provides a central privacy guarantee equivalent to the worst-case local privacy level. Thus, achieving strong central privacy as well as personalized local privacy with a utility-promising model is a challenging problem. In this work, a general framework (APES) is built up to strengthen model privacy under personalized local privacy by leveraging the privacy amplification effect of the shuffle model. To tighten the privacy bound, we quantify the heterogeneous contributions to the central privacy user by user. The contributions are characterized by the ability of generating "echos" from the perturbation of each user, which is carefully measured by proposed methods Neighbor Divergence and Clip-Laplace Mechanism. Furthermore, we propose a refined framework (S-APES) with the post-sparsification technique to reduce privacy loss in high-dimension scenarios. To the best of our knowledge, the impact of shuffling on personalized local privacy is considered for the first time. We provide a strong privacy amplification effect, and the bound is tighter than the baseline result based on existing methods for uniform local privacy. Experiments demonstrate that our frameworks ensure comparable or higher accuracy for the global model. Yixuan Liu 0002, Suyun Zhao, Li Xiong 0001, Hong Chen 0001 |
AAAI | 2 |
| 2023 | Superclass Learning with Representation EnhancementabstractIn many real scenarios, data are often divided into a handful of artificial super categories in terms of expert knowledge rather than the representations of images. Concretely, a superclass may contain massive and various raw categories, such as refuse sorting. Due to the lack of common semantic features, the existing classification techniques are intractable to recognize superclass without raw class labels, thus they suffer severe performance damage or require huge annotation costs. To narrow this gap, this paper proposes a superclass learning framework, called SuperClass Learning with Representation Enhancement(SCLRE), to recognize super categories by leveraging enhanced representation. Specifically, by exploiting the self-attention technique across the batch, SCLRE collapses the boundaries of those raw categories and enhances the representation of each superclass. On the enhanced representation space, a superclassaware decision boundary is then reconstructed. Theoretically, we prove that by leveraging attention techniques the generalization error of SCLRE can be bounded under superclass scenarios. Experimentally, extensive results demonstrate that SCLRE outperforms the baseline and other contrastive-based methods on CIFAR-100 datasets and four high-resolution datasets. Zeyu Gan, Suyun Zhao, Jinlong Kang, Liyuan Shang, Hong Chen 0001, Cuiping Li 0001 |
CVPR | 2 |
| 2023 | Semi-Supervised Learning via Weight-aware Distillation under Class Distribution MismatchabstractSemi-Supervised Learning (SSL) under class distribution mismatch aims to tackle a challenging problem wherein unlabeled data contain lots of unknown categories unseen in the labeled ones. In such mismatch scenarios, traditional SSL suffers severe performance damage due to the harmful invasion of the instances with unknown categories into the target classifier. In this study, by strict mathematical reasoning, we reveal that the SSL error under class distribution mismatch is composed of pseudo-labeling error and invasion error, both of which jointly bound the SSL population risk. To alleviate the SSL error, we propose a robust SSL framework called Weight-Aware Distillation (WAD) that, by weights, selectively transfers knowledge beneficial to the target task from unsupervised contrastive representation to the target classifier. Specifically, WAD captures adaptive weights and high-quality pseudo-labels to target instances by exploring point mutual information (PMI) in representation space to maximize the role of unlabeled data and filter unknown categories. Theoretically, we prove that WAD has a tight upper bound of population risk under class distribution mismatch. Experimentally, extensive results demonstrate that WAD outperforms five state-of-the-art SSL approaches and one standard baseline on two benchmark datasets, CIFAR10 and CIFAR100, and an artificial cross-dataset. The code is available at https://github.com/RUC-DWBI-ML/research/tree/main/WAD-master. Pan Du 0002, Suyun Zhao, Zisen Sheng, Cuiping Li 0001, Hong Chen 0001 |
ICCV | 2 |
| 2023 | Granularity self-information based uncertainty measure for feature selection and robust classification
Shuang An, Qijin Xiao, Changzhong Wang, Suyun Zhao |
Fuzzy Sets Syst. | 4 |
| 2023 | Hadamard Encoding Based Frequent Itemset Mining under Local Differential Privacy
Dan Zhao 0009, Suyun Zhao, Hong Chen 0001, Ruixuan Liu, Cuiping Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | Contrastive Active Learning Under Class Distribution MismatchabstractActive learning(AL) has been successful based on the premise that labeled and unlabeled data come from the same class distribution. However, its performance undergoes a severe deterioration under class distribution mismatch, wherein the unlabeled data contain numerous instances out of the class distribution of labeled data. In this article, we solve this practical yet rarely studied problem by minimizing the AL error, which is formally defined and decomposed as the valid query error and invalid query error. Specifically, the invalid query error is associated with the queries from unknown categories, and the valid query error is attributed to less informative queries from target categories. In light of this discovery, we propose a contrastive AL framework, named ConAL, to simultaneously learn the semantics and distinctiveness of the instances by contrastive techniques, thereby reducing the invalid query error and valid query error, respectively. Theoretically, we prove that the AL error of ConAL has a tight upper bound. Experimentally, ConAL achieves superior performance on two benchmark datasets, CIFAR10 and CIFAR100, and a cross-dataset with class distribution across multi-datasets. Furthermore, we validate that the ConAL technique performs admirably even on the realistic dataset. To the best of our knowledge, ConAL is the first AL work for class distribution mismatch. Pan Du 0002, Suyun Zhao, Shuwen Chai, Hong Chen 0001, Cuiping Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Relative Fuzzy Rough Approximations for Feature Selection and ClassificationabstractFuzzy rough set (FRS) theory is generally used to measure the uncertainty of data. However, this theory cannot work well when the class density of a data distribution differs greatly. In this work, a relative distance measure is first proposed to fit the mentioned data distribution. Based on the measure, a relative FRS model is introduced to remedy the mentioned imperfection of classical FRSs. Then, the positive region, negative region, and boundary region are defined to measure the uncertainty of data with the relative FRSs. Besides, a relative fuzzy dependency is defined to evaluate the importance of features to decision. With the proposed feature evaluation, we propose a feature selection algorithm and design a classifier based on the maximal positive region. The classification principle is that an unlabeled sample will be classified into the class corresponding to the maximal degree of the positive region. Experimental results show the relative fuzzy dependency is an effective and efficient measure for evaluating features, and the proposed feature selection algorithm presents better performance than some classical algorithms. Besides, it also shows the proposed classifier can achieve slightly better performance than the KNN classifier, which demonstrates that the maximal positive region-based classifier is effective and feasible. Shuang An, Enhui Zhao, Changzhong Wang, Ge Guo 0001, Suyun Zhao, Piyu Li |
IEEE Trans. Cybern. | 5 |
| 2023 | Inaccurate-Supervised Learning With Generative Adversarial NetsabstractInaccurate-supervised learning (ISL) is a weakly supervised learning framework for imprecise annotation, which is derived from some specific popular learning frameworks, mainly including partial label learning (PLL), partial multilabel learning (PML), and multiview PML (MVPML). While PLL, PML, and MVPML are each solved as independent models through different methods and no general framework can currently be applied to these frameworks, most existing methods for solving them were designed based on traditional machine-learning techniques, such as logistic regression, KNN, SVM, decision tree. Prior to this study, there was no single general framework that used adversarial networks to solve ISL problems. To narrow this gap, this study proposed an adversarial network structure to solve ISL problems, called ISL with generative adversarial nets (ISL-GANs). In ISL-GAN, fake samples, which are quite similar to real samples, gradually promote the Discriminator to disambiguate the noise labels of real samples. We also provide theoretical analyses for ISL-GAN in effectively handling ISL data. In this article, we propose a general framework to solve PLL, PML, and MVPML, while in the published conference version, we adopt the specific framework, which is a special case of the general one, to solve the PLL problem. Finally, the effectiveness is demonstrated through extensive experiments on various imprecise annotation learning tasks, including PLL, PML, and MVPML. Yabin Zhang 0005, Hairong Lian, Guang Yang 0040, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Fldp: Flexible Strategy For Local Differential PrivacyabstractLocal differential privacy (LDP), a technique applying unbiased statistical estimations instead of real data, is often adopted in data collection. In particular, this technique is used in frequency oracles (FO) because it can protect each user’s privacy and prevent leakage of sensitive information. However, the definition of LDP is so conservative that it requires all inputs to be indistinguishable after perturbation. Indeed, LDP protects each value; however, it is rarely used in practical scenarios owing to its cost in terms of accuracy. In this paper, we address the challenge of providing weakened but flexible protection where each value only needs to be indistinguishable from part of the domain after perturbation. First, we present this weakened but flexible LDP (FLDP) notion which splits the domain. We then prove the association with LDP. Second, we design a Flexible Hadamard Response (FHR) approach for the common FO issue while satisfying FLDP. The proposed approach balances communication cost, computational complexity, and estimation accuracy. Finally, experimental results using practical and synthetic datasets verify the effectiveness and efficiency of our approach. Dan Zhao 0009, Hong Chen 0001, Suyun Zhao, Ruixuan Liu, Cuiping Li 0001 |
ICASSP | 3 |
| 2022 | Collecting Triangle Counts with Edge Relationship Local Differential PrivacyabstractCounting subgraphs in decentralized settings has drawn increasing attention for graph analysis, wherein triangle count is one of the fundamental statistics. However, triangle counts may breach edge privacy, such as sensitive relations of individuals. Protecting edge privacy in triangle counts collection is a challenging problem due to the strong correlations among data from different clients. Decentralized Differential Privacy (DDP), as a possible option, protects edge privacy on correlated data to some extent. However, DDP provides a weak privacy guarantee by only hiding one edge in global. Unlike DDP, Local Differential Privacy (LDP) is a widely adopted standard for data collection which hides multiple data points in global at a time. But the LDP notion does not consider data correlations. With the understanding of these limitations, we introduce Edge Relationship Local Differential Privacy (Edge-RLDP), which provides a strong privacy guarantee as LDP and considers data correlations simultaneously. Based on Edge-RLDP, a baseline framework for triangle counts collection is proposed, as well as an improved two-phase framework, which strikes a better balance between privacy and data utility. Our improved framework fully utilizes the privacy budget by asking each client to only report the count of randomly sampled triangles after measuring the global data correlation. Theoretically, we rigorously prove that our framework satisfies ($\varepsilon, \delta$) -Edge-RLDP. Experimentally, we demonstrate our framework outperforms the state-of-art methods in terms of triangle count accuracy under a stricter privacy definition. Suyun Zhao, Yixuan Liu 0002, Dan Zhao 0009, Hong Chen 0001, Cuiping Li 0001 |
ICDE | 2 |
| 2022 | Exploring Binary Classification Hidden within Partial Label LearningabstractPartial label learning (PLL) is to learn a discriminative model under incomplete supervision, where each instance is annotated with a candidate label set. The basic principle of PLL is that the unknown correct label y of an instance x resides in its candidate label set s, i.e., P(y ∈ s | x) = 1. On which basis, current researches either directly model P(x | y) under different data generation assumptions or propose various surrogate multiclass losses, which all aim to encourage the model-based Pθ(y ∈ s | x)→1 implicitly. In this work, instead, we explicitly construct a binary classification task toward P(y ∈ s | x) based on the discriminative model, that is to predict whether the model-output label of x is one of its candidate labels. We formulate a novel risk estimator with estimation error bound for the proposed PLL binary classification risk. By applying logit adjustment based on disambiguation strategy, the practical approach directly maximizes Pθ(y ∈ s | x) while implicitly disambiguating the correct one from candidate labels simultaneously. Thorough experiments validate that the proposed approach achieves competitive performance against the state-of-the-art PLL methods. Hengheng Luo, Yabin Zhang 0005, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001 |
IJCAI | 3 |
| 2022 | Efficient protocols for heavy hitter identification with local differential privacy
Dan Zhao 0009, Suyun Zhao, Hong Chen 0001, Ruixuan Liu, Cuiping Li 0001, Wenjuan Liang |
Frontiers Comput. Sci. | 2 |
| 2021 | Contrastive Coding for Active Learning under Class Distribution MismatchabstractActive learning (AL) is successful based on the assumption that labeled and unlabeled data are obtained from the same class distribution. However, its performance deteriorates under class distribution mismatch, wherein the un-labeled data contain many samples out of the class distribution of labeled data. To effectively handle the problems under class distribution mismatch, we propose a contrastive coding based AL framework named CCAL. Unlike the existing AL methods that focus on selecting the most informative samples for annotating, CCAL extracts both semantic and distinctive features by contrastive learning and combines them in a query strategy to choose the most informative un-labeled samples with matched categories. Theoretically, we prove that the AL error of CCAL has a tight upper bound. Experimentally, we evaluate its performance on CIFAR10, CIFAR100, and an artificial cross-dataset that consists of five datasets; consequently, CCAL achieves state-of-the-art performance by a large margin with remarkably lower annotation cost. To the best of our knowledge, CCAL is the first work related to AL for class distribution mismatch. Pan Du 0002, Suyun Zhao, Shuwen Chai, Hong Chen 0001, Cuiping Li 0001 |
ICCV | 2 |
| 2021 | A relative uncertainty measure for fuzzy rough feature selection
Shuang An, Jiaying Liu 0012, Changzhong Wang, Suyun Zhao |
Int. J. Approx. Reason. | 4 |
| 2021 | Partial Label Learning via Conditional-Label-Aware Disambiguation
Suyun Zhao, Zhi-Gang Dai, Hong Chen 0001, Cuiping Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2021 | Ensemble selection with joint spectral clustering and structural sparsity
Zhenlei Wang, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Yufeng Shen |
Pattern Recognit. | 2 |
| 2021 | An Accelerator for Rule Induction in Fuzzy Rough TheoryabstractRule-based classifier, that extract a subset of induced rules to efficiently learn/mine while preserving the discernibility information, plays a crucial role in human-explainable artificial intelligence. However, in this era of big data, rule induction on the whole datasets is computationally intensive. So far, to the best of our knowledge, no known method focusing on accelerating rule induction has been reported. This is the first study to consider the acceleration technique to reduce the scale of computation in rule induction. We propose an accelerator for rule induction based on fuzzy rough theory; the accelerator can avoid redundant computation and accelerate the building of a rule classifier. First, a rule induction method based on consistence degree, called consistence-based value reduction (CVR), is proposed and used as basis to accelerate. Second, we introduce a compacted search space termed key set, which only contains the key instances required to update the induced rule, to conduct value reduction. The monotonicity of key set ensures the feasibility of our accelerator. Third, a rule-induction accelerator is designed based on key set, and it is theoretically guaranteed to display the same results as the unaccelerated version. Specifically, the rank preservation property of key set ensures consistency between the rule induction achieved by the accelerator and the unaccelerated method. Finally, extensive experiments demonstrate that the proposed accelerator can perform remarkably faster than the unaccelerated rule-based classifier methods, especially on datasets with numerous instances. Suyun Zhao, Zhi-Gang Dai, Xizhao Wang, Hengheng Luo, Hong Chen 0001, Cuiping Li 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2020 | Partial Label Learning via Generative Adversarial Nets
Yabin Zhang 0005, Guang Yang 0040, Suyun Zhao, Hairong Lian, Hong Chen 0001, Cuiping Li 0001 |
ECAI | 3 |
| 2020 | Discernibility matrix based incremental feature selection on fused decision tables
Lidi Zheng, Yeliang Xiu, Suyun Zhao, Xizhao Wang, Hong Chen 0001, Cuiping Li 0001 |
Int. J. Approx. Reason. | 5 |
| 2020 | Incremental feature selection based on fuzzy rough sets
Suyun Zhao, Xizhao Wang, Hong Chen 0001, Cuiping Li 0001, Eric C. C. Tsang |
Inf. Sci. | 2 |
| 2019 | Local Differential Privacy with K-anonymous for Frequency EstimationabstractData release, such as statistics of data distribution, in many data analysis and machine learning tasks is needed, which poses significant risks of user's privacy. Usually, to preserve privacy of every individual, frequency estimation based on LDP (Local Differential Privacy) is used to replace the real distribution of data. Unfortunately, when an individual sends values multiple times, privacy leakage, i.e., same value problems may occur, along with other performance problems such as memory usage problem. To narrow these gaps, SAnonLDP (Sample Anonymous Local Differential Privacy) is proposed in this paper. We build the SAnonLDP framework by integrating k-anonymous into LDP, which includes four blocks: random grouping; anonymous and Walsh-Fourier transforms; random response; singular value decomposition (SVD). Among them, the second block 'Anonymous and Walsh-Fourier transforms' significantly decreases the communication cost and the memory requirements. The left blocks make up for the loss of information to achieve an acceptable frequency estimation. More important, we verify that this estimation is unbiased by the strict mathematical reasoning. Finally, the numerical experiments demonstrate that SAnonLAP achieves better KL-divergence and estimation error compared to another known privacy model: RAPPOR. Dan Zhao 0009, Hong Chen 0001, Suyun Zhao, Cuiping Li 0001, Ruixuan Liu |
IEEE BigData | 3 |
| 2019 | Privacy Preserving Machine Learning with Limited Information Leakage
Wenyi Tang, Suyun Zhao, Boning Zhao, Yunzhi Xue, Hong Chen 0001 |
NSS | 3 |
| 2019 | An Accelerator of Feature Selection Applying a General Fuzzy Rough Model
Suyun Zhao, Hong Chen 0001, Cuiping Li 0001 |
PAKDD (2) | 2 |
| 2019 | PARA: A positive-region based attribute reduction accelerator
Suyun Zhao, Xizhao Wang, Hong Chen 0001, Cuiping Li 0001 |
Inf. Sci. | 2 |
| 2018 | Fuzzy Rough Based Feature Selection by Using Random Sampling
Zhenlei Wang, Suyun Zhao, Yangming Liu, Hong Chen 0001, Cuiping Li 0001, Sun Xiran |
PRICAI | 2 |
| 2015 | A Novel Approach to Building a Robust Fuzzy Rough ClassifierabstractCurrently, most robust classifiers with parameters focus on the determination of the optimal or suboptimal parameters. There are no research studies or even discussions about robust classifiers on all of the possible parameters. This paper considers the robust rough classifier and finds that the robust rough classifier satisfies a nested topological structure; then, the nested classifier, which reflects the classifier on all of the possible parameters, is proposed. First, some notions, such as the robust discernibility vector, the robust value reduct, and the robust covering vector, are proposed; these notions can reflect the classical corresponding notions on all of the possible parameters. It is more important that these notions share a common characteristic: the nested structure. The nested structure of these notions makes nested classifier theoretically possible. Furthermore, some novel algorithms are designed to compute the robust value reduct, the robust covering degree, and the robust classifier. These algorithms make the nested classifier technologically possible. Finally, numerical experiments demonstrate that the nested classifier is effective and efficient for classification and predication. Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Xiaoyong Du 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2014 | Randomly sampling based fuzzy rough reductionabstractNow most methods of the existing fuzzy rough based attribute reduction can't be applied to large scale data sets because of high time and space consuming. To overcome this problem, we proposed a randomly sampling based method to decrease the computing consumption. First, we propose some pseudo concepts of fuzzy rough sets, such as the pseudo lower approximation and pseudo attribute reduction, which is not exactly, but approximately equal to classical concepts, and then it drastically accelerate computation and significantly reduce space consumption. The numerical experiments show that the random sampling algorithm is significantly efficient both in time and space with little loss of information. Suyun Zhao, Hong Chen 0001 |
SMC | 2 |
| 2013 | Efficient Event Prewarning for Sensor Networks with Multi Microenvironments
Yinglong Li, Hong Chen 0001, Suyun Zhao, Shangfeng Mo |
Euro-Par | 3 |
| 2013 | The Nested Structure in Fuzzy Rough ClassifierabstractCurrently most robust fuzzy rough classifiers with parameters focus on the robustness and less-sensitiveness to noise. No work studies or even discusses about the topological structure of robust fuzzy rough classifiers. This paper finds that the robust rough classifier satisfies a nested topological structure, and then NESTED CLASSIFIER, which reflects the classifier on different parameters, is proposed. First some notions, such as robust discernibility vector, robust value reduct and robust covering vector, are proposed which share the common characteristic: the nested structure. The nested structure of these notions makes the nested classifier theoretically possible. Furthermore, some novel algorithms are designed to compute robust value reduct, robust covering degree and robust classifier. These algorithms make the nested classifier technologically possible. Finally numerical experiments demonstrate that the nested classifier is more efficient than the existing ones. Suyun Zhao, Hong Chen 0001, Cuiping Li 0001 |
SMC | 1 |
| 2013 | Distance-Based Feature Selection from Probabilistic Data
Bin Pei, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001 |
WAIM | 3 |
| 2013 | Big data challenge: a data management perspective
Jinchuan Chen, Yueguo Chen, Xiaoyong Du 0001, Cuiping Li 0001, Jiaheng Lu, Suyun Zhao, Xuan Zhou 0001 |
Frontiers Comput. Sci. | 6 |
| 2013 | FARP: Mining fuzzy association rules from a probabilistic quantitative database
Bin Pei, Suyun Zhao, Hong Chen 0001, Xuan Zhou 0001, Dingjie Chen |
Inf. Sci. | 2 |
| 2013 | Nested structure in parameterized rough reduction
Suyun Zhao, Xizhao Wang, Degang Chen 0002, Eric C. C. Tsang |
Inf. Sci. | 1 |
| 2013 | RFRR: Robust Fuzzy Rough ReductionabstractThis paper proposes a robust method of dimension reduction using fuzzy rough sets, in which the reduction results can reflect the reducts obtained on all of the possible parameters. Here, the reducts being obtained on all of the possible parameters mean that all of the reducts are obtained on different degrees of robustness to handle noise. This method is completely different from the existing methods of fuzzy rough reduction. The differences are shown in three aspects: the concept, the tool, and the algorithm. First, the key concept of attribute reduction is redefined in a new way. That is, the robust fuzzy rough reduct, which is shortened to a robust reduct, is proposed to reflect the classical reducts obtained on all of the possible parameters. The new “robust reduct” is not a crisp subset of condition attributes; rather, it is a fuzzy subset, whose most interesting property is that any cut set of the robust reduct is a classical reduct on a certain parameter. Second, the tool used to measure the discernibility power is different from the existing discernibility measures. In this paper, the robustness of each attribute to handle misclassification and perturbation is considered. By considering both the robustness and the discernibility, a robust fuzzy discernibility matrix is designed. Finally, the algorithms used to find the robust reducts are designed based upon the robust fuzzy discernibility matrix, which is completely different from the existing algorithms used to find the classical reducts. Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Mengyao Zhai, Xiaoyong Du 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2012 | Mining Fuzzy Association Rules from Heterogeneous Probabilistic DatasetsabstractAssociation rule mining (ARM), as a useful method to discover relations between attributes of objects, has been widely studied. The previous methods focused on ARM either from a certain dataset with different type attributes, or from a probabilistic dataset with only Boolean attributes. However, little work on ARM from a probabilistic dataset with coexistence of different type attributes has been mentioned. Such dataset is named Heterogeneous Probabilistic Dataset (HPD), which is prevalent in the real-world applications. This paper develops a generic framework to discover association rules from a HPD. Considering the different type data in the dataset, we first convert a HPD to a probabilistic dataset with fuzzy sets by fuzzification. A novel Shannon-like Entropy is then introduced to measure the information of an item with coexistence of fuzzy uncertainty hidden in different type data and random uncertainty in the transformed dataset. Based on this Shannon-like Entropy, Support and Confidence degrees for such multi-uncertain dataset are defined. Finally, we design an Apriori-like algorithm to mine association rules from a HPD using the above measures. Experimental results show that the proposed algorithm for HPD is feasible and effective. Bin Pei, Suyun Zhao, Hong Chen 0001 |
ICTAI | 3 |
| 2012 | A Novel Algorithm for Finding Reducts With Fuzzy Rough SetsabstractAttribute reduction is one of the most meaningful research topics in the existing fuzzy rough sets, and the approach of discernibility matrix is the mathematical foundation of computing reducts. When computing reducts with discernibility matrix, we find that only the minimal elements in a discernibility matrix are sufficient and necessary. This fact motivates our idea in this paper to develop a novel algorithm to find reducts that are based on the minimal elements in the discernibility matrix. Relative discernibility relations of conditional attributes are defined and minimal elements in the fuzzy discernibility matrix are characterized by the relative discernibility relations. Then, the algorithms to compute minimal elements and reducts are developed in the framework of fuzzy rough sets. Experimental comparison shows that the proposed algorithms are effective. Degang Chen 0002, Lei Zhang 0006, Suyun Zhao, Qinghua Hu, Pengfei Zhu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2012 | Sample Pair Selection for Attribute Reduction with Rough SetabstractAttribute reduction is the strongest and most characteristic result in rough set theory to distinguish itself to other theories. In the framework of rough set, an approach of discernibility matrix and function is the theoretical foundation of finding reducts. In this paper, sample pair selection with rough set is proposed in order to compress the discernibility function of a decision table so that only minimal elements in the discernibility matrix are employed to find reducts. First relative discernibility relation of condition attribute is defined, indispensable and dispensable condition attributes are characterized by their relative discernibility relations and key sample pair set is defined for every condition attribute. With the key sample pair sets, all the sample pair selections can be found. Algorithms of computing one sample pair selection and finding reducts are also developed; comparisons with other methods of finding reducts are performed with several experiments which imply sample pair selection is effective as preprocessing step to find reducts. Degang Chen 0002, Suyun Zhao, Lei Zhang 0006, Yongping Yang, Xiao Zhang 0012 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Fuzzy rough set based attribute reduction for information systems with fuzzy decisions
Qiang He 0003, Congxin Wu, Degang Chen 0002, Suyun Zhao |
Knowl. Based Syst. | 4 |
| 2010 | Local reduction of decision system with fuzzy rough sets
Degang Chen 0002, Suyun Zhao |
Fuzzy Sets Syst. | 2 |
| 2010 | Building a Rule-Based Classifier—A Fuzzy-Rough Set ApproachabstractThe fuzzy-rough set (FRS) methodology, as a useful tool to handle discernibility and fuzziness, has been widely studied. Some researchers studied on the rough approximation of fuzzy sets, while some others focused on studying one application of FRS: attribute reduction (i.e., feature selection). However, constructing classifier by using FRS, as another application of FRS, has been less studied. In this paper, we build a rule-based classifier by using one generalized FRS model after proposing a new concept named as ¿consistence degree¿ which is used as the critical value to keep the discernibility information invariant in the processing of rule induction. First, we generalized the existing FRS to a robust model with respect to misclassification and perturbation by incorporating one controlled threshold into knowledge representation of FRS. Second, we propose a concept named as ¿consistence degree¿ and by the strict mathematical reasoning, we show that this concept is reasonable as a critical value to reduce redundant attribute values in database. By employing this concept, we then design a discernibility vector to develop the algorithms of rule induction. The induced rule set can function as a classifier. Finally, the experimental results show that the proposed rule-based classifier is feasible and effective on noisy data. Suyun Zhao, Eric C. C. Tsang, Degang Chen 0002, Xizhao Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | The Model of Fuzzy Variable Precision Rough SetsabstractThe fuzzy rough set (FRS) model has been introduced to handle databases with real values. However, FRS was sensitive to misclassification and perturbation (here misclassification means error or missing values in classification, and perturbation means small changes of numerical data). The variable precision rough sets (VPRSs) model was introduced to handle databases with misclassification. However, it could not effectively handle the real-valued datasets. Now, it is valuable from theoretical and practical aspects to combine FRS and VPRS so that a powerful tool, which not only can handle numerical data but also is less sensitive to misclassification and perturbation, can be developed. In this paper, we set up a model named fuzzy VPRSs (FVPRSs) by combining FRS and VPRS with the goal of making FRS a special case. First, we study the knowledge representation ways of FRS and VPRS, and then, propose the set approximation operators of FVPRS. Second, we employ the discernibility matrix approach to investigate the structure of attribute reductions in FVPRS and develop an algorithm to find all reductions. Third, in order to overcome the NP-complete problem of finding all reductions, we develop some fast heuristic algorithms to obtain one near-optimal attribute reduction. Finally, we compare FVPRS with RS, FRS, and several flexible RS-based approaches with respect to misclassification and perturbation. The experimental comparisons show the feasibility and effectiveness of FVPRS. Suyun Zhao, Eric C. C. Tsang, Degang Chen 0002 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2008 | On fuzzy approximation operators in attribute reduction with fuzzy rough sets
Suyun Zhao, Eric C. C. Tsang |
Inf. Sci. | 1 |
| 2007 | An approach of attributes reduction based on fuzzy TL rough setsabstractIn fuzzy rough sets a fuzzy T-similarity relation is employed to describe the similarity degree between two objects and to construct lower and upper approximations for arbitrary fuzzy sets. Different triangular norm T identifies different point of view of similarity. Thus a reasonable selection of triangular norm is meaningful to practical applications of fuzzy rough sets. In this paper we discuss the selection of triangular norm and emphasize the well-known Lukasiewicz's triangular norm TLas a reasonable selection. And then we focus on attributes reduction with TL-fuzzy rough sets. First we define attributes reduction with TL-fuzzy rough sets and characterize it by dependency function. Second, the structure of proposed attributes reduction is investigated in detail by the approach of discernibility matrix. Third, an algorithm of computing attributes reduction is designed and finally a comparison with other methods of attributes reduction with fuzzy rough sets is performed using several experiments. The experimental results show that the method of attributes reduction with TL-fuzzy rough sets is feasible and valid. Degang Chen 0002, Eric C. C. Tsang, Suyun Zhao |
SMC | 3 |
| 2007 | Learning fuzzy rules from fuzzy samples based on rough set technique
Xizhao Wang, Eric C. C. Tsang, Suyun Zhao, Degang Chen 0002, Daniel S. Yeung |
Inf. Sci. | 3 |
| 2006 | A Discussion of Attribute Reduction in Fuzzy Rough Sets Using Support Vector MachineabstractThis paper mainly focuses on the attribute reduction in fuzzy rough sets. An algorithm using discernibility matrix to compute all the attribute reductions is developed. After reducing the attributes, we introduce Support Vector Machine (SVM) as a classification technique to test the knowledge representation ability of attribute reduction. The numerical results show that the attribute reduction with fuzzy rough sets contains the same information as the original one. Eric C. C. Tsang, Degang Chen 0002, Suyun Zhao, Qiang He 0003 |
SMC | 3 |