Qifen Yang

dblp:324/8983 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-1639-8020ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Chrono: Efficient Serverless Analytics With Adaptive Fine-Grained Partitioning and Shadow Execution
Zhaorui Wu, Yuhui Deng 0001, Jiande Huang, Qifen Yang, Peng Zhou 0032, Geyong Min
IEEE Trans. Cloud Comput.4
2025 Intervention-Driven Correlation Reduction: A Data Generation Approach for Achieving Counterfactually Fair Predictors
abstract
Achieving counterfactual fairness is a critical objective in advancing fairness research within machine learning. Studies have shown that machine learning models often inherit biases from their training data, leading to unfair decision-making. Fair data generation methods aim to mitigate these biases, ensuring that predictors trained on such data uphold fairness. However, in the context of counterfactual fairness, existing methods for generating fair data are often limited in their applicability and lead to significant performance losses in downstream predictors. To address these issues, this paper proposes a new algorithm for generating counterfactually fair data, allowing predictors trained on this generated data to adhere to counterfactual fairness. We propose a new metric, Intervention-Driven Correlation (IDC), to evaluate the counterfactual fairness of generative models. IDC assesses fairness by applying random interventions to samples and measuring the statistical correlation between the degree of intervention and the outcome of interest. This metric is applicable to both discrete and continuous sensitive attributes and labels. Furthermore, our studies reveal a critical insight: counterfactually fair data does not always guarantee counterfactually fair predictors when deployed in real-world scenarios. We identify the root causes of this issue and propose a robust solution. To bridge this gap, we propose the IDC-Reduction method, which ensures the fairness of downstream predictors by generating counterfactually fair data. Experimentally, our method outperforms existing approaches and achieves counterfactual fairness regardless of the type of downstream predictors.
Dehua Zhou, Bowei Wu, Ke Wang 0068, Qifen Yang, Yuhui Deng 0001, Siu-Ming Yiu
ICDE4
2025 Causal Pathway-Integrated Generative Adversarial Networks for Counterfactually Fair Data Generation
Haoming Mo, Yuhui Deng 0001, Qifen Yang, Jiande Huang, Yi Zhou 0009
ICIC (10)3
2025 FIFA: A Forest-Based Sliding Window Aggregation Scheme for Out-of-Order Data Streams
abstract
Sliding window aggregation is a core operation in data stream analysis that extracts summaries from the most recent data stream. An evict or insert of the window can be handled in$O(1)$for in-order data streams. However, real-world data streams are typically disordered due to network delays. To process out-of-order data streams, existing methods primarily use a tree to maintain the sliding windows. Since the complexity of the tree is related to the window size, the performance of these methods will drop sharply or become unavailable when facing a big window. To overcome the limitation of existing methods, this paper presents Finger B-Trees Forest Aggregation (FIFA). This novel forest-based sliding window aggregation scheme optimally handles out-of-order data streams. At its heart, FIFA uses aggregation forest to extend Finger B-Trees. Specifically, FIFA evenly divides a window into several chunks and constructs a separate tree to maintain each chunk. When an out-of-order item arrives, FIFA first locates the corresponding chunk of the item and then uses Finger B-Trees to insert it into the window efficiently. Chunking reduces the complexity of the tree and the coupling of aggregation results by isolating items within windows. Thus, an insert or evict takes amortized$O(c)$in the worst case, where$c$is the size of each chunk. Finally, extensive experiments based on real-world data demonstrate that FIFA achieves an average 2-fold throughput improvement on out-of-order data streams compared with the state-of-the-art (SOTA) aggregation schemes.
Jiande Huang, Yuhui Deng 0001, Jianjun Li 0012, Lijuan Lu, Qifen Yang, Geyong Min
IEEE Trans. Big Data5
2025 MAFRO: Optimal-Granularity Fuzzy Decision Rule-Based Classification Architecture for Attribute Unlearning
abstract
Recently, many laws and regulations have granted users the right to be forgotten, i.e., the right to require data controllers to delete user data. Various methods for machine unlearning have been proposed to remove individual data points. However, they do not scale to the scenarios where larger groups of features are to be removed. To address this challenge, we propose MAFRO, an optimal-granularity fuzzy decision rule–based classifier that accelerates unlearning via influence functions. Building on granular computing (GrC), MAFRO first selects a minimal reduct of attributes, then constructs fuzzy granules with a Gaussian membership function to extract concise decision rules and realizes unlearning through the influence function. Specifically, instead of training with the full set of attributes, we use the reduct, a minimal subset of attributes that can classify the data with the same accuracy as the full set of attributes. Next, we extract fuzzy rules based on the reduct. Finally, fusing the generated rules establishes the linear model with strongly convex loss functions. In this way, MAFRO can quantify the divergence caused by attribute deleting and update the model without retraining it, thereby adapting the influence of data removal on the model and accelerating the unlearning process. We conduct extensive experiments to evaluate MAFRO on ten typical datasets in terms of performance and unlearning speed. We compare MAFRO with the state-of-theart algorithms. Experimental results demonstrate that MAFRO enhances accuracy by an average of 6.96%, and achieves up to 236× speedup for attribute unlearning tasks.
Jiande Huang, Yuhui Deng 0001, Yi Zhou 0009, Qifen Yang, Geyong Min
IEEE Trans. Fuzzy Syst.4
2025 DBCGM: A Granular Model for Big Data Classification Based on Data Bisection and Cascade Weighted Clustering
Jiande Huang, Yuhui Deng 0001, Yi Zhou 0009, Shujie Pang, Qifen Yang, Geyong Min
IEEE Trans. Knowl. Data Eng.5
2025 Gecko: Efficient Sliding Window Aggregation With Granular-Based Bulk Eviction Over Big Data Streams
abstract
Sliding window aggregation, which extracts summaries from data streams, is a core operation in streaming analysis. Though existing sliding window algorithms that perform single eviction and insertion operations can achieve a worst-case time complexity of$O(1)$for in-order streams, real-world data streams often involve out-of-order data and exhibit burst data characteristics, which pose performance challenges to these sliding window algorithms. To address this challenging issue, we proposeGecko- a novel sliding window aggregation algorithm that supports bulk eviction. Gecko leverages a granular-based eviction strategy for various bulk sizes, enabling efficient bulk eviction while maintaining the performance close to that of in-order stream algorithms for single evictions. For large data bulks, Gecko performs coarse-grained eviction at the chunk level, followed by fine-grained eviction using leftward binary tree aggregation (LTA) as a complementary method. Moreover, Gecko partitions data based on chunks to prevent the impacts of out-of-order data on other chunks, thereby enabling efficient handling of out-of-order data streams. We conduct extensive experiments to evaluate the performance of Gecko. Experimental results demonstrate that Gecko exhibits superior performance over other solutions, which is consistent with theoretical expectations. In real-world data scenarios, Gecko improves the average throughput of the state-of-the-art algorithm b_FiBA by 1.7 times, with a maximum improvement of up to 3.5 times. Gecko also demonstrates the best latency performance among all compared schemes.
Jianjun Li 0012, Yuhui Deng 0001, Jiande Huang, Yi Zhou 0009, Qifen Yang, Geyong Min
IEEE Trans. Knowl. Data Eng.5
2023 HCDC: A novel hierarchical clustering algorithm based on density-distance cores for data sets with varying density
Qifen Yang, Wanyi Gao, Ziyang Li 0007, Shuhua Zhu, Yuhui Deng 0001
Inf. Syst.1
2023 A Robust Learning Membership Scaling Fuzzy C-Means Algorithm Based on New Belief Peak
abstract
Fuzzy C-means clustering (FCM) has been a commonly used algorithm in fuzzy clustering for decades. However, it still faces two problems: how to determine the initial cluster center and how to determine the number of clusters. The recently proposed robust learning fuzzy C-means (RL-FCM) can automatically obtain the optimal number of clusters. However, it assumes that the initial cluster center is the entire dataset, which incurs a significant time cost and involves parameters that are also difficult to determine. Additionally, RL-FCM is unable to handle imbalanced datasets and datasets with a large span of sample attributes. Therefore, we propose a robust learning membership scaling fuzzy C-means algorithm based on new belief peaks (RL-MFCM). Within the framework of the confidence function, the neighbors of the sample points provide evidence for the sample points being cluster centers. Consequently, according to Jiang's combination rule, we consider the new belief peak as the initial cluster center. To avoid excessive interference of the mixing ratio of the cluster to the calculation of membership degree, we employ triangle inequality to improve the influence of the samples in the cluster in the clustering process. We analyze the time complexity of the proposed algorithm and conduct comparative experiments with existing fuzzy clustering algorithms on artificial and real datasets in the article. Experiments demonstrate that our proposed algorithm accurately estimates the number of clusters and exhibits superior clustering performance without needing initialization.
Qifen Yang, Wanyi Gao, Zhenye Yang, Shuhua Zhu, Yuhui Deng 0001
IEEE Trans. Fuzzy Syst.1
2022 An improvement of spectral clustering algorithm based on fast diffusion search for natural neighbor and affinity propagation
Qifen Yang, Ziyang Li 0007, Wanyi Gao, Shuhua Zhu, Yuhui Deng 0001
J. Supercomput.1
2022 Correction to: An improvement of spectral clustering algorithm based on fast diffusion search for natural neighbor and affinity propagation
Qifen Yang, Ziyang Li 0007, Wanyi Gao, Shuhua Zhu, Yuhui Deng 0001
J. Supercomput.1