VLDB 2026 Research / reviewers in the wild / expert
Yong-Feng Ge
dblp:164/9068
· DBLP profile ↗
24ranked-venue papers
19as first author
18since 2021 · last 2026
0000-0002-5955-6295ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 9 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Evolutionary Differential Privacy in Cross-Platform Spatial CrowdsourcingabstractThe development of mobile web services has brought significant attention to spatial crowdsourcing. The uneven distribution of tasks and workers has led to recent research on Cross-Platform Spatial Crowdsourcing (CPSC), aiming for a multi-win situation for platforms, workers, and task requesters. Previous studies on CPSC problems focused on task assignment and worker selection performance, overlooking the importance of privacy preservation. This article addresses the existing challenges of privacy preservation and service quality by formulating a Privacy-Preserving Cross-Platform Spatial Crowdsourcing (PP-CPSC) problem and proves it to be NP-hard. We propose an Evolutionary Differential Privacy (Evo-DP) approach to optimize PP-CPSC. Evo-DP’s evolutionary framework enables efficient and flexible optimization of privacy budget allocation. Within Evo-DP, each solution to the privacy budget allocation is represented as an individual in the population. To approximate the optimal solution, three evolutionary operations—mutation, crossover, and scaling—are employed for population updates, along with a selection process. A hybrid population model is introduced to balance exploration and exploitation abilities. Experimental results demonstrate Evo-DP’s superiority over previous strategies in terms of solution quality, convergence speed, and scalability. Yong-Feng Ge, Hua Wang 0002, Elisa Bertino, Jinli Cao, Yanchun Zhang, Zhonglong Zheng |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2025 | Analysis and Multi-objective Protection of Public Medical Datasets from Privacy and Utility PerspectivesabstractAbstract In this era of big data, seamless distribution of healthcare information is crucial for improving patient care and advancing medical research, necessitating meticulous attention to preserving health data privacy. However, overly stringent protection measures can impede the efficient utilization of invaluable resources for medical research and personalized healthcare, posing a central challenge in balancing privacy protection with effective data utilization. This study aims to explore various methods used to protect the privacy of patients’ health records, and evaluates their advantages and limitations. Additionally, it conducts an in-depth analysis of a public medical dataset concerning privacy protection, assessing the effectiveness of k-anonymity and l-diversity privacy criteria and examining the influence of quasi-identifier (QID) attributes on privacy preservation. The study showcases techniques to achieve privacy standards, including generalization and suppression. Furthermore, it introduces a novel approach that utilizes the genetic algorithm (GA) and a non-dominated sorting technique to maximize both privacy and utility in health data through multi-objective optimization. After examining the results, this paper offers a guide for data owners on selecting attributes for medical data publication and choosing suitable privacy preservation strategies. Through the exploration of the GA and the non-dominated sorting approach, this paper suggests that the proposed GA can offer promising non-dominated solutions to the issue of health data privacy in the era of data-driven healthcare. A combination of these algorithms can enhance privacy protection and provide healthcare professionals and researchers with essential knowledge, ultimately benefiting patient care and ensuring a more secure database system. Samsad Jahan, Yong-Feng Ge, Md. Enamul Kabir, Kate N. Wang 0001 |
Data Sci. Eng. | 2 |
| 2025 | Multiobjective Privacy-Preserving Task Assignment in Spatial CrowdsourcingabstractLocation information is crucial for efficient task assignment in spatial crowdsourcing, but sharing such information raises privacy concerns. Differential privacy (DP) offers a solution by protecting location privacy while preserving data usefulness. Existing DP-based spatial crowdsourcing frameworks have two main limitations: 1) they fail to provide personalized privacy preservation for workers and 2) they prioritize incentive mechanisms (such as utility maximization and cost minimization) while overlooking quality control. To address these limitations, we formulate the multiobjective privacy-preserving task assignment (MP-TA) problem. This problem aims to maximize both incentives and quality while meeting service rate requirements and ensuring personalized privacy protection for workers. Accordingly, we present a three-phase framework comprising worker proposal, candidate worker selection, and task assignment optimization. To generate high-quality eligible solutions for both objectives, we introduce a distributed cooperative co-evolutionary multiobjective memetic algorithm (DCC-MMA) based on sequential subproblem division and knee-driven migration operation. Matching-based crossover, matching-based mutation, and fix operations are designed to enhance search efficiency. Experimental results demonstrate DCC-MMA's superiority in solution quality, convergence speed, and scalability compared to state-of-the-art algorithms. Yong-Feng Ge, Hua Wang 0002, Elisa Bertino, Jinli Cao, Yanchun Zhang |
IEEE Trans. Cybern. | 1 |
| 2025 | Distributed Bandit-Based Cooperative Coevolution for Large-Scale Multi-Objective Data Publishing
Yong-Feng Ge, Hua Wang 0002, Elisa Bertino, Jinli Cao, Yanchun Zhang, Zhonglong Zheng |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | CADIF-OSN: Detecting Cloned Accounts with Missing Profile Attributes on Online Social NetworksabstractThe growth of online social networks (OSNs) has become increasingly significant. Potential cloned accounts on these platforms raise serious concerns due to the risks they pose to user privacy and security. Previous works in the detection of cloned accounts on OSNs do not yield satisfactory results and lack consideration of the impact of missing attributes on the detection process. We propose cloned account detection with imputation framework for online social networks (CADIF-OSN) to accurately find potential cloned accounts on OSNs. This framework enables the accurate identification of potential cloned accounts on OSNs by leveraging their public profile information, even in cases where some of the information may not be accessible. The framework comprises four key components: 1) Fuzzy string matching with Levenshtein Distance that quickly generates suspicious account pairs by matching all the accounts' usernames and screennames; 2) An embedded method Doc2Vec that transforms all existing profile information of accounts into estimable vectors; 3) A HyperImpute model that imputes the missing information; and 4) A deep-forest model that is trained to detect cloned accounts. We evaluated our framework using a Twitter dataset consisting of 3,826 pairs of cloned accounts and 70,000 normal accounts. The evaluation results demonstrate that our framework significantly surpasses existing approaches in terms of Precision and F1-score. Dewei Ning, Yong-Feng Ge, Hua Wang 0002, Changjun Zhou |
CIKM | 2 |
| 2024 | Federated Genetic Algorithm: Two-Layer Privacy-Preserving Trajectory Data PublishingabstractNowadays, trajectory data is widely available and used in various real-world applications such as urban planning, navigation services, and location-based services. However, publishing trajectory data can potentially leak sensitive information about identity, personal profiles, and social relationships, and requires privacy protection. This paper focuses on optimizing Privacy-Preserving Trajectory Data Publishing (PP-TDP) problems, addressing the limitations of existing techniques in the trade-off between privacy protection and information preservation. We propose the Federated Genetic Algorithm (FGA) in this paper, aiming to achieve better local privacy protection and global information preservation. FGA consists of multiple local optimizers and a single global optimizer. The parallel local optimizer enables the local data center to retain the original trajectory data and share only the locally anonymized outcomes. The global optimizer collects the local anonymized outcomes and further optimizes the preservation of information while achieving comprehensive privacy protection. To optimize the discrete-domain PP-TDP problems more efficiently, this paper proposes a grouping-based strategy, an intersection-based crossover operation, and a complement-based mutation operation. Experimental results demonstrate that FGA outperforms its competitors in terms of solution accuracy and search efficiency. Yong-Feng Ge, Hua Wang 0002, Jinli Cao, Yanchun Zhang, Georgios Kambourakis |
GECCO | 1 |
| 2024 | Dynamic-Parameter Genetic Algorithm for Multi-objective Privacy-Preserving Trajectory Data Publishing
Samsad Jahan, Yong-Feng Ge, Hua Wang 0002, Md. Enamul Kabir |
WISE (5) | 2 |
| 2024 | Evolutionary Dynamic Database Partitioning Optimization for Privacy and UtilityabstractDistributed database system (DDBS) technology has shown its advantages with respect to query processing efficiency, scalability, and reliability. Moreover, by partitioning attributes of sensitive associations into different fragments, DDBSs can be used to protect data privacy. However, it is complex to design a DDBS when one has to optimize privacy and utility in a time-varying environment. This paper proposes a distributed prediction-randomness framework for the evolutionary dynamic multiobjective partitioning optimization of databases. In the proposed framework, two sub-populations contain individuals representing database partitioning solutions. One sub-population utilizes a Markov chain-based predictor to predict discrete-domain solutions for database partitioning when the environment changes, and the other sub-population utilizes the random initialization operator to maintain population diversity. In addition, a knee-driven migration operator is utilized to exchange information between two sub-populations. Experimental results show that the proposed algorithm outperforms the competing solutions with respect to accuracy, convergence speed, and scalability. Yong-Feng Ge, Hua Wang 0002, Elisa Bertino, Zhi-hui Zhan, Jinli Cao, Yanchun Zhang, Jun Zhang 0003 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Distributed Cooperative Coevolution of Data Publishing Privacy and TransparencyabstractData transparency is beneficial to data participants’ awareness, users’ fairness, and research work’s reproducibility. However, when addressing transparency requirements, we cannot ignore data privacy. This article defines the multi-objective data publishing (MODP) problem, optimizing data privacy and transparency at the same time. Accordingly, we propose a distributed cooperative coevolutionary genetic algorithm (DCCGA) to optimize the MODP problem. In the population of DCCGA, each individual represents an anonymization solution to MODP. Three modules in DCCGA, i.e., grouping module, cooperative coevolutionary module, and evolving module, are proposed for distributed sub-population update and evaluation, improving DCCGA’s optimization performance and parallel efficiency. Moreover, a matrix-based crossover operator and a matrix-based mutation operator are designed to exchange and adjust anonymization information in the individuals efficiently. Experimental results demonstrate that the proposed DCCGA outperforms the competitors with respect to solution accuracy, convergence speed, and scalability. Besides, we verify the effectiveness of all the proposed components in DCCGA. Yong-Feng Ge, Elisa Bertino, Hua Wang 0002, Jinli Cao, Yanchun Zhang |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | Privacy-preserving data publishing: an information-driven distributed genetic algorithmabstractAbstract The privacy-preserving data publishing (PPDP) problem has gained substantial attention from research communities, industries, and governments due to the increasing requirements for data publishing and concerns about data privacy. However, achieving a balance between preserving privacy and maintaining data quality remains a challenging task in PPDP. This paper presents an information-driven distributed genetic algorithm (ID-DGA) that aims to achieve optimal anonymization through attribute generalization and record suppression. The proposed algorithm incorporates various components, including an information-driven crossover operator, an information-driven mutation operator, an information-driven improvement operator, and a two-dimensional selection operator. Furthermore, a distributed population model is utilized to improve population diversity while reducing the running time. Experimental results confirm the superiority of ID-DGA in terms of solution accuracy, convergence speed, and the effectiveness of all the proposed components. Yong-Feng Ge, Hua Wang 0002, Jinli Cao, Yanchun Zhang, Xiaohong Jiang 0001 |
World Wide Web (WWW) | 1 |
| 2024 | Hierarchical adaptive evolution framework for privacy-preserving data publishingabstractAbstract The growing need for data publication and the escalating concerns regarding data privacy have led to a surge in interest in Privacy-Preserving Data Publishing (PPDP) across research, industry, and government sectors. Despite its significance, PPDP remains a challenging NP-hard problem, particularly when dealing with complex datasets, often rendering traditional traversal search methods inefficient. Evolutionary Algorithms (EAs) have emerged as a promising approach in response to this challenge, but their effectiveness, efficiency, and robustness in PPDP applications still need to be improved. This paper presents a novel Hierarchical Adaptive Evolution Framework (HAEF) that aims to optimizet-closeness anonymization through attribute generalization and record suppression using Genetic Algorithm (GA) and Differential Evolution (DE). To balance GA and DE, the first hierarchy of HAEF employs a GA-prioritized adaptive strategy enhancing exploration search. This combination aims to strike a balance between exploration and exploitation. The second hierarchy employs a random-prioritized adaptive strategy to select distinct mutation strategies, thus leveraging the advantages of various mutation strategies. Performance bencmark tests demonstrate the effectiveness and efficiency of the proposed technique. In 16 test instances, HAEF significantly outperforms traditional depth-first traversal search and exceeds the performance of previous state-of-the-art EAs on most datasets. In terms of overall performance, under the three privacy constraints tested, HAEF outperforms the conventional DFS search by an average of 47.78%, the state-of-the-art GA-based ID-DGA method by an average of 37.38%, and the hybrid GA-DE method by an average of 8.35% in TLEF. Furthermore, ablation experiments confirm the effectiveness of the various strategies within the framework. These findings enhance the efficiency of the data publishing process, ensuring privacy and security and maximizing data availability. Mingshan You, Yong-Feng Ge, Kate N. Wang 0001, Hua Wang 0002, Jinli Cao, Georgios Kambourakis |
World Wide Web (WWW) | 2 |
| 2023 | TLEF: Two-Layer Evolutionary Framework for t-Closeness Anonymization
Mingshan You, Yong-Feng Ge, Kate N. Wang 0001, Hua Wang 0002, Jinli Cao, Georgios Kambourakis |
WISE | 2 |
| 2022 | An Information-Driven Genetic Algorithm for Privacy-Preserving Data Publishing
Yong-Feng Ge, Hua Wang 0002, Jinli Cao, Yanchun Zhang |
WISE | 1 |
| 2022 | DSGA: A Distributed Segment-Based Genetic Algorithm for Multi-Objective Outsourced Database Partitioning
Yong-Feng Ge, Zhi-hui Zhan, Jinli Cao, Hua Wang 0002, Yanchun Zhang, Kuei-Kuei Lai, Jun Zhang 0003 |
Inf. Sci. | 1 |
| 2022 | MDDE: multitasking distributed differential evolution for privacy-preserving database fragmentation
Yong-Feng Ge, Maria E. Orlowska, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
VLDB J. | 1 |
| 2021 | Set-Based Adaptive Distributed Differential Evolution for Anonymity-Driven Database FragmentationabstractAbstract By breaking sensitive associations between attributes, database fragmentation can protect the privacy of outsourced data storage. Database fragmentation algorithms need prior knowledge of sensitive associations in the tackled database and set it as the optimization objective. Thus, the effectiveness of these algorithms is limited by prior knowledge. Inspired by the anonymity degree measurement in anonymity techniques such as k-anonymity, an anonymity-driven database fragmentation problem is defined in this paper. For this problem, a set-based adaptive distributed differential evolution (S-ADDE) algorithm is proposed. S-ADDE adopts an island model to maintain population diversity. Two set-based operators, i.e., set-based mutation and set-based crossover, are designed in which the continuous domain in the traditional differential evolution is transferred to the discrete domain in the anonymity-driven database fragmentation problem. Moreover, in the set-based mutation operator, each individual’s mutation strategy is adaptively selected according to the performance. The experimental results demonstrate that the proposed S-ADDE is significantly better than the compared approaches. The effectiveness of the proposed operators is verified. Yong-Feng Ge, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
Data Sci. Eng. | 1 |
| 2021 | Knowledge transfer-based distributed differential evolution for dynamic database fragmentation
Yong-Feng Ge, Maria E. Orlowska, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
Knowl. Based Syst. | 1 |
| 2021 | Distributed Memetic Algorithm for Outsourced Database FragmentationabstractData privacy and utility are two essential requirements in outsourced data storage. Traditional techniques for sensitive data protection, such as data encryption, affect the efficiency of data query and evaluation. By splitting attributes of sensitive associations, database fragmentation techniques can help protect data privacy and improve data utility. In this article, a distributed memetic algorithm (DMA) is proposed for enhancing database privacy and utility. A balanced best random distributed framework is designed to achieve high optimization efficiency. In order to enhance global search, a dynamic grouping recombination operator is proposed to aggregate and utilize evolutionary elements; two mutation operators, namely, merge and split, are designed to help arrange and create evolutionary elements; a two-dimension selection approach is designed based on the priority of privacy and utility. Furthermore, a splicing-driven local search strategy is embedded to introduce rare utility elements without violating constraints. Extensive experiments are carried out to verify the performance of the proposed DMA. Furthermore, the effectiveness of the proposed distributed framework and novel operators is verified. Yong-Feng Ge, Wei-jie Yu 0001, Jinli Cao, Hua Wang 0002, Zhi-hui Zhan, Yanchun Zhang, Jun Zhang 0003 |
IEEE Trans. Cybern. | 1 |
| 2020 | Distributed Differential Evolution for Anonymity-Driven Vertical Fragmentation in Outsourced Data Storage
Yong-Feng Ge, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
WISE (2) | 1 |
| 2019 | A benefit-driven genetic algorithm for balancing privacy and utility in database fragmentationabstractIn outsourcing data storage, privacy and utility are significant concerns. Techniques such as data encryption can protect the privacy of sensitive information but affect the efficiency of data usage accordingly. By splitting attributes of sensitive associations, database fragmentation can protect data privacy. In the meantime, data utility can be improved through grouping data of high affinity. In this paper, a benefit-driven genetic algorithm is proposed to achieve a better balance between privacy and utility for database fragmentation. To integrate useful fragmentation information in different solutions, a matching strategy is designed. Two benefit-driven operators for mutation and improvement are proposed to construct valuable fragments and rearrange elements. The experimental results show that the proposed benefit-driven genetic algorithm is competitive when compared with existing approaches in database fragmentation. Yong-Feng Ge, Jinli Cao, Hua Wang 0002, Jiao Yin 0003, Wei-jie Yu 0001, Zhi-hui Zhan, Jun Zhang 0003 |
GECCO | 1 |
| 2018 | Competition-Based Distributed Differential EvolutionabstractDifferential evolution (DE) is a simple and efficient evolutionary algorithm for global optimization. In distributed differential evolution (DDE), the population is divided into several sub-populations and each sub-population evolves independently for enhancing algorithmic performance. Through sharing elite individuals between sub-populations, effective information is spread. However, the information exchanged through individuals is still too limited. To address this issue, a competition-based strategy is proposed in this paper to achieve comprehensive interaction between sub-populations. Two operators named opposition-invasion and cross-invasion are designed to realize the invasion from good performing sub-populations to bad performing subpopulations. By utilizing opposite invading sub-population, the search efficiency at promising regions is improved by opposition-invasion. In cross-invasion, information from both invading and invaded sub-populations is combined and population diversity is maintained. Moreover, the proposed algorithm is implemented in a parallel master-slave manner. Extensive experiments are conducted on 15 widely used large-scale benchmark functions. Experimental results demonstrate that the proposed competition-based DDE (DDE-CB) could achieve competitive or even better performance compared with several state-of-the-art DDE algorithms. The effect of proposed competition-based strategy cooperation with well-known DDE variants is also verified. Yong-Feng Ge, Wei-jie Yu 0001, Zhi-hui Zhan, Jun Zhang 0003 |
CEC | 1 |
| 2018 | Distributed Differential Evolution Based on Adaptive Mergence and Split for Large-Scale OptimizationabstractNowadays, large-scale optimization problems are ubiquitous in many research fields. To deal with such problems efficiently, this paper proposes a distributed differential evolution with adaptive mergence and split (DDE-AMS) on subpopulations. The novel mergence and split operators are designed to make full use of limited population resource, which is important for large-scale optimization. They are adaptively performed based on the performance of the subpopulations. During the evolution, once a subpopulation finds a promising region, the current worst performing subpopulation will merge into it. If the merged subpopulation could not continuously provide competitive solutions, it will be split in half. In this way, the number of subpopulations is adaptively adjusted and better performing subpopulations obtain more individuals. Thus, population resource can be adaptively arranged for subpopulations during the evolution. Moreover, the proposed algorithm is implemented with a parallel master-slave manner. Extensive experiments are conducted on 20 widely used large-scale benchmark functions. Experimental results demonstrate that the proposed DDE-AMS could achieve competitive or even better performance compared with several state-of-the-art algorithms. The effects of DDE-AMS components, adaptive behavior, scalability, and parameter sensitivity are also studied. Finally, we investigate the speedup ratios of DDE-AMS with different computation resources. Yong-Feng Ge, Wei-jie Yu 0001, Ying Lin 0001, Yue-Jiao Gong, Zhi-hui Zhan, Weineng Chen, Jun Zhang 0003 |
IEEE Trans. Cybern. | 1 |
| 2016 | Enhancing distributed differential evolution with a space-driven topologyabstractDifferential evolution (DE) is a simple and efficient evolutionary algorithm for global optimization. In distributed differential evolution (DDE), the population is divided into several sub-populations and each sub-population evolves independently for enhancing population diversity as well as algorithmic performance. Sub-populations in DDE share their elite individuals with neighborhood through a predefined migration topology. However, the construction of traditional migration topologies does not consider the position information of sub-populations in the search space. The position information is helpful in controlling the degree of diversity between the sub-populations and their migrated individuals. A proper degree of diversity could promote the balance between exploration and exploitation for DDE algorithms. To achieve this target, a dynamic space-driven migration topology is proposed in this paper. The proposed topology is constructed and updated according to the distances between sub-populations. Based on this proposed topology, some sub-populations receive diverse individuals from neighborhood far away while others communicate with neighborhood nearby. Numerical experiments have been performed on 13 diverse test functions. Results verify the advantage of DDE with the proposed migration topology compared to those with classic topologies. Yong-Feng Ge, Wei-jie Yu 0001, Jingjing Li 0002, Zhiwen Yu 0002, Jun Zhang 0003 |
CEC | 1 |
| 2015 | Reconstructing Cross-Cut Shredded Text Documents: A Genetic Algorithm with Splicing-Driven ReproductionabstractIn this work we focus on reconstruction of cross-cut shredded text documents (RCCSTD), which is of high interest in the fields of forensics and archeology. A novel genetic algorithm, with splicing-driven crossover, four mutation operators, and a row-oriented elitism strategy, is proposed to improve the capability of solving RCCSTD in complex space. We also design a novel and comprehensive objective function based on both edge and empty vector-based splicing error to guarantee that the correct reconstruction always has the lowest cost value. Experiments are conducted on six RCCSTD scenarios, with experimental results showing that the proposed algorithm significantly outperforms the previous best-known algorithms for this problem. Yong-Feng Ge, Yue-Jiao Gong, Wei-jie Yu 0001, Xiaomin Hu, Jun Zhang 0003 |
GECCO | 1 |