EDBT 2026 Demo / reviewers in the wild / expert
Qingying Yu
dblp:147/7244
· DBLP profile ↗
32ranked-venue papers
8as first author
22since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 8 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | DuGTRL: Dual-view grid-based trajectory representation learning framework integrating spatiotemporal semantics
Yamei Liu, Qingying Yu, Chuanming Chen, Xiaoyao Zheng, Yonglong Luo |
Expert Syst. Appl. | 2 |
| 2026 | Attention dynamic graph convolutional network for traffic flow prediction
Chenhui Wei, Chuanming Chen, Dongmei Pan, Qingying Yu, Xiaoyao Zheng, Yonglong Luo |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Similar yet different: Robust and transferable face privacy protection via adversarial identity editing
Xiaoyao Zheng, Liangmin Guo, Qingying Yu, Yonglong Luo |
Knowl. Based Syst. | 5 |
| 2026 | Federated Recommendation Model Based on Personalized Attention and Privacy-Preserving Dynamic GraphabstractGraph Neural Networks (GNNs) have been widely adopted in recommendation systems. When integrated into a federated learning framework, GNNs can enhance the model’s expressive capability. However, challenges arise in personalized representation and graph expansion due to the heterogeneity and locality of user data in federated recommendation systems. To address these challenges, we propose a federated recommendation model based on personalized attention and privacy-preserving dynamic graphs. The method first matches neighbor users for each selected client. Subsequently, it counts the interaction frequencies of items for both local and neighbor users to construct personalized weights, which captures the unique characteristics of different users. Additionally, we designs a method for constructing privacy-preserving dynamic graphs. In each round of federated training, the selected client adds pseudo-interaction items to its own interaction subgraph, perturbing the real interactions. After completing local training, the noisy interaction subgraph is incorporated into the global graph to capture higher-order connectivity information among users while safeguarding their interaction privacy. We conduct extensive experiments on three benchmark datasets, and the results demonstrate that the proposed PADG method achieves superior performance while effectively protecting privacy. Xiaoyao Zheng, Shukai Ye, Ming Zheng, Liangmin Guo, Qingying Yu, Yonglong Luo |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2026 | Location Privacy Protection Method Based on Local Differential Privacy in Crowdsensing With Approximately Accurate Task AllocationabstractWith the widespread adoption of smartphones and other mobile intelligent devices, Mobile Crowd Sensing (MCS) is widely used. Typically, the real locations of the workers and tasks must be submitted to the service platform to complete the task allocation. Therefore, the protection of location information has become a key factor in influencing user participation. To address the issue of location information leakage, we propose a location information protection method based on local differential privacy, which can protect the location privacy of workers and tasks while generating approximately accurate task allocation results. Firstly, we divide the region into$k$*$k$grids and merge girds with a similar dispersion to form clusters. Then, this paper utilizes the inverse sampling of the cumulative distribution function (CDF) of the flipped Huber distribution to generate a personalized noise location set for each cluster. Furthermore, the exponential mechanism is used to select the obfuscated location for each user. Finally, the platform selects workers based on the perturbed location to complete the task allocation. Theoretical analysis shows that our mechanism satisfies differential privacy and achieves an approximately accurate task allocation. Experimental results demonstrate that, compared to existing methods, this method exhibits superior performance across different datasets and effectively balances the utility of data and the protection of location privacy. Yutao Huang, Tianjiao Ni, Qingying Yu, Yonglong Luo |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Density-Aware Personalized Differential Privacy for Multi-Objective Task Allocation in Mobile CrowdsensingabstractWith the rapid advancement of Mobile Crowdsensing (MCS) technology, its impact on daily life continues to grow, making it indispensable to modern society. However, user data collection and analysis pose significant privacy risks. Although existing privacy-preserving task allocation schemes incorporate some basic adjustments for personalized noise, they fail to dynamically adapt based on user distribution and primarily focus on single-objective constraints. To bridge this gap, we propose a Density-Aware Personalized Differential Privacy for Multi-Objective Task Allocation in Mobile Crowdsensing (PDPMTA) scheme that adaptively adjusts noise intensity based on the distribution density of user data in feature space, then obfuscates sensitive information using differential privacy techniques. PDPMTA introduces a hybrid optimization strategy combining Simulated Annealing (SA) with NSGA-II, where simulated annealing is periodically applied to population subsets to balance exploration and exploitation, achieving more effective convergence toward the Pareto front. Experimental results on the real-world datasets confirm the scheme's effectiveness in optimizing travel distance, platform costs, and cost-efficiency. Zhichao Fang, Tianjiao Ni, Qingying Yu, Yonglong Luo |
ICPADS | 5 |
| 2025 | Multi-Scale Dual-Domain Attention Network for Traffic Flow PredictionabstractIn modern urban traffic management, accurate traffic flow prediction contributes to travel decision optimization, signal control, emergency response and resource scheduling, which improves road efficiency, reduces congestion and promotes sustainable urban planning. However, traffic flow data are characterized by complex spatio-temporal dependence, multi-scale periodicity, and highly dynamic changes. Existing studies mostly focus on single spatio-temporal domain modeling and ignore frequency domain information fusion, which makes it difficult to comprehensively capture the potential laws of traffic flow. To this end, a multi-scale spatio-temporal fusion Transformer prediction model is proposed, which systematically integrates frequency-domain analysis with spatio-temporal dependent modeling. The model contains three parts: (1) spatial-temporal adaptive neighbor selection algorithm, which dynamically supplements topological information based on spatio-temporal correlation to enhance the efficiency of inter-subdivisional information transfer; (2) frequency-domain feature coupling module, which fuses the frequency-domain and spatial-domain features by fast Fourier transform to enhance the ability of temporal pattern sensing; and (3) spatio-temporal-frequency-domain dual-attention encoder, which combines the linear-attention mechanism to efficiently capture the longrange spatio-temporal dependencies. Experimental results on several real traffic datasets show that the model significantly outperforms existing methods in terms of mean absolute error, root mean square error and mean absolute percentage error, and demonstrates stronger robustness in complex spatio-temporal patterns and abnormal fluctuation scenarios. Chenhui Wei, Chuanming Chen, Ming Zheng, Tianjiao Ni, Qingying Yu |
ICPADS | 6 |
| 2025 | UFIDSF: An undersampling approach based on feature importance and double side filter for imbalanced data classification
Ming Zheng, Liangchen Hu, Qingying Yu, Xiaoyao Zheng |
Future Gener. Comput. Syst. | 5 |
| 2025 | A Small-Scale Restricted Double Auction Mechanism Based on Local Differential PrivacyabstractAuctions have been widely applied in resource allocation due to their fairness and efficiency. For instance, platforms receive requests from service requesters and utilize auction theory to select suitable service providers. Existing studies typically assume that winners are determined based on bidders’ true valuations by allowing arbitrary transactions between requesters and providers, which can lead to serious valuation privacy leakage issues and limitations in application scenarios. Although some research has addressed these concerns using differential privacy techniques, they mostly rely on a trusted platform, and the introduction of noise results in utility loss, making them unsuitable for restricted auction contexts. To overcome these limitations, we propose a restricted double auction mechanism based on local differential privacy. Specifically, we extract the characteristics of the valuation data and constrain the noise addition probability density function based on the data features. Then we design a novel exponential selection mechanism that ensures that the relative positions of the obfuscated bids remain unchanged compared to the original valuations, while satisfying ε-local differential privacy. Furthermore, we develop an auction matching mechanism that maintains properties such as truthfulness under restricted allocation. The simulation results demonstrate that the proposed bid obfuscation mechanism ensures that the relative positions of the interfered bids remain unchanged while incurring low time overhead. Compared to existing mechanisms, our restrictive auction mechanism can generate greater social welfare while reducing the risk of valuation privacy leakage. Yutao Huang, Tianjiao Ni, Qingying Yu, Yonglong Luo |
IEEE Internet Things J. | 5 |
| 2025 | Density Peak Clustering Algorithm Based on Data Field Theory and Grid Similarity
Qingying Yu, Gege Shi, Dongsheng Xu 0004, Chuanming Chen, Yonglong Luo |
J. Comput. Sci. Technol. | 1 |
| 2025 | MMformer: Transformer-Based Trajectory Map-Matching Model for Large-Scale Road NetworksabstractNumerous trajectory data mining algorithms rely on rich map information for improved effectiveness, necessitating essential trajectory preprocessing steps, such as map matching. Current deep models are restricted to small road networks and cannot adapt to or learn large-scale, variable, and noisy trajectories. From a data-driven perspective, we effectively applied and developed a transformer function and proposed a transformer-based map matching (MMformer) for large-scale road networks. This model does not require maps and trajectories to be gridded, and directly runs on the vectors of trajectory points. The decoder module can learn, gather, and store the connections of road segments; therefore, the outputs of the decoder have a graph-like inductive bias. Trajectory point vector, including coming direction, enhances encoder and decoder performance. After experimenting with different embedding modules, the trajectory point vectors were embedded using a single linear layer. Extensive experiments demonstrated that MMformer can perform the map-matching task for large-scale road networks, and its map-matching accuracy on large-scale road networks is 10% higher than that of existing models on small road networks. Xiaoping Luo, Qingying Yu, Yonglong Luo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Trajectory outlier detection method based on group divisionabstractTrajectory-outlier detection can be used to discover the fraudulent behaviour of taxi drivers during operations. Existing detection methods typically consider each trajectory as a whole, resulting in low accuracy and slow speed. In this study, a trajectory outlier detection method based on group division is proposed. First, the urban vector region is divided into a series of grids of fixed size, and the grid density is calculated based on the urban road network. Second, according to the grid density, the grids were divided into high- and low-density grids, and the code sequence for each trajectory was obtained using grid coding and density. Third, the trajectory dataset is divided into several groups based on the number of low-density grids through which each trajectory passes. Finally, based on the high-density grid sequences, a regular subtrajectory dataset was obtained within each trajectory group, which was used to calculate the trajectory deviation to detect outlying trajectories. Based on experimental results using real trajectory datasets, it has been found that the proposed method performs better at detecting abnormal trajectories than other similar methods. Chuanming Chen, Dongsheng Xu 0004, Xiaoyao Zheng, Qingying Yu |
Intell. Data Anal. | 7 |
| 2024 | Which standard classification algorithm has more stable performance for imbalanced network traffic data?
Ming Zheng, Qingying Yu, Liangmin Guo, Fulong Chen 0002 |
Soft Comput. | 5 |
| 2024 | Lighter Sequential Recommendation Algorithm With Time Interval Awareness AugmentationabstractSequential recommendation models analyze users’ historical interactions to predict the next item they will en gage with. In order to better capture users’ dynamic interest preferences, most existing sequential recommendation models that introduce heterogeneous time intervals lead to increased model complexity, which raises computational costs and training difficulty. This is particularly evident in long sequential data, where the model need to handle a large variety of different time intervals. Additionally, accurately modeling the impact of long time intervals on user behavior remains a significant challenge. To address these issues, we propose a lightweight sequential recommendation algorithm with time interval awareness augmen tation (TALSAN). This model introduces a novel uniform data augmentation operator to improve the distribution of original data samples and employs a time-aware self-attention layer to model user interactions, maintaining the continuity of the original sequence. By integrating temporal context with posi tional features, TALSAN constructs a streamlined self-attention network for predicting user behavior. Comparative testing on datasets such as ML-100K, ML-1M, Amazon Beauty, Amazon Toys, and Amazon Fashion demonstrates the model’s superiority over existing baselines. Our results confirm that TALSAN not only mitigates cold start issues but also enhances the ability to learn user preferences, leading to improved prediction accuracy. Xiaoyao Zheng, Shengfei Jiang, Zhenghua Chen, Qingying Yu, Liangmin Guo, Yonglong Luo |
IEEE Trans. Serv. Comput. | 6 |
| 2023 | Parallel Pattern Matching over Brotli Compressed Network TrafficabstractPattern matching is a crucial technique for network traffic detection applications. As a fundamental computation model used by pattern matching, the finite state automata execute sequential matching due to the state dependence among transitions. Meanwhile, most services tend to compress their data to improve transmission or storage efficiency. The increased compressed data challenges the straightforward method of matching the whole decompressed data and incurs data dependence among the compression encodings. The related approaches either leverage techniques to break the state dependence of matching uncompressed data or accelerate matching compressed data in a single-threaded manner without considering the state and data dependence. None of them can perform parallel matching over compressed data. This paper provides PETALS, a parallel pattern matching method over Brotli compressed network traffic. PETALS partitions the original compressed traffic into fixed- length blocks for parallel matching and patches the broken compression encodings crossing blocks to break the data dependence. Then, it merges the compressed traffic matching method into path fusion, an enumerative parallelization of finite state automata, to present parallel matching over compressed traffic. Evaluation using real-world network traffic and regular expressions shows that PETALS can raise the speedup from 1.53x to 3.53x of the state-of-the-art parallelization schemes on a 56-cores machine. Xiuwen Sun, Guangzheng Zhang, Qingying Yu, Jie Cui 0004, Hong Zhong 0001 |
TrustCom | 4 |
| 2023 | Trajectory personalization privacy preservation method based on multi-sensitivity attribute generalization and local suppressionabstractFast-developing mobile location-aware services generate an enormous volume of trajectory data while adding value to people’s lives. However, trajectory data contains not only location information, but also sensitive personal information. If the original trajectory data is published directly, it could result in serious privacy leaks. Most of the existing privacy-preserving trajectory publishing methods only protect the location information or set the same privacy preservation levels for all moving objects. To meet the users’ personalized privacy requirements and ensure the utility of trajectory location and sensitive information, we propose a trajectory personalized privacy preservation method based on multi-sensitivity attribute generalization and local suppression. First, we set different security levels for each trajectory by calculating the correlation between sensitive attributes to establish a sensitive attribute classification tree. Second, we generalized sensitive attributes based on privacy preservation levels for each trajectory, the trajectory data still at risk of privacy leakage after generalization was locally suppressed. Finally, an anonymized trajectory dataset was generated. Experimental results on real datasets demonstrated that our method could improve data availability while preserving privacy. Qingying Yu, Zhenxing Xiao, Shan Gong, Chuanming Chen |
Intell. Data Anal. | 1 |
| 2023 | Efficient regular expression matching over hybrid dictionary-based compressed data
Xiuwen Sun, Da Mo, Chunhui Ye, Qingying Yu, Jie Cui 0004, Hong Zhong 0001 |
J. Netw. Comput. Appl. | 5 |
| 2023 | Improved path planning algorithm for mobile robots
Xiaoyu Duan, Pingan Xu, Xiaoyao Zheng, Qingying Yu, Yonglong Luo |
Soft Comput. | 6 |
| 2022 | High-Frequency Trajectory Map Matching Algorithm Based on Road Network TopologyabstractAccurately mapping the raw global position system (GPS) trajectories to the road network is the basis for studying the application of trajectory data. This study proposes a novel off-line map matching algorithm based on road network topology, to address the problems of low execution efficiency and poor matching accuracy of selective look-ahead map matching (SLAMM) algorithm. First, the noise points of the trajectory data are removed by data preprocessing. Second, the algorithm searches for critical samples in the trajectory data and segments the data accordingly. Then, the adjacent road segments around the transition node corresponding to the critical sample are selected as candidate arcs. Finally, the segmented trajectory data are matched to the road network by constructing an error ellipse. The algorithm fully considers the topology of the road network and the characteristics of high-frequency trajectory data. The experimental results, using Beijing trajectory data to perform matching on an actual road network environment, show that the proposed algorithm is more efficient and robust than other map matching algorithms for high-frequency trajectories. Qingying Yu, Zhen Ye 0003, Chuanming Chen, Yonglong Luo |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Using information entropy and a multi-layer neural network with trajectory data to identify transportation modesabstractResidents’ trajectory data denote their instantaneous locations along their movements. Mobility research that applies trajectory mining techniques to identify the transportation modes of these movements can inform urban transportation planning. Herein, we propose a five-step approach with information entropy and a multi-layer neural network to identify transportation modes from trajectory data. First, this approach extracts the motion features at each time-stamped location based on foundation geospatial data and spatiotemporal trajectory data, including the speed, acceleration, change of direction, rate of change in direction, and distance from each basic transportation facility. The second step uses information entropy to identify the features that play key roles in identifying transportation modes. The third step weighs each attribute in the feature vector consisting of the selected features and normalizes it to prepare it as input data. The fourth step constructs, trains, and tests a multi-layer neural network with seven-fold cross-validation. The final step includes a post-processing method to optimize the identification result. We use F-measure metric to evaluate the performance. Experimental results on a real trajectory dataset show that the proposed approach can identify the transportation mode at each time-stamped location and outperforms existing transportation-mode identification methods in terms of accuracy and stability. Qingying Yu, Yonglong Luo, Dongxia Wang 0004, Chuanming Chen |
Int. J. Geogr. Inf. Sci. | 1 |
| 2021 | Personalized trajectory privacy-preserving method based on sensitive attribute generalization and location perturbationabstractTrajectory data may include the user’s occupation, medical records, and other similar information. However, attackers can use specific background knowledge to analyze published trajectory data and access a user’s private information. Different users have different requirements regarding the anonymity of sensitive information. To satisfy personalized privacy protection requirements and minimize data loss, we propose a novel trajectory privacy preservation method based on sensitive attribute generalization and trajectory perturbation. The proposed method can prevent an attacker who has a large amount of background knowledge and has exchanged information with other attackers from stealing private user information. First, a trajectory dataset is clustered and frequent patterns are mined according to the clustering results. Thereafter, the sensitive attributes found within the frequent patterns are generalized according to the user requirements. Finally, the trajectory locations are perturbed to achieve trajectory privacy protection. The results of theoretical analyses and experimental evaluations demonstrate the effectiveness of the proposed method in preserving personalized privacy in published trajectory data. Chuanming Chen, Wenshi Lin, Shuanggui Zhang, Zitong Ye, Qingying Yu, Yonglong Luo |
Intell. Data Anal. | 5 |
| 2021 | UFFDFR: Undersampling framework with denoising, fuzzy c-means clustering, and representative sample selection for imbalanced data classification
Ming Zheng, Tong Li 0004, Xiaoyao Zheng, Qingying Yu, Chuanming Chen, Changlong Lv |
Inf. Sci. | 4 |
| 2020 | A privacy-preserving density peak clustering algorithm in cloud computingabstractSummary Aiming at preventing the privacy disclosure of sensitive information, issues related to privacy protection in cloud computing have attracted the interest of researchers. To protect the privacy of users during clustering in a cloud computing environment, we present a privacy‐preserving density peak clustering (PPDPC) algorithm that neither discloses personal privacy information nor leaks the cluster centers. Our scheme contains two steps of density peak clustering: First, a cloud service provider calculates the cluster centers without knowing each participant's private data and without disclosing any cluster center information to the other participants, and second, participant allocation is secure and every participant is prevented from identifying the other members of the same cluster. Security analysis and comparison experiments show that the proposed PPDPC algorithm not only obtains good accuracy with respect to density peak clustering but also resists collusion attacks even if the cloud service provider is collaborating with all except one participant. Both theoretical analysis and experimental results confirm the security and accuracy of our method. Shang Ci, Xiaoyao Zheng, Qingying Yu, Yonglong Luo |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | A Framework of Abnormal Behavior Detection and Classification Based on Big Trajectory Data for Mobile NetworksabstractBig trajectory data feature analysis for mobile networks is a popular big data analysis task. Due to the large coverage and complexity of the mobile networks, it is difficult to define and detect anomalies in urban motion behavior. Some existing methods are not suitable for the detection of abnormal urban vehicle trajectories because they use the limited single detection techniques, such as determining the common patterns. In this study, we propose a framework for urban trajectory modeling and anomaly detection. Our framework takes into account the fact that anomalous behavior manifests the overall shape of unusual locations and trajectories in the spatial domain as well as the way these locations appear. Therefore, this study determines the peripheral features required for anomaly detection, including spatial location, sequence, and behavioral features. Then, we explore sports behaviors from the three types of features and build a taxi trajectory model for anomaly detection. Anomaly detection, including sports behaviors, are (i) detour behavior detection using an algorithm for global router anomaly detection of trajectories having a pair of same starting and ending points; this method is based on the isolation forest algorithm; (ii) local speed anomaly detection based on the DBSCAN algorithm; and (iii) local shape anomaly detection based on the local outlier factor algorithm. Using a real-life dataset, we demonstrate the effectiveness of our methods in detecting outliers. Furthermore, experiments show that the proposed algorithms perform better than the classical algorithm in terms of high accuracy and recall rate; thus, the proposed methods can accurately detect drivers’ abnormal behavior. Yonglong Luo, Qingying Yu, Xuejing Li, Zhenqiang Sun |
Secur. Commun. Networks | 3 |
| 2019 | Trajectory similarity clustering based on multi-feature distance measurement
Qingying Yu, Yonglong Luo, Chuanming Chen, Shigang Chen |
Appl. Intell. | 1 |
| 2019 | TPPG: Privacy-preserving trajectory data publication based on 3D-Grid partitionabstractThe issue of privacy preservation is receiving more and more attention when publishing trajectory data. In this paper, we study the challenges of published trajectory data anonymization. Most existing anonymization methods directly delete the trajectories or locations violating specific constraints , it is likely to cause a large loss of information. To address the problem, this paper proposes a trajectory privacy preservation method based on 3D-Grid partition in order to reduce information loss in the process of trajectory anonymization. This method first divides the trajectory region into several spatio-temporal units (denoted as 3D-cells), and then conducts location exchange or suppression in each spatio-temporal unit. Based on the trajectory data partition, within each 3D-cell, the proposed method exchanges locations among trajectories or removes very few locations of some sub-trajectories which do not meet the conditions rather than the whole trajectory. Our method considers three scenarios of trajectory distribution and measures trajectory similarity based on time, orientation, spatial locations and other features of trajectory. After the reconstruction of the related anonymous sub-trajectories, an anonymized trajectory dataset is obtained. Theoretical analysis and experimental results show that, compared to other methods, the proposed algorithm effectively preserves trajectory data privacy and improves the anonymous results of trajectory data in terms of accuracy and availability. Chuanming Chen, Yonglong Luo, Qingying Yu, Guiyin Hu |
Intell. Data Anal. | 3 |
| 2019 | Hierarchical interpolation point anonymity for trajectory privacy protectionabstractThe traditional trajectory privacy protection algorithm approaches the task as a single-layer problem. Taking a perspective in harmony with an approach more characteristic of human thinking, in which complex problems are solved hierarchically, we propose a two-level hierarchical granularity model f or this problem. The first level of the proposed model is a coarse-grained layer, in which the original dataset is divided into groups. The second level is a fine-grained layer, where problems are solved in each group instead of on the original dataset, which reduces complexity and computation while improving efficiency. On the basis of this hierarchical model, we propose the interpolation trajectory-anonymous privacy protection algorithm with temporal and spatial granularity constraints. In addition, we propose interpolation-based modified Hausdorff distance on adjacent segment (IMHD_AS), which provides a smaller clustering area and better data utility than the traditional Euclidean distance, as the trajectory similarity criterion for clustering within each group. Further, we theoretically prove that the proposed algorithm outperforms the traditional algorithm in terms of data distortion and anonymity cost and verify its efficacy experimentally. Compared with the classic anonymity algorithm, the maximum information loss and the anonymity cost are reduced by up to 21.04% and 28.32%, respectively. Zepei Zhang, Yonglong Luo, Qingying Yu |
Intell. Data Anal. | 4 |
| 2018 | Trajectory outlier detection approach based on common slices sub-sequence
Qingying Yu, Yonglong Luo, Chuanming Chen |
Appl. Intell. | 1 |
| 2018 | Probabilistic optimal projection partition KD-Tree k-anonymity for data publishing privacy protectionabstractData needs to be released to the relevant decision makers and researchers. Privacy protection should be carried out first because it contains personal sensitive information. The k-anonymity algorithm is an important privacy protection algorithm, and partitioning is one of its key methods. To reduce the computational complexity and low speed of existing privacy-preserving algorithms for high-dimensional data publishing, a probabilistic optimal projection partition k-dimensional (KD)-tree k-anonymity algorithm is proposed. First, some attribute dimensions are probabilistically selected from the global domain. Then, for these dimensions, the partition coefficient is calculated and the optimal partition point is determined. Furthermore, an improved KD-tree structure is introduced in which a node is a collection rather than a data point. The proposed KD-tree node is divided into left and right child nodes by the hyper-plane passing through the dividing point and perpendicular to the optimal dimension. The proposed algorithm is validated by a theoretical analysis and comparison experiments. The results show that the proposed algorithm can reduce the average generalization range by 11% to 22% compared to traditional k-anonymity. This enables better division and better dataset availability. Moreover, the runtime is reduced by 8% to 32% compared to globally optimal projection partitioning k-anonymity. Yonglong Luo, Yefeng Jiang, Wenli Wu, Qingying Yu |
Intell. Data Anal. | 5 |
| 2016 | Outlier-eliminated k-means clustering algorithm based on differential privacy preservation
Qingying Yu, Yonglong Luo, Chuanming Chen, Xintao Ding |
Appl. Intell. | 1 |
| 2016 | Neighborhood relevant outlier detection approach based on information entropyabstractOutlier detection is an interesting issue in data mining and machine learning. In this paper, to detect outliers, an information-entropy-based k-nearest neighborhood relevant outlier factor algorithm is proposed that is combined with Shannon information theory and the triangle pruning strategy. The algorithm accounts for the data points whose k-nearest neighbors are distributed on the edge of the range within the designated radius. In particular, the neighborhood influence on each point is considered to address the problem of information concealment and submergence. Information entropy is used to calculate the weights to distinguish the importance of each attribute. Then, based on the attribute weights, the improved pruning strategy reduces the computational complexity of the subsequent procedures by removing some inliers and obtaining the outlier candidate dataset. Finally, according to the weighted distance between the objects in the candidate dataset and those in the original dataset, the algorithm calculates the dissimilarity between each object and its k-nearest neighbors. The data points with the top $r$ dissimilarity are regarded as the outliers. Experimental results show that, compared to existing methods, the proposed approach improves pruning and detection rates while maintaining the coverage rate. Qingying Yu, Yonglong Luo, Chuanming Chen, Weixin Bian |
Intell. Data Anal. | 1 |
| 2014 | Fingerprint ridge orientation field reconstruction using the best quadratic approximation by orthogonal polynomials in two discrete variables
Weixin Bian, Yonglong Luo, Deqin Xu, Qingying Yu |
Pattern Recognit. | 4 |