Guangjun Wu

dblp:21/4046 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Computer networks · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Digital Scapegoat: An Incentive Deception Model for Resisting Unknown APT Stealing Attacks on Critical Data Resource
abstract
It is a challenging problem to resist unknown advanced persistent threats (APTs) on stealing data resources in an information system of critical infrastructures, because APT attackers have very specific objectives and compromise the system stealthily and slowly. We observe that it is a necessary condition for APT attackers to achieve their campaigns via controlling unknown Trojans to access and exfiltrate critical files. We present a theoretical model called Digital Scapegoat (abbreviated as DS-IDep) that constructs an Incentive Deception defense schema to hijack the attacker’s access to critical files and redirect it to avatar files without awareness. We propose a FlipIDep Game model (GF) and a Markov Game model (GM) to characterize completely the payoffs, equilibria, and best strategies from the perspective of the attacker and the defender respectively. We also design an exponential risk propagation model to evaluate the ability of DS-IDep to eliminate stealing impact when the risk is propagated between states. Theoretically, we can achieve the objective of stealing impact elimination (LK0.7) and the probability of an attack operation bypassing the defense surface is less than 0.1 (r* × μ <0.1) under Stackelberg strategies. We develop a kernel-level incentive deception defense surface according to the theoretical parameters of the DS-IDep. The experimental results show that DS-IDep can resist APT stealing attacks from unknown Trojans. We also evaluate the DS-IDep in five well-known software applications. It demonstrates that DS-IDep can address unknown attacks from compromised software with less than 10% performance overhead.
Xiao-chun Yun, Guangjun Wu, Qige Song, Zixian Tang, Zhenyu Cheng 0001
IEEE Trans. Inf. Forensics Secur.2
2024 Understanding Atomics and Memory Ordering Issues in Real-World Rust Software
abstract
Rust is designed as a systems programming language that aims to provide safety guarantees and performance efficiency. In practice, programmers usually use atomic correlations to share data across threads. For example, by using atomic operations to correlate with non-atomic addresses, they can design lock-free data structures for efficient concurrency. Although atomic operations are used in safe code, memory ordering misuses can still lead to atomic concurrency bugs and performance loss.In this paper, we conduct the first empirical study of atomic operations and memory ordering usage in Rust, manual inspection of 2883 atomic usages in real-world applications, including 15 thread bugs and 150 performance issues. We also study their usage scenarios, performance comparisons and issue fixes to provide a better understanding on Rust’s memory ordering misuses and guide better code practices in the future.We design AtomVChecker, an automated static analyzer to detect memory ordering misuses. we evaluate our tool on four widely-used concurrent libraries, it can automatically analyze 228 atomic correlations with 80% accuracy. Based on the atomic correlation analysis, AtomVChecker finds a total of 51 performance loss issues in 9 Rust packages, with all of them recently confirmed by the project maintainer based on our reports.
Tengfei Tu, Su-Juan Qin, Guangjun Wu, Fei Gao 0001, Mingchao Wan
ISSRE4
2024 AC-DNN: An Adaptive Compact DNNs Architecture for Collaborative Learning Among Heterogeneous Smart Devices
abstract
With the rapid development of the Internet of Things (IoT), a massive number of smart devices are deployed in industry and critical infrastructures. Nowadays, IoT smart devices have drawn increasing attention for collaborative learning tasks, e.g., persistent monitoring and online recognition. In this article, we present an adaptive compact deep neural network (DNN) approach (termed as AC-DNN) to tackle the challenging problem of unreliable transmission for collaborative learning tasks among heterogeneous smart devices. We introduce a cross-platform model weight encoding, decoding, and dispatching architecture to accommodate to differential smart devices and improve the reliability of intermediate model transmission via encapsulating binary model weights into self-contained transactions. To decrease encoding and decoding overhead, we design a quantile-based histogram sketch to compress the intermediate model. We conduct extensive evaluations to test our AC-DNN framework and deploy the AC-DNN on federated learning testbed FedAvg. We evaluate our approach functionality using different DNN architectures, such as convolutional neural network and ResNet and compare their effectiveness within the different network structures. The experiments reveal that our approach can improve the reliability of collaborative learning tasks among smart devices. Meanwhile, we can achieve nearly 70% weight compression compared to the original model size with minimal loss of accuracy. Our approach facilitates the deployment of a DNN-like network among discrete mobile smart devices for deep and persistent learning tasks.
Guangjun Wu, Fengxin Liu, Qige Song, Zixian Tang
IEEE Internet Things J.1
2024 Federated Learning Enabled Credit Priority Task Processing for Transportation Big Data
abstract
Due to the epidemic COVID-19 spread and Intelligent Transportation System (ITS) development, investigators are now to conduct their research over the generated Transportation Big Data (TBD) in many critical areas, such as medical supplies, food supplies, as well as logistics supplies. At present, Vehicular Edge Computing (VEC) is an emerging paradigm to integrate resources from vehicles, road-site units, base stations, and cloud center to promote the performance of TBD tasks scheduling and running. In this paper, we design a three-layered TBD task processing architecture with a federated learning mechanism for credit priority-based task scheduling and running. In our design, we consider the efficiency of task offloading and misbehavior attack problems simultaneously. We propose a vehicular federated learning framework combined with Multi-Layer Perceptron (MLP) credit measurement, which can preserve the privacy of vehicles and obtain the related features for vehicular credit prediction. We also propose a task offloading algorithm to solve the optimization problem for credit priority task offloading between edge computing servers and vehicles. The proposed solution can prioritize tasks and assign sufficient resources for reliable and active task requesters. Experimental results expose that the proposed mechanism outperforms the state-of-the-art solutions when considering efficiency and attack simultaneously for TBD tasks scheduling and running.
Guangjun Wu, Jun Li 0085, Zhaolong Ning, Yong Wang 0032, Binbin Li 0001
IEEE Trans. Intell. Transp. Syst.1
2023 UCWSC: A unified cross-modal weakly supervised classification ensemble framework
abstract
In recent years, Internet data has grown exponentially, but due to the lack of labels, the data that can be used is still relatively small. To solve this problem, research on weak supervision has emerged. However, common weakly supervised research often focuses on either single-modal data or multi-modal data research, which cannot be compatible with both types of data at the same time. Motivated by this observation, we propose a unified cross-modal weakly supervised classification ensemble framework (UCWSC) to tackle this issue. Especially, our proposed framework is based on high-order feature information of different modes. First, We introduce a feature fusion method based on high-order features to increase the amount of acquired information. Then we propose a modified Feature MixMatch algorithm with learning from feature representations. We propose feature fusion and decision fusion methods for weakly supervised classification of multi-modal data with voting and weighting mechanisms as discriminators to obtain the final classification results, respectively. We demonstrate the compatibility of these techniques, our classification accuracy can reach around 99% on the Wikipedia dataset and 78% on the MVSA-Multiple dataset.
Huiyang Chang, Binbin Li 0003, Guangjun Wu, Haiping Wang 0003, Siyu Jia, Zisen Qi, Xiaohua Jiang
CSCWD3
2022 Edge Federated Learning for Social Profit Optimality: A Cooperative Game Approach
Wenyuan Zhang 0002, Guangjun Wu, Yongfei Liu, Binbin Li 0003
CollaborateCom (1)2
2022 Federated Learning-Based Intrusion Detection on Non-IID Data
Yongfei Liu, Guangjun Wu, Wenyuan Zhang 0002, Jun Li 0085
ICA3PP2
2022 BMKS: A Blockchain Based Multi-Keyword Search Scheme for Medical Data Sharing
abstract
In recent years, electronic medical records(EMRs) sharing has played a vital role in formulating optimized treatment plans, providing data sets for researchers, and accelerating the development of biomedical science. However, This unprecedented era of technological confluence poses significant data security and privacy challenges. To solve these problems, we propose a blockchain-based multi-keyword search scheme for medical data sharing called BMKS, focusing on ensuring medical data confidentiality and realizing secure data retrieval. In BMKS, we introduce a two-level search scheme to achieve efficient and verifiable keyword searches. The introduction of the improved Bloom filter significantly improves query efficiency. In addition, we take advantage of blockchain technology to realize ciphertext search and pre-decryption, which reduces user's decryption overhead. Moreover, blockchain's transparency and tamper-preventing characteristics help record the access control process in a traceable and auditable way. The performance evaluation and security analysis show that BMKS is comprehensively safe, efficient, and practical.
Guangjun Wu, Bingqing Zhu, Jun Li 0085
ISCC1
2022 Fourier Enhanced MLP with Adaptive Model Pruning for Efficient Federated Recommendation
Zhengyang Ai, Guangjun Wu, Binbin Li 0001, Yong Wang 0032, Chuantong Chen
KSEM (3)2
2022 Towards Better Personalization: A Meta-Learning Approach for Federated Recommender Systems
Zhengyang Ai, Guangjun Wu, Zisen Qi, Yong Wang 0032
KSEM (2)2
2022 Privacy-Preserving Deep Learning in Internet of Healthcare Things with Blockchain-Based Incentive
Wenyuan Zhang 0002, Guangjun Wu, Jun Li 0085
KSEM (3)3
2022 Blockchain-Enabled Privacy-Preserving Access Control for Data Publishing and Sharing in the Internet of Medical Things
abstract
Recently, the rapid developments in the Internet of Medical Things (IoMT) enable smart devices to generate and transmit massive personal electronic medical records (EMRs). However, there are many sensitive attributes in an EMR, which could be accessed by external or internal unauthorized users for malicious purposes. In this article, we present a triple subject purpose-based access control (TS-PBAC) model, which is compatible with a blockchain-enabled reliable transaction network, and design an individual-centric security and privacy-preserving mechanism for access control with different purposes and roles in IoMT scenarios. Specifically, we design hierarchical purpose tree (HPT) and related policies to guarantee the legality of an external user with different purposes. To improve the privacy for sensitive attributes against an internal attacker, we design a local differential privacy (LDP)-based policy and role-based access control scheme in an edge computing paradigm to grant fine-granularity rights for authorized users. In addition, we introduce mutual evaluation metrics to evaluate data quality from a patient-and-medical-service level in an open anonymous network, only using logs kept in the blockchain. We test our approach by real-world EMRs with 100000 patients. The experimental results show that the proposed privacy-preserving scheme can better protect patient’s privacy than traditional access control policies in IoMT environments, and can make reliable and stable access control decisions between data publishers and data requesters with different purposes.
Guangjun Wu, Zhaolong Ning, Jun Li 0085
IEEE Internet Things J.1
2022 A Sketching Approach for Obtaining Real-Time Statistics Over Data Streams in Cloud
abstract
Many applications of complex event processing (CEP) in Cloud can tolerate analytical errors to some extent, and it provides us an opportunity to optimize real-time analytics using methods of approximate query processing over big data streams. In this article, we present a novel rules-based sampling technique, which supports to construct sketch over one-pass and high-speed asynchronous data streams and provides accurate answers for different types of analytical queries. Moreover, we propose two methods of distributed sketching implementation, i.e., D-AQP$_b$and D-AQP$_i$, to make our approach to be compatible with batch processing and interactive processing architectures respectively, and be appropriate for stream processing systems in Cloud. Experimental results with real-world and synthetic datasets indicate that our approach can obtain more accurate estimates and improve two times of system throughput when compared with state-of-the-art Hadoop-based approximate engine BlinkDB. When compared with current batch processing systems Spark and stream processing system Spark-Streaming, our methods of D-AQP$_b$and D-AQP$_i$can achieve 2 and 4 orders of magnitude improvement on query response time respectively.
Guangjun Wu, Xiao-chun Yun, Yong Wang 0032, Binbin Li 0001, Yong Liu 0018
IEEE Trans. Cloud Comput.1
2022 Joint Optimization of Task Offloading and Resource Allocation Based on Differential Privacy in Vehicular Edge Computing
abstract
In the Internet of Vehicles (IoVs), task offloading is necessary to ensure the low-response delay due to the limitation of vehicular computational capacity. Task offloading involving social behavior can improve the utilization of computational resources in IoVs. To offload tasks effectively, the connected vehicles (CVs) need to upload context information, such as speed and location to road side unit (RSU) and base station (BS), which brings dramatic threat and risk for CVs’ privacy security. To solve the above-mentioned issue, we propose a privacy-preserving vehicular edge computing (PP-VEC) system architecture in this article. In the PP-VEC, the vehicular tasks can be offloaded to RSUs and adjacent CVs with adequate computing resources. Privacy mechanism disturbs the context information of CVs based on differential privacy technology before uploading it to the BS for offloading decisions to protect the CVs’ privacy. This article adopts the local differential privacy algorithm based on histogram algorithm and proposes a K-neighbor joint optimization of task offloading and resource allocation algorithm (K-NJTA) to optimize the global delay of task execution. We demonstrate the effectiveness of the proposed methods by simulation experiments. The results demonstrate K-NJTA on task execution delay and our local differential privacy algorithm can protect CVs’ privacy while has less effect on the task offloading algorithm due to the context information distribution.
Jun Li 0085, Guangjun Wu, Handi Chen, Shihui Sun
IEEE Trans. Comput. Soc. Syst.3
2022 Privacy-Preserved Electronic Medical Record Exchanging and Sharing: A Blockchain-Based Smart Healthcare System
abstract
The digitization of Electronic Medical Record (EMR) provides potential access to a wealth of medical information, but also presents new challenges in privacy-preserved EMR exchanging and sharing. In this paper, we propose a blockchain-based smart healthcare system with fine-grained privacy protection for reliable data exchanging and sharing among different users. We design a blockchain-enabled dynamic access control framework combined with Local Differential Privacy (LDP) strategies to provide the attribute-based privacy protection in transaction workflow. We design four types of smart contracts in the framework to meet the requirements of anonymous transaction, dynamic access control, beneficial matching decision, and evaluation of published data in an open network. To satisfy fine-grained privacy protection, we classify sensitive attributes of EMRs into different levels and set differential privacy budgets to randomize attributes before data publishing. Also, we design data quality function to depict the disturbance incurred by LDP-based privacy preferences at the requester view, and present appropriate many-to-many matching decisions among participants for beneficial transactions. Finally, we develop a prototype system and test our approach using 200,000 real-world EMRs. Experimental results show that the proposed privacy-preserved scheme can make stable and reliable transactions between EMR publishers and requesters. The prototype system achieves individual-centric privacy configuration at the patient site, while providing error-guaranteed statistics at the requester site. Additionally, the access control policies, logs of anonymous transaction are kept in the blockchain to provide system-level traceability.
Guangjun Wu, Zhaolong Ning, Bingqing Zhu
IEEE J. Biomed. Health Informatics1
2020 Parallel Belief Propagation Optimized by Coloring on GPUs
Junteng Hou, Chengxiang Si, Guangjun Wu
ICA3PP (1)4
2020 Parallel SCC Detection Based on Reusing Warps and Coloring Partitions on GPUs
Junteng Hou, Guangjun Wu, Bingnan Ma
ICA3PP (1)3
2020 MG-Hybrid: A Strongly Connected Components Detection Algorithm using Multiple GPUs
abstract
Detection of strongly connected component (SCC) on the GPU has become a fundamental operation to accelerate graph computing. Existing SCC detection methods on multiple GPUs introduce massive unnecessary data transformation between multiple GPUs. In this paper, we propose a novel distributed SCC detection approach using multiple GPUs plus CPU. Our approach includes three key ideas: (1) segmentation and labeling over large-scale datasets; (2) collecting and merging the segmented SCCs; and (3) running tasks assignment over multiples GPUs and CPU. We implement our approach under a hybrid distributed architecture with multiple GPUs plus CPU. Our approach can achieve device-level optimization and can be compatible with the state-of-the-art algorithms. We conduct extensive theoretical and experimental analysis to demonstrate efficiency and accuracy of our approach. The experimental results expose that our approach can achieves 11.2×, 1.2×, 1.2× speedup for SCC detection using NVIDIA K80 compared with Tarjan's, FB-Trim, and FB-Hybrid algorithms respectively.
Junteng Hou, Guangjun Wu, Bingnan Ma, Chengxiang Si, Siyu Jia
ISCAS3
2019 Accelerating Real-Time Tracking Applications over Big Data Stream with Constrained Space
Guangjun Wu, Xiao-chun Yun, Ge Fu, Chao Li 0062, Yong Liu 0018, Binbin Li 0001, Yong Wang 0032
DASFAA (1)1
2018 Dynamic Count-Min Sketch for Analytical Queries Over Continuous Data Streams
abstract
The methods of approximate query processing have been proposed for analytics over high-speed data streams, which compact continuous streams into a space-constrained sketch and provide reliable estimates for different queries. Count-Min (CM) is the state-of-the-art sketching structure supporting many queries with error-guaranteed estimates under limited space. However, we need to create a counter table beforehand in CM according to the size of data streams, while it is usually unpredictable for dynamic data streams. In this paper, we proposed an approach, called Dynamic Count-Min sketch (DCM), which is appropriate for dynamic data set and can provide accurate estimates for point query and self-join size query. Our approach constitutes incremental CM sketches and allocates space in a pay-as-you-go manner. Our mathematical analysis and substantial experiments both show that our approach is appropriate for data sets with dynamic or skewed inputs and can provide error-guaranteed estimates with less space compared to CM.
Guangjun Wu, Hong Zhang 0023, Bingnan Ma
HiPC2
2017 Supporting Real-Time Analytic Queries in Big and Fast Data Environments
Guangjun Wu, Xiao-chun Yun, Chao Li 0062, Yipeng Wang 0001, Xiaoyu Zhang 0002, Siyu Jia, Guangyan Zhang
DASFAA (2)1
2017 A nonparametric approach to the automated protocol fingerprint inference
Yipeng Wang 0001, Xiao-chun Yun, Yongzheng Zhang 0002, Guangjun Wu
J. Netw. Comput. Appl.5
2015 Update vs. upgrade: Modeling with indeterminate multi-class active learning
Xiaoyu Zhang 0002, Xiaobin Zhu 0001, Xiao-chun Yun, Guangjun Wu, Yipeng Wang 0001
Neurocomputing5
2015 FastRAQ: A Fast Approach to Range-Aggregate Queries in Big Data Environments
abstract
Range-aggregate queries are to apply a certain aggregate function on all tuples within given query ranges. Existing approaches to range-aggregate queries are insufficient to quickly provide accurate results in big data environments. In this paper, we propose FastRAQ-a fast approach to range-aggregate queries in big data environments. FastRAQ first divides big data into independent partitions with a balanced partitioning algorithm, and then generates a local estimation sketch for each partition. When a range-aggregate query request arrives, FastRAQ obtains the result directly by summarizing local estimates from all partitions. FastRAQ has O(1) time complexity for data updates and O(N/P×B) time complexity for range-aggregate queries, where N is the number of distinct tuples for all dimensions, P is the partition number, and B is the bucket number in the histogram. We implement the FastRAQ approach on the Linux platform, and evaluate its performance with about 10 billions data records. Experimental results demonstrate that FastRAQ provides range-aggregate query results within a time period two orders of magnitude lower than that of Hive, while the relative error is less than 3 percent within the given confidence interval.
Xiao-chun Yun, Guangjun Wu, Guangyan Zhang, Keqin Li 0001
IEEE Trans. Cloud Comput.2
2014 MMD: An Approach to Improve Reading Performance in Deduplication Systems
abstract
The approach of data deduplication has been widely used in backup systems and primary storage such as virtual machine platform. However, the reading speed in those systems suffers due to chunk fragmentation in deduplication. So it has become an important problem to improve reading performance in deduplication systems. In this paper, firstly we propose a new storage method using multiple disks to boost reading performance, which is called MMD. MMD takes advantage of the multiple parallelized disks, each of which is used as independent logical device. Then we present a deduplication model based on MMD, which focuses on optimization of data layout on disks to improve reading speed. Two I/O scheduling algorithms in that model are discussed, which aim at assigning the containers in deduplication systems to appropriate disks. Experiments show that MMD can achieve an obvious reading performance improvement than RAID in deduplication systems.
Chao Li 0062, Xiao-chun Yun, Guangjun Wu
NAS5
2008 Design and Implementation of Multi-Version Disk Backup Data Merging Algorithm
abstract
Multi-version data management in disk backup and recovery is to manage the temporal attribute of backuped data. It can support to retrieve timestamp (time slice) disk data according to different query type. Exiting multi-version data management algorithms have two shortcomings. First, they are inefficient in multi-time point data query and updating which are adopted by data backup and recovery usually. Second, they use centralized data indexes which are not suitable for backup data management. To overcome these limitations, Backup Data Merging (BDM) algorithm is proposed in this paper, which uses distributed storage structure according to disk data format. By range operation, BDM algorithm can generate timestamp (time slice) data index dynamically. By comparing with traditional algorithms, BDM algorithm achieves high performance in storage utilization and query efficiency.
Guangjun Wu, Xiao-chun Yun
WAIM1