Han Wu 0001

dblp:13/1864-1 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-8870-5694ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Adaptive Mini-Batching Strategy for Reliable Streaming Data Delivery in Real-Time
abstract
Modern applications show an increasing demand for continuously processing massive data streams in real time. Mini-batching technique is commonly used for transporting streaming data across these applications. However, although a larger mini-batch size increases throughput, it also raises the end-to-end latency and easily violate the latency constraint required by real-time applications. While existing work mostly studies this problem at the computation stage, this work explores it in the streaming data transportation stage. We show that selecting a proper mini-batch size is essential for efficient streaming data delivery. We further identify the shortcomings of existing latency measurements and introduce new Quality of Service (QoS) metrics: latency violation rate, timely throughput, message loss, and duplicate rate. To address the challenge of mini-batching under varying network conditions, we first develop prediction models for the proposed QoS metrics and then adaptively adjust the mini-batch size based on these predictions. In our experiments, the random forest regressor achieves an$R^{2}$of 0.99 for performance metrics, and the multilayer perceptron achieves an MAE below 0.02 for reliability metrics. Using these predictions, the proposed adaptive strategy continuously updates the mini-batch size according to the observed network state. In a Kafka testbed with network packet loss rate approaching 25%, our strategy improves timely throughput by up to 30% compared with empirical mini-batch size selection. The results confirm the effectiveness of the adaptive mini-batching approach.
Han Wu 0001, Zhihao Shang, Huaming Wu, Katinka Wolter
IEEE Trans. Reliab.1
2025 GTV: Generating Tabular Data via Vertical Federated Learning
abstract
Synthetic data has emerged as a promising avenue for privacy-preserving data sharing. However, constructing synthetic data generators necessitates access to the real dataset, posing challenges, particularly when data features are disparately distributed across different organizations. Vertical Federated Learning (VFL) is a collaborative approach to training machine learning models among distinct tabular data holders, such as financial institutions, who possess disjoint features for the same group of customers. In this paper, we introduce the GTV framework for Generating Tabular Data via Vertical Federated Learning and demonstrate that VFL can be successfully used to implement GANs for distributed tabular data in a privacy-preserving manner, with performance close to centralized GANs which assume shared data. We make design choices with respect to the distribution of GAN generator and discriminator models, and we introduce a training-with-shuffling technique so that no party can reconstruct training data from the GAN conditional vector. The paper presents (1) an implementation of GTV, (2) a detailed quality evaluation of GTV-generated synthetic data, (3) an examination of the GTV framework for different data distributions and number of clients, and (4) an analysis of GTV’s robustness against Membership Inference Attacks with different settings of Differential Privacy, for a range of datasets with diverse distribution characteristics. Our results demonstrate that GTV can consistently generate high-fidelity synthetic tabular data of comparable quality to that generated by a centralized GAN algorithm. The difference in machine learning utility can be as low as 2.7%, even under extremely imbalanced data distributions across clients. Code is available at: https://github.com/zhao-zilong/gtv
Zilong Zhao 0001, Han Wu 0001, Aad P. A. van Moorsel, Lydia Y. Chen
DSN2
2025 Impacts of Physical-Layer Information on Epidemic Spreading in Cyber-Physical Networked Systems
abstract
Since Granell et al. proposed a multiplex network for information and epidemic propagation, researchers have explored how information propagation affects epidemic dynamics. However, the role of individuals acquiring information through physical interactions has received relatively less attention. In this work, we introduce a novel source of information: physical-layer information, and derive the epidemic outbreak threshold using the Microscopic Markov Chain Approach (MMCA). Our simulation results indicate that the outbreak threshold derived from the MMCA is consistent with the Monte Carlo (MC) simulation results, thereby confirming the accuracy of the theoretical model. Furthermore, we find that the physical-layer information effectively increases the population’s awareness density and the infection threshold$\beta _{c}$, while reducing the population’s infection density, thereby suppressing the spreading of the epidemic. Another interesting finding is that when the density of 2-simplex information is relatively high, the 2-simplex plays a role similar to pairwise interaction, significantly enhancing the population’s awareness density and effectively preventing large-scale epidemic outbreaks. In addition, our model works equally well for cyber-physical systems with similar interaction mechanisms, while we simulate and validate it in a real grid system.
Xianglai Yuan, Yichao Yao, Han Wu 0001, Minyu Feng
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots
abstract
In this research, we investigate the influence of utilizing artificial intelligence (AI)-generated photographs on malicious bots that engage in disinformation, fraud, reputation manipulation, and other types of malicious activity on social networks. Our research aims to compare the performance metrics of social bots that employ AI photos with those that use other types of photographs. To accomplish this, we analyzed a dataset with 13 748 measurements of 11 423 bots from the VK social network and identified 73 cases where bots employed generative adversarial network (GAN)-photos and 84 cases where bots employed diffusion or transformers photos. We conducted a qualitative comparison of these bots using metrics such as price, survival rate, quality, speed, and human trust. Our study findings indicate that bots that use AI-photos exhibit less danger and lower levels of sophistication compared to other types: AI-enhanced bots are less expensive, less popular on exchange platforms, of inferior quality, less likely to be operated by humans, and, as a consequence, faster and more susceptible to being blocked by social networks. We also did not observe any significant difference between GAN-based and diffusion/transformers-based bots, indicating that diffusion/transformers models did not contribute to increased bot sophistication compared to GAN models. Our contributions include a proposed methodology for evaluating the impact of photos on bot sophistication, along with a publicly available dataset for other researchers to study and analyze bots. Our research findings suggest a contradiction to theoretical expectations: in practice, bots using AI-generated photos pose less danger.
Maxim Kolomeets, Han Wu 0001, Lei Shi 0003, Aad P. A. van Moorsel
IEEE Trans. Comput. Soc. Syst.2
2023 Complex online harms and the smart home: A scoping review
abstract
Technological advances in the smart home have created new opportunities for supporting digital citizens’ well-being and facilitating their empowerment but have enabled new types of complex online harms to develop. Recent statistics have indicated that ‘smart’ technology ownership increases yearly, driven by lower costs and increased accessibility. Research on smart homes has also grown, focusing on technology perspectives at the expense of a user-centric approach sensitive to the smart home’s harms, risks, and vulnerabilities. This scoping review addresses the information gap by underscoring the scope of literature that exists regarding complex online harms, vulnerabilities, and risks associated with smart home technologies and citizens’ agency. The goal is to understand the state of knowledge, gaps in the literature, and areas for future study. The importance and originality of this paper lie in its interdisciplinary review and approach. It is hoped that this research will contribute to a deeper understanding of complex online harms in the smart home. Three online databases were utilised to identify papers published between 2017 and 2022, from which we selected 235 publications written in English that addressed harms, risks, vulnerabilities, and agency in the smart home context. This allowed us to map contemporary literature to reveal significant gaps in our understanding of the complex online harms affecting smart home users and identify opportunities for further research. This review identified emerging themes of ‘risks’, ‘vulnerabilities’, and ‘harms’ in that order of frequency within the literature on smart homes. The usage of terms is skewed towards computing science and information security, which comprised the majority of the literature at 54.6%. Human–computer interaction papers contributed 24.4%, while social sciences accounted for 16.2%. Risks, harms and vulnerabilities within smart home ecosystems and IoTs are ongoing issues with complexities that necessitate research. Privacy, security, and well-being are key themes that embody the scope of complex harms affecting smart home devices in the broad literature. This review establishes disciplinary research gaps, especially in user-centred perspectives, due to a heavy technology focus in the existing literature. Therefore, further research is needed to address emergent risks, harms and vulnerabilities of smart home devices and understand how user agency and autonomy can complement the design, interface, and socio-technical aspects of smart home systems.
Shola Olabode, Rebecca Owens, Viana Nijia Zhang, Jehana Copilah-Ali, Maxim Kolomeets, Han Wu 0001, Shrikant Malviya, Karolina Markeviciute, Tasos Spiliotopoulos, Cristina Neesham, Lei Shi 0003, Deborah Chambers
Future Gener. Comput. Syst.6
2022 CBlockSim: A Modular High-Performance Blockchain Simulator
abstract
To avoid the inconvenience of the deployment of large-scale blockchains, blockchain simulators are used to facilitate blockchain design and implementation. We evaluate state-of-the-art simulators and find that they suffer from low performance and scalability. To build a more general and faster blockchain simulator, we extend an existing blockchain simulator. We add a network module integrated with a network topology generation algorithm and a block propagation algorithm to simulate the block propagation efficiently. We design a binary transaction pool structure and adopt bitwise operations to accelerate the simulation and reduce memory usage. Moreover, we modularize the simulator based on five primary blockchain processes. Significant blockchain elements are implemented in individual modules and can be combined flexibly to simulate different types of blockchains. Experiments demonstrate that the new simulator reduces the simulation time by an order of magnitude and improves scalability, enabling us to simulate more than ten thousand nodes.
Xuyang Ma, Han Wu 0001, Du Xu, Katinka Wolter
ICBC2
2022 Federated Learning for Tabular Data: Exploring Potential Risk to Privacy
abstract
Federated Learning (FL) has emerged as a potentially powerful privacy-preserving machine learning method-ology, since it avoids exchanging data between participants, but instead exchanges model parameters. FL has traditionally been applied to image, voice and similar data, but recently it has started to draw attention from domains including financial services where the data is predominantly tabular. However, the work on tabular data has not yet considered potential attacks, in particular attacks using Generative Adversarial Networks (GANs), which have been successfully applied to FL for non-tabular data. This paper is the first to explore leakage of private data in Federated Learning systems that process tabular data. We design a Generative Adversarial Networks (GANs)-based attack model which can be deployed on a malicious client to reconstruct data and its properties from other participants. As a side-effect of considering tabular data, we are able to statistically assess the efficacy of the attack (without relying on human observation such as done for FL for images). We implement our attack model in a recently developed generic FL software framework for tabular data processing. The experimental results demonstrate the effectiveness of the proposed attack model, thus suggesting that further research is required to counter GAN-based privacy attacks.
Han Wu 0001, Zilong Zhao 0001, Lydia Y. Chen, Aad P. A. van Moorsel
ISSRE1
2021 Constrained Multiobjective Optimization for IoT-Enabled Computation Offloading in Collaborative Edge and Cloud Computing
abstract
Internet-of-Things (IoT) applications are becoming more resource-hungry and latency-sensitive, which are severely constrained by limited resources of current mobile hardware. Mobile cloud computing (MCC) can provide abundant computation resources, while mobile-edge computing (MEC) aims to reduce the transmission latency by offloading complex tasks from IoT devices to nearby edge servers. It is still challenging to satisfy the quality of service with different constraints of IoT devices in a collaborative MCC and MEC environment. In this article, we propose three constrained multiobjective evolutionary algorithms (CMOEAs) for solving IoT-enabled computation offloading problems in collaborative edge and cloud computing networks. First of all, a constrained multiobjective computation offloading model considering time and energy consumption is established in the mobile environment. Inspired by the push and pull search framework, three CMOEAs are developed by combing the advantages of population-based search algorithms with flexible constraint handling mechanisms. On one hand, three popular and challenging constrained benchmark suites are selected to test the performance of the proposed algorithms by comparing them to the other seven state-of-the-art CMOEAs. On the other hand, a multiserver multiuser multitask computation offloading experimental scenario with a different number of IoT devices is used to evaluate the performance of three proposed algorithms and other compared algorithms as well as representative offloading schemes. The experimental results of the benchmark suites and computation offloading problems demonstrate the effectiveness and superiority of the proposed algorithms.
Guang Peng, Huaming Wu, Han Wu 0001, Katinka Wolter
IEEE Internet Things J.3
2020 Learning to Reliably Deliver Streaming Data with Apache Kafka
abstract
The rise of streaming data processing is driven by mass deployment of sensors, the increasing popularity of mobile devices, and the rapid growth of online financial trading. Apache Kafka is often used as a real-time messaging system for many stream processors. However, efficiently running Kafka as a reliable data source is challenging, especially in the case of real-time processing with unstable network connection. We find that changing configuration parameters can significantly impact the guarantee of message delivery in Kafka. Therefore the key to solving the above problem is to predict the reliability of Kafka given various configurations and network conditions. We define two reliability metrics to be predicted, the probability of message loss and the probability of message duplication. Artificial neural networks (ANN) are applied in our prediction model and we select some key parameters, as well as network metrics as the features. To collect sufficient training data for our model we build a Kafka testbed based on Docker containers. With the neural network model we can predict Kafka's reliability for different application scenarios given various network environments. Combining with other metrics that a streaming application user may care for, a weighted key performance indicator (KPI) of Kafka is proposed for selecting proper configuration parameters. In the experiments we propose a rough dynamic configuration scheme, which significantly improves the reliability while guaranteeing message timeliness.
Han Wu 0001, Zhihao Shang, Katinka Wolter
DSN1
2020 Evolutionary Large-scale Sparse Multi-objective Optimization for Collaborative Edge-cloud Computation Offloading
Guang Peng, Huaming Wu, Han Wu 0001, Katinka Wolter
IJCCI3
2020 A Reactive Batching Strategy of Apache Kafka for Reliable Stream Processing in Real-time
abstract
Modern stream processing systems need to process large volumes of data in real-time. Various stream processing frameworks have been developed and messaging systems are widely applied to transfer streaming data among different applications. As a distributed messaging system with growing popularity, Apache Kafka processes streaming data in small batches for efficiency. However, the robustness of Kafka's batching method against variable operating conditions is not known. In this paper we study the impact of the batch size on the performance of Kafka. Both configuration parameters, the spatial and temporal batch size, are considered. We build a Kafka testbed using Docker containers to analyze the distribution of Kafka's end-to-end latency. The experimental results indicate that evaluating the mean latency only is unreliable in the context of real-time systems. In the experiments where network faults are injected, we find that the batch size affects the message loss rate in the presence of an unstable network connection. However, allocating resources for message processing and delivery that will violate the reliability requirements implemented as latency constraints of a real-time system is inefficient To address these challenges we propose a reactive batching strategy. We evaluate our batching strategy in both good and poor network conditions. The results show that the strategy is powerful enough to meet both latency and throughput constraints even when network conditions are variable.
Han Wu 0001, Zhihao Shang, Guang Peng, Katinka Wolter
ISSRE1