VLDB 2026 Research / reviewers in the wild / expert
Weizhe Zhang
dblp:z/WeizheZhang
· DBLP profile ↗
202ranked-venue papers
21as first author
154since 2021 · last 2026
0000-0003-4783-876XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 8 first-author · 40 since 2021Computer networks · 44 · 2 first-author · 36 since 2021Security and privacy · 25 · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 16 · 13 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaTurbo: A Scalable FPGA-based Engine for MegaFlow Classifier in Open vSwitchabstractOpen vSwitch (OVS) is a key component in cloud and data center networks, yet its MegaFlow classifier imposes significant CPU overhead. Existing SmartNIC-based acceleration approaches for the MegaFlow classifier typically employ simplistic hardware offloading techniques, which exhibit limited scalability for dynamic, large-scale flow tables. Motivated by these challenges, we argue that a hardware accelerator specifically tailored for the MegaFlow classifier is necessary, forming the basis of our FPGA-based solution, MegaTurbo. The core innovations of MegaTurbo are threefold: (1) a scalable and hardware-friendly decision-tree based packet classification algorithm, specifically optimized for the structure of MegaFlow rules; (2) a novel hardware architecture incorporating multiple pipelined matching engines, designed to process multiple decision trees generated by the software algorithm in parallel; and (3) a heterogeneous framework composed of CPU and FPGA, which can work together to support online rule updates, with little and bounded impact on rule searching. Experimental results on a Xilinx Virtex UltraScale+ FPGA demonstrate that MegaTurbo achieves a sustained classification throughput of 500 MPPS while supporting dynamic rule updates at 300-500 KUPS on 100K-scale rulesets. These results not only validate the effectiveness of our domain-specific co-design approach, but also highlight the potential of FPGA-based SmartNICs to address the performance bottlenecks of software switches in large-scale cloud and data center networks. Zhongxian Liang, Wenjun Li 0004, Yao Xin, Ying Wan 0001, Hui Li 0022, Weizhe Zhang |
FPGA | 8 |
| 2026 | Contribution-Aware Coalition Federated Learning in Edge-Assisted Healthcare Monitoring Systems
Hualong Wu, Weizhe Zhang, Desheng Wang 0002 |
ICC | 3 |
| 2026 | HypCC: A Hyperbolic Congestion Control Approach for Heterogeneous Flows in Datacenters
Heran Yang, Weizhe Zhang |
ICC | 3 |
| 2026 | PBSketch: Finding Periodic Burst Items in Data StreamsabstractDetecting periodic burst (PB) items in data streams is crucial for applications like rate limiting but remains unexplored. % While combining existing sketch algorithms offers a baseline, it suffers from significant inaccuracy and inefficiency. In this paper, we propose PBSketch, the first dedicated sketch algorithm designed for detecting PB items in real time. Its key techniques mainly include: 1) a two-stage hierarchical structure that efficiently maintains potential burst items and discards those without potential; 2) a fine-grained PB selection mechanism during window processing, coupled with the Window Smoothing Processing optimization to amortize performance overhead and eliminate processing spikes. % We provide its error bounds through rigorous theoretical analysis. Our extensive experiments show that PBSketch outperforms the baseline solution in accuracy and speed. By deploying it on an FPGA platform, the throughput is further significantly improved. Moreover, it effectively optimizes a practical application of rate limiting, clearly improving performance with almost negligible overhead. Zhuochen Fan, Zhongxian Liang, Zirui Liu 0002, Dayu Wang, Dong Wen 0004, Wenjun Li 0004, Tong Yang 0003, Yuzhou Liu 0001, Weizhe Zhang |
KDD (1) | 9 |
| 2026 | FlowTurbo: From Best-Effort to Hit-Driven MegaFlow Hardware Offloading in Open vSwitchabstractOffloading fast-path MegaFlows in Open vSwitch to hardware accelerators is a common approach for accelerating packet forwarding in modern cloud data centers. However, due to the limited capabilities of current hardware accelerators, existing solutions still rely on coarse-grained, best-effort offloading, which struggles with dynamic, large-scale traffic and results in inefficient resource utilization and limited performance gains. We present FlowTurbo, a self-adaptive, system-level offloading approach that implements hit-driven MegaFlow hardware offloading by jointly optimizing software rule scheduling and hardware rule lookup. The core innovations of FlowTurbo are threefold: (1) a traffic-aware, hit-driven MegaFlow offloading framework that selectively migrates hotspot wildcard rules to hardware; (2) a domain-specific, hardware-friendly sketch for MegaFlow rules that tracks rule hotness and enables the scheduler to make timely and precise offloading decisions; and (3) a domain-specific, algorithm-hardware co-designed packet classification accelerator that supports both line-rate rule matching and online rule updates. We implemented FlowTurbo on Open vSwitch and prototyped its hardware accelerator on a Xilinx Alveo U200. Evaluation using multiple real-world traffic traces shows that FlowTurbo achieves an average acceleration coverage of 89.4%, and the hardware accelerator delivers a maximum throughput of 400 MOPS while consuming only 3.3% of FPGA logic resources. Zhongxian Liang, Wenjun Li 0004, Yao Xin, Tong Yang 0003, Gaogang Xie, Weizhe Zhang |
SIGCOMM | 12 |
| 2026 | JitterSketch: Finding Jittery Flows in Network StreamsabstractIn the modern internet, with the proliferation of real-time applications such as online gaming and video conferencing, the timely detection of network jitter has become a critical task in network measurement. Network jitter is defined as the abrupt fluctuations in packet inter-arrival times within network flows, which severely degrade the Quality of Service for these applications. Traditional jitter detection methods primarily focus on macro-level end-to-end or hop-by-hop latency variations, neglecting the fine-grained jitter that occurs within specific flows. In this paper, we present JitterSketch, the first sketch-based algorithm specifically designed for detecting jittery flows. JitterSketch employs a novel three-stage structure to efficiently filter out infrequent and stable flows, thereby identifying and reporting the jittery flows that have the most significant impact on network quality. Extensive experiments demonstrate that JitterSketch achieves an improvement of up to 50 percentage points in both recall and precision rates compared to baseline solutions, while maintaining high processing throughput. Furthermore, we deployed JitterSketch in a QoS simulation system, where it yielded significant improvements in QoS. Zhongxian Liang, Qilong Shi, Xiyan Liang, Wenjun Li 0004, Tong Yang 0003, Yangyang Wang 0001, Mingwei Xu 0001, Weizhe Zhang |
WWW | 9 |
| 2026 | A Multi-Level Acceleration Scheme for AI Model Training on ARM Architecture ProcessorsabstractABSTRACT With the widespread application of reinforcement learning and deep learning on edge devices, training neural networks on ARM architecture processors has become an urgent demand. However, existing mainstream deep learning frameworks are not sufficiently optimized for training workloads on ARM CPUs, resulting in low training efficiency. To address this problem, this paper proposes and implements a multi‐level AI training acceleration scheme based on the open‐source C++ library mlpack, targeting the slow training speed and high resource consumption of Convolutional Neural Networks (CNNs) on ARM platforms. The scheme accelerates training by deeply optimizing the im2col algorithm to convert convolutions into efficient matrix multiplications, utilizing ARM NEON SIMD instructions to optimize linear operators, and integrating an FP64/FP16 mixed‐precision training strategy with dynamic loss scaling. Experimental results on LeNet‐5 and VGG11‐style CNNs show substantial performance gains over the original mlpack and mainstream frameworks. On an NVIDIA Jetson AGX Orin, our implementation achieves up to 7.3 speedup over the original mlpack baseline and up to 11.3 × and 5.69 × end‐to‐end speedups over PyTorch and TensorFlow, respectively, while still delivering multi‐fold reductions in training time on a low‐resource Raspberry Pi platform. In DQN‐based reinforcement learning for Atari Breakout, our solution attains a 4.98 × end‐to‐end speedup over the PyTorch single‐threaded baseline and maintains 2.37 × and 4.23 × advantages over 4‐threaded PyTorch and TensorFlow implementations. Ablation studies confirm the complementary nature of the proposed optimizations, with convolutional, linear, and mixed‐precision components jointly contributing to the overall speedup and enabling an attractive performance–accuracy trade‐off for ARM‐based edge computing. Mingdong Xie, Meng Hao 0002, Weizhe Zhang, NingCheng Wang |
Concurr. Comput. Pract. Exp. | 4 |
| 2026 | How to bridge spatial and temporal heterogeneity in link prediction? A contrastive method
Yu Tai, Weizhe Zhang |
Inf. Sci. | 4 |
| 2026 | Performance Prediction of Concurrent DNN Training Tasks in GPU Spatial Sharing EnvironmentsabstractGPU sharing is commonly employed in GPU clusters to improve utilization, with spatial sharing being one of the most widely adopted techniques. However, spatial sharing can lead to resource interference, making task execution times difficult to predict. Predictable execution times for each task are crucial in GPU cluster management and task scheduling. In this article, we propose a performance predictor for multi-DNN training tasks in GPU spatial sharing environments. We first conduct experiments on spatial sharing for multiple DNN workloads on a single GPU, demonstrating that concurrent execution of multiple tasks improves overall performance and GPU resource utilization compared to serial execution. By analyzing warp stall reasons collected during task execution, we investigate the interference for computation and memory resources under MPS on GPUs. Finally, we design a performance predictor that predicts the execution time of a target DNN training task when it runs concurrently with other tasks under GPU spatial sharing via MPS. The predictor is capable of predicting the execution time of each task for previously unseen combinations of DNN training tasks. Extensive evaluations on modern GPUs show that compared to other baseline methods, our approach exhibits higher prediction accuracy, as well as improved stability and robustness. Experiments on multiple GPU architectures, as well as at higher concurrency levels, further demonstrate that our method possesses strong generalization and scalability. We also conducted a performance analysis under diverse workload pattern and a case study to validate the practical applicability of our predictor in real scheduling environments. Sichao Chen, Desheng Wang 0002, Weizhe Zhang, Meng Hao 0002, Yu-Chu Tian |
ACM Trans. Archit. Code Optim. | 3 |
| 2026 | PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power ConstraintsabstractThe proliferation of deep learning inference services in power-constrained environments necessitates GPU management strategies that maximize throughput within strict power envelopes. Existing approaches often treat frequency scaling and resource partitioning as orthogonal problems or rely on static hardware assumptions, leading to suboptimal energy efficiency. This article presents PctoDL , a power-aware scheduling system that maximizes aggregate inference throughput by jointly optimizing spatial resource partitioning, batch size, and SM/memory frequency settings. To address the throughput–power tradeoff in power-constrained multi-tenant inference, PctoDL couples resource partitioning with coordinated frequency control under a fixed power cap. It combines a physics-informed iterative greedy partitioning algorithm, a thermodynamic model-predictive controller for runtime frequency regulation, and an online joint optimization mechanism for adaptive refinement. On the NVIDIA RTX 3080 Ti platform, PctoDL improves average throughput over BatchDVFS by 108.41%, with a peak gain of 262.74%. On the NVIDIA A100 platform, it delivers an average gain of 19.74% and a maximum gain of 57.03%. Compared with Morak’s coarse-grained partitioning approach, PctoDL achieves average/peak gains of 79.05%/137.93% on the RTX 3080 Ti and 26.33%/70.21% on the A100. Meng Hao 0002, Zikun Wu, Xueyang Tian, Siyu Yang 0002, Guotong Guo, Yiming Wang 0010, Farui Wang, Desheng Wang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 11 |
| 2026 | GreenDLS: An Energy-Efficient and SLO-Aware Deep Learning Serving SystemabstractThe growing demand for deploying deep learning (DL) models, particularly large language models (LLMs), has made it imperative to optimize GPU energy consumption while meeting service-level objectives (SLOs). Significant energy use and carbon dioxide (CO2) emissions from GPU-based inference tasks contribute substantially to the environmental footprint of the DL deployment. Existing approaches primarily rely on batching and dynamic voltage and frequency scaling (DVFS) to optimize service performance or throughput, but often overlook memory frequency adjustments and holistic energy optimization under dynamic workloads. This study presents GREENDLS, a DL serving system that integrates deep reinforcement learning (DRL) with offline prediction models to optimize energy consumption while adhering to inference latency SLOs, achieving significant energy savings. GREENDLS dynamically adjusts batch size, GPU streaming multiprocessor (SM) frequency, and GPU memory frequency based on inference request rates. It also accounts for GPU energy consumption during idle phases, such as batch filling, enabling multi-parameter and fine-grained energy optimization. Compared to the Clipper system, GREENDLS achieves energy savings of up to 45.57% on the RTX 3080Ti and 39.44% on the Tesla V100S. When compared to the EAIS system, which only combines batching with GPU SM frequency adjustment, GREENDLS achieves energy savings of up to 36.44% on the RTX 3080Ti and 15.44% on the Tesla V100S. Against the method proposed by Yu et al., GREENDLS attains energy savings of up to 46.34% on the RTX 3080Ti and 33.74% on the Tesla V100S. In LLM inference tasks using Qwen, GREENDLS reduces average energy consumption by 40.79% compared to Clipper, 10.42% compared to EAIS, and 42.42% compared to Yu et al. These results clearly demonstrate that GREENDLS more effectively optimizes energy consumption compared to traditional methods that rely primarily on batching or a combination of batching and DVFS, while still ensuring SLO compliance. Meng Hao 0002, Xueyang Tian, Siyu Yang 0002, Yiming Wang 0010, Desheng Wang 0002, Weizhe Zhang |
IEEE Trans. Computers | 10 |
| 2026 | AudioJailbreak: Jailbreak Attacks Against End-to-End Large Audio-Language ModelsabstractJailbreak attacks to Large audio-language models (LALMs) are studied recently, but they exclusively focused on the attack scenario where the adversary can fully manipulate user prompts (named strong adversary) and limited in effectiveness, applicability, and practicability. In this work, we first conduct an extensive evaluation showing that advanced text jailbreak attacks cannot be easily ported to end-to-end LALMs via text-to-speech (TTS) techniques. We then propose AUDIOJAILBREAK, a novel audio jailbreak attack, featuring (1) asynchrony: the jailbreak audios do not need to align with user prompts in the time axis by crafting suffixal jailbreak audios; (2) universality: a single jailbreak perturbation is effective for different prompts by incorporating multiple prompts into the perturbation generation; (3) stealthiness: the malicious intent of jailbreak audios is concealed by proposing various intent concealment strategies; and (4) over-the-air robustness: the jailbreak audios remain effective when being played over the air by incorporating reverberation into the perturbation generation. In contrast, all prior audio jailbreak attacks cannot offer asynchrony, universality, stealthiness, and/or over-the-air robustness. Moreover, AUDIOJAILBREAK is also applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary). Extensive experiments with thus far the most LALMs demonstrate the high effectiveness of AUDIOJAILBREAK, in particular, it can jailbreak openAI's GPT-4o-Audio and bypass Meta's Llama-Guard-3 safeguard, in the weak adversary scenario. We highlight that our work peeks into the security implications of audio jailbreak attacks against LALMs, and realistically fosters improving their robustness, especially for the newly proposed weak adversary. Guangke Chen, Fu Song, Zhe Zhao 0007, Xiaojun Jia, Yang Liu 0003, Yanchen Qiao, Weizhe Zhang, Weiping Tu, Yuhong Yang 0001, Bo Du 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | Accelerating Secure Machine Learning Training on GPUs With Pipeline ParallelismabstractThe proliferation of data-driven machine learning (ML) applications makes privacy and data security risks increasingly prominent. Secure multi-party computation (MPC) offers a privacy-preserving ML method by enabling joint model training without data disclosure, but its reliance on complex cryptography adds significant delay and resource demands, challenging its practicality in large-scale applications. To address the low GPU utilization within the MPC framework, we propose an innovative privacy-preserving ML training optimization framework that exploits pipeline parallelism. Through a detailed analysis of the underlying principles of MPC-based training, we identify computation and communication as the primary bottlenecks for linear and non-linear computations, respectively. Drawing from traditional ML optimization strategies, we design a sub-network partitioning and pipeline parallelism method, specifically tailored for MPC training. This method not only allows for simultaneous training computations across different network layers, but also strategically overlaps computation with communication to enhance GPU utilization and reduce latency. Additionally, we develop a distributed communication mechanism to further improve communication efficiency. We integrate our framework into two distinct, state-of-the-art secure training frameworks: CryptGPU and Piranha. Compared to their original versions, our enhancements boost training speeds by up to 51%, significantly increasing GPU utilization, with negligible impact on model convergence speed and accuracy. Meng Hao 0002, Mingdong Xie, Weizhe Zhang, Linxuan Wang, Xueyang Tian, Desheng Wang 0002 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Privacy-Preserving Federated Learning Scheme With Mitigating Model Poisoning Attacks: Vulnerabilities and CountermeasuresabstractThe privacy-preserving federated learning schemes based on the setting of two honest-but-curious and non-colluding servers offer promising solutions in terms of security and efficiency. However, our investigation reveals that these schemes still suffer from privacy leakage when considering model poisoning attacks from malicious users. Specifically, we demonstrate that the privacy-preserving computation process for defending against model poisoning attacks inadvertently leaks privacy to one of the honest-but-curious servers, enabling it to access users' gradients in plaintext. To address this issue, we propose an enhanced privacy-preserving and Byzantine-robust federated learning (PBFL) framework that simultaneously achieves privacy, robustness, and efficiency. Central to our design is a novel Byzantine-tolerant aggregation strategy that defends against both conventional and adaptive poisoning attacks. It integrates normalization judgment, cosine similarity computation, and adaptive user weighting, with a dual-scoring trust mechanism and outlier suppression for stealthy attacks. In addition, we develop two privacy-preserving subroutines, namely secure normalization judgment and secure cosine similarity measurement, which operate over encrypted gradients using a trapdoor fully homomorphic encryption (FHE) scheme, ensuring both confidentiality and robust aggregation correctness. Theoretical analyses confirm that our scheme guarantees security, convergence, and efficiency even with malicious users and one malicious server. Extensive experiments demonstrate that our method effectively breaks prior privacy attacks, maintains high accuracy under diverse poisoning strategies, and significantly reduces computation and communication overhead compared to state-of-the-art PBFL schemes. Jiahui Wu 0001, Tiecheng Sun, Haiyan Wang 0009, Weizhe Zhang |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Toward Generalizable Deepfake Detection via Forgery-Aware Audio-Visual Adaptation: A Variational Bayesian ApproachabstractThe widespread application of AIGC contents has brought not only unprecedented opportunities, but also potential security concerns, e.g., audio-visual deepfakes. Therefore, it is of great importance to develop an effective and generalizable method for multi-modal deepfake detection. Typically, the audio-visual correlation learning could expose subtle cross-modal inconsistencies, e.g., audio-visual misalignment, which serve as crucial clues in deepfake detection. In this paper, we reformulate the correlation learning with variational Bayesian estimation, where audio-visual correlation is approximated as a Gaussian distributed latent variable, and thus develop a novel framework for deepfake detection, i.e., Forgery-aware Audio-Visual Adaptation with Variational Bayes (FoVB). Specifically, given the prior knowledge of pre-trained backbones, we adopt two core designs to estimate audio-visual correlations effectively. First, we exploit various difference convolutions and a high-pass filter to discern local and global forgery traces from both modalities. Second, with the extracted forgery-aware features, we estimate the latent Gaussian variable of audio-visual correlation via variational Bayes. Then, we factorize the variable into modality-specific and correlation-specific ones with orthogonality constraint, allowing them to better learn intra-modal and cross-modal forgery traces with less entanglement. Extensive experiments demonstrate that our FoVB outperforms other state-of-the-art methods in various benchmarks. Fan Nie, Jiangqun Ni, Jian Zhang 0086, Bin Zhang 0048, Weizhe Zhang, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | XPCH: A Cross-Chain Payment Protocol via Connecting the Payment Channel Hubs
Yuanming Shao, Yuming Feng 0002, Weizhe Zhang, Bin Xiao 0001, Jianhuan Wang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Detection and Mitigation Data Poisoning Attacks in Multimodal Online Federated Learning
Heqiang Wang, Xiaoxiong Zhong, Hualong Wu, Fangming Liu, Weizhe Zhang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | Dual-Thresholded Heatmap-Guided Proposal Clustering and Negative Certainty Supervision With Enhanced Base Network for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) has attracted significant attention in recent years, as it does not require box-level annotations. State-of-the-art methods generally adopt a multi-module network, which employs WSDDN as the multiple instance detection network module and uses multiple instance refinement modules to refine performance. However, these approaches suffer from three key limitations. First, existing methods tend to generate pseudo GT boxes that either focus only on discriminative parts, failing to capture the whole object, or cover the entire object but fail to distinguish between adjacent intra-class instances. Second, the foundational WSDDN architecture lacks a crucial background class representation for each proposal and exhibits a large semantic gap between its branches. Third, prior methods discard ignored proposals during optimization, leading to slow convergence. To address these challenges, we propose the Dual-thresholded heAtmap-guided proposal clustering and Negative Certainty supervision with Enhanced base network (DANCE) method for WSOD. Specifically, we first devise a heatmap-guided proposal selector (HGPS) algorithm, which utilizes dual thresholds on heatmaps to pre-select proposals, enabling pseudo GT boxes to both capture the full object extent and distinguish between adjacent intra-class instances. We then construct a weakly supervised basic detection network (WSBDN), which augments each proposal with a background class representation and uses heatmaps for pre-supervision to bridge the semantic gap between matrices. At last, we introduce a negative certainty supervision (NCS) loss on ignored proposals to accelerate convergence. Extensive experiments on the challenging PASCAL VOC and MS COCO datasets demonstrate the effectiveness and superiority of our method. Our code is publicly available at https://github.com/gyl2565309278/DANCE. Yuelin Guo, Haoyu He 0001, Zitong Huang, Renhao Lu, Lu Shi 0002, Weizhe Zhang |
IEEE Trans. Image Process. | 8 |
| 2026 | Filtering and Accelerating: A Unified Framework for High-Performance Persistence EstimationabstractEfficient data stream processing, particularly for persistence estimation, is crucial in handling high-velocity data streams characterized by skewed distributions of item frequencies. Unlike more straightforward frequency metrics, persistence captures items' recurrence across multiple time windows, posing a significant challenge to existing single-structure sketches where high-persistence and low-persistence items collide. To address this, we introduce the Hypersistent Sketch, a unified framework for high-performance estimation built on two decoupled mechanisms: filtering and accelerating. The filtering component, a Cold Filter, directly addresses the skewed nature of data streams. It separates hot items from the majority of cold ones, which allows for differential treatment. The accelerating component, a Burst Filter, then optimizes the processing of hot items. It significantly improves throughput by preventing repeated insertions within a single window. We demonstrate its generality by applying it to various state-of-the-art sketches (e.g., On-Off, Waving, P-Sketch), showing it consistently enhances their original performance. We also deploy our framework on Redis platforms, demonstrating the framework’s broad applicability and scalability. Qilong Shi, Weiqiang Xiao, Nianfu Wang, Wenjun Li 0004, Tong Yang 0003, Zhijun Li 0002, Weizhe Zhang, Mingwei Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2026 | Multimodal Online Federated Learning With Modality Missing in Internet of ThingsabstractThe Internet of Things (IoT) ecosystem generates vast amounts of multimodal data from heterogeneous sources such as sensors, cameras, and microphones. As edge intelligence continues to evolve, IoT devices have progressed from simple data collection units to nodes capable of executing complex computational tasks. This evolution necessitates the adoption of distributed learning strategies to effectively handle multimodal data in an IoT environment. Furthermore, the real-time nature of data collection and limited local storage on edge devices in IoT call for an online learning paradigm. To address these challenges, we introduce the concept of Multimodal Online Federated Learning (MMO-FL), a novel framework designed for dynamic and decentralized multimodal learning in IoT environments. Building on this framework, we further account for the inherent instability of edge devices, which frequently results in missing modalities during the learning process. We conduct a comprehensive theoretical analysis under both complete and missing modality scenarios, providing insights into the performance degradation caused by missing modalities. To mitigate the impact of modality missing, we propose the Prototypical Modality Mitigation (PMM) algorithm, which leverages prototype learning to effectively compensate for missing modalities. Experimental results on two multimodal datasets further demonstrate the superior performance of PMM compared to benchmarks. Heqiang Wang, Xiang Liu 0004, Xiaoxiong Zhong, Lixing Chen, Fangming Liu, Weizhe Zhang |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Denoising and Adaptive Online Vertical Federated Learning for Sequential Multi-Sensor Data in IIoTabstractWith the advancement of computational capabilities in edge devices such as intelligent sensors in the Industrial Internet of Things (IIoT), these sensors evolving beyond simple data collection to support complex computational tasks. This advancement provides new opportunities for adopting distributed learning approaches in IIoT. In this study, we focus on enhancing learning performance in an industrial assembly line scenario where multiple distributed sensors sequentially collect real-time data with distinct feature spaces. However, existing research lacks an online distributed learning framework tailored for such IIoT settings. To address this gap, we propose the Denoising and Adaptive Online Vertical Federated Learning (DAO-VFL) algorithm, a novel algorithm that leverages the computing potential of edge sensors while addressing key challenges such as communication overhead and data privacy. DAO-VFL effectively manages continuous data streams and adapts to shifting learning objectives. Furthermore, it can address critical challenges prevalent in industrial environment, such as communication noise and heterogeneity of sensor capabilities. To support the proposed algorithm, we provide a comprehensive theoretical analysis, highlighting the effects of noise reduction and adaptive local iteration decisions on the regret bound. Experimental results on two real-world datasets further demonstrate the superior performance of DAO-VFL compared to benchmarks. Heqiang Wang, Xiaoxiong Zhong, Fangming Liu, Weizhe Zhang |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Edge-Assisted Real-Time Dynamic 3D Point Cloud Rendering for Multi-Party Mobile Virtual RealityabstractMulti-party Mobile Virtual Reality (MMVR) enables multiple mobile users to share virtual scenes for an immersive multimedia experience in scenarios such as gaming, social interaction, and industrial mission collaboration. Dynamic 3D Point Cloud (DPCL) is an emerging representation form of MMVR that can be consumed as a free-viewpoint video with 6 degrees of freedom. With limited on-device resources, it is a challenge to achieve a satisfying frame rate for DPCL rendering, which makes edge-assisted rendering a practical solution. However, repeated loading of DPCL scenes with a substantial amount of metadata introduces a significant redundancy overhead that cannot be overlooked when enabling multiple edge servers to support the rendering requirements of user groups. In this paper, we design PoClVR, an edge-assisted DPCL rendering system for MMVR applications, which introduces an object-level splitting mode to alleviate performance bottlenecks caused by redundant loading. In addition, PoClVR dynamically selects the splitting mode and scheduling decisions to adapt to varying task requirements and available computational resources, thereby improving overall system efficiency. To evaluate the performance of PoClVR, we implement and deploy a realistic prototype system and also conduct large-scale trace-driven simulations. The experimental results show that PoClVR can reduce resource usage by up to approximately 49.3% under different task requirements and resource conditions, while decreasing bottleneck performance degradation by up to 77.3%. Ximing Wu, Kongyange Zhao, Xu Chen 0004, Teng Liang, Weizhe Zhang |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Support Vector Machine Aided Semi-intelligent Performance Enhancement of CMOS Gate-voltage Bootstrapped Sampling SwitchabstractAs a forefront core functional module, a bootstrapped sampling switch (BSS) built-in an analog-digital converter (ADC) contributes to realize a high-precision signal sampling in bioelectrical sensing systems. A novel algorithm-based automatic approach to guide the optimal design of a high-performance bootstrapped sampling switch is proposed. Within the first manual topology optimization, a complementary sampling transfer-gate is designed to suppress the clock feed-through effect, and a dynamic body bias control module and an improved bootstrapped closed-loop-path are constructed to effectively improve the linearity. In the secondary core algorithm-based performance solving stage, a support vector regression (SVR) machine is adopted to further explore the best design tradeoff between dynamic noise feature and power dissipation according to the model-training-based optimal solution of the design parameters. SMIC 180 nm/1.8 V standard CMOS technology is employed to implement the front/back-end design of the proposed BSS circuit, and pre-/post-layout simulations for feature verification are performed. Final experimental results show that, targeting a test benchmark signal of 100 Hz and 1.2 V p-p in 51.2 kHz sampling frequency, after mixed-optimization-based ENOB and SNR can be reached up to 12.969 bits and 103.046 dB, respectively. Similar to the other two key dynamic performance indexes, SFDR and THD are also improved to 70.905 dBc and -69.508 dB, respectively. In the design case comparison, with a higher power supply voltage of 1.8 V and kHz-order frequency, the average power consumption is only approximately 0.188 μW. These feature results demonstrated the significant effectiveness of the proposed SVR algorithm aided artificial optimization approach on bootstrapped sampling switch design, and the comprehensive specification of switches can meet the demand of the specific application of highly accurate human bioelectrical signal sampling. Bo Liu 0031, Pengshuai Dong, Weizhe Zhang, Qingduan Meng, Jun Wang 0064 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2026 | MigrRDMA: Enabling RDMA Live Migration in the SoftwareabstractLive migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA has been widely adopted in data centers, and has attracted both academia and industry for years. However, live migration of RDMA is not supported in today’s data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, it relies on extra hardware support. This paper proposes MigrRDMA, a software-based RDMA live migration system. MigrRDMA provides a software indirection layer to achieve transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA’s indirection layer focuses on keeping the RDMA states on the migration source and destination identical from the perspective of applications. We implemented the MigrRDMA prototype over Mellanox RNICs. Our evaluation shows that MigrRDMA adds little downtime when migrating a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 2% ∼ 9% extra overhead in the data path. When migrating Hadoop tasks, MigrRDMA only incurs an extra 3-second job completion time. Weizhe Zhang, Fengyuan Ren |
IEEE Trans. Netw. | 2 |
| 2026 | Defining and Detecting the Defects of Large Language Model-Based Autonomous AgentsabstractArtificial intelligence (AI) agents are systems capable of perceiving their environment, autonomously planning and executing tasks. Recent advancements in Large Language Models (LLMs) have introduced a transformative paradigm for AI agents, enabling them to interact with external resources and tools through prompt techniques. This advancement has significantly extended the capabilities of LLMs, positioning LLM-based AI Agents as an important research area. In such agents, the workflow integrates developer-written code, which manages framework construction and logic control, with LLM-generated natural language that enhances dynamic decision-making and interaction. However, inconsistencies between LLM outputs and developer logic can lead to defects, such as tool invocation failures. These issues introduce specific risks, leading to various defects in LLM-based AI Agents, including service interruptions and incorrect output. Despite the importance of these issues, there is a lack of systematic work that focuses on analyzing LLM-based AI Agents to uncover defects in their code. To address this gap, we present the first study focused on identifying and detecting defects in LLM Agents. We collected and analyzed 14,754 relevant developer reports from StackOverflow and GitHub. We further filtered 2,604 valid posts to define and classify eight types of agent code defects. Then, we designed a static analysis tool, named Agentable, to detect these defects. Agentable leverages Code Property Graphs (CPGs) and LLMs to analyze Agent workflows by efficiently identifying specific code patterns and analyzing natural language descriptions. To evaluate Agentable, we constructed two datasets: AgentSet, which consists of 84 real world Agent projects, and AgentTest, which contains 78 Agent projects specifically designed to include various types of defects. Our evaluation shows that Agentable achieves a precision of 88.79% on the real-world agent dataset and a recall of 91.03% on the manually labeled defect dataset. Furthermore, our analysis identifies 889 defects in real-world agent projects, highlighting the prevalence of these issues in practice. Kaiwen Ning, Jiachi Chen, Wei Li 0121, Zexu Wang, Yuming Feng 0002, Weizhe Zhang, Zibin Zheng |
IEEE Trans. Software Eng. | 7 |
| 2025 | LPIA: Label Preference Inference Attack Against Federated Graph Learning
Jiaxue Bai, Lu Shi 0002, Yang Liu 0039, Weizhe Zhang |
ACISP (3) | 4 |
| 2025 | LossControl: Defending Membership Inference Attacks by Controlling the LossabstractMachine learning models are vulnerable to membership inference attacks (MIAs), where adversaries attempt to predict whether specific samples are part of the model’s training set. Previous studies have demonstrated a strong correlation between the distinguishability of training and testing loss distributions and the model’s susceptibility to MIAs. Motivated by existing results, we propose a novel training framework called LossControl, which focuses on manipulating loss to mitigate privacy leaks. In LossControl, we first utilize Soft-label Training to replace the general learning process, which facilitates model training while improving generalization. Next, we monitor overfitting samples during the training process and prevent further loss reduction by applying our designed Loss Ascent to these samples without sacrificing model performance. Through extensive evaluations across four diverse datasets (including images, medical data, and transaction records), our method consistently outperforms defense mechanisms against state-of-the-art attacks and achieves optimal model performance in most experiments, demonstrating LossControl’s superior resilience against MIAs and its ability to strike an unparalleled balance between privacy and utility. Renhao Lu, Weizhe Zhang, Haoyu He 0001 |
ICASSP | 5 |
| 2025 | Hypersistent Sketch: Enhanced Persistence Estimation via Fast Item SeparationabstractEfficient data stream processing, particularly for persistence estimation, is crucial in handling high-velocity data streams characterized by skewed distributions of item frequencies. Unlike more straightforward frequency metrics, persistence captures items' recurrence across multiple time windows, requiring nuanced processing approaches. In response, we introduce the Hypersistent Sketch, an algorithm that significantly enhances persistence estimation through innovative filtering techniques. Our design incorporates a Cold Filter to address the skewed nature of data streams where a few high-frequency (hot) items dominate. This filter allows for differential treatment by using smaller counters for most low-frequency (cold) items, thus conservatively allocating memory resources that would otherwise be sized uniformly based on hot items. However, the Cold Filter can reduce throughput due to its segregative processing. To mitigate this, we implement a Burst Filter, which optimizes the processing of hot items. The Burst Filter significantly improves throughput by preventing repeated insertions within a single window—where persistence increases by at most one—and deferring the insertion until the window's end. Comparative evaluations demonstrate that the Hypersistent Sketch outperforms existing solutions like the On-Off Sketch, offering up to 3 times improved throughput while maintaining competitive accuracy and substantially reducing memory usage in handling large-scale data streams. Qilong Shi, Weiqiang Xiao, Nianfu Wang, Wenjun Li 0004, Zhijun Li 0002, Weizhe Zhang, Mingwei Xu 0001 |
ICDE | 7 |
| 2025 | Circumventing Backdoor Space via Weight SymmetryabstractDeep neural networks are vulnerable to backdoor attacks, where malicious behaviors are implanted during training. While existing defenses can effectively purify compromised models, they typically require labeled data or specific training procedures, making them difficult to apply beyond supervised learning settings. Notably, recent studies have shown successful backdoor attacks across various learning paradigms, highlighting a critical security concern. To address this gap, we propose Two-stage Symmetry Connectivity (TSC), a novel backdoor purification defense that operates independently of data format and requires only a small fraction of clean samples. Through theoretical analysis, we prove that by leveraging permutation invariance in neural networks and quadratic mode connectivity, TSC amplifies the loss on poisoned samples while maintaining bounded clean accuracy. Experiments demonstrate that TSC achieves robust performance comparable to state-of-the-art methods in supervised learning scenarios. Furthermore, TSC generalizes to self-supervised learning frameworks, such as SimCLR and CLIP, maintaining its strong defense capabilities. Our code is available at https://github.com/JiePeng104/TSC. Jie Peng 0009, Hengji Dong, Weizhe Zhang, Haoyu He 0001 |
ICML | 6 |
| 2025 | ServerlessLego: An Elastic Serverless Framework Assembling Model Building Blocks to Provide SLO-Aware Inference ServicesabstractInference of large language models (LLMs) is common in cloud environments. As the elastic resource management capabilities and the flexible pay-as-you-go billing model offered by serverless, LLM inference services are increasingly migrated to serverless platforms. However, the increasing size of LLMs in recent years has introduced a new cold start issue for serverless frameworks, which in turn impacts their scalability under dynamic workloads. To address these issues, we propose ServerlessLego, an elastic serverless computing framework. ServerlessLego partitions LLMs into layers, then groups and deploys them to different instances, and loads these groups in parallel. These instances perform a subscription-based pipeline. To address dynamically request loads, ServerlessLego models the incoming request patterns and the inference time of running requests, providing an SLO-Aware instance scheduling. Experiments show that ServerlessLego reduces the cold start time of serverless frameworks by 58.15 % and improves throughput by 43.39 % compared to the baseline for dynamic workloads. Moreover, ServerlessLego can horizontally schedule instance based on request SLOs and arrival rates. Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Yuming Feng 0002 |
ICPADS | 3 |
| 2025 | Aegis Sketch: High-Throughput and Accurate Top-$k$ Elephant Flows Detection in Large-Scale Parallel Network Traffic ProcessingabstractIn large-scale parallel network traffic processing systems, detecting Top-$k$elephant flows is essential for real-time monitoring and traffic management. However, under the dual constraints of high update rates and limited memory, existing approaches struggle to balance throughput and accuracy. Many fail to exploit the heavy-tailed distribution of network traffic, resulting in frequent hash collisions and irreversible eviction errors that severely limit detection performance. Current methods fall into two categories: counter-based approaches, which maintain a candidate set (e.g., a min-heap) for high accuracy but suffer from high synchronization overhead and unrecoverable evictions, and sketch-based approaches, which are naturally parallelizable but prone to accuracy loss under hash collisions. To address these challenges, we propose Aegis Sketch, a novel framework for high-throughput and accurate Top-$k$elephant flow detection in large-scale parallel network traffic processing. Aegis Sketch incorporates an ordered storage mechanism that fundamentally reduces hash collisions without incurring additional structural overhead, thereby significantly improving throughput and accuracy. In addition, a multi-party competitive replacement policy prioritizes the preservation of true elephant flows during contention, effectively mitigating the accuracy loss caused by irreversible evictions. Experiments on real-world traffic traces show that Aegis Sketch outperforms existing methods such as Elastic Sketch and OneSketch in throughput, detection accuracy, and memory efficiency, achieving up to 1.35 times higher throughput and 22 % higher precision. These results demonstrate its effectiveness and efficiency for large-scale parallel network traffic processing. Desheng Wang 0002, Weizhe Zhang |
ICPADS | 4 |
| 2025 | HyDLR: Load-Aware Dynamic Rescheduling for Deep Learning Hybrid DeploymentabstractResource contention, driven by traffic surges from online services, presents a significant challenge in hybrid clusters where latency-sensitive and best-effort deep learning tasks are colocated. To address this, we propose HyDLR, a dynamic, loadaware hybrid deployment scheduling method that dynamically reallocates offline tasks to ensure Quality of Service (QoS) for online services while enhancing overall resource utilization. The bursty nature and stringent QoS demands of online tasks, coupled with the fluctuating resource footprints of offline tasks, can lead to severe resource pressure on nodes and undermine system stability. HyDLR first designs a load-aware rescheduling policy that dynamically identifies resource hotspots by monitoring metrics such as CPU satisfaction degree, memory, and GPU memory utilization. It then leverages eviction and task migration to optimize workload distribution. Furthermore, a two-stage filtering algorithm, guided by a multi-objective optimization model, targets system-wide load balancing and minimal rescheduling overhead. By incorporating a dynamically adjusted priority queue and a cost-feedback mechanism, HyDLR improves scheduling efficiency without compromising stability. Experimental results demonstrate that HyDLR significantly reduces the frequency of task migrations while achieving a well-balanced system load. The rate of cascading rescheduling events is kept below 3%, demonstrating superior performance over existing approaches. This work offers an effective solution for resource management in complex, hybrid deployment scenarios, laying a foundation for more efficient data center scheduling and demonstrating strong potential for practical adoption. Desheng Wang 0002, Shuo Si, Sichao Chen, Weizhe Zhang |
ICPADS | 5 |
| 2025 | DynGPU: A Dynamic GPU Sharing Framework for Enhanced Resource Utilization and Task Scheduling in Concurrent DNN TrainingabstractTraining deep neural networks (DNNs) is a common task in GPU clusters. However, in practical cluster environments, multiple concurrent DNN training tasks often fail to fully leverage GPU resources, resulting in suboptimal GPU utilization. Furthermore, existing GPU sharing frameworks primarily rely on static scheduling and frequently overlook task deadlines, leading to task delays and inefficient scheduling. To address these issues, we propose a dynamic GPU sharing framework (DynGPU) that intercepts GPU kernel executions to perform resource scheduling in multi-task environments. DynGPU incorporates a dynamic task priority adjustment mechanism that adapts task priorities in real time based on task progress, historical data, and remaining time to deadlines. By guaranteeing resources for high-priority tasks while maximizing resource allocation for low-priority tasks, DynGPU reduces resource contention and improves system throughput, enabling more timely task completions. Experiments show that, compared to dedicated GPU execution, DynGPU can reserve up to 97.5 % of throughput for high-priority tasks. Compared to state-of-the-art baselines, DynGPU achieves up to an 8.4 % improvement in task completion time. Zhiji Yu, Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Meng Hao 0002, Yu-Chu Tian |
ICPADS | 3 |
| 2025 | Definition and Detection of Centralization Defects in Smart ContractsabstractIn recent years, security incidents stemming from centralization defects in smart contracts have led to substantial financial losses. A centralization defect refers to any error, flaw, or fault in a smart contract's design or development stage that introduces a single point of failure. Such defects allow a specific account or user to disrupt the normal operations of smart contracts, potentially causing malfunctions or even complete project shutdowns. Despite the significance of this issue, most current smart contract analyses overlook centralization defects, focusing primarily on other types of defects. To address this gap, our paper introduces six types of centralization defects in smart contracts by manually analyzing 597 Stack Exchange posts and 117 audit reports. For each defect, we provide a detailed description and code examples to illustrate its characteristics and potential impacts. Additionally, we introduce a tool named CDRipper (Centralization Defects Ripper) designed to identify the defined centralization defects. Specifically, CDRipper constructs a permission dependency graph (PDG) and extracts the permission dependencies of functions from the source code of smart contracts. It then detects the sensitive operations in functions and identifies centralization defects based on predefined patterns. We conduct a large-scale experiment using CDRipper on 244,424 real-world smart contracts and evaluate the results based on a manually labeled dataset. Our findings reveal that 82,446 contracts contain at least one of the six centralization defects, with our tool achieving an overall precision of 93.7%. Zewei Lin, Jiachi Chen, Jiajing Wu, Weizhe Zhang, Zibin Zheng |
ICSE | 4 |
| 2025 | Smartreco: Detecting Read-Only Reentrancy via Fine-Grained Cross-DApp AnalysisabstractDespite the increasing popularity of Decentralized Applications (DApps), they are suffering from various vulnerabilities that can be exploited by adversaries for profits. Among such vulnerabilities, Read-Only Reentrancy (called ROR in this paper), is an emerging type of vulnerability that arises from the complex interactions between DApps. In the recent three years, attack incidents of ROR have already caused around 30M USD losses to the DApp ecosystem. Existing techniques for vulnerability detection in smart contracts can hardly detect Read-Only Reentrancy attacks, due to the lack of tracking and analyzing the complex interactions between multiple DApps. In this paper, we propose SmartReco, a new framework for detecting Read-Only Reentrancy vulnerability in DApps through a novel combination of static and dynamic analysis (i.e., fuzzing) over smart contracts. The key design behind SmartReco is threefold: (1) SmartReco identifies the boundary between different DApps from the heavy-coupled cross-contract interactions. (2) SmartReco performs fine-grained static analysis to locate points of interest (i.e., entry functions) that may lead to ROR. (3) SmartReco utilizes the on-chain transaction data and performs multi-function fuzzing (i.e., the entry function and victim function) across different DApps to verify the existence of ROR. Our evaluation of a manual-labeled dataset with 45 RORs shows that SmartReco achieves a precision of 88.64 % and a recall of 86.67 %. In addition, SmartReco successfully detects 43 new RORs from 123 popular DApps. The total assets affected by such RORs reach around 520,000 USD. Zibin Zheng, Yuhong Nan, Mingxi Ye, Kaiwen Ning, Yu Zhang 0036, Weizhe Zhang |
ICSE | 7 |
| 2025 | An Efficient Collaborative Algorithm for Building-Wide Mobile Edge ComputingabstractThe rapid proliferation of multi-mobile devices in smart building environments has intensified the demand for efficient computation offloading strategies in mobile edge computing. To address the challenges of task offloading, communication, and computation resource allocation while also considering the mobility of mobile devices and task priorities, we design a target server query strategy based on resource matching. This strategy accommodates the varying resource requirements of different task types and avoids increasing algorithm complexity. Based on this strategy, we propose the Greedy-Based Collaborative Algorithm to minimize the average execution time of tasks. Simulation results demonstrate that the proposed algorithm outperforms baseline algorithms and that the energy consumption of mobile devices remains acceptable. Hualong Wu, Weizhe Zhang, Desheng Wang 0002 |
INDIN | 3 |
| 2025 | SSR: Safeguarding Staking Rewards by Defining and Detecting Logical Defects in DeFi StakingabstractDecentralized Finance (DeFi) staking is one of the most prominent applications within the DeFi ecosystem, where DeFi projects enable users to stake tokens on the platform and reward participants with additional tokens. However, logical defects in DeFi staking could enable attackers to claim unwarranted rewards by manipulating reward amounts, repeatedly claiming rewards, or engaging in other malicious actions. To mitigate these threats, we conducted the first study focused on defining and detecting logical defects in DeFi staking. Through the analysis of 64 security incidents and 144 audit reports, we identified six distinct types of logical defects, each accompanied by detailed descriptions and code examples. Building on this empirical research, we developed SSR (Safeguarding Staking Reward), a static analysis tool designed to detect logical defects in DeFi staking contracts. SSR utilizes a large language model (LLM) to extract fundamental information about staking logic and constructs a DeFi staking model. It then identifies logical defects by analyzing the model and the associated semantic features. We constructed a ground truth dataset based on known security incidents and audit reports to evaluate the effectiveness of SSR. The results indicate that SSR achieves an overall precision of 92.31%, a recall of 87.92%, and an F1-score of 88.85%. Additionally, to assess the prevalence of logical defects in real-world smart contracts, we compiled a large-scale dataset of 15,992 DeFi staking contracts. SSR detected that 3,557 (22.24%) of these contracts contained at least one logical defect. Zewei Lin, Jiachi Chen, Zexu Wang, Yuming Feng 0002, Weizhe Zhang, Zibin Zheng |
ASE | 6 |
| 2025 | Finding Insecure State Dependency in DApps via Multi-Source Tracing and Semantic EnrichmentabstractDecentralized Applications (DApps) serve as the gateway to utilizing blockchain technology. As their prevalence continues to grow, DApps are becoming increasingly interconnected. For instance, a DApp does not need to manage the prices of various tokens internally, as it can retrieve this information from other DApps that provide more up-to-date data. However, such deep reliance also introduces more attack surfaces, posing greater risks to both DApps and their users. In this paper, we refer to the security threat arising from the interdependence of DApps as Insecure State Dependency (ISD). Public reports indicate that ISD has led to losses exceeding 340 million USD.Existing ISDs are mostly found by extensive manual auditing and lucky incidents, as automated discovery of such issues is extremely difficult. More specifically, it is by no means trivial to (1) achieve precise data tracking in the intertwined and invisible interactions of DApps, (2) obtain fine-grained semantic information in low semantic bytecode. In this paper, we propose a novel framework, called InsFinder, for detecting ISD in DApps. Specifically, InsFinder consists of three unique modules to overcome the aforementioned challenges. (1) InsFinder employs dynamic cross-DApp taint analysis to achieve accurate multi-source data tracking in heavily coupled DApp interactions. (2) InsFinder uses source mapping to map bytecode identifiers into meaningful source code, such as variable names or statements, enabling a deeper understanding of bytecode. (3) InsFinder implements fine-grained access control and static analysis for ISD entry point detection. Evaluation on a manually annotated dataset with 93 real-world ISDs shows that InsFinder successfully detects 72 of them, achieving a precision of 84.7% and a recall of 77.4%. Furthermore, InsFinder successfully uncovers 165 previously unreported ISDs across 122 DApp projects. These ISDs collectively impact over 2 million USD. Yuhong Nan, Wei Li 0121, Kaiwen Ning, Zewei Lin, Zitong Yao, Yuming Feng 0002, Weizhe Zhang, Zibin Zheng |
ASE | 8 |
| 2025 | PSSketch: Finding Persistent and Sparse Flow with High Accuracy and EfficiencyabstractFinding persistent sparse (PS) flow is critical to early warning of various threats. Previous works have predominantly focused on either heavy or persistent flows, with limited attention given to PS flows. Although some recent studies pay attention to PS flows, they struggle to establish an objective criterion due to insufficient data-driven observations, resulting in reduced accuracy. In this paper, we define a new criterion ''anomaly boundary'' to distinguish PS flows from regular flows. Specifically, a flow whose persistence exceeds a threshold will be protected, while a protected flow with a density lower than a threshold is reported as a PS flow. We then introduce PSSketch, a high-precision layered sketch, to find PS flows. PSSketch employs variable-length bitwise counters, where the first layer tracks the frequency and persistence of all flows, and the second layer protects potential PS flows and records overflow counts from the first layer. Some optimizations have also been implemented to reduce memory consumption further and improve accuracy. The experiments show that PSSketch reduces memory consumption by 1-2 orders of magnitude compared to the strawman solution combined with existing work. Compared with SOTA solutions for finding PS flows, it outperforms up to 2.94x higher in F1 score and reduces ARE by 1-2 orders of magnitude. Meanwhile, PSSketch achieves a higher throughput than these solutions. Qilong Shi, Xiyan Liang, Han Wang 0022, Wenjun Li 0004, Ziling Wei, Weizhe Zhang, Shuhui Chen |
KDD (2) | 7 |
| 2025 | EVRM: Elastic Virtual Resource Management framework for cloud virtual instances
Desheng Wang 0002, Weizhe Zhang, Zhiji Yu, Yu-Chu Tian, Keqin Li 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | LRD-Raft: Log Replication Decouple for Efficient and Secure Consensus in Consortium-Blockchain-Based IoTabstractCurrently, consortium blockchain has been used in Internet of Things (IoT) systems to ensure secure data sharing across organizations. Consortium blockchain typically uses the Raft algorithm because of its high performance. However, in geo-distributed IoT environments, the single-point overhead problem of leader nodes affects the security and efficiency of consensus due to the high latency and frequency of client requests. To address this challenge, we propose a novel solution: log replication decoupling raft (LRD-Raft), which enables follower nodes to participate in log replication as well. This scheme reduces the overhead of the leader node by delegating part of the log replication task to the follower node. We also propose an adaptive coding protocol that dynamically adjusts the erasure code parameters according to the number of cluster healthy nodes to save the cluster’s network traffic. We evaluated LRD-Raft in different network latency environments and different cluster sizes, and the experimental results show that LRD-Raft has a more extensive performance system compared to Raft in high network latency and large data block transmission environments, and also has a certain level of resistance to DoS attacks, which improves the performance and security of the consensus mechanism. Heru Yang, Yuming Feng 0002, Weizhe Zhang |
IEEE Internet Things J. | 3 |
| 2025 | Towards JPEG-Resistant Image Forgery Detection and Localization Via Self-Supervised Domain AdaptationabstractWith wide applications of image editing tools, forged images (splicing, copy-move, removal and etc.) have been becoming great public concerns. Although existing image forgery localization methods could achieve fairly good results on several public datasets, most of them perform poorly when the forged images are JPEG compressed as they are usually done in social networks. To tackle this issue, in this paper, a self-supervised domain adaptation network, which is composed of a backbone network with Siamese architecture and a compression approximation network (ComNet), is proposed for JPEG-resistant image forgery detection and localization. To improve the performance against JPEG compression, ComNet is customized to approximate the JPEG compression operation through self-supervised learning, generating JPEG-agent images with general JPEG compression characteristics. The backbone network is then trained with domain adaptation strategy to localize the tampering boundary and region, and alleviate the domain shift between uncompressed and JPEG-agent images. Extensive experimental results on several public datasets show that the proposed method outperforms or rivals to other state-of-the-art methods in image forgery detection and localization, especially for JPEG compression with unknown QFs. Yuan Rao 0002, Jiangqun Ni, Weizhe Zhang, Jiwu Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | SEPPDL: A Secure and Efficient Privacy-Preserving Deep Learning Inference Framework for Autonomous DrivingabstractThe autonomous driving system necessitates using privacy-preserving deep learning (PPDL) technologies as the safety assurance for its extensive application. However, existing PPDL solutions depend on intricate protocol designs for robust security. Although leveraging advanced dedicated hardware platforms can significantly improve inference efficiency, the PPDL frameworks that make the best use of hardware platform computility are scarce. Thus, balancing efficiency and security in PPDL remains an open question. This study presents SEPPDL, a secure tripartite inference framework for deep learning based on secret-sharing to balance privacy security and computational efficiency. We reduce the communication and calculation time by designing a deep learning quantization representation scheme, two new computational protocols, and a computation library that utilizes the integer computation units of the GPU. The experimental results show that compared with state-of-the-art PPDL frameworks, the SEPPDL framework reduces the communication and computation delay in the model inference to 1/2 and 1/3 of the existing optimal frameworks while maintaining the accuracy of the model inference. Meanwhile, the SEPPDL framework achieves a 10-fold performance improvement in a lightweight model. As the model scale increases, the performance of the SEPPDL-based model even achieves an 86-fold improvement compared to VGG16. Wang Bobo, Meng Hao 0002, Weizhe Zhang |
ACM Trans. Auton. Adapt. Syst. | 6 |
| 2025 | Deep Learning Workload Mapping Optimization on Jetson PlatformsabstractTo improve the performance and energy efficiency of deep learning (DL) applications, recent edge computing platforms have built-in heterogeneous accelerators, such as general-purpose graphics processing units (GPUs) and neural processing units (NPUs). For example, widely used NVIDIA Jetson platforms contain CPU, GPU, and deep learning accelerator (DLA), a type of NPU. It is non-trivial to map DL workloads to suitable accelerators to improve performance, energy efficiency, or even both. This article presents JDIMO, 1 a Jetson-aware deep-learning inference workload mapping optimization framework, to simultaneously improve energy efficiency and performance. JDIMO first measures energy-performance data of the fundamental nodes and the sub-networks with energy-efficiency improvement potential according to the topology structure of a DL network. Then, under the guidance of an analytical energy-performance model, the framework exploits an algorithm based on the variable-length sliding window to find the optimal mapping configuration and the optimal number of CUDA streams. We evaluate JDIMO by applying it to seven DL applications on a Jetson Orin NX (16GB) platform. JDIMO saves 47.5% EDP (energy delay product) and 22.6% energy and improves 138.3% QPS (queries per second) on average compared to the DLA-possible configuration. JDIMO saves 22.5% EDP and 12.6% energy and improves 13.5% QPS on average compared to JEDI, the most similar work to ours. Meanwhile, JDIMO also reduces 93.8% optimization time on average compared to JEDI. Farui Wang, Meng Hao 0002, Siyu Yang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | Dynamic Power Management Through Multi-agent Deep Reinforcement Learning for Heterogeneous SystemsabstractPower management and optimization play a significant role in modern computer systems, from battery-powered devices to servers running in data centers. Existing approaches for power capping fail to meet the requirements presented by dynamic workloads, and the situation becomes even more severe, given the divergent energy efficiency of workloads on heterogeneous hardware platforms. Adaptively optimizing energy consumption for dynamic workloads presents a great challenge to heterogeneous systems. To tackle this challenge, we present a machine learning based method to improve system-level power efficiency. We employ multi-agent deep reinforcement learning (MADRL) to automatically explore the relationship between long-term performance and the power budget for workloads of different types on classic CPU-GPU heterogeneous platforms. Our framework equips each device with an agent, enabling decentralized control over its power budget while maintaining centralized coordination to maximize the running time of applications within a power cap. We evaluate our approach against state-of-the-art methods on CPU-GPU platforms. Experimental results show that our method improves performance by an average of 8.5%. Additionally, our method is significantly more stable compared to the state-of-the-art heuristic approach. Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Weizhi Kong, Yuan Wen |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Partitioned Scheduling and Analysis for a Typed DAG Task on Heterogeneous Multi-CoresabstractHeterogeneous multi-core architectures are gaining popularity in recent years as they combine the benefits of different processors, resulting in improved execution capacity and energy efficiency. However, analyzing response times and allocating resources for the typed directed acyclic graph (DAG) task, which has complex execution logic, on heterogeneous multi-core systems poses significant challenges. Major approaches may yield overly pessimistic worst-case response time (WCRT) estimates in certain scenarios while failing to adequately address critical structural characteristics inherent to typed DAG tasks. To address these limitations, this article explores the WCRT analysis and core allocations for the typed DAG task under partitioned scheduling. In this work, we first delve into the characteristics of the topology structure of the typed DAG task and propose a novel WCRT upper bound to enhance the accuracy of WCRT analysis. Then, a subtask allocation strategy is presented, which enables an effectively utilization of the resources of multi-cores. Finally, the performance of the proposed analysis algorithm and allocation strategy are tested by implementing a verification system on a real heterogeneous multi-core platform. Experimental results demonstrate that our proposed WCRT analysis algorithm exhibits substantial improvements of 38.7% and 37.43% in the theoretical analysis performance and actual analysis accuracy, respectively. Similarly, our proposed core allocation strategy improves the theoretical and the actual execution efficiency of the system by 10.6% and 7.41%, respectively. These results substantiate the practical value of our enhanced WCRT derivation methodology and allocation scheme in improving system resource utilization efficiency. Yehan Ma, Mingdong Xie, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | HEngine: A High Performance Optimization Framework on a GPU for Homomorphic EncryptionabstractHomomorphic encryption (HE) represents an encryption technology that allows for direct computation on encrypted data without requiring decryption. However, the substantial computational complexity and significant latency associated with HE has impeded its broader adoption in practical applications. To address these challenges, we propose a GPU-based acceleration framework, namely HEngine, tailored for homomorphic encryption tasks. Specifically, we first propose a warp shuffle-based optimization method for two key phases, i.e., inverse Chinese Remainder Theorem (ICRT) and number theoretic transformation (NTT), to mitigate synchronization overhead in homomorphic encryption. Secondly, we propose to fuse the NTT kernel with the inner product kernel to address the imbalance between memory access and computation. Thirdly, considering the potential difference in the amount of tasks of users in the real-world, we design two different encoding methods for small batch and large batch inference tasks to improve computational efficiency. Finally, experiments demonstrate that our proposed framework achieves a 218× speedup on homomorphic multiplication tasks compared with the CPU-based SEAL library. In addition, for convolutional neural network inference tasks on shallow network structures, our proposed framework achieves amortized inference performance at the millisecond level and sub-millisecond level on small batch and large batch data, respectively. For convolutional neural network inference tasks on deeper network structures (i.e., ResNet-20), our proposed framework achieves second-level inference. Meng Hao 0002, Weizhe Zhang, Desheng Wang 0002 |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | ESAFL: Efficient Secure Additively Homomorphic Encryption for Cross-Silo Federated LearningabstractCross-silo federated learning (FL) enables multiple clients to collaboratively train a machine learning model without sharing training data, but privacy in FL remains a major challenge. Techniques using homomorphic encryption (HE) have been designed to solve this but bring their own challenges. Many techniques using single-key HE (SKHE) require clients to fully trust each other to prevent privacy disclosure between clients. However, fully trusted clients are hard to ensure in practice. Other techniques using multi-key HE (MKHE) aim to protect privacy from untrusted clients but lead to the disclosure of training results in public channels by untrusted third parties, e.g., the public cloud server. Besides, MKHE has higher computation and communication complexity compared with SKHE. We present a new FL protocol ESAFL that leverages a novel efficient and secure additively HE (ESHE) based on the hard problem of ring learning with errors. ESAFL can ensure the security of training data between untrusted clients and protect the training results against untrusted third parties. In addition, theoretical analyses present that ESAFL outperforms current techniques using MKHE in computation and communication, and intensive experiments show that ESAFL achieves approximate$204\times$−$953\times$and$11\times$−$14\times$training speedup while reducing the communication burden by$77\times$−$109\times$and$1.25\times$−$2\times$compared with the state-of-the-art FL models using SKHE. Jiahui Wu 0001, Weizhe Zhang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | LS2: Boosting Hidden Separation for Backdoor Defense With Learning Speed-Driven Label Smoothing
Jie Peng 0009, Haoyu He 0001, Hengji Dong, Weizhe Zhang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | DPDeno: A Post-Processing Framework for Releasing Differentially Private Spatio-Temporal Mobility FeaturesabstractThe spatio-temporal (ST) mobility patterns derived from trajectory data are crucial for applications such as location-based services and urban analytics. However, releasing these mobility features raises significant privacy concerns, as they may expose sensitive personal location information. Differential privacy (DP) is widely used to safeguard individual privacy during data releases, but existing methods for releasing ST features often suffer from utility loss because their high dimensionality requires injecting substantial noise to meet privacy guarantees. Several recent approaches attempt to address this issue by reducing noise in differentially private spatio-temporal (DPST) features, but they either discard valuable information while compressing noisy data representations or rely solely on restrictive road network topology constraints, resulting in only modest utility improvements. In this paper, we present DPDeno, a post-processing framework designed to significantly enhance the utility of DPST features. First, DPDeno generates synthetic trajectory datasets using public information (e.g., road network data) and applies existing DP methods to create paired DPST (noisy) and ST (clean) features. It then trains a spatio-temporal graph autoencoder (STGAE), which models each feature as a graph, with road segments as nodes and transitions over time as edges. By minimizing node- and edge-level reconstruction losses between the noisy and clean pairs, STGAE learns to refine DPST inputs toward the structural consistency of their clean counterparts, thereby improving their practical utility. The trained model is then used to post-process real DPST features. Importantly, DPDeno preserves the original DP guarantee, as STGAE is trained solely on synthetic data generated from public sources without accessing any private information. Experimental results on two real-world trajectory datasets show that DPDeno significantly improves both the statistical accuracy and practical usability of released mobility features. Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Renyu Yang, Weizhe Zhang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Vulnerabilities of NSPFL: Privacy-Preserving Federated Learning With Data Integrity AuditingabstractThe secure and privacy-preserving federated learning scheme, NSPFL, aims to safeguard data privacy while also auditing data integrity. The solution provided by this scheme is highly novel. However, NSPFL has significant design shortcomings in terms of both privacy protection and data integrity verification. This work identifies specific issues within NSPFL and proposes effective countermeasures. Furthermore, our proposed solution can serve as a general approach for privacy-preserving multiparty computations, safeguarding privacy while enhancing efficiency. Jiahui Wu 0001, Tiecheng Sun, Weizhe Zhang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | GeoRecover: Recovery From Poisoning Attacks for LDP-Enabled Spatial Density AggregationabstractThe spatial density distribution collected and aggregated from users’ trajectory data is vital for location-based services like regional popularity analysis and congestion measurement. However, spatial density aggregation poses privacy concerns since trajectory data usually originate from users. Local differential privacy (LDP) addresses these concerns by allowing users to perturb their data before reporting it. Yet, LDP is vulnerable to poisoning attacks where attackers manipulate data from malicious users. Recent studies attempt to defend against such attacks in LDP-enabled frequency estimation but suffer from inaccurate data recovery due to empirical presets of malicious user proportions and inaccurate malicious data estimation. These issues worsen in spatial density aggregation, as high-dimensional trajectory data help conceal malicious information. In this work, we propose GeoRecover, a method to defend against poisoning attacks in LDP-enabled spatial density aggregation by addressing previous limitations. GeoRecover designs an adaptive model to unify these attacks. Under this model, GeoRecover estimates the proportion of malicious users using statistical differences between genuine and malicious data and learns malicious data statistics through LDP properties. This allows GeoRecover to recover accurate spatial density distribution by subtracting malicious users’ contributions. Evaluations on two real-world datasets show GeoRecover outperforms state-of-the-art methods in recovery accuracy, defense capability, and practical performance. Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Weizhe Zhang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake DetectionabstractWith the advancement of deepfake generation techniques, the importance of deepfake detection in protecting multimedia content integrity has become increasingly obvious. Recently, temporal inconsistency clues have been explored to improve the generalizability of deepfake video detection. According to our observation, the temporal artifacts of forged videos in terms of motion information usually exhibits quite distinct inconsistency patterns along horizontal and vertical directions, which could be leveraged to improve the generalizability of detectors. In this paper, a transformer-based framework forDiffusion Learning ofInconsistencyPattern (DIP) is proposed, which exploits directional inconsistencies for deepfake video detection. Specifically, DIP begins with a spatiotemporal encoder to represent spatiotemporal information. A directional inconsistency decoder is adopted accordingly, where direction-aware attention and inconsistency diffusion are incorporated to explore potential inconsistency patterns and jointly learn the inherent relationships. In addition, the SpatioTemporal Invariant Loss (STI Loss) is introduced to contrast spatiotemporally augmented sample pairs and prevent the model from overfitting nonessential forgery artifacts. Extensive experiments on several public datasets demonstrate that our method could effectively identify directional forgery clues and achieve state-of-the-art performance. Fan Nie, Jiangqun Ni, Jian Zhang 0086, Bin Zhang 0048, Weizhe Zhang |
IEEE Trans. Multim. | 5 |
| 2025 | HeavyFinder: A Lightweight Network Measurement Framework for Detecting High-Frequency Elements in Skewed Data StreamsabstractSkewed data streams are characterized by uneven distributions in which a small fraction of elements occur with much higher frequency than others. The detection of these high-frequency elements presents significant practical challenges, particularly under stringent memory constraints, as existing detection techniques have typically relied on predefined thresholds that require significant memory usage. However, this approach is highly inefficient since not all elements require equal storage space. To address these limitations, we introduce HeavyFinder (HF), a novel lightweight network measurement architecture designed to detect high-frequency elements in skewed data. HF employs a threshold-free update strategy that enables dynamic adaptation to variable data, thereby providing greater flexibility for tracking high-frequency elements without requiring fixed thresholds. Furthermore, an included memory-light strategy enables high accuracy for non-uniform distributions, even with limited memory allocation. Experimental results showed that HF significantly improved performance in four query tasks, producing an accuracy of 99.81% when identifying the top-k elements. The average absolute error (AAE) was also reduced to 10-4 using only 100KB of memory, which was significantly lower than that of conventional methods. Lin Yao 0001, Weizhe Zhang |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | Data Offloading for Edge-Enabled Smart City Services via Location-Functionality Correlation AnalysisabstractMobile edge computing significantly reduces the delay for smart city services by offloading the related datasets on the edge server. The correlation among smart city datasets regarding their location and functionality plays a crucial role in determining the data offloading strategy. The more closely related these datasets are, the higher the probability they will be requested simultaneously. However, most existing methods often ignore the impact of location-functionality correlation on the datasets' access frequencies. If datasets with high correlation are offloaded on separate edge servers, it will increase cross-server delay. Therefore, we proposed a method to analyze location-functionality correlation among smart city datasets and optimize data offloading strategy. Then, we formalize a data offloading model to optimize the total expected profits, which is an NP-hard problem. To solve this problem, we propose a pruned Q-learning for data offloading algorithm that learns the location-functionality correlation among the datasets. To validate its effectiveness, we first conduct simulation experiments to validate the offloading strategy under both complex service scenarios and large-scale datasets. Furthermore, we implement a data offloading architecture based on SuperMap iServer. The experimental results demonstrate that our algorithm reduces the delay and energy consumption by 34.65% and 25.54%, respectively, compared to the baseline algorithms. Qingyang Fan, Weizhe Zhang, Chen Ling 0006 |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | Unity is Strength: Enhancing Precision in Reentrancy Vulnerability Detection of Smart Contract Analysis ToolsabstractReentrancy is one of the most notorious vulnerabilities in smart contracts, resulting in significant digital asset losses. However, many previous works indicate that current Reentrancy detection tools suffer from high false positive rates. Even worse, recent years have witnessed the emergence of new Reentrancy attack patterns fueled by intricate and diverse vulnerability exploit mechanisms. Unfortunately, current tools face a significant limitation in their capacity to adapt and detect these evolving Reentrancy patterns. Consequently, ensuring precise and highly extensible Reentrancy vulnerability detection remains critical challenges for existing tools. To address this issue, we propose a tool named ReEP, designed to reduce the false positives for Reentrancy vulnerability detection. Additionally, ReEP can integrate multiple tools, expanding its capacity for vulnerability detection. It evaluates results from existing tools to verify vulnerability likelihood and reduce false positives. ReEP also offers excellent extensibility, enabling the integration of different detection tools to enhance precision and cover different vulnerability attack patterns. We perform ReEP to eight existing state-of-the-art Reentrancy detection tools. The average precision of these eight tools increased from the original 0.5% to 73% without sacrificing recall. Furthermore, ReEP exhibits robust extensibility. By integrating multiple tools, the precision further improved to a maximum of 83.6%. These results demonstrate that ReEP effectively unites the strengths of existing works, enhances the precision of Reentrancy vulnerability detection tools. Zexu Wang, Jiachi Chen, Peilin Zheng, Yu Zhang 0036, Weizhe Zhang, Zibin Zheng |
IEEE Trans. Software Eng. | 5 |
| 2025 | C2DP: CLIP-conditioned knowledge distillation for membership inference privacy protection
Zimeng Jia, Lun Xin, Renhao Lu, Weizhe Zhang |
World Wide Web (WWW) | 8 |
| 2024 | Bubble Sketch: A High-performance and Memory-efficient Sketch for Finding Top-k Items in Data StreamsabstractSketch algorithms are crucial for identifying top-k items in large-scale data streams. Existing methods often compromise between performance and accuracy, unable to efficiently handle increasing data volumes with limited memory. We present Bubble Sketch, a compact algorithm that excels in both performance and accuracy. Bubble Sketch achieves this by (1) Recording only full keys of hot items, significantly reducing memory usage, and (2) Using threshold relocation to resolve conflicts, enhancing detection accuracy. Unlike traditional methods, Bubble Sketch eliminates the need for a Min-Heap, ensuring fast processing speeds. Experiments show Bubble Sketch outperforms the other seven algorithms compared, with the highest throughput and precision, and surpasses HeavyKeeper in accuracy by up to two orders of magnitude. Qilong Shi, Yuxi Liu 0017, Hanyue Zheng, Yao Xin, Wenjun Li 0004, Tong Yang 0003, Yangyang Wang 0001, Yang Xu 0010, Weizhe Zhang, Mingwei Xu 0001 |
CIKM | 10 |
| 2024 | How to Bridge Graph and Sequence Patterns in Session-Based Recommendation? A Self-Supervised MethodabstractSession-based Recommendation aims to reveal the item distribution patterns in anonymous session sequences. Most existing approaches model the distribution patterns by utilizing either sequential or structural information individually to absorb different pattern knowledge, which can only model the distinct one-sided facet of item distribution in sessions, thus leading to suboptimal performance. Self-supervised learning provides a natural solution as a bridge to fill the gap between different learning paradigms in session-based recommendations, which remains unexplored. In this paper, we regard the distinct learning paradigm as an individual channel and then integrate the sequential and graphical channels with a contrastive bridge architecture. We name the novel framework DC-Rec, for Dual Channel Recommendation, to model the comprehensive session characteristics. Extensive experiments conducted on two real-world datasets demonstrate that our model consistently outperforms the state-of-the-art recommendation methods. Yu Tai, Sheng Yin, Weizhe Zhang |
ICASSP | 7 |
| 2024 | BitMatcher: Bit-level Counter Adjustment for SketchesabstractSketch has been widely used in the field of large-scale data stream processing. However, common fixed-counter algorithms such as Count-Min Sketch have to allocate larger counters, which wastes a lot of memory due to the high skewness of real-world data streams. To reduce memory usage, we propose to dynamically adjust the counter size that matches the distribution of the data stream. We introduce BitMatcher, a fast global-adjusting algorithm that automatically adjusts the counter to the appropriate size to match the data stream. During stream processing, BitMatcher identifies items hashed into a bucket based on isolated fingerprints. If it overflows, BitMatcher changes the flag bits in the bucket and dynamically increases or shrinks the size of some counters in a fine-grained manner. BitMatcher can also relocate a cold item in the bucket with the idea of cuckoo hashing to preserve the potential hot item while achieving global load balancing. Through the above way of dealing with overflow caused by skewed data, BitMatcher precisely manipulates allocated bits and maximizes memory utilization. The experiments show that BitMatcher has high throughput and can outperform SOTA by up to 4 orders of magnitude in terms of accuracy. We also deployed BitMatcher on several platforms, showing its software and hardware scalability. Qilong Shi, Chengjun Jia, Wenjun Li 0004, Zaoxing Liu, Tong Yang 0003, Jianan Ji, Gaogang Xie, Weizhe Zhang, Minlan Yu |
ICDE | 8 |
| 2024 | Escape Cache Traps by Rate Feedback for Ndn Real-Time Video StreamingabstractIn-network caching is one of the most important characteristic of Named Data Networking (NDN). However, while replacing producers in responding to interest requests, caching data packets also shields consumers from perceiving the bottleneck bandwidth of the transmission path between the producer and the consumer. Therefore, when the content source switches from the cache node to the producer due to data exhaustion, the consumer can not adjust the requesting rate accordingly, and may lead to the serious bufferbloat or packet loss - we call it as Cache Trap. We found that Cache Trap occurs commonly in streaming services and the state-of-art NDN congestion control schemes cannot achieve efficient and stable quality of service when it happens. To escape Cache Trap, this paper proposes an explicit rate feedback congestion control algorithm, named as RFCC. RFCC leverages NDN routers' ability of encapsulating customized information in data packets to send link state information to consumers. Specifically, when responding to interest packets, RFCC nodes estimate data throughput received from the producer and insert this information into the returned data packets. The consumer perceives the change of content source according to the hopcount tag in data packet, and then adjusts the sending rate of interest packet based on the explicit rate information. We have implemented RFCC in both real-world NDN live video streaming and NDNsim simulation platforms, and compared it with the state-of-arts congestion control algorithms in a variety of scenarios. The experimental results show that when Cache Trap occurs, RFCC maintains a stable QoE in live video streaming, reduces 50% delay jitters compared with DPCCP and achieves$2.4 \times$throughput compared with PCON. Zhaohua Zhu, Yongrui Chen 0001, Linggang Li, Zhijun Li 0002, Weizhe Zhang, Yu Zhang 0036 |
ICNP | 5 |
| 2024 | An MTD-driven Hybrid Defense Method Against DDoS Based on Markov Game in Multi-controller SDN-enabled IoT NetworksabstractThe widespread deployment of low-cost, vulnerable IoT devices allows attackers to exploit them to generate botnets and launch distributed denial-of-service (DDoS) attacks, which has become a serious security challenge for ensuring quality of service (QoS). For cost-effective defense against DDoS, we propose a novel hybrid defense method that includes proactive moving target defense (MTD) and passive security control to resist DDoS threats at different stages in IoT networks in this paper. We construct a multi-stage Markov game model to portray the game as a competition between the attacker and the defender for the control duration of the attack surface, and design an optimal defense strategy algorithm. In particular, we introduce a new parameter of action execution interval expectation in the game and add node importance evaluation in the reward quantification so that the optimal action execution interval of each defense technique can be output. We also consider the possibility that advanced attackers may launch DDoS on the SDN controller in the game. The experimental results demonstrate that our proposed method can defend against DDoS cost-effectively and ensure the QoS in IoT networks with acceptable overhead. Yuming Feng 0002, Weizhe Zhang, Zijun Feng, Xiaoxiong Zhong, Fangming Liu |
IWQoS | 2 |
| 2024 | FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake DetectionabstractNowadays, the abuse of AI-generated content (AIGC), especially the facial images known as deepfake, on social networks has raised severe security concerns, which might involve the manipulations of both visual and audio signals. For multimodal deepfake detection, previous methods usually exploit forgery-relevant knowledge to fully finetune Vision transformers (ViTs) and perform cross-modal interaction to expose the audio-visual inconsistencies. However, these approaches may undermine the prior knowledge of pretrained ViTs and ignore the domain gap between different modalities, resulting in unsatisfactory performance. To tackle these challenges, in this paper, we propose a new framework, i.e., Forgery-aware Audio-distilled Multimodal Learning (FRADE), for deepfake detection. In FRADE, the parameters of pretrained ViT are frozen to preserve its prior knowledge, while two well-devised learnable components, i.e., the Adaptive Forgery-aware Injection (AFI) and Audio-distilled Cross-modal Interaction (ACI), are leveraged to adapt forgery relevant knowledge. Specifically, AFI captures high-frequency discriminative features on both audio and visual signals and injects them into ViT via the self-attention layer. Meanwhile, ACI employs a set of latent tokens to distill audio information, which could bridge the domain gap between audio and visual modalities. The ACI is then used to well learn the inherent audio-visual relationships by cross-modal interaction. Extensive experiments demonstrate that the proposed framework could outperform other state-of-the-art multimodal deepfake detection methods under various circumstances. Fan Nie, Jiangqun Ni, Jian Zhang 0086, Bin Zhang 0048, Weizhe Zhang |
ACM Multimedia | 5 |
| 2024 | Modified Counter: A Fast and Dynamic Structure for Locating High-Frequency Items in Data StreamsabstractThe demand for high-speed data stream processing has increased as networks have become more complex. Specific challenges include managing large-scale streams and addressing issues caused by data skewness and the presence of anomalies. As such, this paper introduces Modified Counter (MC), a novel data structure in which conventional XOR operations are replaced by a unique one-item lookup process wherein each item hash is partitioned into a quotient for bin selection and a remainder for precise entry location. This technique significantly reduces hash collisions, thereby improving retrieval accuracy and fingerprint indexing times, which increases both processing speed and spatial efficiency. Counter structures and entry sorting are further optimized by selectively evicting low-frequency items and prioritizing high-frequency data, which is beneficial for skewed distributions and high-conflict scenarios. As a result, MC offers unprecedented load balancing and adaptability, particularly for high-speed data streams. Validation experiments showed that MC outperformed conventional structures, consistently achieving the highest throughput for frequency queries, the lowest error rates, and the highest precision for heavy hitter queries, which are critical for high-speed data processing. Notably, an insertion rate of 19.4 Mbps was achieved with the IP dataset. The average absolute error was approximately 10-1.5with a memory of only 0.1 MB, improving to below 10-2.5with 1.0 MB, representing an improvement of ~2 orders of magnitude over conventional models. These results suggest MC to be a valuable structure for reducing memory dependency by more efficiently locating high-frequency items in high-speed data streams. Lin Yao 0004, Weizhe Zhang |
MSN | 3 |
| 2024 | LightFinder: Finding Persistent Items with Small Memory
Weiqiang Xiao, Weizhe Zhang |
NPC (1) | 3 |
| 2024 | Fast Memory Disaggregation with SwiftSwap
Xiangwei Zhang, Desheng Wang 0002, Weizhe Zhang, Zhiji Yu, Meng Hao 0002 |
NPC (1) | 3 |
| 2024 | Optimizing depthwise separable convolution on DCUabstractAbstract The integration of Large Language Models (LLMs) with Convolutional Neural Networks (CNNs) is significantly advancing the development of large models. However, the computational cost of large models is high, necessitating optimization for greater efficiency. One effective way to optimize the CNN is the use of depthwise separable convolution (DSC), which decouples spatial and channel convolutions to reduce the number of parameters and enhance efficiency. In this study, we focus on porting and optimizing DSC kernel functions from the GPU to the Deep Computing Unit (DCU), a computing accelerator developed in China. For depthwise convolution, we implement a row data reuse algorithm to minimize redundant data loading and memory access overhead. For pointwise convolution, we extend our dynamic tiling strategy to improve hardware utilization by balancing resource allocation among blocks and threads, and we enhance arithmetic intensity through a channel distribution algorithm. We implement depthwise and pointwise convolution kernel functions and integrate them into PyTorch as extension modules. Experiments demonstrate that our optimized kernel functions outperform the MIOpen library on the DCU, achieving up to a 3.59 $$\times$$ × speedup in depthwise convolution and up to a 3.54 $$\times$$ × speedup in pointwise convolution. These results highlight the effectiveness of our approach in leveraging the DCU’s architecture to accelerate deep learning operations. Meng Hao 0002, Weizhe Zhang, Gangzhao Lu, Xueyang Tian, Siyu Yang 0002, Mingdong Xie, Chenyu Yuan, Desheng Wang 0002 |
CCF Trans. High Perform. Comput. | 3 |
| 2024 | Adaptive asynchronous federated learning
Renhao Lu, Weizhe Zhang, Qiong Li 0001, Xiaoxiong Zhong, Desheng Wang 0002, Zenglin Xu, Mamoun Alazab |
Future Gener. Comput. Syst. | 2 |
| 2024 | Decentralized AI-Based Task Distribution on Blockchain for Cloud Industrial Internet of Things
Amir Javadpour 0001, Arun Kumar Sangaiah, Weizhe Zhang, Ankit Vidyarthi, Sayyed Hamid Reza Ahmadi 0001 |
J. Grid Comput. | 3 |
| 2024 | Cooperative Service Caching in Vehicular Edge Computing Networks Based on Transportation Correlation AnalysisabstractIn vehicular edge computing, the vehicular services are cached and replaced among Roadside Units (RSU) to minimize delay in delay-sensitive services. However, without prior knowledge of these service preferences, the cached services suffer a low hit ratio and a high network delay. Before the appearance of requests from vehicles, most existing vehicular service caching methods design the service-sharing mechanisms among RSUs to optimize the hit ratio, but they ignore the influence of vehicle mobility on service-sharing. Moreover, in vehicular services replacement, the differences between actual and historical vehicle trajectories negatively impact the service-sharing mechanisms constructed before, which existing methods ignored. Thus, this paper addresses the vehicular service caching problem based on the surrounding function-features of vehicles and transportation correlations. Firstly, we use the surrounding function-features of vehicles to estimate the service preference and ensure the hit ratio of cached services. Secondly, we formulate the vehicular service caching problem as a constrained optimization problem and design a cooperation mechanism between RSUs, which is based on transportation correlations of vehicle trajectories. Thirdly, we propose two vehicular service caching methods based on Gibbs sampling to optimize the network delay of vehicular services and deal with the negative influence of vehicle mobility. Simulation in real datasets from Shenzhen, China shows that our methods have a better network delay and hit ratio than existing baselines by 15.26% and 5.6%, respectively. Chen Ling 0006, Weizhe Zhang, Qingyang Fan, Zebang Feng, Desheng Wang 0002 |
IEEE Internet Things J. | 2 |
| 2024 | Two-Stage Client Selection for Federated Learning Against Free-Riding Attack: A Multiarmed Bandits and Auction-Based ApproachabstractUtilizing the federated learning (FL) technique, data owners can collaboratively train artificial intelligence models, retaining all training data on their premises to minimize the potential for personal data breaches. However, self-interested users (e.g., free riders) bring new challenges that hinder the development of FL techniques. To this end, we propose a two-stage client selection scheme comprising a multiarmed bandit (MAB)-based candidate client selection method and an auction-based training client selection method. Specifically, our client selection scheme initially formulates the FL system into an MAB system, where clients are the arms and the server is the player. Then, we quantify the similarity between a local model and the server side, which is the designed metric for model aggregation and reward computation updating based on the fuzzy mathematical strategy. Next, based on the Thompson Sampling strategy, the server can intelligently determine the reward of each client, and clients with more significant rewards have the chance for local model training. With an auction method, the server can determine the training clients to reduce the training cost while maximizing each client’s revenue. Extensive experiments on real-world data sets demonstrate that the proposed scheme outperforms representative FL schemes (i.e., FedAvg, FedProx, FedMax, and MFL) regarding the model’s convergence rate and cost in FL systems with free riders. Renhao Lu, Weizhe Zhang, Qiong Li 0001, Xiaoxiong Zhong, Desheng Wang 0002, Lu Shi 0002, Yuelin Guo |
IEEE Internet Things J. | 2 |
| 2024 | Minimizing Service Latency Through Image-Based Microservice Caching and Randomized Request Routing in Mobile Edge ComputingabstractIn the context of mobile edge computing (MEC), the traditional method of requesting microservices from a central cloud can result in increased delay for users due to the physical distance between the user and the cloud server. To address this issue, MEC advocates for placing servers closer to the users at the edge of the network. However, this approach is constrained by the storage capacity and computing resources of edge servers (ESs). Therefore, it is crucial to devise a strategy for processing user requests that minimizes the average request delay. To address this problem, This article formulates the microservice caching problem as an image-based microservice placement and task request routing problem. We model the problem as an integer linear programming problem with multicondition constraints. Considering the limited resources of ESs, we propose a microservice placement algorithm called approximate algorithm based on randomized task request routing. The proposed algorithm is designed to provide near-optimal solutions in polynomial time, leveraging Chernoff’s theorem. Our approach is evaluated through comparisons with two existing algorithms: 1) the image-pull-based microservice cache request algorithm and 2) the greedy-based microservice cache and request routing algorithm. The results demonstrate that our algorithm exhibits superior performance compared to existing methods. Desheng Wang 0002, Weizhe Zhang, Guanqing Lou |
IEEE Internet Things J. | 3 |
| 2024 | Topic-aware Masked Attentive Network for Information Cascade PredictionabstractPredicting information cascades holds significant practical implications, including applications in public opinion analysis, rumor control, and product recommendation. Existing approaches have generally overlooked the significance of semantic topics in information cascades or disregarded the dissemination relations. Such models are inadequate in capturing the intricate diffusion process within an information network inundated with diverse topics. To address such problems, we propose a neural-based model using Topic-Aware Masked Attentive Network for Information Cascade Prediction (ICP-TMAN) to predict the next infected node of an information cascade. First, we encode the topical text into user representation to perceive the user-topic dependency. Next, we employ a masked attentive network to devise the diffusion context to capture the user-context dependency. Finally, we exploit a deep attention mechanism to model historical infected nodes for user embedding enhancement to capture user-history dependency. The results of extensive experiments conducted on three real-world datasets demonstrate the superiority of ICP-TMAN over existing state-of-the-art approaches. Yu Tai, Yuanming Shao, Weizhe Zhang, Arun Kumar Sangaiah |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2024 | An IoT Device Identification Method Using Extracted Fingerprint From Sequence of Traffic Grayscale ImagesabstractWith the widespread deployment and application of various types of IoT devices, preventing illegal intrusion and impersonation attacks of IoT devices has become an important security challenge. Device identification helps to limit the behavior of suspicious devices and enhances the security of the device access process. In this paper, we propose a novel deep learning-based automatic fingerprint extraction model that addresses low efficiency and complexity of traditional feature engineering process, which are often rely on expert experience. The proposed model integrates advanced modules such as Depthwise Separable Convolution (DSC) and Gated Recurrent Unit (GRU), as well as architectures of inverted residuals and linear bottlenecks to enhance the performance of fingerprint extraction. After converting the raw device traffic into the sequence of traffic grayscale images, the model can analyze spatial and temporal features from them to generate highly distinguishable device fingerprints automatically. Additionally, we also achieve fast fingerprint search based on Hierarchical Navigable Small World (HNSW) to support device identification. Our proposed method can not only indicate deviations in device behavior from expected specifications, but also identify unknown and unreliable IoT devices. The experimental results show that our method has excellent performance and more comprehensive identification capabilities in multiple dimensions. Yuming Feng 0002, Yu Zhang 0036, Weizhe Zhang, Desheng Wang 0002 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | On the Security of Verifiable and Oblivious Secure Aggregation for Privacy-Preserving Federated LearningabstractRecently, to resist privacy leakage and aggregation result forgery in federated learning (FL), Wang et al. proposed a verifiable and oblivious secure aggregation protocol for FL, called VOSA. They claimed that VOSA was aggregate unforgeable and verifiable under a malicious aggregation server and gave detailed security proof. In this article, we show that VOSA is insecure, in which local gradients/aggregation results and their corresponding authentication tags/proofs can be tampered with without being detected by the verifiers. After presenting specific attacks, we analyze the reason for this security issue and give a suggestion to prevent it. Jiahui Wu 0001, Weizhe Zhang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Generating Location Traces With Semantic- Constrained Local Differential PrivacyabstractValuable information and knowledge can be learned from users’ location traces and support various location-based applications such as intelligent traffic control, incident response, and COVID-19 contact tracing. However, due to privacy concerns, no authority could simply collect users’ private location traces for mining or even publishing. To echo such concerns, local differential privacy (LDP) enables individual privacy by allowing each user to report a perturbed version of their data. Unfortunately, when applied to location traces, LDP cannot preserve the semantics in the context of location traces because it treats all locations (i.e., various points of interest) as equally sensitive. This results in a low utility of LDP mechanisms for collecting location traces. In this paper, we address the challenge of collecting and sharing location traces with valuable semantics while providing sufficient privacy protection for participating users. We first propose semantic-constrained local differential privacy (SLDP), a new privacy model to provide a provable mathematical privacy guarantee while preserving desirable semantics. Then, we design a location trace perturbation mechanism (LTPM) that users can use to perturb their traces in a way that satisfies SLDP. Finally, we propose a private location trace synthesis (PLTS) framework in which users use LTPM to perturb their traces before sending them to the collector, who aggregates the users’ perturbed data to generate location traces with valuable semantics. Extensive experiments on three real-world datasets demonstrate that our PLTS outperforms existing state-of-the-art methods by at least 21% in a range of real-world applications, such as spatial visiting queries and frequent pattern mining, under the same privacy leakage. Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Weizhe Zhang, Jie Xu 0007 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | On the Security of "LSFL: A Lightweight and Secure Federated Learning Scheme for Edge Computing"abstractZhang et al. (2023) recently proposed a secure federated learning (FL) scheme named LSFL to guarantee Byzantine robustness while protecting privacy in FL. In this work, we show that LSFL breaches privacy it claimed. Specifically, we demonstrate that the secure Byzantine robustness procedure of LSFL exposes significant information of all participant models and data to a semi-honest server, thereby damaging privacy. Then, we analyze the reason for this security issue and give a suggestion to prevent privacy breaches in LSFL. Jiahui Wu 0001, Weizhe Zhang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Multi-Source and Multi-modal Deep Network Embedding for Cross-Network Node ClassificationabstractIn recent years, to address the issue of networked data sparsity in node classification tasks, cross-network node classification (CNNC) leverages the richer information from a source network to enhance the performance of node classification in the target network, which typically has sparser information. However, in real-world applications, labeled nodes may be collected from multiple sources with multiple modalities (e.g., text, vision, and video). Naive application of single-source and single-modal CNNC methods may result in sub-optimal solutions. To this end, in this article, we propose a model called Multi-source and Multi-modal Cross-network Deep Network Embedding (M 2 CDNE) for cross-network node classification. In M 2 CDNE, we propose a deep multi-modal network embedding approach that combines the extracted deep multi-modal features to make the node vector representations network invariant. In addition, we apply dynamic adversarial adaptation to assess the significance of marginal and conditional probability distributions between each source and target network to make node vector representations label discriminative. Furthermore, we devise to classify nodes in the target network through the related source classifier and aggregate different predictions utilizing respective network weights, corresponding to the discrepancy between each source and target network. Extensive experiments performed on real-world datasets demonstrate that the proposed M 2 CDNE significantly outperforms the state-of-the-art approaches. Weizhe Zhang, Yan Wang 0002, Lin Jing |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Backdoor Two-Stream Video Models on Federated LearningabstractVideo models on federated learning (FL) enable continual learning of the involved models for video tasks on end-user devices while protecting the privacy of end-user data. As a result, the security issues on FL, e.g., the backdoor attacks on FL and their defense have increasingly become the domains of extensive research in recent years. The backdoor attacks on FL are a class of poisoning attacks, in which an attacker, as one of the training participants, submits poisoned parameters and thus injects the backdoor into the global model after aggregation. Existing backdoor attacks against videos based on FL only poison RGB frames, which makes it that the attack could be easily mitigated by two-stream model neutralization. Therefore, it is a big challenge to manipulate the most advanced two-stream video model with a high success rate by poisoning only a small proportion of training data in the framework of FL. In this paper, a new backdoor attack scheme incorporating the rich spatial and temporal structures of video data is proposed, which injects the backdoor triggers into both the optical flow and RGB frames of video data through multiple rounds of model aggregations. In addition, the adversarial attack is utilized on the RGB frames to further boost the robustness of the attacks. Extensive experiments on real-world datasets verify that our methods outperform the state-of-the-art backdoor attacks and show better performance in terms of stealthiness and persistence. Jie Peng 0009, Weizhe Zhang, Jiangqun Ni, Arun Kumar Sangaiah, Aniello Castiglione |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Recursive Multi-Tree Construction With Efficient Rule Sifting for Packet Classification on FPGAabstractAs a programmable accelerator, SmartNIC provides more opportunities for algorithmic packet classification. Our aim in this work is to achieve both line-speed rule search and efficient rule update, two highly desired metrics for SDN data plane. We leverage the parallelism offered by the FPGA in SmartNIC following an algorithm/hardware co-design paradigm. Particularly, we first design an algorithm that constructs multiple trees for the rule set with a recursive rule sifting process. Unlike traditional space-cutting-based multi-tree construction, our rule sifting mechanism breaks the space constraints of rule-to-tree mapping and enables bounded height on each tree, thus providing the potential of bounded worst-case and line-speed performance. We then design a flexible hardware architecture with multiple systolic arrays that can be implemented in parallel on FPGA. Each systolic array works as a coarse-grained pipeline, and the multiple trees constructed earlier will be mapped onto these pipeline stages. This hardware-software mapping enables bounded worst-case rule searching. Additionally, incremental rule update is achieved simply by traversing the pipeline in one pass, with little and bounded impact on rule searching. Experimental results show that our design achieves an average classification throughput of 600.8/147.5 MPPS and an update throughput of 8.2/5.9 MUPS for 10k/100k-scale 5-tuple and OpenFlow rule sets. Yao Xin, Wenjun Li 0004, Chengjun Jia, Yang Xu 0010, Bin Liu 0001, Zhihong Tian 0001, Weizhe Zhang |
IEEE/ACM Trans. Netw. | 8 |
| 2024 | Multi-Attribute Auction-Based Grouped Federated LearningabstractFederated Learning empowers data owners to collectively train an artificial intelligence model without exposing data. However, the heterogeneous resources and the self-interested users bring new challenges hindering the development of federated learning. To this end, we propose a Multi-attribute Auction-based Grouped Federated Learning scheme, called MAGFL, comprising a grouped federated learning framework and a multi-attribute auction-based group selection strategy. Initially, our grouped federated learning framework clusters clients into groups according to local characteristics. Then, we propose a quality assessment method to assess the quality of each group based on a fuzzy approach. Furthermore, the FL server distributes economic rewards to training clients to motivate more clients to join the FL system, which is likened to a multi-attribute auction market where each group agent bids for training opportunities. Moreover, we design a novel global model update method with added Adam (i.e., Adaptive Moment Estimation) operations into the global update stage, which can fully utilize the local and global update direction to accelerate the convergence rate of scheme MGAFL. Extensive experiments on real-world datasets demonstrate that the proposed scheme outperforms representative federated learning schemes (i.e., FedAvg, FedProx, and FedAvg-Adam) regarding the model's convergence rate and capacity to deal with heterogeneous systems. Renhao Lu, Yan Wang 0002, Qiong Li 0001, Xiaoxiong Zhong, Weizhe Zhang |
IEEE Trans. Serv. Comput. | 7 |
| 2024 | CRPWarner: Warning the Risk of Contract-Related Rug Pull in DeFi Smart ContractsabstractIn recent years, Decentralized Finance (DeFi) has grown rapidly due to the development of blockchain technology and smart contracts. As of March 2023, the estimated global cryptocurrency market cap has reached approximately $949 billion. However, security incidents continue to plague the DeFi ecosystem, and one of the most notorious examples is the “Rug Pull” scam. This type of cryptocurrency scam occurs when the developer of a particular token project intentionally abandons the project and disappears with investors’ funds. Despite only emerging in recent years, Rug Pull events have already caused significant financial losses. In this work, we manually collected and analyzed 103 real-world rug pull events, categorizing them based on their scam methods. Two primary categories were identified:Contract-relatedRug Pull (through malicious functions in smart contracts) andTransaction-relatedRug Pull (through cryptocurrency trading without utilizing malicious functions). Based on the analysis of rug pull events, we propose CRPWarner (short forContract-relatedRugPull RiskWarner) to identify malicious functions in smart contracts and issue warnings regarding potential rug pulls. We evaluated CRPWarner on 69 open-source smart contracts related to rug pull events and achieved a 91.8% precision, 85.9% recall, and 88.7% F1-score. Additionally, when evaluating CRPWarner on 13,484 real-world token contracts on Ethereum, it successfully detected 4168 smart contracts with malicious functions, including zero-day examples. The precision of large-scale experiments reaches 84.9%. Zewei Lin, Jiachi Chen, Jiajing Wu, Weizhe Zhang, Yongjuan Wang, Zibin Zheng |
IEEE Trans. Software Eng. | 4 |
| 2024 | DRLCAP: Runtime GPU Frequency Capping With Deep Reinforcement LearningabstractPower and energy consumption is the limiting factor of modern computing systems. As the GPU becomes a mainstream computing device, power management for GPUs becomes increasingly important. Current works focus on GPU kernel-level power management, with challenges in portability due to architecture-specific considerations. We presentDRLCap, a general runtime power management framework intended to support power management across various GPU architectures. It periodically monitors system-level information to dynamically detect program phase changes and model the workload and GPU system behavior. This elimination from kernel-specific constraints enhances adaptability and responsiveness. The framework leverages dynamic GPU frequency capping, which is the most widely used power knob, to control the power consumption.DRLCapemploys deep reinforcement learning (DRL) to adapt to the changing of program phases by automatically adjusting its power policy through online learning, aiming to reduce the GPU power consumption without significantly compromising the application performance. We evaluateDRLCapon three NVIDIA and one AMD GPU architectures. Experimental results show thatDRLCapimproves prior GPU power optimization strategies by a large margin. On average, it reduces the GPU energy consumption by 22% with less than 3% performance slowdown on NVIDIA GPUs. This translates to a 20% improvement in the energy efficiency measured by the energy-delay product (EDP) over the NVIDIA default GPU power management strategy. For the AMD GPU architecture,DRLCapsaves energy consumption by 10%, on average, with a 4% percentage loss, and improves energy efficiency by 8%. Yiming Wang 0010, Meng Hao 0002, Weizhe Zhang, Qiuyuan Tang, Zheng Wang 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2024 | Model-Free GPU Online Energy OptimizationabstractGPUs play a central and indispensable role as accelerators in modern high-performance computing (HPC) platforms, enabling a wide range of tasks to be performed efficiently. However, the use of GPUs also results in significant energy consumption and carbon dioxide (CO2) emissions. This article presents MF-GPOEO, a model-free GPU online energy efficiency optimization framework. MF-GPOEO leverages a synthetic performance index and a PID controller to dynamically determine the optimal clock frequency configuration for GPUs. It profiles GPU kernel activity information under different frequency configurations and then compares GPU kernel execution time and gap duration between kernels to derive the synthetic performance index. With the performance index and measured average power, MF-GPOEO can use the PID controller to try different frequency configurations and find the optimal frequency configuration under the guidance of user-defined objective functions. We evaluate the MF-GPOEO by running it with 74 applications on an NVIDIA RTX3080Ti GPU. MF-GPOEO delivers a mean energy saving of 26.2% with a slight average execution time increase of 3.4% compared with NVIDIA's default clock scheduling strategy. Farui Wang, Meng Hao 0002, Weizhe Zhang, Zheng Wang 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2023 | Accelerated Genetic Algorithm with Population Control for Energy-Aware Virtual Machine Placement in Data Centers
Yu-Chu Tian, Maolin Tang, You-Gan Wang, Jiong Jin, Weizhe Zhang |
ICONIP (2) | 7 |
| 2023 | A Representation Learning Link Prediction Approach Using Line Graph Neural Networks
Yu Tai, Weizhe Zhang |
PRCV (9) | 5 |
| 2023 | Multi-modal Graph and Sequence Fusion Learning for Recommendation
Yu Tai, Weizhe Zhang |
PRCV (1) | 6 |
| 2023 | Multi-behavior Enhanced Graph Neural Networks for Social Recommendation
Anfeng Huang, Yu Tai, Weizhe Zhang |
PRCV (10) | 6 |
| 2023 | Defending Against Data Poisoning Attacks: From Distributed Learning to Federated LearningabstractAbstract Federated learning (FL), a variant of distributed learning (DL), supports the training of a shared model without accessing private data from different sources. Despite its benefits with regard to privacy preservation, FL’s distributed nature and privacy constraints make it vulnerable to data poisoning attacks. Existing defenses, primarily designed for DL, are typically not well adapted to FL. In this paper, we study such attacks and defenses. In doing so, we start from the perspective of DL and then give consideration to a real-world FL scenario, with the aim being to explore the requisites of a desirable defense in FL. Our study shows that (i) the batch size used in each training round affects the effectiveness of defenses in DL, (ii) the defenses investigated are somewhat effective and moderately influenced by batch size in FL settings and (iii) the non-IID data makes it more difficult to defend against data poisoning attacks in FL. Based on the findings, we discuss the key challenges and possible directions in defending against such attacks in FL. In addition, we propose detect and suppress the potential outliers(DSPO), a defense against data poisoning attacks in FL scenarios. Our results show that DSPO outperforms other defenses in several cases. Weizhe Zhang, Andrew C. Simpson, Yang Liu 0039, Zoe Lin Jiang |
Comput. J. | 2 |
| 2023 | Efficient intrusion detection toward IoT networks using cloud-edge collaboration
Yixiao Xu, Bangzhou Xin, Weizhe Zhang |
Comput. Networks | 7 |
| 2023 | An Energy-optimized Embedded load balancing using DVFS computing in Cloud Data centers
Amir Javadpour 0001, Arun Kumar Sangaiah, Pedro Pinto 0001, Forough Ja'fari, Weizhe Zhang, Ali Majed Hossein Abadi, Sayyed Hamid Reza Ahmadi 0001 |
Comput. Commun. | 5 |
| 2023 | Enhanced resource allocation in distributed cloud using fuzzy meta-heuristics optimization
Arun Kumar Sangaiah, Amir Javadpour 0001, Pedro Pinto 0001, Samira Rezaei, Weizhe Zhang |
Comput. Commun. | 5 |
| 2023 | Dependable federated learning for IoT intrusion detection against poisoning attacks
Weizhe Zhang |
Comput. Secur. | 5 |
| 2023 | FMSA: a meta-learning framework-based fast model stealing attack technique against intelligent network intrusion detection systemsabstractAbstract Intrusion detection systems are increasingly using machine learning. While machine learning has shown excellent performance in identifying malicious traffic, it may increase the risk of privacy leakage. This paper focuses on implementing a model stealing attack on intrusion detection systems. Existing model stealing attacks are hard to implement in practical network environments, as they either need private data of the victim dataset or frequent access to the victim model. In this paper, we propose a novel solution called Fast Model Stealing Attack (FMSA) to address the problem in the field of model stealing attacks. We also highlight the risks of using ML-NIDS in network security. First, meta-learning frameworks are introduced into the model stealing algorithm to clone the victim model in a black-box state. Then, the number of accesses to the target model is used as an optimization term, resulting in minimal queries to achieve model stealing. Finally, adversarial training is used to simulate the data distribution of the target model and achieve the recovery of privacy data. Through experiments on multiple public datasets, compared to existing state-of-the-art algorithms, FMSA reduces the number of accesses to the target model and improves the accuracy of the clone model on the test dataset to 88.9% and the similarity with the target model to 90.1%. We can demonstrate the successful execution of model stealing attacks on the ML-NIDS system even with protective measures in place to limit the number of anomalous queries. Kaisheng Fan, Weizhe Zhang, Guangrui Liu |
Cybersecur. | 2 |
| 2023 | An intelligent sustainable efficient transmission internet protocol to switch between User Datagram Protocol and Transmission Control Protocol in IoT computingabstractAbstract Today, Internet of things (IoT), Cloud and Fog networks have spread out around the world. The more these networks grow, the more their energy consumption comes to attention. Many efforts have been made during recent years to decrease this energy consumption, mainly focused on utilizing low‐power devices. Green algorithms are recently proposed to reduce energy consumption by modifying the structure of many algorithms employed in the network and its protocols. This paper proposes a new green reliability algorithm for Transmission Control Protocol/Internet Protocol (TCP/IP protocol) in Fog computing. The proposed algorithm does not require extensive TCP/IP protocol changes or relevant hardware. It is based on transferring less number of packets in the network by using the advantage of differences between TCP and User Datagram Protocol (UDP). TCP and User Datagram Protocol (UDP) are different in nature as the number of total packets in UDP is half that of TCP. As a result, the number of complete packets in UDP is half that of TCP. The proposed method is built around the loss of some packets in applications, such as voice and online video, does not severely degrade the end results. Therefore, the UDP protocol can substitute TCP in such situations. The criterion to switch between the two is the minimum acceptable Quality of Service (QoS) of the overall network. In other words, the UDP protocol will be used as long as QoS requirements are met. The switching process between UDP and TCP is dynamic, optimized by estimating network noise in the period. Additionally, we evaluated the proposed method based on several QoS functions, including delay, throughput, and energy usage. Shadi Mahmoodi Khaniabadi, Amir Javadpour 0001, Mehdi Gheisari, Weizhe Zhang, Yang Liu 0039, Arun Kumar Sangaiah |
Expert Syst. J. Knowl. Eng. | 4 |
| 2023 | A Collaborative Stealthy DDoS Detection Method Based on Reinforcement Learning at the Edge of Internet of ThingsabstractThe weaknesses of Internet of Things (IoT) devices leads to vulnerabilities easily, which can be exploited by criminals to launch Distributed Denial-of-Service (DDoS) attacks, becoming a major security hazard. Nowadays, the rapid development of the IoT makes the IoT-based DDoS attacks have the characteristics of wide distribution, large scale, and more stealthy that brings greater challenges for the DDoS detection. In this article, we conduct our research based on the edge side of IoT for providing earlier detection capability and more efficient resource utilization. We propose a novel reinforcement learning-based collaborative DDoS detection method and design a lightweight unsupervised classifier based on statistics. We deploy the classifiers in IoT edge gateways to detect anomalies by analyzing network traffic features in time. In order to deal with the dynamic changes of the IoT environment, we use the soft actor–critic (SAC) reinforcement learning model deployed on the edge server to adjust the parameter configuration of the underlying unsupervised classifier dynamically, which can ensure excellent detection effect for different types of IoT devices. In addition, a collaborative aggregation module is designed in the edge server to share the observation state and historical experience, which has a unique collaborative reward mechanism for the reinforcement learning model to fully mobilize the collaborative work capability. The experiments on public data set and constructed real-world testbed demonstrate that our proposed method has excellent detection performance and especially it can also discover stealthy IoT-based DDoS attacks accurately. Yuming Feng 0002, Weizhe Zhang, Shujun Yin, Yang Xiang 0001, Yu Zhang 0036 |
IEEE Internet Things J. | 2 |
| 2023 | Paired Swarm Optimized Relational Vector Learning for FDI Attack Detection in IoT-Aided Smart GridabstractIoT-aided smart grid heavily depends on the most innovative communication technologies that could make the grid system susceptible to false data injection attacks (FDIAs). The main objective of the FDI attackers remains in damaging or corrupting the state estimation strategy in the smart grid resulting in blackouts and/or to influence the electricity market. With a number of features involved in the smart grid system, FDIA detection is said to be complicated. By the conventional bad data detection systems, the FDIA detection accuracy and validation made by the receiver operating characteristic (ROC) curve were marginally acceptable. However, due to the time complexity and overhead incurred, the detection of FDIA is a hot research topic. In this work, we design an efficient FDIA detection method by coupling cooperative paired swarm optimization and relational vector learning techniques (CPSO-RVL), to address the above-said issues. First, the cooperative paired particle swarm optimization model is proposed to attain an appropriate feature for improving the computational efficiency of FDIA detection. Next, with the obtained significant features, the relational vector learning-based FDIA detection model is designed for robust classification between FDIA and non-FDIA with minimum overhead. The extensive experiments show that the proposed method outperforms existing baseline approaches by 16% and 34% in terms of computation time and computation overhead, respectively. Sumarga Kumar Sah Tyagi, Deepak Kumar Jain 0001, Yi-Cheng Tu, Weizhe Zhang |
IEEE Internet Things J. | 5 |
| 2023 | Predicting information diffusion using the inter- and intra-path of influence transitivity
Yu Tai, Weizhe Zhang, Yan Wang 0002 |
Inf. Sci. | 3 |
| 2023 | SVScanner: Detecting smart contract vulnerabilities via deep semantic extraction
Hengyan Zhang, Weizhe Zhang, Yuming Feng 0002, Yang Liu 0039 |
J. Inf. Secur. Appl. | 2 |
| 2023 | Accelerated computation of the genetic algorithm for energy-efficient virtual machine placement in data centersabstractAbstract Energy efficiency is a critical issue in the management and operation of cloud data centers, which form the backbone of cloud computing. Virtual machine (VM) placement has a significant impact on energy-efficiency improvement for virtualized data centers. Among various methods to solve the VM-placement problem, the genetic algorithm (GA) has been well accepted for the quality of its solution. However, GA is also computationally demanding, particularly in the computation of its fitness function. This limits its application in large-scale systems or specific scenarios where a fast VM-placement solution of good quality is required. Our analysis in this paper reveals that the execution time of the standard GA is mostly consumed in the computation of its fitness function. Therefore, this paper designs a data structure extended from a previous study to reduce the complexity of the fitness computation from quadratic to linear one with respect to the input size of the VM-placement problem. Incorporating with this data structure, an alternative fitness function is proposed to reduce the number of instructions significantly, further improving the execution-time performance of GA. Experimental studies show that our approach achieves 11 times acceleration of GA computation for energy-efficient VM placement in large-scale data centers with about 1500 physical machines in size. Yu-Chu Tian, You-Gan Wang, Weizhe Zhang |
Neural Comput. Appl. | 4 |
| 2023 | A blockchain-based privacy-preserving advertising attribution architecture: Requirements, design, and a prototype implementationabstractAbstract In the era of digital marketing, advertisements have become an indispensable part. One of the central challenges is advertising attribution which explains the amount of contribution every publisher has with the conversions. However, through observation, we have found that current advertising platforms attribution, advertisers attribution, or third‐party platforms attribution all have the problems of trust, data leakage, and data forgery. To fill the gap, our work's main contribution is combining blockchain with advertising attribution to propose an architecture for improving the privacy‐preserving degree and amount. In the proposed architecture, publishers, and advertisers can store real‐time data on a blockchain. The attribution results are credible because blockchain is decentralized, tamper‐proof, and traceable. We combine privacy set intersection and zero‐knowledge proof technology to increase the privacy of flowing data. In addition, we describe a preliminary prototype in which publishers, advertisers, and advertising platforms can get the corresponding attribution details. To show its effectiveness, we analyze it from different perspectives, including communication cost, attribution accuracy, and time cost. The results show that our communication cost has significantly reduced compared to the recent studies. Yang Liu 0039, Liangjie Lin, Weizhe Zhang, Xuan Wang 0002, Mehdi Gheisari, Hamid Esmaeili Najafabadi |
Softw. Pract. Exp. | 4 |
| 2023 | A transformer-based approach for improving app review response generationabstractAbstract Mobile apps are becoming an integral part of people's daily life by providing various functionalities, such as messaging and gaming. App developers try their best to ensure user experience during app development and maintenance to improve the rating of their apps on app platforms and attract more user downloads. Previous studies indicated that responding to users' reviews tends to change their attitude towards the apps positively. Users who have been replied are more likely to update the given ratings. However, reading and responding to every user review is not an easy task for developers since it is common for popular apps to receive tons of reviews every day. Thus, automation tools for review replying are needed. To address the need above, the paper introduces a Transformer‐based approach, named TRRGen, to automatically generate responses to given user reviews. TRRGen extracts apps' categories, rating, and review text as the input features. By adapting a Transformer‐based model, TRRGen can generate appropriate replies for new reviews. Comprehensive experiments and analysis on the real‐world datasets indicate that the proposed approach can generate high‐quality replies for users' reviews and significantly outperform current state‐of‐art approaches on the task. The manual validation results on the generated replies further demonstrate the effectiveness of the proposed approach. Weizhe Zhang, Cuiyun Gao 0001, Michael R. Lyu |
Softw. Pract. Exp. | 1 |
| 2023 | Smart-DNN+: A Memory-efficient Neural Networks Compression Framework for the Model InferenceabstractDeep Neural Networks (DNNs) have achieved remarkable success in various real-world applications. However, running a Deep Neural Network (DNN) typically requires hundreds of megabytes of memory footprints, making it challenging to deploy on resource-constrained platforms such as mobile devices and IoT. Although mainstream DNNs compression techniques such as pruning, distillation, and quantization can reduce the memory overhead of model parameters during DNN inference, they suffer from three limitations: (i) low model compression ratio for the lightweight DNN structures with little redundancy, (ii) potential degradation in model inference accuracy, and (iii) inadequate memory compression ratio is attributable to ignoring the layering property of DNN inference. To address these issues, we propose a lightweight memory-efficient DNN inference framework called Smart-DNN+, which significantly reduces the memory costs of DNN inference without degrading the model quality. Specifically, ① Smart-DNN+ applies a layerwise binary-quantizer with a remapping mechanism to greatly reduce the model size by quantizing the typical floating-point DNN weights of 32-bit to the 1-bit signs layer by layer. To maintain model quality, ② Smart-DNN+ employs a bucket-encoder to keep the compressed quantization error by encoding the multiple similar floating-point residuals into the same integer bucket IDs. When running the compressed DNN in the user’s device, ③ Smart-DNN+ utilizes a partially decompressing strategy to greatly reduce the required memory overhead by first loading the compressed DNNs in memory and then dynamically decompressing the required materials for model inference layer by layer. Experimental results on popular DNNs and datasets demonstrate that Smart-DNN+ achieves lower 0.17%–0.92% memory costs at lower runtime overheads compared with the states of the art without degrading the inference accuracy. Moreover, Smart-DNN+ potentially reduces the inference runtime up to 2.04× that of conventional DNN inference workflow. Donglei Wu, Weihao Yang, Xiangyu Zou, Wen Xia, Zhenbo Hu, Weizhe Zhang, Binxing Fang |
ACM Trans. Archit. Code Optim. | 7 |
| 2023 | A Multi-Objective Virtual Network Migration Algorithm Based on Reinforcement LearningabstractVirtual network migration (VNM) helps improve network performance by remapping a subset of virtual nodes or links to physical infrastructure, aligning the resource allocation to the virtual network's changing conditions. However, existing VNM methods neglect integrating multiple objectives that affect network performance, such as energy, communication, migration, and service level agreement violation (SLAV). It is challenging to make VNM decisions to optimize the overall objective in a large-scale cloud environment. This article establishes a multi-objective optimization model and proposes a multi-objective VNM algorithm called MiOvnm. The MiOvnm employs the double deep$Q$-learning approach to cope with ample state space. It also applies an action selection method called actfilter to deal with large-scale action space. The MiOvnm finds the migration action with optimal potential reward from the candidate action set. Simulation results demonstrate the superiority of our MiOvnm to the state-of-the-art methods. More specifically, MiOvnm reduces average SLAV, communication cost, and total cost by 24.32%, 4.95%, and 12.45%, respectively. Furthermore, evaluation results in a real-world OpenStack platform reveal that making full use of computation and network resources, the MiOvnm reduces the completion time of computation- and network-intensive benchmarks by 11.35% and 10.31%, respectively, with a total cost reduction of 26.02%. Desheng Wang 0002, Weizhe Zhang, Junren Lin, Yu-Chu Tian |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Adversarial ELF Malware Detection Method Using Model InterpretationabstractRecent research shows that executable and linkable format (ELF) malware detection models based on deep learning are vulnerable to adversarial attacks. The most commonly used method in previous work is adversarial training to defend adversarial examples. Nevertheless, it is inefficient and only effective for specific adversarial attacks. Given that the perturbation byte insertion positions of existing adversarial malware generation methods are relatively fixed, we propose a new method to detect adversarial ELF malware. Using model interpretation techniques, we analyze the decision-making basis of the malware detection model and extract the features of adversarial examples. We further use anomaly detection techniques to identify adversarial examples. As an add-on module of the malware detection model, the proposed method does not require modifying the original model and does not need to retrain the model. Evaluating results show that the method can effectively defend the adversarial attacks against the malware detection model. Yanchen Qiao, Weizhe Zhang, Zhicheng Tian, Laurence T. Yang, Yang Liu 0039, Mamoun Alazab |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Hierarchical Clustering Based on Dendrogram in Sustainable Transportation SystemsabstractEach group in a data-driven automobile network has its cluster head. A group can communicate with each other and members of other groups once it has been founded. Vehicles belonging to each group near the other group allow intergroup communication. Because nodes in automotive networks move so quickly, routing in these networks is a complex problem to solve. Each cluster in hierarchical clustering can be partitioned into multiple sub-clusters. Put another way, and the data is stored in a cluster, which is then divided into more clusters. The data is stored directly in separate clusters in non-hierarchical approaches. A dendrogram is a type of hierarchical tree. We anticipate increasing information sharing in clusters by properly clustering vehicles on the road and establishing clusters of the desired size in the relevant dendrogram. We can select clusters of the necessary extent and compare the Quality of Service (QoS) network’s outcomes by breaking the dendrogram at different levels. The findings reveal that the suggested method outperforms AIVISN in delay, PDR, overhead, and Drooped packets compared to AIVISN, 7.12%, 12.21%,8.32%, and 7.34%, respectively. Arun Kumar Sangaiah, Amir Javadpour 0001, Forough Ja'fari, Weizhe Zhang, Shadi Mahmoodi Khaniabadi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | A Model-Based Method for Enabling Source Mapping and Intrusion Detection on Proprietary Can BusabstractWith the deep integration of the Internet of Things (IoT) technology and the increase of computational power and memory, vehicles can also serve as the infrastructures for Intelligent Transportation System (ITS), e.g., as fog nodes. However, when connecting vehicles to the internet, alongside with the benefits it brings, it also opens many new challenges such as security attacks. Controller Area Network (CAN) is one of the main in-vehicle communication protocols in modern cars. Its lack of sender verification mechanism makes CAN particularly vulnerable to cyber-attacks including masquerade attack. Fingerprinting Electronic Control Units (ECUs) based on hardware characteristics has been proved feasible and effective on defending CAN buses. However, most state-of-the-art works exploited the supervised learning algorithm to identify the transmitter based on the signal characteristics. This makes the decision process hard to understand, and it also limits the deployment on proprietary CAN bus without prior knowledge. To solve this, we design a novel clock-skew-based approach capable of pinpointing the sender and detecting intrusion on proprietary CAN bus. We take a single CAN frame as the object for measurement, and adjust the measuring process such that our approach can be independent of the transmission time of frames. Based on the statistical analysis of data from real vehicles, we propose a novel box-plot algorithm based on score mechanism to filter the raw data. Finally, the clock skews are estimated and accumulated to build a linear model for representing the transmitter ECU. The evaluation results on one CAN prototype and two production vehicles show that our approach is able to well identify and differentiate ECUs on the bus without prior knowledge. The data processed by the proposed box-plot algorithm can describe the hardware characteristics of ECUs precisely. We also show the ability of our approach to protecting the CAN bus against the masquerade attack. Jia Zhou 0003, Guoqi Xie, Haibo Zeng 0001, Weizhe Zhang, Laurence T. Yang, Mamoun Alazab, Renfa Li |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | MSDS: A Novel Framework for Multi-Source Data Selection Based Cross-Network Node ClassificationabstractIn this paper, we study the problem of multi-source cross-network node classification, which aims to classify unlabeled nodes in a target network by leveraging the knowledge learned from the rich labeled nodes in multiple source networks. The existing multi-source transfer learning approaches generally fail to model the structural information of networks, and the current cross-network node classification models mainly neglect that not all source networks can boost the task performance in the target network. Thus, none can be directly applied to the multi-source cross-network node classification task. To this end, in this paper, we propose a novel multi-source data selection (MSDS) based framework for cross-network node classification, which integrates multi-source transfer learning with network embedding to learn label-discriminative and network-invariant node representations. In MSDS, we first propose the multi-source network data selection, which applies three distances to jointly select the transferable source networks to well alleviate the problem of suboptimal solution or even negative transfer. In addition, we devise a new feature information alignment technique to make node vector representations network-invariant. Moreover, we incorporate aggregated structural information and feature information to make node representations label-discriminative. Extensive experiments on real-world datasets demonstrate that the proposed approaches outperform the state-of-the-art non-transfer and single-source transfer approaches in terms of classification accuracy. Weizhe Zhang, Yan Wang 0002, Zhaonian Zou |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | CoopCon: Cooperative Hybrid Congestion Control Scheme for Named Data NetworkingabstractCongestion control is a key technology for guaranteeing quality-of-service (QoS) in Named Data Networking (NDN). Hybrid congestion control has gradually developed into the mainstream method in NDN congestion control, which capitalizes on the advantages of receiver adjusting rate and router diverting traffic to deal with congestion. However, it has to address how to effectively coordinate consumers and routers to prevent transport performance degradation caused by repeated and excessive control. In this paper, a hybrid congestion control scheme named CoopCon is proposed, which fully gives the cooperation between consumers and routers to control the congestion adaptively. Moreover, the optimal path is used resiliently by CoopCon to enhance the robustness of multipath forwarding and multicast data delivery in NDN. The proposed CoopCon is implemented in ndnSIM. And simulation results show that CoopCon consistently achieves higher total throughput than existing work. In particular, the total throughput of consumers deployed with CoopCon is 19.4% higher than that of consumers deployed with PCON in the BRITE-generated topology. Additionally, CoopCon also achieves the best fairness in a dumbbell topology, with a fairness index of even 0.96. Zhuo Li 0009, Xingdi Shen, Hao Xun, Weizhe Zhang, Peng Luo 0004 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2023 | Auction-Based Cluster Federated Learning in Mobile Edge Computing SystemsabstractFederated Learning (FL), allowing data owners to conduct model training without sending their raw data to third-party servers, can enhance data privacy in Mobile Edge Computing (MEC) which brings data processing closer to the data sources. However, the heterogeneity of local data and constrained local resources in MEC bring new challenges hindering the development of FL. To this end, we propose an Auction-based Cluster Federated Learning scheme, called ACFL, comprising a clustered FL framework and an auction-based client selection strategy. Our clustered FL framework first introduces a mean-shift clustering algorithm to FL, which can intelligently cluster clients according to their local data distribution. Then, we select clients from each cluster using an auction mechanism to participate in FL training, which can mitigate the impact of data heterogeneity on model convergence and balance energy consumption. Moreover, we prove the proposed clustered FL framework converges at a sublinear rate. Extensive experiments conducted on real-world datasets demonstrate that the proposed FL scheme outperforms the conventional FL schemes in terms of convergence rate and energy balance. Renhao Lu, Weizhe Zhang, Yan Wang 0002, Qiong Li 0001, Xiaoxiong Zhong, Desheng Wang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | TDTA: Topology-Based Real-Time DAG Task Allocation on Identical Multiprocessor PlatformsabstractModern real-time systems contain complex workloads, which are usually modeled as directed acyclic graph (DAG) tasks and deployed on multiprocessor platforms. The complex execution logic of DAG tasks results in excessive schedulability analysis overhead, and the current DAG task allocation strategy cannot efficiently utilize processor resources (inner parallelization of DAG tasks). In this article, an invalid-edge deletion (IED) method is proposed to reduce the execution complexity of the DAG tasks while guaranteeing the correctness of the execution logic. Besides, we bound the number of complete paths for DAG tasks, which re-limits the searching space of the schedulability analysis. Then, a topology-based DAG tasks allocation (TDTA) strategy is developed, which reduces the interference caused by higher-priority DAG tasks to enable the full utilization of the processor resources. The experimental results show that the IED method effectively reduces the overhead of DAG task analysis, and the performance of the TDTA strategy is better than the performance of other state-of-the-art strategies. Weizhe Zhang, Nan Guan, Yehan Ma |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | PDA-GNN: propagation-depth-aware graph neural networks for recommendation
Yu Tai, Weizhe Zhang |
World Wide Web (WWW) | 6 |
| 2022 | DNN Real-Time Collaborative Inference Acceleration with Mobile Edge ComputingabstractThe collaborative inference approach splits the Deep Neural Networks (DNNs) model into two parts. It runs collaboratively on the end device and cloud server to minimize inference latency and protect data privacy, especially in the 5G era. The scheme of DNN model partitioning depends on the network bandwidth size. However, in the context of dynamic mobile networks, resource-constrained devices cannot efficiently execute complex model partitioning algorithms to obtain optimal partitioning in real-time. In this paper, to overcome this challenge, we first formulate the model partitioning problem as a Min-cut problem to seek the optimal partition. Second, we propose a Collaborative Inference method based on model Compression named CIC. CIC enhances the efficiency of the execution of model partitioning algorithms on resource-constrained end devices by reducing the algorithm's complexity. CIC generates a splitting model based on the inherent characteristics of the DNN model and the platform resources. The splitting models are independent of the network environment, generated offline, and constantly used in the current environment. CIC has excellent compressibility, and even DNN models with hundreds of layers can be rapidly partitioned on resource-constrained devices. Experimental results show that our method is significantly more effective than existing solutions, speeding up model partitioning decision time by up to 100x, reducing inference latency by up to 2.6x, and increasing throughput by up to 3.3x in the best case. Yan Li 0075, Weizhe Zhang |
IJCNN | 4 |
| 2022 | Lie group manifold analysis: an unsupervised domain adaptation approach for image classification
Weizhe Zhang, Yawen Bai |
Appl. Intell. | 3 |
| 2022 | VulnerGAN: a backdoor attack through vulnerability amplification against machine learning-based network intrusion detection systems
Guangrui Liu, Weizhe Zhang, Kaisheng Fan, Shui Yu 0001 |
Sci. China Inf. Sci. | 2 |
| 2022 | Adversarial malware sample generation method based on the prototype of deep learning detector
Yanchen Qiao, Weizhe Zhang, Zhicheng Tian, Laurence T. Yang, Yang Liu 0039, Mamoun Alazab |
Comput. Secur. | 2 |
| 2022 | MTGK: Multi-source cross-network node classification via transferable graph knowledge
Weizhe Zhang, Yawen Bai |
Inf. Sci. | 3 |
| 2022 | You are what the permissions told me! Android malware detection based on hybrid tactics
Huanran Wang, Weizhe Zhang |
J. Inf. Secur. Appl. | 2 |
| 2022 | Traffic flow control using multi-agent reinforcement learning
Ahmad Zeynivand, Amir Javadpour 0001, S. Bolouki, Arun Kumar Sangaiah, Forough Ja'fari, Pedro Pinto 0001, Weizhe Zhang |
J. Netw. Comput. Appl. | 7 |
| 2022 | Evading generated-image detectors: A deep dithering approach
Hao Xie 0002, Jiangqun Ni, Jian Zhang 0086, Weizhe Zhang, Jiwu Huang |
Signal Process. | 4 |
| 2022 | Corrigendum to 'Evading generated-image detectors: A deep dithering approach' [Signal Processing 197(2022) 108558]
Hao Xie 0002, Jiangqun Ni, Jian Zhang 0086, Weizhe Zhang, Jiwu Huang |
Signal Process. | 4 |
| 2022 | Improving Interference Analysis for Real-Time DAG Tasks Under Partitioned SchedulingabstractReal-time systems with strict timing constraints have been widely applied in many fields. The Directed acyclic graph (DAG) task model has been widely studied and applied to model real-time systems with partial parallelism and precedence constraints in each task. Our paper focuses on the worst-case response time (WCRT) analysis of DAG tasks under partitioned scheduling on multiprocessors. We investigate a parallel structure named$Str$, which helps obtain more accurate analysis results, and propose a new offline scheduling analysis algorithm named reducing repetitive calculation (RRC). Experiments with synthetic workload are conducted to compare the results calculated by RRC and the state-of-the-art, as well as the observed average response time on a real embedded system. Results show that RRC has better performance in terms of analysis accuracy. Weizhe Zhang, Nan Guan, Yue Tang 0001 |
IEEE Trans. Computers | 2 |
| 2022 | Blockchain-Based Decentralized Public Auditing for Cloud StorageabstractPublic auditing schemes for cloud storage systems have been extensively explored with the increasing importance of data integrity. A third-party auditor (TPA) is introduced in public auditing schemes to verify the integrity of outsourced data on behalf of users. To resist malicious TPAs, many blockchain-based public verification schemes have been proposed. However, existing auditing schemes rely on a centralized TPA, and they are vulnerable to tempting auditors who may collude with malicious blockchain miners to produce biased auditing results. In this article, we propose a blockchain-based decentralized public auditing (BDPA) scheme by utilizing a decentralized blockchain network to undertake the responsibility of a centralized TPA, and also mitigate the influence of tempting auditors and malicious blockchain miners by taking the concept of decentralized autonomous organization (DAO). A detailed security analysis shows that BDPA can preserve data integrity against tempting auditors and malicious blockchain miners. A comprehensive performance evaluation demonstrates that BDPA is feasible and scalable. Jiangang Shu, Xing Zou, Xiaohua Jia, Weizhe Zhang, Ruitao Xie |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | A Novel Video Steganographic Scheme Incorporating the Consistency Degree of Motion VectorsabstractIn this letter, a novel steganographic scheme in motion vector domain (MV) for H.264 video is presented, which can significantly improve the security performance against the newly emerged powerful multi-domain feature set MVC (motion vector consistency). By taking into account both the consistency degree of motion vectors for sub-blocks within a macroblock (MB) or sub-macroblock (sub-MB), and the MV statistics, the corresponding distortion function called dMVC is proposed. The proposed dMVC is also shown to be capable of integrating with existing methods to resist the joint steganalytic attacks of both MVC feature and local optimality features, e.g., NPELO, in the framework of minimal distortion embedding. Compared with other state-of-the-art MV-based steganographic schemes, experimental results on YUV sequences at various embedding rates and QPs show that the proposed method gains significant performance improvement while maintaining good coding efficiency. Ying Liu 0062, Jiangqun Ni, Weizhe Zhang, Jiwu Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Consortium Blockchain-Based Access Control Framework With Dynamic Orderer Node Selection for 5G-Enabled Industrial IoTabstract5G-enabled Industrial Internet of Things (IIoT) deployment will bring more severe security and privacy challenges, which puts forward higher requirements for access control. Blockchain-based access control method has become a promising security technology, but it still faces high latency in consensus process and weak adaptability to dynamic changes in network environment. This article proposes a novel access control framework for 5G-enabled IIoT based on consortium blockchain. We design three types of chaincodes for the framework named policy management chaincode (PMC), access control chaincode (ACC), and credit evaluation chaincode (CEC). The PMC and ACC are deployed on the same data channel to implement the management of access control policies and the authorization of access. The CEC deployed on another channel is used to add behavior records collected from IIoT devices and calculate the credit value of IIoT domain. Specifically, we design a two-step credit-based Raft consensus mechanism, which can select the orderer nodes dynamically to achieve fast and reliable consensus based on historical behavior records stored in the ledger. Furthermore, we implement the proposed framework on a real-world testbed and compare it with the framework based on practical Byzantine fault tolerance consensus. The experiment results show that our proposed framework can maintain lower consensus cost time with 100 ms level and achieves four to five times throughput with lower hardware resource consumption and communication consumption. Besides, our design also improves the security and robustness of the access control process. Yuming Feng 0002, Weizhe Zhang, Xiapu Luo, Bin Zhang 0048 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Two-Phase Industrial Manufacturing Service Management for Energy Efficiency of Data CentersabstractData-driven industrial manufacturing services are proliferating. They use large amounts of data generated from Industrial-Internet-of-Things (IIoT) devices for intelligent services to end-service-users. However, cloud data centers hosting these services consume a huge amount of energy, resulting in a high operational cost. To address this issue, an energy-efficient resource allocation framework is proposed in this article for cloud services. It operates in two phases. First, a multithreshold-based host CPU utilization classification scheme is developed to classify hosts into four groups for improved CPU resource allocation. It is designed through analyzing CPU utilization data by using the least median squares regression technique. Thereby, the scheme limits search space, thus reducing time complexity. In the second phase, with a metaheuristic search, an energy- and thermal-aware resource allocation method is developed to find an energy-efficient host for allocating resources to services. From real data center workload traces, extensive experiments show that our framework outperforms existing baseline approaches with 6.9%, 33.75%, and 34.1% on average in terms of temperature, energy consumption, and service-level-agreement violation, respectively. Weizhe Zhang, Yu-Chu Tian, Sumarga Kumar Sah Tyagi, Ibrahim A. Elgendy, Omprakash Kaiwartya |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Efficient Lane-Level Map Building via Vehicle-Based CrowdsourcingabstractBy providing rich context of lane information on roads, lane-level maps play a vital role in intelligent transportation systems. Since Global Positioning Systems (GPS) have been widely applied to vehicles, vehicle-based crowdsourcing offers an economical way to the lane-level map building by collecting and analyzing the GPS trajectories of vehicles. However, existing works cannot directly extract lane-level road information from raw and interleaved crowdsourcing trajectories, and moreover they are time-consuming and inaccurate. In this article, we propose a lane-level map building scheme, which can directly extract lane-level road information from raw crowdsourcing GPS trajectories with both efficiency and accuracy improvement. Consider the global similarity between trajectories, we design an efficient trajectory segmentation and clustering algorithm based on improved discrete Fréchet distance and entropy theory, which can directly and accurately deal with the interleaved and messy trajectories. To improve the efficiency, we employ the Least Square Estimate (LSE) to constrain Gaussian Mixture Model (GMM) and design an efficient and accurate lane-level road information extraction algorithm. Comprehensive comparative experiments and performance evaluation on a real-world trajectory dataset show that the proposed scheme outperforms the state-of-the-art works in terms of both efficiency and accuracy. Jiangang Shu, Songlei Wang, Xiaohua Jia, Weizhe Zhang, Ruitao Xie, Hejiao Huang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Secure Task Offloading in Blockchain-Enabled Mobile Edge Computing With Deep Reinforcement LearningabstractMobile Edge Computing (MEC) is a promising and fast-developing paradigm that provides cloud services at the edge of the network. MEC enables IoT devices to offload and execute their real-time applications at the proximity of these devices with low latency. Such applications include efficient manufacture inspection, virtual/augmented reality, image recognition, Internet of Vehicles (IoV), and e-Health. However, task offloading experiences security and privacy attacks such as data tampering, private data leakage, data replication, etc. To this end, in this paper, we propose a new blockchain-based framework for secure task offloading in MEC systems with guaranteed performance in terms of execution delay and energy consumption. First, blockchain technology is introduced as a platform to achieve data confidentiality, integrity, authentication, and privacy of task offloading in MEC. Second, we formulate an integration model of resource allocation and task offloading for a multi-user with multi-task MEC systems to optimize the energy and time cost. This is an NP-hard problem because of the curse-of-dimensionality and dynamic characteristics challenges of the considered scenario. Therefore, a deep reinforcement learning-based algorithm is developed to derive the close-optimal task offloading decision efficiently. Theoretical analysis and experimental results demonstrate that the proposed framework is resilient to several task offloading security attacks and it can save about 22.2% and 19.4% of system consumption with respect to the local and edge execution scenarios. Moreover, the benchmark analysis proves that the framework consumes few resources in terms of memory and disk usage, CPU utilization, and transaction throughput. Ahmed Samy, Ibrahim A. Elgendy, Haining Yu, Weizhe Zhang, Hongli Zhang 0001 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2022 | Malware Classification Based on Multilayer Perception and Word2Vec for IoT SecurityabstractWith the construction of smart cities, the number of Internet of Things (IoT) devices is growing rapidly, leading to an explosive growth of malware designed for IoT devices. These malware pose a serious threat to the security of IoT devices. The traditional malware classification methods mainly rely on feature engineering. To improve accuracy, a large number of different types of features will be extracted from malware files in these methods. That brings a high complexity to the classification. To solve these issues, a malware classification method based on Word2Vec and Multilayer Perception (MLP) is proposed in this article. First, for one malware sample, Word2Vec is used to calculate a word vector for all bytes of the binary file and all instructions in the assembly file. Second, we combine these vectors into a 256x256x2-dimensional matrix. Finally, we designed a deep learning network structure based on MLP to train the model. Then the model is used to classify the testing samples. The experimental results prove that the method has a high accuracy of 99.54%. Yanchen Qiao, Weizhe Zhang, Xiaojiang Du, Mohsen Guizani |
ACM Trans. Internet Techn. | 2 |
| 2022 | Improving Quality of Service in 5G Resilient Communication with the Cellular Structure of SmartphonesabstractRecent studies in information computation technology (ICT) are focusing on Next-generation networks, SDN (Software-defined networking), 5G, and 6G. Optimal working mode for device-to-device (D2D) communication is aimed at improving the quality of service with the frequency spectrum structure is of research areas in 5G. D2D communication working modes are selected to meet both the predefined system conditions and provide maximum throughput for the network. Due to the complexity of the direct solutions, we formulated the problem as an optimization problem and found the optimal working modes under different parameters of the system through extensive simulations. After determining the links’ optimal modes, we calculated the network throughput; because of selecting the best working modes, we obtained the highest throughput. A major finding from this research is that D2D communication pairs are more inclined to use full-duplex (FD) mode in short distances to meet system requirements, and so most communications take place in FD mode at these distances. According to these results, using FD communication at short distances offers better conditions and Quality of service (QoS) than QoS-D2D method. Arun Kumar Sangaiah, Amir Javadpour 0001, Pedro Pinto 0001, Forough Ja'fari, Weizhe Zhang |
ACM Trans. Sens. Networks | 5 |
| 2022 | Optimizing Depthwise Separable Convolution Operations on GPUsabstractThe depthwise separable convolution is commonly seen in convolutional neural networks (CNNs), and is widely used to reduce the computation overhead of a standard multi-channel 2D convolution. Existing implementations of depthwise separable convolutions target accelerating model training with large batch sizes with a large number of samples to be processed at once. Such approaches are inadequate for small-batch-sized model training and the typical scenario of model inference where the model takes in a few samples at once. This article aims to bridge the gap of optimizing depthwise separable convolutions by targeting the GPU architecture. We achieve this by designing two novel algorithms to improve the column and row reuse of the convolution operation to reduce the number of memory operations performed on the width and the height dimensions of the 2D convolution. Our approach employs a dynamic tile size scheme to adaptively distribute the computational data across GPU threads to improve GPU utilization and to hide the memory access latency. We apply our approach on two GPU platforms: an NVIDIA RTX 2080Ti GPU and an embedded NVIDIA Jetson AGX Xavier GPU, and two data types: 32-bit floating point (FP32) and 8-bit integer (INT8). We compared our approach against cuDNN that is heavily tuned for the NVIDIA GPU architecture. Experimental results show that our approach delivers over 2× (up to 3×) performance improvement over cuDNN. We show that, when using a moderate batch size, our approach averagely reduces the end-to-end training time of MobileNet and EfficientNet by 9.7 and 7.3 percent respectively, and reduces the end-to-end inference time of MobileNet and EfficientNet by 12.2 and 11.6 percent respectively. Gangzhao Lu, Weizhe Zhang, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Online Power Management for Multi-Cores: A Reinforcement Learning Based ApproachabstractPower and energy is the first-class design constraint for multi-core processors and is a limiting factor for future-generation supercomputers. While modern processor design provides a wide range of mechanisms for power and energy optimization, it remains unclear how software can make the best use of them. This article presents a novel approach for runtime power optimization on modern multi-core systems. Our policy combines power capping and uncore frequency scaling to match the hardware power profile to the dynamically changing program behavior at runtime. We achieve this by employing reinforcement learning (RL) to automatically explore the energy-performance optimization space from training programs, learning the subtle relationships between the hardware power profile, the program characteristics, power consumption and program running times. Our RL framework then uses the learned knowledge to adapt the chip's power budget and uncore frequency to match the changing program phases for any new, previously unseen program. We evaluate our approach on two computing clusters by applying our techniques to 11 parallel programs that were not seen by our RL framework at the training stage. Experimental results show that our approach can reduce the system-level energy consumption by 12 percent, on average, with less than 3 percent of slowdown on the application performance. By lowering the uncore frequency to leave more energy budget to allow the processor cores to run at a higher frequency, our approach can reduce the energy consumption by up to 17 percent while improving the application performance by 5 percent for specific workloads. Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Dynamic GPU Energy Optimization for Machine Learning Training WorkloadsabstractGPUs are widely used to accelerate the training of machine learning workloads. As modern machine learning models become increasingly larger, they require a longer time to train, leading to higher GPU energy consumption. This paper presents GPOEO, an online GPU energy optimization framework for machine learning training workloads. GPOEO dynamically determines the optimal energy configuration by employing novel techniques for online measurement, multi-objective prediction modeling, and search optimization. To characterize the target workload behavior, GPOEO utilizes GPU performance counters. To reduce the performance counter profiling overhead, it uses an analytical model to detect the training iteration change and only collects performance counter data when an iteration shift is detected. GPOEO employs multi-objective models based on gradient boosting and a local search algorithm to find a trade-off between execution time and energy consumption. We evaluate the GPOEO by applying it to 71 machine learning workloads from two AI benchmark suites running on an NVIDIA RTX3080Ti GPU. Compared with the NVIDIA default scheduling strategy, GPOEO delivers a mean energy saving of 16.2% with a modest average execution time increase of 5.1%. Farui Wang, Weizhe Zhang, Shichao Lai, Meng Hao 0002, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Repeatable Multi-Dimensional Virtual Network Embedding in Cloud Service PlatformabstractVirtual network embedding (VNE) can effectively deploy virtual networks (VNs) onto shared substrate network (SN) resources. However, with the consistent changing scalability and diversity demands of VNs, traditional VNE methods prove to be a challenging task for current cloud service platforms. Thus, we model a repeatable multi-dimensional virtual network embedding (RMD-VNE) problem for implementing multi-dimensional virtual networks (MD-VNs) that involves real servers, virtual machines, containers, and network simulators. The MD-VN is preprocessed and embedded via a heuristic method denoted asReMiDvne. Following its transformation for the containers and simulation networks, the MD-VN topology undergoes a process of coarsening, partitioning, and uncoarsening.ReMiDvnethen applies a topology-aware repeatable embedding solution to complete the embedding stage. Experimental results demonstrate thatReMiDvneoutperforms seven baseline approaches through small-, 1,000- and 10,000-scale VNE simulation experiments. Remarkably,ReMiDvneimproves the average rates of acceptance ratio, revenue, and revenue-cost ratio by up to 40.45, 40.45, and 299.03 percent, respectively, and reduces the average rate of cost by up to 64.16 percent. Furthermore, real-world VNE experiments are conducted based on the OpenStack platform. The results reveal the ability ofReMiDvneto efficiently reduce communication costs by up to 45.93 and 63.43 percent for download and upload, respectively. Weizhe Zhang, Desheng Wang 0002, Shui Yu 0001, Yan Wang 0002 |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | HFL-DP: Hierarchical Federated Learning with Differential PrivacyabstractFederated learning (FL) is a framework of distributed machine learning, which aims to protect data privacy by transferring parameters instead of private data from local clients. Compared with the typical cloud-client architecture, applying FL on a cloud-edge-client hierarchical architecture could train the model faster and achieve better communication-computation trade-offs. However, hierarchical federated learning (HFL) still suffers from privacy leakage by analyzing uploaded parameters from clients or edge servers. To address this problem, we propose a privacy-preserving scheme based on the theory of local differential privacy (LDP), where adding the noise to the shared model parameters before uploading them to edge and cloud servers. According to our analysis by the moment accounting, the proposed algorithm can realize the strict differential privacy guarantee for the layers of clients and edge servers with adjustable privacy protection levels. We evaluate its performance based on the image classification tasks, and the result demonstrates that our theoretical analyses are consistent with simulations. Lu Shi 0002, Jiangang Shu, Weizhe Zhang, Yang Liu 0039 |
GLOBECOM | 3 |
| 2021 | MHCNC: A Novel Framework for Multi-Source Heterogeneous Cross-Network Node ClassificationabstractIn this study, we consider the issue of multi-source heterogeneous cross-network node classification, which employs the plentiful labeled information from multiple source networks to assist in classifying unlabeled nodes in a target network. Traditional single-source cross-network node classification methods mainly focus on the scenario where the source and target networks share the same feature spaces. The current multi-source transfer learning methods are generally unable to model the network's structural information. Thus, both of them cannot be applied to the problem of multi-source heterogeneous cross-network node classification. In this study, we introduce a multi-source heterogeneous cross-network node classification (MHCNC) framework, which integrates multi-source heterogeneous transfer learning with graph convolutional neural network to learn network-invariant and label-discriminative node representations. In MHCNC, we first devise the heterogeneous feature transformation that transforms the multiple source feature spaces onto the target network to obtain new feature representations to mitigate the distribution discrepancy between networks. In addition, we incorporate an inductive learning-based convolutional neural network to classify the unlabeled nodes in a target network. Moreover, we propose a novel model fusion strategy based on the classification weight of each model. Comprehensive experiments conducted on real-world social networks verify that our algorithms outperform the most advanced methods in terms of classification accuracy. Weizhe Zhang |
GLOBECOM | 3 |
| 2021 | Smart-DNN: Efficiently Reducing the Memory Requirements of Running Deep Neural Networks on Resource-constrained PlatformsabstractDeep neural networks (DNNs) have gained considerable attention in various real-world applications due to their strong performance in representation learning. However, running a DNN needs tremendous memory resources, which significantly restricts DNN from being applicable on resource-constrained platforms (e.g., IoT, mobile devices, etc.). Lightweight DNNs can accommodate the characteristics of mobile devices, but the hardware resources of mobile or IoT devices are extremely limited, and the resource consumption of lightweight models needs to be further reduced. However, the current neural network compression approaches (i.e., pruning, quantization, knowledge distillation, etc.) works poorly on the lightweight DNNs, which are already simplified. In this paper, we present a novel framework called Smart-DNN, which can efficiently reduce the memory requirements of running DNNs on resource-constrained platforms. Specifically, we slice a neural network into several segments and use SZ error-bounded lossy compression to compress each segment separately while keeping the network structure unchanged. When running a network, we first store the compressed network into memory and then partially decompress the corresponding part layer by layer. According to experimental results on four popular lightweight DNNs (usually used in resource-constrained platforms), Smart-DNN achieves memory saving of 1/10∼1/5, while slightly sacrificing inference accuracy and unchanging the neural network structure with accepted extra runtime overhead. Zhenbo Hu, Xiangyu Zou, Wen Xia, Weizhe Zhang, Donglei Wu |
ICCD | 5 |
| 2021 | A Novel Android Malware Detection Method Based on Visible User InterfaceabstractMachine learning has been increasingly adopted to detect Android malwares. Most existing studies depend on features in code space such as information flows and API calls. Malware variants would engage these models in a never-ending war. Inspired by the observation that some variants share similar or even identical user interfaces (UIs), this paper explores employing visible UI screenshot as the indicator to build a novel Android malware detection method. To achieve this vision, we built the first Android Application Screenshot Dataset (AnASD) consisting of more than twenty thousand UI screenshots produced by both benign applications and malwares. A thorough analysis was conducted to characterize the dataset, especially the UI difference between benign applications and malwares. Then a set of state of the art deep learning classifiers on AnASD were trained and evaluated. The results of both sim-ilarity measurement and classification performance proved the feasibility to detect Android malwares based on user interfaces. To facilitate the research community, the dataset is free available at https://doi.org/10.6084/m9.figshare.14445768. Shuaishuai Tan, Zhiyi Tian, Xiaoxiong Zhong, Shui Yu 0001, Weizhe Zhang, Guozhong Dong |
TrustCom | 5 |
| 2021 | Kalman prediction-based virtual network experimental platform for smart living
Desheng Wang 0002, Weizhe Zhang, Yang Xiang 0003, Yu-Chu Tian |
Comput. Commun. | 2 |
| 2021 | Optimal priority assignment for messages on controller area network with maximum system robustnessabstractSummary Controller Area Network (CAN) is broadly used for in‐vehicle networking. Each message on the CAN bus is given a unique identifier as its priority. When more than one node starts to transmit at the same time, the message with the highest priority will be sent. The schedulability of the message set is significantly affected by the priority assignment algorithms employed. To ensure that the CAN system tolerate additional interference, we propose system robustness, a concept to represent the total robustness of the message set. Maximum system robustness of message set problem is NP‐hard. We find that assigning a higher priority to the message with small transmission time or low utilization may lead to more system robustness. Based on this, this paper presents four optimal priority assignment approximation algorithms using the Audsley's priority assignment method: TMPA, UMPA, GTMPA, and GUMPA. TMPA and UMPA are heuristic algorithms. GTMPA and GUMPA are greedy algorithms. Experimental results show that, compared with the existing optimal algorithms, TMPA, GTMPA, and GUMPA can improve the system robustness of CAN message sets effectively. Enci Bai, Weizhe Zhang |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | Node-Fusion: Topology-aware virtual network embedding algorithm for repeatable virtual network mapping over substrate nodesabstractSummary Cloud computing has become a new Internet application model, where network virtualization is recognized as an important technology for allowing multiple heterogeneous virtual networks (VNs) to coexist on a shared substrate network (SN). As demands in cloud computing increase, the scale of VN greatly increases as well, and providing an end‐to‐end SN to embed VNs in terms of scale is difficult. To utilize SN resources fully, we devise a topology‐aware Node‐Fusion algorithm, which is different from the traditional virtual network embedding (VNE) algorithms, for repeatable VNE over substrate nodes problem. We rank the resource of nodes through a novel solution by considering the CPU and bandwidth of adjacent link capacity and the number of adjacent links of each node as resources, and rank a node on the basis of resources. Furthermore, we embed several virtual nodes into the same substrate node together in accordance with Node‐Fusion interconnection value during the node mapping process, which can greatly improve the success ratio of the subsequent link mapping phase. Evaluation results confirm that Node‐Fusion outperforms traditional classical heuristics (Link‐opt, Node‐opt, and ORSTA), which are modified to fit into our model, with regard to acceptance ratio, long‐term revenue, long‐term cost, and revenue‐cost ratio. Desheng Wang 0002, Weizhe Zhang, Chuanyi Liu |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | An Evolutionary Study of IoT MalwareabstractRecent years have witnessed lots of attacks targeted at the widespread Internet of Things (IoT) devices and malicious activities conducted by compromised IoT devices. After some notorious IoT malware released their source code, many new variants emerge, which are usually more powerful and stealthy. Although numerous existing studies have analyzed some exposed families, there is a lack of systematic study to make full use of them, which can be a fundamental step for provenance, triage, labeling, lineage analysis, and authorship attribution. The key challenge of conducting an IoT malware evolutionary study is how to collect sufficient and accurate information about malware and identify the relationships among them. In this article, we take the first step to investigate the IoT malware evolution by leveraging the information from two sources that complement each other. First, we crawl online articles about IoT malware and employ natural language processing techniques to extract the features of malware samples and their relationships with other malware family, which allow us to form the basic lineage graph. Second, we collect real malware samples through our widely deployed honeypots and design a new classifier to group them into families and identify lineage relationships among them. Such results are used to enhance the basic lineage graph. Eventually, we construct the final lineage graph for 72 IoT malware families by correlating the information from the aforementioned sources, which can help the research community better understand and fight IoT malware now and in the future. Our study has been incorporated into the threat awareness system of NSFOCUS company. Huanran Wang, Weizhe Zhang, Peng Liu 0005, Xiapu Luo, Yang Liu 0039, Yan Li 0075, Wenmao Liu, Runzi Zhang, Xing Lan |
IEEE Internet Things J. | 2 |
| 2021 | Secure and Optimized Load Balancing for Multitier IoT and Edge-Cloud Computing SystemsabstractMobile-edge computing (MEC) has emerged as a new computing paradigm with great potential to alleviate resource limitations attributed to mobile device users (MDUs) by offloading intensive computations to ubiquitous MEC server. However, most of the current offloading policies allow MDUs to transmit their tasks to the same connected small base stations (sBSs), which invariably increases latency and limits performance gain due to overload. Moreover, the security issue mitigating sensitive communication of information is not adequately addressed. Therefore, in this study, in addition to proposing a joint load balancing and computation offloading (CO) technique for MEC systems, we introduce a new security layer to circumvent potential security issues. First, a load balancing algorithm for efficient redistribution of MDUs among sBSs is proposed. In addition, a new advanced encryption standard (AES) cryptographic technique suffused with electrocardiogram (ECG) signal-based encryption and decryption key is presented as a security layer to safeguard the vulnerability of data during the transmission. Furthermore, an integrated model of load balancing, CO and security is formulated as a problem whose goal is to decrease the time and energy demands of the system. Detailed experimental results prove that our model with and without the additional security layers can save about 68.2% and 72.4% of system consumption compared to the local execution. Weizhe Zhang, Ibrahim A. Elgendy, Mohamed Hammad, Abdullah M. Iliyasu, Xiaojiang Du, Mohsen Guizani, Ahmed A. Abd El-Latif 0001 |
IEEE Internet Things J. | 1 |
| 2021 | CL-ADMM: A Cooperative-Learning-Based Optimization Framework for Resource Management in MECabstractWe consider the problem of the intelligent and efficient resource management framework in mobile-edge computing (MEC), which can reduce delay and energy consumption, and features distributed optimization and efficient congestion avoidance. In this article, we present a cooperative learning framework for resource management in MEC from an alternating direction method of multipliers (ADMMs) perspective, named the CL-ADMM framework. First, computing a task requires both the user personal data and corresponding program that processes it, to efficiently cache program in a group, a novel program popularity estimation scheme is proposed, which is based on a semi-Markov process model. Then, a greedy program cooperative caching mechanism is established, which can effectively reduce delay and energy consumption. Second, to address group congestion, a dynamic task migration scheme based on improved cooperative Q-learning is proposed, which can effectively reduce delay and alleviate congestion. Third, to minimize delay and energy consumption for resource allocation in a group, we formulate it as an optimization problem with a large number of variables, and then exploit a novel ADMM-based scheme to solve this problem, which can reduce the complexity of the problem with a new set of auxiliary variables, these subproblems are all convex problems that can be solved by using a primal-dual approach, which guarantees its convergence. Finally, we prove its convergence by using the Lyapunov theory. The numerical results demonstrate the effectiveness of the CL-ADMM framework in reducing delay and energy consumption in MEC. Xiaoxiong Zhong, Xinghan Wang 0001, Li Li 0015, Yuanyuan Yang 0001, Yang Qin 0001, Tingting Yang 0001, Bin Zhang 0048, Weizhe Zhang |
IEEE Internet Things J. | 8 |
| 2021 | EOM-NPOSESs: Emergency Ontology Model Based on Network Public Opinion Spread ElementsabstractThe construction of an emergency ontology model plays an important role in emergency management, which is an important basis for emergency public opinion management and decision-making. Integration of network public opinion spread elements into the emergency ontology model is crucial for realizing knowledge sharing in the field of emergency and public opinion responses. In this study, we crawl a large amount of emergency data from different data sources and construct an emergency dataset. Based on this dataset, we analyze the public opinion elements of emergencies and propose an emergency ontology model based on network public opinion spread elements (EOM-NPOSESs). Thereafter, we consider the coronavirus disease (COVID-19) emergency as an example to construct the EOM-NPOSESs. Finally, we design some strategies to realize rule reasoning and present the COVID-19 emergency application based on the constructed EOM-NPOSESs and the geographic information system platform. The results demonstrate that EOM-NPOSESs can not only describe the semantic relationship between emergencies and emergency elements but also perform semantic logical reasoning on different emergencies. Guozhong Dong, Weizhe Zhang, Haowen Tan, Shuaishuai Tan |
Secur. Commun. Networks | 2 |
| 2021 | An Efficient and Secured Framework for Mobile Cloud ComputingabstractSmartphone devices are widely used in our daily lives. However, these devices exhibit limitations, such as short battery lifetime, limited computation power, small memory size and unpredictable network connectivity. Therefore, numerous solutions have been proposed to mitigate these limitations and extend the battery lifetime with the use of the offloading technique. In this paper, a novel framework is proposed to offload intensive computation tasks from the mobile device to the cloud. This framework uses an optimization model to determine the offloading decision dynamically based on four main parameters, namely, energy consumption, CPU utilization, execution time, and memory usage. In addition, a new security layer is provided to protect the transferred data in the cloud from any attack. The experimental results showed that the framework can select a suitable offloading decision for different types of mobile application tasks while achieving significant performance improvement. Moreover, different from previous techniques, the framework can protect application data from any threat. Ibrahim A. Elgendy, Weizhe Zhang, Chuan-Yi Liu, Ching-Hsien Hsu |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | Correction to "An Efficient and Secured Framework For Mobile Cloud Computing"abstractPresents corrections to author affiliation information in the above named paper. Ibrahim A. Elgendy, Weizhe Zhang, Chuan-Yi Liu, Ching-Hsien Hsu |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | Efficient JPEG Batch Steganography Using Intrinsic Energy of Image ContentsabstractBatch steganography aims at properly allocating a large payload to multiple covers, so as to keep the whole covert communication at a satisfactory level of security. JPEG is currently one of the most widely used formats for image storage and transmission. This paper presents an efficient JPEG batch steganographic scheme, which allocates the payload in a linear manner w.r.t. a new heuristic measure - the intrinsic energy of JPEG image contents, in which more concerns are with the high frequency components, and the proposed measure could also be easily generalized to cover selection in batch steganographic applications. And a calibration strategy is elaborately designed to balance the security level when JPEG covers of various QFs are involved in JPEG batch steganography. In this way, the proposed scheme can effectively resolve the problem that the statistical undetectability fluctuates dramatically w.r.t. the size and quality factor when the batch set is involved with various image parameters, and consequently maintains the overall security of the practical JPEG batch steganographic system. Experimental results show that the proposed method exhibits security performance superior or comparable to the state-of-the-art batch schemes while maintaining a low computational cost. Xianglei Hu, Jiangqun Ni, Weizhe Zhang, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Collaborative Intrusion Detection for VANETs: A Deep Learning-Based Distributed SDN ApproachabstractVehicular Ad hoc Network (VANET) is an enabling technology to provide a variety of convenient services in intelligent transportation systems, and yet vulnerable to various intrusion attacks. Intrusion detection systems (IDSs) can mitigate the security threats by detecting abnormal network behaviours. However, existing IDS solutions are limited to detect abnormal network behaviors under local sub-networks rather than the entire VANET. To address this problem, we utilize deep learning with generative adversarial networks and explore distributed SDN to design a collaborative intrusion detection system (CIDS) for VANETs, which enables multiple SDN controllers jointly train a global intrusion detection model for the entire network without directly exchanging their sub-network flows. We prove the correctness of our CIDS in both IID (Independent Identically Distribution) and non-IID situations, and also evaluate its performance through both theoretical analysis and experimental evaluation on a real-world dataset. Detailed experimental results validate that our CIDS is efficient and effective in intrusion detection for VANETs. Jiangang Shu, Weizhe Zhang, Xiaojiang Du, Mohsen Guizani |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Automatic translation of data parallel programs for heterogeneous parallelism through OpenMP offloading
Farui Wang, Weizhe Zhang, Meng Hao 0002, Gangzhao Lu, Zheng Wang 0001 |
J. Supercomput. | 2 |
| 2021 | Fine-Grained Powercap Allocation for Power-Constrained Systems Based on Multi-Objective Machine LearningabstractPower capping is an important solution to keep the system within a fixed power constraint. However, for the over-provisioned and power-constrained systems, especially the future exascale supercomputers, powercap needs to be reasonably allocated according to the workloads of compute nodes to achieve trade-offs among performance, energy and powercap. Thus it is necessary to model performance and energy and to predict the optimal powercap allocation strategies. Existing power allocation approaches have insufficient granularity within nodes. Modeling approaches usually model performance and energy separately, ignoring the correlation between objectives, and do not expose the Pareto-optimal powercap configurations. Therefore, this article combines the powercap with uncore frequency scaling and proposes an approach to predict the Pareto-optimal powercap configurations on the power-constrained system for input MPI and OpenMP parallel applications. Our approach first uses the elaborately designed micro-benchmarks and a small number of existing benchmarks to build the training set, and then applies a multi-objective machine learning algorithm which combines the stacked single-target method with extreme gradient boosting to build multi-objective models of performance and energy. The models can be used to predict the optimal processor and memory powercap settings, helping compute nodes perform fine-grained powercap allocation. When the optimal powercap configuration is determined, the uncore frequency scaling is used to further optimize the energy consumption. Compared with the reference powercap configuration, the predicted optimal configurations predicted by our method can achieve an average powercap reduction of 31.35 percent, an average energy reduction of 12.32 percent, and average performance degradation of only 2.43 percent. Meng Hao 0002, Weizhe Zhang, Yiming Wang 0010, Gangzhao Lu, Farui Wang, Athanasios V. Vasilakos |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Multitier Service Migration Framework Based on Mobility Prediction in Mobile Edge ComputingabstractMobile edge computing (MEC) pushes computing resources to the edge of the network and distributes them at the edge of the mobile network. Offloading computing tasks to the edge instead of the cloud can reduce computing latency and backhaul load simultaneously. However, new challenges incurred by user mobility and limited coverage of MEC server service arise. Services should be dynamically migrated between multiple MEC servers to maintain service performance due to user movement. Tackling this problem is nontrivial because it is arduous to predict user movement, and service migration will generate service interruptions and redundant network traffic. Service interruption time must be minimized, and redundant network traffic should be reduced to ensure service quality. In this paper, the container live migration technology based on prediction is studied, and an online prediction method based on map data that does not rely on prior knowledge such as user trajectories is proposed to address this challenge in terms of mobility prediction accuracy. A multitier framework and scheduling algorithm are designed to select MEC servers according to moving speeds of users and latency requirements of offloading tasks to reduce redundant network traffic. Based on the map of Beijing, extensive experiments are conducted using simulation platforms and real‐world data trace. Experimental results show that our online prediction methods perform better than the common strategy. Our system reduces network traffic by 65% while meeting task delay requirements. Moreover, it can flexibly respond to changes in the user’s moving speed and environment to ensure the stability of offload service. Weizhe Zhang |
Wirel. Commun. Mob. Comput. | 3 |
| 2021 | Blockchain-Based DNS Root Zone Management Decentralization for Internet of ThingsabstractDomain Name System (DNS) is a widely used infrastructure for remote control and batch management of IoT devices. As a critical Internet infrastructure, DNS is structured as a tree‐like hierarchy with single root zone authority at the top, which puts the operation of DNS at risk from single point of failure. The current root zone management is lack of transparency and accountability, since only the root zone file is published as the final outcome of operations inside the root zone authority. Towards distributed root zone operation in DNS, this paper presents a blockchain‐based root operation architecture—RootChain, composed of multiple root servers. On the basis of maintaining the single root authority for top‐level domain (TLD), RootChain decentralizes TLD data publication by empowering delegated TLD authorities to publish authenticated data directly. The transparency and accountability of root zone operation are attained by smart‐contracting the whole life cycle of TLD operation and logging all operations on the chain. RootChain is transparent to recursive/stub resolver and DNS/DNSSEC‐compatible. A proof‐of‐concept prototype of RootChain has been implemented with Hyperledger Fabric and evaluated by experiments. Yu Zhang 0036, Zhongda Xia, Zhongze Wang, Weizhe Zhang, Hongli Zhang 0001, Binxing Fang |
Wirel. Commun. Mob. Comput. | 6 |
| 2021 | Joint computation offloading and task caching for multi-user and multi-task MEC systems: reinforcement learning-based algorithms
Ibrahim A. Elgendy, Weizhe Zhang, Brij B. Gupta, Ahmed A. Abd El-Latif 0001 |
Wirel. Networks | 2 |
| 2020 | Optimizing GPU Memory Transactions for Convolution OperationsabstractConvolution computation is a common operation in deep neural networks (DNNs) and is often responsible for performance bottlenecks during training and inferencing. Existing approaches for accelerating convolution operations aim to reduce computational complexity. However, these strategies often increase the memory footprint with extra memory accesses, thereby leaving much room for performance improvement. This paper presents a novel approach to optimize memory access for convolution operations, specifically targeting GPU execution. Our approach leverages two optimization techniques to reduce the number of memory operations for convolution operations performed on the width and height dimensions. For convolution computations on the width dimension, we exploit shuffle instructions to exchange the overlapped columns of the input for reducing the number of memory transactions. For convolution operations on the height dimension, we multiply each overlapped row of the input with multiple rows of a filter to compute multiple output elements to improve the data locality of row elements. We apply our approach to 2D and multi-channel 2D convolutions on an NVIDIA 2080Ti GPU. For 2D convolution, our approach delivers over faster performance than the state-of-the-art image processing libraries. For multi-channel 2D convolutions, we obtain up to speedups over the quickest algorithm of cuDNN. We apply our approach to 2D and multi-channel 2D convolutions on an NVIDIA 2080Ti GPU. For 2D convolution, our approach delivers over 2× faster performance than the state-of-the-art image processing libraries. For multi-channel 2D convolutions, we obtain up to 1.3× speedups over the quickest algorithm of cuDNN. Gangzhao Lu, Weizhe Zhang, Zheng Wang 0001 |
CLUSTER | 2 |
| 2020 | Malware Classification Method Based on Word Vector of Bytes and Multilayer PerceptionabstractThe traditional machine learning-based malware classification methods are mainly based on feature engineering. In order to improve accuracy, many features will be extracted from malware files in these methods. That brings a high complexity to the classification. To solve this issue, this paper proposes a malware classification method based on the word vector of bytes in the malware sample and Multilayer Perception (MLP). A malware sample consists of large number of bytes with values ranging from 0x00 to 0xFF. Therefore, every malware sample could be considered as a document written by bytes. And this document could be divided into sentences based on padding or meaningless bytes. In this paper, first, we use Word2Vec to calculate a 256 dimensions word vector for each byte. Second, we combine them into a matrix in ascending order. Third, we use MLP to train the model on the training samples. Finally, we use the trained model to classify the testing samples. The experimental results show that the method has a high accuracy of 98.89%. Yanchen Qiao, Bin Zhang 0048, Weizhe Zhang |
ICC | 3 |
| 2020 | HomoPAI: A Secure Collaborative Machine Learning Platform based on Homomorphic EncryptionabstractHomomorphic Encryption (HE) allows encrypted data to be processed without decryption, which could maximize the protection of user privacy without affecting the data utility. Thanks to strides made by cryptographers in the past few years, the efficiency of HE has been drastically improved, and machine learning on homomorphically encrypted data has become possible. Several works have explored machine learning based on HE, but most of them are restricted to the outsourced scenario, where all the data comes from a single data owner. We propose HomoPAI, an HE-based secure collaborative machine learning system, enabling a more promising scenario, where data from multiple data owners could be securely processed. Moreover, we integrate our system with the popular MPI framework to achieve parallel HE computations. Experiments show that our system can train a logistic regression model on millions of homomorphically encrypted data in less than two minutes. Cheng Hong 0001, Hunter Qu, Weizhe Zhang |
ICDE | 7 |
| 2020 | Delta-DNN: Efficiently Compressing Deep Neural Networks via Exploiting Floats SimilarityabstractDeep neural networks (DNNs) have gained considerable attention in various real-world applications due to the strong performance on representation learning. However, a DNN needs to be trained many epochs for pursuing a higher inference accuracy, which requires storing sequential versions of DNNs and releasing the updated versions to users. As a result, large amounts of storage and network resources are required, significantly hampering DNN utilization on resource-constrained platforms (e.g., IoT, mobile phone). Zhenbo Hu, Xiangyu Zou, Wen Xia, Sian Jin, Dingwen Tao, Yang Liu 0039, Weizhe Zhang, Zheng Zhang 0006 |
ICPP | 7 |
| 2020 | Double-Wing Mixture of Experts for Streaming Recommendations
Shoujin Wang, Yan Wang 0002, Hongwei Liu 0002, Weizhe Zhang |
WISE (2) | 5 |
| 2020 | The QoS and privacy trade-off of adversarial deep learning: An evolutionary game approach
Zhe Sun 0005, Lihua Yin, Chao Li 0027, Weizhe Zhang, Ang Li 0005, Zhihong Tian 0001 |
Comput. Secur. | 4 |
| 2020 | Ransomware classification using patch-based CNN and self-attention network on embedded N-grams of opcodes
Bin Zhang 0048, Wentao Xiao, Xi Xiao 0001, Arun Kumar Sangaiah, Weizhe Zhang, Jiajia Zhang 0001 |
Future Gener. Comput. Syst. | 5 |
| 2020 | Multi-metric domain adaptation for unsupervised transfer learningabstractUnsupervised domain adaptation aims to learn a classifier for the unlabelled target domain by leveraging knowledge from a labelled source domain. This study presents a novel domain adaptation framework from global and local transfer perspectives, referred to as multi‐metric domain adaptation (MMDA) for unsupervised transfer learning. At the global level, MMDA minimises the marginal and within‐class distances and maximises the between‐class distance between domains while maintaining the features of the source domain to improve the cross‐domain adaptability. At the local level, MMDA exploits both in‐ and cross‐domain manifold structures embedded in data samples to increase the discriminative ability. The authors learn a coupled transformation that projects the source and target domain data onto respective subspace where the statistical and geometrical divergences are reduced simultaneously. They formulate global and local adaptation methods in an optimisation problem and derive an analytic solution to the objective function. Extensive experiments demonstrate that MMDA shows improvements in classification accuracy compared with several existing state‐of‐the‐art methods. Yawen Bai, Weizhe Zhang |
IET Image Process. | 5 |
| 2020 | An IoT Honeynet Based on Multiport Honeypots for Capturing IoT AttacksabstractInternet of Things (IoT) devices are vulnerable against attacks because of their limited network resources and complex operating systems. Thus, a honeypot is a good method of capturing malicious requests and collecting malicious samples but is rarely used on the IoT. Accordingly, this article implements three kinds of honeypots to capture malicious behaviors. First, on the basis of the CVE-2017–17215 vulnerability, we implement a medium-high interaction honeypot that can simulate a specific series of router UPnP services. It has functions, such as service simulation, log recording, malicious sample download, and service self-check. Second, given the limited details available for the simulated UPnP service and to help the honeypot respond to unrecognizable malicious requests, we use the actual IoT device firmware that matches the vulnerability to build a high-interaction honeypot. In addition, we investigate the most exposed SOAP service ports and design corresponding multiport honeypot to improve the capacity of the honeynet, providing a hybrid service from a real device and simulating honeypots. The Docker in the honeynet, which reduces the volume of the honeypot and realizes the rapid deployment of the honeynet, encapsulates all these honeypots. Moreover, the honeynet control center is simultaneously designed to distribute commands and transfer files to each physical node in the honeynet. We implemented the proposed honeynet system and deployed it in practice. We have successfully caught many unknown malicious attacks excluded in the VT, which proved the effectiveness of the proposed framework. Weizhe Zhang, Bin Zhang 0048, Zeyu Ding 0003 |
IEEE Internet Things J. | 1 |
| 2020 | Trustworthy Enhancement for Cloud Proxy based on Autonomic ComputingabstractAiming to improve Internet content accessing capacity of the system, cloud proxy platforms are used to improve the visiting performance in network export environment. Limited by complexity of cloud proxy system, trustworthy guarantee of cloud system becomes a difficult problem. Considering the self-government of autonomic computing, it could enhance cloud system trustworthy and avoids system management security and reliable problems brought by complex construction. Based on the idea of self-supervisory, a mechanism to enhance security of cloud system was proposed in this paper. First, a trustworthy autonomous enhancement framework for virtual machines was proposed. Second, a method to extract linear relationship of monitoring items in the virtual machine based on ARX model was put forward. According to the mapping relation between monitoring items and system modules, an abnormal module positioning technology based on Naive Bayes classifier was developed to realize self-sensing of abnormal system conditions. Finally, security threats of virtual machines including malicious dialogue and buffer memory of hot attacks were tested through experiments. Results showed that the proposed trustworthy enhancement mechanism of virtual machines based on autonomic computing could achieve trustworthy enhancement of virtual machines effectively and provide an effective safety protection for the cloud system. Weizhe Zhang, Chuanyi Liu, Honglei Sun |
IEEE Trans. Cloud Comput. | 2 |
| 2020 | OODT: Obstacle Aware Opportunistic Data Transmission for Cognitive Radio Ad Hoc NetworksabstractIn recent years, a large number of smart devices will be connected in Internet of Things (IoT) using an ad hoc network, which needs more frequency spectra. The cognitive radio (CR) technology can improve spectrum utilization in an opportunistic communication manner for IoT, forming a promising paradigm known as cognitive radio ad hoc networks, CRAHNs. However, dynamic spectrum availability and mobile devices/persons make it difficult to develop an efficient data transmission scheme for CRAHNs under an obstacle environment. Opportunistic routing can leverage the broadcast nature of wireless channels to enhance network performance. Inspired by this, in this paper, we propose an Obstacle aware Opportunistic Data Transmission scheme (OODT) in CRAHNs from a computational geometry perspective, considering energy efficiency and social features. In the proposed scheme, we exploit a new routing metric, which is based on an obstacle avoiding algorithm using a polygon boundary 1-searcher technology, and an auction model for selecting forwarding candidates. In addition, we prove that the candidate selection problem is NP-hard and propose a heuristic algorithm for candidate selection. The simulation results show that the proposed scheme can achieve better performance than existing schemes. Xiaoxiong Zhong, Li Li 0015, Yuanping Zhang, Bin Zhang 0048, Weizhe Zhang, Tingting Yang 0001 |
IEEE Trans. Commun. | 5 |
| 2020 | Speeding Up the Schedulability Analysis and Priority Assignment of Sporadic Tasks Under Uniprocessor FPNSabstractFixed-priority non-preemptive scheduling (FPNS) is widely used in practice because of its simplicity and predictability. This article aims to enhance the efficiency of the schedulability analysis and priority assignment of sporadic tasks under uniprocessor FPNS. To speed-up the schedulability analysis, we first improve the state-of-the-art worst-case response time analysis for uniprocessor fixed-priority non-preemptive scheduling. In addition, we present two special conditions under which the worst-case response time of a task can be analyzed from its first job, which further improves the efficiency of the analysis. To accelerate the priority assignment, we present two priority-assignment algorithms based on the improved Audsley's algorithm: improved Audsley-based longest deadline first (IA-LDF) and improved Audsley-based longest worst-case execution time first (IA-LCF). The numerical experiments show that IA-LDF and IA-LCF can lead to 31.2% and 36% decrease in runtime compared to longest deadline first (LDF) and longest worst-case execution time first (LCF), respectively. Weizhe Zhang, Enci Bai, Jing Li 0025 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Efficient and Secure Multi-User Multi-Task Computation Offloading for Mobile-Edge Computing in Mobile IoT NetworksabstractMobile edge computing (MEC) is a new paradigm to alleviate resource limitations of mobile IoT networks through computation offloading with low latency. This article presents an efficient and secure multi-user multi-task computation offloading model with guaranteed performance in latency, energy, and security for mobile-edge computing. It does not only investigate offloading strategy but also considers resource allocation, compression and security issues. Firstly, to guarantee efficient utilization of the shared resource in multi-user scenarios, radio and computation resources are jointly addressed. In addition, JPEG and MPEG4 compression algorithms are used to reduce the transfer overhead. To fulfill security requirements, a security layer is introduced to protect the transmitted data from cyber-attacks. Furthermore, an integrated model of resource allocation, compression, and security is formulated as an integer nonlinear problem with the objective of minimizing the weighted sum of energy under a latency constraint. As this problem is considered as NP-hard, linearization and relaxation approaches are applied to transform the problem into a convex one. Finally, an efficient offloading algorithm is designed with detailed processes to make the computation offloading decision for computation tasks of mobile users. Simulation results show that our model not only saves about 46% of system overhead consumption in comparison with local execution but also scale well for large-scale IoT networks. Ibrahim A. Elgendy, Weizhe Zhang, Yiming Zeng 0001, Yu-Chu Tian, Yuanyuan Yang 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | Performance Optimization for Relative-Error-Bounded Lossy Compression on Scientific DataabstractScientific simulations in high-performance computing (HPC) environments generate vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for postanalysis. Unlike traditional data reduction schemes such as deduplication or lossless compression, not only can error-controlled lossy compression significantly reduce the data size but it also holds the promise to satisfy user demand on error control. Pointwise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications with lossy compression since error control can adapt to the error bound in the dataset automatically. Pointwise relative-error-bounded compression is complicated and time consuming. We develop efficient precomputation-based mechanisms based on the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error-bounded compression with excellent compression ratios. In addition, we reduce traversing operations for Huffman decoding, significantly accelerating the decompression process in SZ. Experiments with eight well-known real-world scientific simulation datasets show that our solution can improve the compression and decompression rates (i.e., the speed) by about 40 and 80 p, respectively, in most of cases, making our designed lossy compression strategy the best-in-class solution in most cases. Xiangyu Zou, Tao Lu 0014, Wen Xia, Xuan Wang 0002, Weizhe Zhang, Haijun Zhang 0002, Sheng Di, Dingwen Tao, Franck Cappello |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | DAMBA: Detecting Android Malware by ORGB AnalysisabstractWith the rapid development of smart devices, mobile phones have permeated many aspects of our life. Unfortunately, their widespread popularization attracted endless attacks that are serious threats for users. As the mobile system with the largest market share, Android has already become the hardest hit for years. To Detect Android Malware by ORGB Analysis, in this paper, we present DAMBA, a novel prototype system based on a C/S architecture. DAMBA extracts the static and dynamic features of apps. For further analyses, we propose TANMAD algorithm, a two-step Android malware detection algorithm, which reduces the range of possible malware families, and then utilizes subgraph isomorphism matching for malware detection. The key novelty of this paper is the modeling of object reference information by constructing directed graphs, which is called object reference graph birthmarks (ORGB). To achieve better efficiency and accuracy, in this paper, we present several optimization strategies for hybrid analysis. DAMBA is evaluated on a large real-world dataset of 2239 malicious and 1000 popular benign apps. The detection accuracy reaches 100% in most cases, and the average detection time is less than 5 s. Experimental results show that DAMBA outperforms the well-known detector, McAfee, which is based on signature recognition. In addition, DAMBA is demonstrated to resist the known malware attacks and their variants efficiently, as well as malware that uses obfuscation techniques. Weizhe Zhang, Huanran Wang, Peng Liu 0005 |
IEEE Trans. Reliab. | 1 |
| 2020 | An adaptive heuristic for managing energy consumption and overloaded hosts in a cloud data center
Weizhe Zhang, Keqin Li 0001, Chuanyi Liu, Muhammad Shafiq 0003, Nabin Kumar Karn |
Wirel. Networks | 2 |
| 2019 | RLS-VNE: Repeatable Large-Scale Virtual Network Embedding over Substrate NodesabstractEmbedding multiple virtual networks (VNs) on a shared substrate network (SN), known as virtual network embedding (VNE), is a challenging problem in cloud platforms. VNE methods can provide strategies to deploy VNs onto SN resources. However, as the scale of VN greatly increases, traditional VNE methods are time-consuming and waste link resource. Meanwhile, traditional VNE methods assign each virtual node of the same VN to different substrate nodes, whereas it is hard to provide larger scale SN to provision the VN. In order to efficiently embed large-scale VNs, multiple virtual nodes from the same VN need to share the same substrate node. We therefore model a repeatable large-scale virtual network embedding (RLSVNE) problem in this study, provisioning large-scale VNs, and propose a heuristic method (Rlsvne) to handle RLS-VNE. Rlsvne pre-processes the VN topology before embedding. In the pre-processing stage, the VN topology is processed through graph coarsening, partitioning, and uncoarsening. After the pre-processing, Rlsvne accomplishes an embedding stage with a topology- aware repeatable embedding solution. 1,000 and 10,000-scale VNE experiments are conducted to demonstrate our Rlsvne. The evaluation results demonstrate that our Rlsvne outperforms three modified heuristics. Rlsvne shows improved performance in reducing substrate cost and fully utilizing substrate resources, achieving high acceptance ratio and revenue values. Desheng Wang 0002, Weizhe Zhang, Shui Yu 0001 |
GLOBECOM | 2 |
| 2019 | Accelerating Relative-error Bounded Lossy Compression for HPC datasets with Precomputation-Based MechanismsabstractScientific simulations in high-performance computing (HPC) environments are producing vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for post-analysis. Unlike the traditional data reduction schemes (such as deduplication or lossless compression), not only can error-controlled lossy compression significantly reduce the data size but it can also hold the promise to satisfy user demand on error control. Point-wise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications in the lossy compression, since error control can adapt to the precision in the dataset automatically. Point-wise relative error bounded compression is complicated and time consuming. In this work, we develop efficient precomputation-based mechanisms in the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error bounded compression with excellent compression ratios. In addition, our mechanisms also help reduce traversing operations for Huffman decoding, and thus significantly accelerate the decompression process in SZ. Experiments with four well-known real-world scientific simulation datasets show that our solution can improve the compression rate by about 30% and decompression rate by about 70% in most of cases, making our designed lossy compression strategy the best choice in class in most cases. Xiangyu Zou, Tao Lu 0014, Wen Xia, Xuan Wang 0002, Weizhe Zhang, Sheng Di, Dingwen Tao, Franck Cappello |
MSST | 5 |
| 2019 | An Ant Colony System for energy-efficient dynamic Virtual Machine Placement in data centers
Fares Alharbi, Yu-Chu Tian, Maolin Tang, Weizhe Zhang, Chen Peng 0001, Minrui Fei |
Expert Syst. Appl. | 4 |
| 2019 | Resource allocation and computation offloading with data security for mobile edge computing
Ibrahim A. Elgendy, Weizhe Zhang, Yu-Chu Tian, Keqin Li 0001 |
Future Gener. Comput. Syst. | 2 |
| 2019 | Performance modeling for MPI applications with low overhead fine-grained profiling
Gangzhao Lu, Weizhe Zhang, Laurence T. Yang |
Future Gener. Comput. Syst. | 2 |
| 2019 | Automatic generation of benchmarks for I/O-intensive parallel applications
Meng Hao 0002, Weizhe Zhang, Marc Snir, Laurence T. Yang |
J. Parallel Distributed Comput. | 2 |
| 2019 | Linear and dynamic programming algorithms for real-time task scheduling with task duplication
Weizhe Zhang, Yawei Liu, Allen Chen |
J. Supercomput. | 1 |
| 2018 | Classification of Holoscopic 3D Micro-Gesture Images and VideosabstractThis paper presents an empirical study on applying convolutional neural networks (CNNs) to the Holoscopic Micro-Gesture Recognition Challenge 2018 (HoMGR 2018[1]). Based on some neural networks trained on large scale datasets such as ImageNet, we are able to get fine-tuned models to work well on the HoMGR 2018 data. Result shows that resolution of inputs is critical for model accuracy. On test sets, the accuracy is 0.867 for frames based challenge and 0.82 for videos based challenge. Weizhe Zhang, Jie Shao 0012 |
FG | 1 |
| 2018 | A novel lossless recovery algorithm for basic matrix-based VSS
Shen Wang 0004, Jianzhi Sang, Weizhe Zhang |
Multim. Tools Appl. | 4 |
| 2018 | Random grid-based threshold visual secret sharing with improved visual quality and lossless recovery ability
Shen Wang 0004, Xuehu Yan, Weizhe Zhang |
Multim. Tools Appl. | 4 |
| 2018 | Demadroid: Object Reference Graph-Based Malware Detection in AndroidabstractSmartphone usage has been continuously increasing in recent years. In addition, Android devices are widely used in our daily life, becoming the most attractive target for hackers. Therefore, malware analysis of Android platform is in urgent demand. Static analysis and dynamic analysis methods are two classical approaches. However, they also have some drawbacks. Motivated by this, we present Demadroid, a framework to implement the detection of Android malware. We obtain the dynamic information to build Object Reference Graph and propose λ -VF2 algorithm for graph matching. Extensive experiments show that Demadroid can efficiently identify the malicious features of malware. Furthermore, the system can effectively resist obfuscated attacks and the variants of known malware to meet the demand for actual use. Huanran Wang, Weizhe Zhang |
Secur. Commun. Networks | 3 |
| 2017 | Network-aware virtual machine migration in an overcommitted cloud
Weizhe Zhang, Shuo Han 0007, Huixiang Chen 0001 |
Future Gener. Comput. Syst. | 1 |
| 2017 | Automatic Memory Control of Multiple Virtual Machines on a Consolidated ServerabstractThrough virtualization, multiple virtual machines (VMs) can coexist and operate on one physical machine. When virtual machines compete for memory, the performances of applications deteriorate, especially those of memory-intensive applications. In this study, we aim to optimize memory control techniques using a balloon driver for server consolidation. Our contribution is three-fold: (1) We design and implement an automatic control system for memory based on a Xen balloon driver. To avoid interference with VM monitor operation, our system works in user mode; therefore, the system is easily applied in practice. (2) We design an adaptive global-scheduling algorithm to regulate memory. This algorithm is based on a dynamic baseline, which can adjust memory allocation according to the memory used by the VMs. (3) We evaluate our optimized solution in a real environment with 10 VMs and well-known benchmarks ( DaCapo and Phoronix Test Suites). Experiments confirm that our system can improve the performance of memory-intensive and disk-intensive applications by up to 500 and 300 percent, respectively. This toolkit has been released for free download as a GNU General Public License v3 software. Weizhe Zhang, Hu-Cheng Xie, Ching-Hsien Hsu |
IEEE Trans. Cloud Comput. | 1 |
| 2017 | Profile-based dynamic application assignment with a repairing genetic algorithm for greener data centers
Meera Vasudevan, Yu-Chu Tian, Maolin Tang, Erhan Kozan, Weizhe Zhang |
J. Supercomput. | 5 |
| 2017 | MeReg: Managing Energy-SLA Tradeoff for Green Mobile Cloud ComputingabstractMobile cloud computing (MCC) provides various cloud computing services to mobile users. The rapid growth of MCC users requires large-scale MCC data centers to provide them with data processing and storage services. The growth of these data centers directly impacts electrical energy consumption, which affects businesses as well as the environment through carbon dioxide (CO2) emissions. Moreover, large amount of energy is wasted to maintain the servers running during low workload. To reduce the energy consumption of mobile cloud data centers, energy-aware host overload detection algorithm and virtual machines (VMs) selection algorithms for VM consolidation are required during detected host underload and overload. After allocating resources to all VMs, underloaded hosts are required to assume energy-saving mode in order to minimize power consumption. To address this issue, we proposed an adaptive heuristics energy-aware algorithm, which creates an upper CPU utilization threshold using recent CPU utilization history to detect overloaded hosts and dynamic VM selection algorithms to consolidate the VMs from overloaded or underloaded host. The goal is to minimize total energy consumption and maximize Quality of Service, including the reduction of service level agreement (SLA) violations. CloudSim simulator is used to validate the algorithm and simulations are conducted on real workload traces in 10 different days, as provided by PlanetLab. Weizhe Zhang |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | Android platform-based individual privacy information protection system
Weizhe Zhang, Xiong Li 0002, Naixue Xiong, Athanasios V. Vasilakos |
Pers. Ubiquitous Comput. | 1 |
| 2016 | DwarfCode: A Performance Prediction Tool for Parallel ApplicationsabstractWe present DwarfCode, a performance prediction tool for MPI applications on diverse computing platforms. The goal is to accurately predict the running time of applications for task scheduling and job migration. First, DwarfCode collects the execution traces to record the computing and communication events. Then, it merges the traces from different processes into a single trace. After that, DwarfCode identifies and compresses the repeating patterns in the final trace to shrink the size of the events. Finally, a dwarf code is generated to mimic the original program behavior. This smaller running benchmark is replayed in the target platform to predict the performance of the original application. In order to generate such a benchmark, two major challenges are to reduce the time complexity of trace merging and repeat compression algorithms. We propose an O(mpn) trace merging algorithm to combine the traces generated by separate MPI processes, where m denotes the upper bound of tracing distance, p denotes the number of processes, and n denotes the maximum of event numbers of all the traces. More importantly, we put forward a novel repeat compression algorithm, whose time complexity is O(nlogn). Experimental results show that DwarfCode can accurately predict the running time of MPI applications. The error rate is below 10 percent for compute and communication intensive applications. This toolkit has been released for free download as a GNU General Public License v3 software. Weizhe Zhang, Albert Mo Kim Cheng, Jaspal Subhlok |
IEEE Trans. Computers | 1 |
| 2016 | Optimization strategy of Hadoop small file storage for big data in healthcare
Zhonghui Du, Weizhe Zhang, Allen Chen |
J. Supercomput. | 3 |
| 2016 | Communication optimization for RDMA-based science data transmission tools
Weizhe Zhang, Meng Hao 0002 |
J. Supercomput. | 1 |
| 2016 | Exploring large-scale small file storage for search engines
Weizhe Zhang, Gangzhao Lu, Qizhen Zhang 0003, Chuanliang Yu |
J. Supercomput. | 1 |
| 2015 | Xen-based virtual honeypot system for smart device
Weizhe Zhang, Tai-Hoon Kim |
Multim. Tools Appl. | 1 |
| 2014 | Identifying and evaluating the internet opinion leader community based on k-clique clustering
Weizhe Zhang, Boran Cao |
Neural Comput. Appl. | 1 |
| 2013 | A performance prediction scheme for computation-intensive applications on cloudabstractAs cloud computing services are gaining popularity, many organizations are considering migrating their large-scale computing applications to cloud. Different cloud service providers (CSPs) may have different computing platforms and billing methods. Most cloud customers don't know which CSP is more suitable for their applications and how much computing resource should be purchased. To address this issue, in this paper, we present a performance prediction scheme that allows a cloud customer to accurately predict computing resource (e.g., running time) for an application. The proposed scheme identifies application's control flow and scaling blocks, constructs a miniature version program to run in local machines, and then replays it in cloud to get the performance ratio between local and cloud. Our real-network experiments show that the scheme can achieve high prediction accuracy with low overhead. Hongli Zhang 0001, Xiaojiang Du, Weizhe Zhang |
ICC | 5 |
| 2010 | Ontology emergence from folksonomiesabstractThe folksonomies built from the large-scale social annotations made by collaborating users are perfect data sources for bootstrapping Semantic Web applications. In this paper, we develop an ontology induction approach to harvest the emergent semantics from the folksonomies. We propose a latent subsumption hierarchy model to uncover the implicit structure of tag space and develop our ontology induction approach on basis of this model. We identify tag subsumptions with a set-theoretical approach and model the tag space as a tag subsumption graph. While turning this graph into a concept hierarchy, we address the problem of inconsistent subsumptions and propose a random walk based tag generality ranking procedure to settle it. We propose an agglomerative hierarchical clustering algorithm utilizing the result of tag generality ranking to generate the concept hierarchy. We conduct experiments on the Delicious dataset. The results of both qualitative and quantitative evaluation demonstrate the effectiveness of the proposed approach. Kaipeng Liu 0001, Binxing Fang, Weizhe Zhang |
CIKM | 3 |
| 2010 | A General Distributed Object Locating Architecture in the Internet of ThingsabstractThis paper proposes a novel platform for object locating application in the Internet of Things environment. In this platform, objects and inquirers access and query locations using uniform service entry interfaces in heterogeneous services. To build a virtual storage system, services entries integrate enterprise database clusters and a DHT peer-to-peer network built with inquirers' devices. The DHT network is originally designed for accurate object locating, to enable fuzzy object locating we construct a hierarchical storage overlay network based on the DHT network. This LBS platform simplifies the object locating operation for ordinary inquirers greatly, moreover it provides huge virtual computing and storage resources for small companies and individual developers. Wenmao Liu, Lihua Yin, Weizhe Zhang, Hongli Zhang 0001 |
ICPADS | 3 |
| 2010 | Scale-Adaptable Recrawl Strategies for DHT-Based Distributed Web Crawling System
Weizhe Zhang, Hongli Zhang 0001, Binxing Fang |
NPC | 2 |
| 2009 | A Forwarding-Based Task Scheduling Algorithm for Distributed Web Crawling over DHTsabstractDistributed Web crawling (DWC) over DHTs is proposed to solve the bottlenecks in the traditional Web crawling. The core of this kind of system is its fully distributed task scheduling mechanism in which the crawlers are treated as peers and the crawlees are treated as resources maintained by the peers. A system model based on the content addressable network (CAN) can further optimize the scheduling mechanism by exploiting the network proximity of the crawlers and the crawlees. In this paper, we propose a new method for CAN in order to achieve load balancing in the CAN-based DWC system. The method not only keeps the load balancing among peers but also keeps the distance between peers and resources very short in our simulations. The shortened peer-resource distance fulfills the need of shortening crawler-crawlee latencies. Weizhe Zhang, Hongli Zhang 0001, Binxing Fang |
ICPADS | 2 |
| 2009 | Measurement and Analysis of BitTorrent AvailabilityabstractOf the many peer-to-peer (P2P) systems in existence, BitTorrent is one of the most appealing that has managed to attract millions of users in the past few years. In this paper we present an extensive study of BitTorrent availability through measurement and analysis. Previous studies are limited to tracker without distributed hash tables (DHT), which have been used widely already. We first discuss the difference of measurement methods, and then develop a new parallel method based on threshold with bounding measurement errors. Afterwards, we focus on three issues: the availability of tracker and DHT, the availability of pieces and the variability of availability over time. Our analysis provides several new findings: (1) Overall dynamics of peers are similar across different index. (2) Compared to tracker, DHT does not have better performance as expected. (3) Seeds contribute most of piece replicas. When they are of offline, the availability of pieces will decrease dramatically. (4) The replicas of all pieces are almost the same without very rare pieces and the distribution of pieces is equal and effective. (5) The variability of availability shows a typical life cycle pattern over time, which means it is difficult for users to obtain files in the latter half of stage. In summary, this paper advances our understanding of availability by comparing different index, exploring the distribution of pieces and analyzing the fluctuation of availability. Hongli Zhang 0001, Weizhe Zhang |
ICPADS | 3 |
| 2006 | Multisite co-allocation algorithms for computational gridabstractEfficient multisite job scheduling facilitates the cooperation of multi-domain massively parallel processor systems in a computing grid environment. However, co-allocation, heterogeneity, adaptability, and scalability emerge as tough challenges for the design of multisite job scheduling models and algorithms. This paper presents a new multisite job scheduling schema based on the multisite job scheduling model and the performance model for a heterogeneous grid environment. There are three key components: resource selection, reservation, and backfilling. The optimal and greedy-heuristic adaptive resource selection strategies are introduced. The conservative and easy backfilling are incorporated into the backfilling procedure. Experiments indicate that the scheduler and the algorithm are effective and perform better than a non-adaptive algorithm. Weizhe Zhang, Albert Mo Kim Cheng, Mingzeng Hu |
IPDPS | 1 |
| 2006 | Multisite co-allocation scheduling algorithms for parallel jobs in computing grid environments
Weizhe Zhang, Binxing Fang, Mingzeng Hu, Hongli Zhang 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2005 | Load Balance Heuristics for Synchronous Iterative Applications on Heterogeneous Cluster SystemsabstractHeterogeneous computing systems are emerging as a computing infrastructure that will enable the use of distributed heterogeneous clusters for a variety of challenging applications. The actual challenge is the load balance for tightly-coupled applications. In this paper, we focus on the important subclass of tightlycoupled applications, synchronous iterative applications and formally define their load balance problem. Two novel static meta heuristic algorithms are proposed for the load distribution: a genetic tabu hybrid search (GTHS) algorithm and a host clustering based iterative search (HCIS) algorithm, when different communication computation ratios are considered. To this end, the analysis and experiment results demonstrate the effectiveness of heuristic algorithms. Weizhe Zhang, Mingzeng Hu, Hongli Zhang 0001 |
PDCAT | 1 |
| 2004 | Multisite Resource Selection and Scheduling Algorithm on Computational GridabstractSummary form only given. Multisite resource selection and task scheduling is paid more and more attention under the grid environment nowadays with the emergence of high speed WAN. Especially in a distributed computational grid, multisite resource selection and scheduling can significantly reduce the average execution time of grand applications with a limited demand in communication. Therefore we present a CGRS resource selection algorithm based on a new density-based grid resource-clustering algorithm, and implement it in our scalable environment. The analysis and experimental results show the correctness and effectiveness of our resource selection algorithm. Weizhe Zhang, Binxing Fang, Hongli Zhang 0001, Mingzeng Hu |
IPDPS | 1 |