VLDB 2026 Research / reviewers in the wild / expert
Juan Fang 0004
dblp:69/5289-4
· DBLP profile ↗
37ranked-venue papers
15as first author
32since 2021 · last 2027
0000-0002-4542-8727ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 9 first-author · 12 since 2021Computer networks · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | DOGD: Distributed Offloading and Graph-Driven Decision framework for parallel DNN inference optimization in edge networks
Juan Fang 0004, Ziyi Teng, Naixue Xiong |
Future Gener. Comput. Syst. | 2 |
| 2026 | Coordinated Resource Management for Energy-Efficient DNN Inference on Heterogeneous Edge Devices
Yuening Wang, Juan Fang 0004, Ran Zhai, Qi Ming, Anca Jurcut |
Euro-Par (2) | 2 |
| 2026 | MAE: Collaborative inference acceleration with efficient DNN partitioning and resource allocation in resource-constrained edge computing
Juan Fang 0004, Yaxin An, Ziyi Teng, Xiaoning Zhai, Heng Tang, Huijie Chen |
Comput. Networks | 1 |
| 2026 | Multiscale Semantic Compression for Robust Collaborative CNN Inference in Low-SNR Environments: An Attention-Enhanced UNet AutoencoderabstractIn collaborative inference scenarios, semantic communication replaces raw data transmission by conveying task-oriented semantic features to improve bandwidth efficiency. However, under noisy wireless channels, the combined effects of semantic compression distortion and channel noise lead to severe information loss, resulting in degraded inference accuracy. To address this issue, this paper proposes a Multi-scale Semantic Compression Collaborative Inference (MSCCI) framework that achieves efficient, stable inference performance under high compression ratios and elevated noise levels. Specifically, a UNet-based encoder extracts multi-scale semantic features on an IoT device. These features are then integrated into a unified stream using a novel semantic fusion compression strategy, thereby substantially reducing communication overhead. The edge server decoder decompresses features and recovers image semantics via progressive upsampling and multi-scale semantic restoration. For noisy wireless channels, the framework incorporates Squeeze-and-Excite (SE) attention for dynamic feature channel weighting and residual connections for enhanced low-SNR robustness. Experimental results demonstrate that our collaborative inference framework for semantic communication outperforms state-of-the-art algorithms, and the approach’s effectiveness and robustness are verified across various channel conditions. Juan Fang 0004, Heng Tang, Ziyi Teng, Huijie Chen |
IEEE Internet Things J. | 1 |
| 2026 | Multiagent Collaborative Inference Optimization for Large-Scale DNNs in IoT Edge Systems
Juan Fang 0004, Heng Tang, Xiaolin Li 0012 |
IEEE Internet Things J. | 1 |
| 2026 | A privacy-preserving information sharing scheme in online social networks
Yehong Luo, Nafei Zhu, Jingsha He, Anca Jurcut, Yuzi Yi, Xiangjun Ma, Juan Fang 0004 |
J. Inf. Secur. Appl. | 7 |
| 2026 | Adaptive arbitration mechanisms for heterogeneous NoCs under diverse load scenarios
Juan Fang 0004, Yiding Li, Yuening Wang, Zekai Jin |
J. Supercomput. | 1 |
| 2026 | Co-design of traffic-aware dynamic VC partitioning and congestion-aware routing in CPU-GPU heterogeneous NoCs
Juan Fang 0004, Haoyu Cheng, Yuening Wang, Juncheng Chen |
J. Supercomput. | 1 |
| 2026 | DOJS: A Distributed Online Joint Scheme to Optimize Cost in Mobile Edge NetworksabstractEdge computing deploys computing and storage resources at the network edge, thereby providing services closer to terminal users. However, in edge networks, the mobility of terminals, the diversity of requests, and the dynamic nature of wireless channels pose significant challenges for efficiently allocating limited wireless and caching resources among multiple terminal devices. To address the issues of unbalanced network load and high caching costs caused by resource allocation in edge networks, we propose a Distributed Online Joint Optimization Scheme (DOJS). Specifically, we design a joint optimization scheme, referred to as DOJS, which combines centralized user association at the cloud with distributed cache placement at the base stations. This scheme analyzes the impact of terminal device association policies on caching costs and develops a caching cost model that integrates the activity level and content request probability of terminal devices. Based on this model, the relationship between user association selection and caching costs is analyzed, and a Game Theory-based User Association (GTUA) selection algorithm is proposed. In order to adapt to the dynamic characteristics of terminal-user requests in mobile edge networks, we develop a dynamic cache update method LS-TD3, which combines Long Short-Term Memory (LSTM) and Twin Delayed Deep Deterministic policy gradient (TD3). Specifically, we integrate the LSTM layer into the policy model framework of reinforcement learning to better predict the content popularity from dynamic time data, thus improving the accuracy of cache decision making. To further reduce computational complexity and enhance overall system performance, we employ a distributed optimization strategy to improve the dynamic caching decision process. Extensive experimental results demonstrate the superiority of the proposed algorithm in achieving inter-node load balancing and minimizing caching costs. Ziyi Teng, Juan Fang 0004, Naixue Xiong |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2026 | Practical Efficient Deployment and Updating for Microservice With Dependencies in Multi-Access Edge ComputingabstractAs mobile edge computing technology advances rapidly, latency-sensitive and resource-intensive applications are being offloaded to edge servers to enhance Quality of Service (QoS) for users. Traditional monolithic architectures, however, struggle to meet the escalating service and traffic requirements of distributed users due to their inherent inflexibility. In response to these challenges, microservices architecture, characterized by scalability and flexibility, has been adopted for dynamic deployment at the network edge. However, the deployment of these lightweight, dependency-rich components in a way that minimally impacts the makespan and maximizes quality of service is complex. Current studies often overlook the deployment of microservices with specific dependencies within constrained environments of edge server clusters and communication links. This paper introduces practical and effective strategies for the deployment and updating of microservices, tailored to various application contexts. Initially, two scenarios are analyzed: one constrained by bandwidth with unlimited storage, and the other by storage with unlimited bandwidth. For each scenario, optimal solutions are developed using a novel enhanced graph construction method. The study progresses to a more intricate scenario involving comprehensive constraints on storage, computation, and communication resources. An optimized deployment method is proposed, utilizing main path embedding followed by an innovative simulated annealing algorithm for iterative refinement. This method is validated by demonstrating that the main path coincides with the critical path. Furthermore, the dynamic reallocation of edge resources is explored through a critical path-based updating algorithm that optimizes microservice locations to reduce overall makespan. Extensive experiments demonstrate that our strategies outperform existing representative benchmark approaches in terms of overall performance and microservice deployment efficiency. Shuaibing Lu, Jie Wu 0001, Zhi Cai, Jackson Yang, Shuyang Zhou, Juan Fang 0004 |
IEEE Trans. Serv. Comput. | 8 |
| 2026 | RLRM: Reinforcement Learning-Based Routing for Ring-Augmented Mesh NoCabstractModern multicore processors increasingly rely on network-on-chip (NoC) architectures to support high-bandwidth and low-latency communication. Traditional mesh-based NoC topologies suffer from uneven load distribution, with central router nodes frequently becoming congestion hot spots. The existing approaches mainly focus on introducing bypass links or independent ring interconnects. However, bypass links can only alleviate congestion between fixed routers, and routers not included in the ring still rely on conventional mesh forwarding, limiting the overall performance improvement. To address these challenges, this article proposes a reinforcement learning-based routing for ring-augmented mesh NoC (RLRM) to improve communication efficiency in multicore processors. First, we design a ring-augmented$8\times 8$mesh topology that integrates horizontal and vertical ring interconnects to achieve balanced load distribution and mitigate congestion hotspots. Second, we equip each router with a topology-aware reinforcement learning (RL) agent that dynamically selects paths based on congestion conditions, thereby proactively avoiding performance bottlenecks. Extensive experiments demonstrate that the proposed solution provides a scalable and efficient architecture for high-performance computing and multicore NoC systems. Experimental results under realistic benchmark traffic patterns show that RLRM reduces average latency by 22.78% compared to conventional mesh and static hybrid routing schemes. Juan Fang 0004, Qi Ming |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | A Multi-strategy Communication Optimization and Adaptive Model Splitting Scheme for Federated Split Learning
Juan Fang 0004, Ziyi Teng, Xiaoning Zhai |
ICA3PP (2) | 1 |
| 2025 | A Topology-Aware GNN Learning Approach for Energy Optimization in Multihop LoRa NetworksabstractEnergy optimization is crucial for extending battery life and reducing maintenance costs in long-range (LoRa) Internet of Things (IoT) networks. Traditional optimization methods usually need excessive computational complexity, limiting the practicability. This drives the recent development of machine learning (ML)-based optimization, such as deep neural networks (DNNs) and reinforcement learning (RL), in wireless local area networks (WLANs). However, compared to WLANs, LoRa owns a more complicated network topology due to the engagement of multi-hop, which is difficult for the existing ML methods to learn. Motivated by this, we propose a topology-aware graph neural network (GNN) learning method, which is specially tailored to tackle the energy optimization problem in multi-hop LoRa networks. By leveraging each node’s topological position to adaptively determine the optimal message-passing depth, the model better integrates the neighborhood information with node feature representations, enhancing the prediction of transmission parameter and overall energy efficiency. Also, a closed-form model of collision probability is derived for the nodes in LoRa, to measure the energy consumption due to retransmissions. Simulation results show that against traditional optimization methods such as game theory, the proposed method can reduce the runtime by five orders of magnitude, with an energy optimization gap below 13%. Compared to existing GNN-based methods, it achieves up to 50% lower energy consumption 20% fewer outage probability, at a similar level of inference time. Huapeng Yang, Xiping Wu, Han Ji 0001, Zhangqin Huang, Juan Fang 0004 |
IEEE Internet Things J. | 5 |
| 2025 | DRCD: a regional-contention-driven arbitration policy for CPU-GPU heterogeneous systemsabstractIn CPU–GPU heterogeneous systems, there exists intense resource contention between CPUs and GPUs. Traditional resource arbitration policies fail to account for the heterogeneity of cores, leading to inefficient network resource utilization for the CPU, which negatively impacts its performance. In heterogeneous networks, the degree of resource contention varies across different regions. This paper first uses reinforcement learning to analyze the message feature weights relied upon for resource arbitration in different network regions. To achieve more efficient resource allocation, a regional-contention-driven arbitration policy is proposed. The simulation results show that, compared to traditional arbitration policy, the overall network latency is reduced by 7.99%, and CPU performance is improved by 11.42%. Furthermore, a dynamic regional-contention-driven arbitration policy is proposed, which further reduces the overall network latency by 10.47% and increases CPU performance by 16.79% compared to traditional arbitration policy. Juan Fang 0004, Haoyu Cheng, Yuening Wang, Ran Zhai |
J. Supercomput. | 1 |
| 2025 | Joint DNN Partitioning and Task Offloading Based on Attention Mechanism-Aided Reinforcement LearningabstractThe rapid advancement of artificial intelligence applications has resulted in the deployment of a growing number of deep neural networks (DNNs) on mobile devices. Given the limited computational capabilities and small battery capacity of these devices, supporting efficient DNN inference presents a significant challenge. This paper considers the joint design of DNN model partitioning and offloading under high-concurrent tasks scenarios. The primary objective is to accelerate DNN task inference and reduce computational delay. Firstly, we propose an innovative adaptive inference framework that partitions DNN models into interdependent sub-tasks through a hierarchical partitioning method. Secondly, we develop a delay prediction model based on a Random Forest (RF) regression algorithm to estimate the computational delay of each sub-task on different devices. Finally, we designed a high-performance DNN partitioning and task offloading method based on an attention mechanism-aided Soft Actor-Critic (AMSAC) algorithm. The bandwidth allocation for each user is determined by the attention mechanism based on the characteristics of the DNN tasks, and the Soft Actor-Critic algorithm is used for adaptive layer-level partitioning and offloading of the DNN model, reducing collaborative inference delay. Extensive experiments demonstrate that our proposed AMSAC algorithm effectively reduces DNN task inference latency cost and improves service quality. Juan Fang 0004, Ziyi Teng |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2025 | Online Elastic Resource Provisioning With QoS Guarantee in Container-Based Cloud ComputingabstractIn cloud data centers, the exponential growth of data places increasing demands on computing, storage, and network resources, especially in multi-tenant environments. While this growth is crucial for ensuring Quality of Service (QoS), it also introduces challenges such as fluctuating resource requirements and static container configurations, which can lead to resource underutilization and high energy consumption. This article addresses online resource provisioning and efficient scheduling for multi-tenant environments, aiming to minimize energy consumption while balancing elasticity and QoS requirements. To address this, we propose a novel optimization framework that reformulates the resource provisioning problem into a more manageable form. By reducing the original multi-constraint optimization to a container placement problem, we apply the interior-point barrier method to simplify the optimization, integrating constraints directly into the objective function for efficient computation. We also introduce elasticity as a key parameter to balance energy consumption with autonomous resource scaling, ensuring that resource consolidation does not compromise system flexibility. The proposed Energy-Efficient and Elastic Resource Provisioning (EEP) framework comprises three main modules: a distributed resource management module that employs vertical partitioning and dynamic leader election for adaptive resource allocation; a prediction module using$\omega$-step prediction for accurate resource demand forecasting; and an elastic scheduling module that dynamically adjusts to tenant scaling needs, optimizing resource allocation and minimizing energy consumption. Extensive experiments across diverse cloud scenarios demonstrate that the EEP framework significantly improves energy efficiency and resource utilization compared to established baselines, supporting sustainable cloud management practices. Shuaibing Lu, Jie Wu 0001, Jackson Yang, Xinyu Deng, Zhi Cai, Juan Fang 0004 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2025 | Reinforcement Learning-Driven Adaptive Prefetch Aggressiveness Control for Enhanced Performance in Parallel System ArchitecturesabstractIn modern parallel system architectures, prefetchers are essential to mitigating the performance challenges posed by long memory access latencies. These architectures rely heavily on efficient memory access patterns to maximize system throughput and resource utilization. Prefetch aggressiveness is a central parameter in managing these access patterns; although increased prefetch aggressiveness can enhance performance for certain applications, it often risks causing cache pollution and bandwidth contention, leading to significant performance degradation in other workloads. While many existing prefetchers rely on static or simple built-in aggressiveness controllers, a more flexible, adaptive approach based on system-level feedback is essential to achieving optimal performance across parallel computing environments. In this paper, we introduce an Adaptive Prefetch Aggressiveness Control (APAC) framework that leverages Reinforcement Learning (RL) to dynamically manage prefetch aggressiveness in parallel system architectures. The APAC controller operates as an RL agent, which optimizes prefetch aggressiveness by dynamically responding to system feedback on prefetch accuracy, timeliness, and cache pollution. The agent receives a reward signal that reflects the impact of each adjustment on both performance and memory bandwidth, learning to adapt its control strategy based on workload characteristics. This data-driven adaptability makes APAC particularly well-suited for parallel architectures, where efficient resource management across cores is essential to scaling system performance. Our evaluation with the ChampSim simulator demonstrates that APAC effectively adapts to diverse workloads and system configurations, achieving performance gains of 6.73$\%$in multi-core systems compared to traditional Feedback Directed Prefetching (FDP). By improving memory bandwidth utilization, reducing cache pollution, and minimizing inter-core interference, APAC significantly enhances prefetching performance in multi-core processors. These results underscore APAC’s potential as a robust solution for performance optimization in parallel system architectures, where efficient resource management is paramount for scaling modern processing environments. Huijing Yang, Juan Fang 0004, Yumin Hou, Xing Su 0001, Naixue Xiong |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Attention Mechanism-Aided Deep Reinforcement Learning for Dynamic Edge CachingabstractThe dynamic mechanism of joint proactive caching and cache replacement, which involves placing content items close to cache-enabled edge devices ahead of time until they are requested, is a promising technique for enhancing traffic offloading and relieving heavy network loads. However, due to limited edge cache capacity and wireless transmission resources, accurately predicting users’ future requests and performing dynamic caching is crucial to effectively utilizing these limited resources. This paper investigates joint proactive caching and cache replacement strategies in a general mobile edge computing (MEC) network with multiple users under a cloud-edge-device collaboration architecture. The joint optimization problem is formulated as a markov decision process (MDP) problem with an infinite range of average network load costs, aiming to reduce network load traffic while efficiently utilizing the limited available transport resources. To address this issue, we design an Attention Weighted Deep Deterministic Policy Gradient (AWD2PG) model, which uses attention weights to allocate the number of channels from server to user, and applies deep deterministic policies on both user and server sides for Cache decision-making, so as to achieve the purpose of reducing network traffic load and improving network and cache resource utilization. We verify the convergence of the corresponding algorithms and demonstrate the effectiveness of the proposed AWD2PG strategy and benchmark in reducing network load and improving hit rate. Ziyi Teng, Juan Fang 0004, Huijing Yang, Huijie Chen, Wei Xiang 0001 |
IEEE Internet Things J. | 2 |
| 2024 | TB-TBP: a task-based adaptive routing algorithm for network-on-chip in heterogenous CPU-GPU architecturesabstractAbstract With the rapid development of heterogeneous network-on-chip (NoC), a vast amount of shared resources are integrated into NoC. Intense resource competition exists between CPUs and GPUs, leading to congestion and a decrease in overall network performance. Reasonable node placement can minimize network conflicts at the topology level. This paper first discusses the placement of shared last-level cache and memory controller, then selects a more rational placement method and optimizes the path. To solve the hot spots problem in center placement method, a task-based routing algorithm is designed to plan the path. Simulation results demonstrate that, compared to the traditional routing algorithm, the overall network latency is reduced by 9%, and the CPU performance is improved by 13.6%. Furthermore, a dynamic task-based routing algorithm is proposed. Compared to the static task routing algorithm, the overall network latency is reduced by 2.08%, and the CPU performance is improved by 4.09%. Juan Fang 0004, Zhichao Wei, Yumin Hou |
J. Supercomput. | 1 |
| 2024 | RL-CoPref: a reinforcement learning-based coordinated prefetching controller for multiple prefetchersabstractAbstract Modern processors employ data prefetchers to alleviate the impact of long memory access latency. However, current prefetchers are designed for specific memory access patterns, which perform poorly on mixed applications with multiple memory access patterns. To address these issues, RL-CoPref, a reinforcement learning (RL)-based coordinated prefetching controller for multiple prefetchers, is proposed in this paper. RL-CoPref takes diverse program context information as the input, learns to maximize cumulative rewards, and evaluates prefetch quality based on prefetch hits/misses and memory bandwidth utilization. It can dynamically adjust the prefetch activation and prefetch degree, enabling multiple prefetchers to complement each other on mixed applications. Our extensive evaluation, utilizing the ChampSim simulator, demonstrates that RL-CoPref can effectively adapt to various workloads and system configurations, optimizing prefetch control. On average, RL-CoPref achieves 76.15% prefetch coverage, having 35.50% IPC improvement, outperforming state-of-the-art individual prefetchers by 5.91–16.54% and outperforming SBP, a state-of-the-art (non-RL) prefetch controller, by 4.64%. Huijing Yang, Juan Fang 0004, Xing Su 0001, Zhi Cai, Yuening Wang |
J. Supercomput. | 2 |
| 2024 | Dependency-Aware Dynamic Task Offloading Based on Deep Reinforcement Learning in Mobile-Edge ComputingabstractThe rapid advancement of mobile edge computing (MEC) networks has enabled the augmentation of the computational power of mobile devices (MDs) by offloading computationally intensive tasks to resource-rich edge nodes. This paper discusses the decision-making process for task offloading and resource allocation among multiple mobile devices connected to a base station. The primary objective is to minimize the time taken to complete tasks while simultaneously reducing energy consumption on the device under a time-varying wireless fading channel. This objective is formulated as an energy-efficiency cost (EEC) minimization problem, which cannot be solved by conventional methods. To address this challenge, we propose a dynamic offloading decision algorithm of dependent tasks (DODA-DT) that adjusts local task execution based on edge node status. The proposed algorithm facilitates fair competition among all devices for edge resources. Additionally, we use a deep reinforcement learning (DRL) algorithm based on an actor-critic learning structure to train the system to quickly identify near-optimal solutions. Numerical simulations demonstrate that the proposed algorithm effectively reduces the total cost of the task in comparison to previous algorithms. Juan Fang 0004, Dezheng Qu, Huijie Chen |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2024 | Combining Lyapunov Optimization and Deep Reinforcement Learning for D2D Assisted Heterogeneous Collaborative Edge CachingabstractThe problem of shared node selection and cache placement in wireless networks is challenging due to the difficulty of finding low-complexity optimal solutions. This paper proposes a new approach combining Lyapunov optimization and reinforcement learning (LoRL) to address content sharing in heterogeneous mobile edge computing (MEC) networks with base station (BS) and device-to-device (D2D) communication. Device in this network can choose to establish D2D links with neighboring devices for content sharing or send requests directly to the base station for content. Content access and energy consumption of shared nodes are modeled as a queuing system. The goal is to assign content sharing nodes to stabilize all queues while maximizing D2D sharing gain and minimizing latency, even in the presence of unknown network state distribution and user sharing costs. The proposed approach enables edge device to independently select associated nodes and make caching decisions, thereby minimizing time-averaged network costs and stabilizing the queuing system. Experimental results show that the proposed algorithm converges to the optimal policy and outperforms other policies in terms of total queue backlog trade-off and network cost. Ziyi Teng, Juan Fang 0004 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | QoS-Aware Online Service Provisioning and Updating in Cost-Efficient Multi-Tenant Mobile Edge ComputingabstractThe vigorous development of IoT technology has spawned a series of applications that are delay-sensitive or resource-intensive. Mobile edge computing is an emerging paradigm that provides services between end devices and traditional cloud data centers to users. However, with the continuously increasing investment of demands, it is nontrivial to maintain a higher quality-of-service (QoS) under the erratic activities of mobile users. In this paper, we investigate the service provisioning and updating problem under the multiple-users scenario by improving the performance of services with long-term cost constraints. We first decouple the original long-term optimization problem into a per-slot deterministic one by using Lyapunov optimization. Then, we propose two service updating decision strategies by considering the trajectory prediction conditions of users. Based on that, we design an online strategy by utilizing the committed horizon control method looking forward to multiple slots predictions. We prove the performance bound of our online strategy theoretically in terms of the trade-off between delay and cost. Extensive experiments demonstrate the superior performance of the proposed algorithm. Shuaibing Lu, Jie Wu 0001, Pengfan Lu, Ning Wang 0018, Juan Fang 0004 |
IEEE Trans. Serv. Comput. | 6 |
| 2023 | A Fine-Grained Cross-Chain Spectrum Sharing Mechanism Based on OracleabstractThe dramatically increased wireless communication needs make non-renewable spectrum resources extremely scarce and costly. Consortium blockchain realizes trusted spectrum sharing among untrusted spectrum owners. Yet, most existing studies ignore spectrum sharing among blockchains, which greatly reduce spectrum utilization. In the paper, we focus on cross-chain spectrum sharing. We propose a Fine-grained Cross-chain Spectrum Sharing mechanism based on Oracle (FCSSO) to realize trusted and efficient cross-chain spectrum transactions. To guarantee benefits of spectrum owners, we design a fine-grained time partition method to decide spectrum renting time in transactions. The method reduces the waste of owners' available spectrum time caused by spectrum handoff. Extensive simulation proves the positive impact of the proposed fine-grained time partition method, and FCSSO outperforms two representative cross-chain mechanisms from two aspects: spectrum owners' benefits and spectrum utilization. Mengjie Cao, Qian Wang 0015, Xiaojiang Du, Juan Fang 0004, Bei Gong, Mohsen Guizani |
GLOBECOM | 4 |
| 2023 | Profit-driven Optimization of Server Deployment and Service Placement in Multi-User Mobile Edge ComputingabstractEdge computing has emerged as a promising paradigm to fulfill the escalating demands of latency-sensitive and computationally intensive applications. In this context, efficient server deployment and service placement have become imperative to optimize performance and increase platform profit. In this paper, we investigate the problem of server deployment and service placement in a multi-user scenario, aiming to enhance the profit of Mobile Network Operators (MNOs) while considering constraints related to distance thresholds, resource limitations, and connectivity requirements. Then, we propose a novel two-stage method to decouple the problem, breaking down server deployment and service placement into two distinct yet interconnected stages. In stage I, the server deployment is formulated as a combinatorial optimization problem within the framework of a Markov Decision Process (MDP), where the state space, action space, and penalty function are defined to effectively model the issue. We propose the SDQ algorithm to establish a relatively stable server deployment strategy. In stage II, the service placement is formulated as a constrained integer linear programming problem. We propose the SPIB-TDB algorithm to optimize service placement. Extensive experimentation validates the exceptional performance of our proposed algorithms in enhancing the profit of MNOs. Juan Fang 0004, Shuaibing Lu |
ICPADS | 1 |
| 2023 | A perceptual and predictive batch-processing memory scheduling strategy for a CPU-GPU heterogeneous systemabstractWhen multiple central processing unit (CPU) cores and integrated graphics processing units (GPUs) share off-chip main memory, CPU and GPU applications compete for the critical memory resource. This causes serious resource competition and has a negative impact on the overall performance of the system. We describe the competition for shared-memory resources in a CPU-GPU heterogeneous multi-core architecture, and a shared-memory request scheduling strategy based on perceptual and predictive batch-processing is proposed. By sensing the CPU and GPU memory request conditions in the request buffer, the proposed scheduling strategy estimates the GPU latency tolerance and reduces mutual interference between CPU and GPU by processing CPU or GPU memory requests in batches. According to the simulation results, the scheduling strategy improves CPU performance by 8.53% and reduces mutual interference by 10.38% with low hardware complexity. Juan Fang 0004, Huijing Yang, Yixiang Xu, Xing Su 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2023 | Dual-branch cross-dimensional self-attention-based imputation model for multivariate time seriesabstractIn real-world scenarios, partial information losses of multivariate time series degrade the time series analysis. Hence, the time series imputation technique has been adopted to compensate for the missing values. Existing methods focus on investigating temporal correlations, cross-variable correlations, and bidirectional dynamics of time series, and most of these methods rely on recurrent neural networks (RNNs) to capture temporal dependency. However, the RNN-based models suffer from the common problems of slow speed and high complexity when dealing with long-term dependency. While some self-attention-based models without any recurrent structures can tackle long-term dependency with parallel computing, they do not fully learn and utilize correlations across the temporal and cross-variable dimensions. To address the limitations of existing methods, we propose a novel so-called dual-branch cross-dimensional self-attention-based imputation (DCSAI) model for multivariate time series, which is capable of performing global and auxiliary cross-dimensional analyses when imputing the missing values. In particular, this model contains masked multi-head self-attention-based encoders aligned with auxiliary generators to obtain global and auxiliary correlations in two dimensions, and these correlations are then combined into one final representation through three weighted combinations. Extensive experiments are presented to show that our model performs better than other state-of-the-art benchmarkers on three real-world public datasets under various missing rates. Furthermore, ablation study results demonstrate the efficacy of each component of the model. Le Fang 0001, Wei Xiang 0001, Yuan Zhou 0006, Juan Fang 0004, Lianhua Chi, ZongYuan Ge |
Knowl. Based Syst. | 4 |
| 2023 | eX-ViT: A Novel explainable vision transformer for weakly supervised semantic segmentationabstractRecently vision transformer models have become prominent models for a multitude of vision tasks. These models, however, are usually opaque with weak feature interpretability, making their predictions inaccessible to the users. While there has been a surge of interest in the development of post-hoc solutions that explain model decisions, these methods can not be broadly applied to different transformer architectures, as rules for interpretability have to change accordingly based on the heterogeneity of data and model structures. Moreover, there is no method currently built for an intrinsically interpretable transformer, which is able to explain its reasoning process and provide a faithful explanation. To close these crucial gaps, we propose a novel vision transformer dubbed the eXplainable Vision Transformer (eX-ViT), an intrinsically interpretable transformer model that is able to jointly discover robust interpretable features and perform the prediction. Specifically, eX-ViT is composed of the Explainable Multi-Head Attention (E-MHA) module, the Attribute-guided Explainer (AttE) module with the self-supervised attribute-guided loss. The E-MHA tailors explainable attention weights that are able to learn semantically interpretable representations from tokens in terms of model decisions with noise robustness. Meanwhile, AttE is proposed to encode discriminative attribute features for the target object through diverse attribute discovery, which constitutes faithful evidence for the model predictions. Additionally, we have developed a self-supervised attribute-guided loss for our eX-ViT architecture, which utilizes both the attribute discriminability mechanism and the attribute diversity mechanism to enhance the quality of learned representations. As a result, the proposed eX-ViT model can produce faithful and robust interpretations with a variety of learned attributes. To verify and evaluate our method, we apply the eX-ViT to several weakly supervised semantic segmentation (WSSS) tasks, since these tasks typically rely on accurate visual explanations to extract object localization maps. Particularly, the explanation results obtained via eX-ViT are regarded as pseudo segmentation labels to train WSSS models. Comprehensive simulation results illustrate that our proposed eX-ViT model achieves comparable performance to supervised baselines, while surpassing the accuracy and interpretability of state-of-the-art black-box methods using only image-level labels. Wei Xiang 0001, Juan Fang 0004, Yi-Ping Phoebe Chen, Lianhua Chi |
Pattern Recognit. | 3 |
| 2023 | Resource provisioning in collaborative fog computing for multiple delay-sensitive usersabstractAbstract Fog computing is an emerging paradigm that supplies storage, computation, and networking resources between traditional cloud data centers and end devices. This article focuses on the resource provisioning problem in collaborative fog computing for multiple delay‐sensitive users. Our goal is to implement a resource provisioning strategy for network operators to minimize the total monetary cost by considering the deadline and capacity constraints. Two scenarios are considered: unlimited‐processor fog nodes (UPFN) and limited‐processor fog nodes (LPFN). In either scenario, we prove that the resource provisioning problem is NP‐hard. First, we consider the UPFN scenario that the processors of fog nodes are unlimited and users' requests can be ideally processed in parallel. Two algorithms are proposed which greedily delete fog nodes based on the local or global collaborative influences until there is no feasible provisioning to guarantee the deadline of users. Then we extend the resource provisioning problem to a more realistic and complicated scenario LPFN in which the scheduling delay cannot be ignored. Two types of tasks are considered. One is the arbitrarily divided tasks, and a near‐optimal solution bounded by has been found. m is the number of fog nodes, and is the upper bound on the Lipschitz constant of the delay function. Another one is the application‐driven tasks, and we propose a heuristic algorithm. Extensive experiments validate the efficiency of the proposed algorithms. Shuaibing Lu, Jie Wu 0001, Ning Wang 0018, Yubin Duan, Jiayue Zhang, Juan Fang 0004 |
Softw. Pract. Exp. | 7 |
| 2022 | Online Service Provisioning and Updating in QoS-aware Mobile Edge ComputingabstractThe vigorous development of IoT technology has spawned a series of applications that are delay-sensitive or resource-intensive. Mobile edge computing is an emerging paradigm which provides services between end devices and traditional cloud data centers to users. However, with the continuously increasing investment of demands, it is nontrivial to maintain a higher quality-of-service (QoS) under the erratic activities of mobile users. In this paper, we investigate the service provisioning and updating problem under the multiple-users scenario by improving the performance of services with long-term cost constraints. We first decouple the original long-term optimization problem into a per-slot deterministic one by using Lyapunov optimization. Then, we propose two service updating decision strategies by considering the trajectory prediction conditions of users. Based on that, we design an online strategy by utilizing the committed horizon control method looking forward to multiple slots predictions. We prove the performance bound of our online strategy theoretically in terms of the trade-off between delay and cost. Extensive experiments demonstrate the superior performance of the proposed algorithm. Shuaibing Lu, Jie Wu 0001, Pengfan Lu, Jiamei Shi, Ning Wang 0018, Juan Fang 0004 |
MSN | 6 |
| 2022 | A novel explainable neural network for Alzheimer's disease diagnosis
Wei Xiang 0001, Juan Fang 0004, Yi-Ping Phoebe Chen, Ruifeng Zhu |
Pattern Recognit. | 3 |
| 2021 | IoT Application Modules Placement and Dynamic Task Processing in Edge-Cloud ComputingabstractIn today's era of Internet of Things (IoT), efficient and real-time processing of massive data generated by IoT device has become the primary issue for traditional cloud computing network architectures. As a supplement of cloud computing, edge computing enhances the real-time performance of service completion by offloading services to edge servers closer to the terminal device for execution, while reducing power consumption and computing load in the cloud. In this article, we propose the following solutions to resolve the different requests of the IoT device: in an “edge-cloud” heterogeneous network environment, create a mapping scheme between application modules and basic resource equipment, considering the two factors of tolerant task latency and system power consumption. In the application step-by-step execution process, heuristic dynamic task processing algorithm is used to reduce the task latency time. Experiments with the “iFogSim” simulator show that, application service quality is significantly improved and system power consumption is greatly reduced, which compared with the stable application module placement strategy and the static task scheduling strategy. Juan Fang 0004, Aonan Ma |
IEEE Internet Things J. | 1 |
| 2020 | Towards cost-efficient resource provisioning with multiple mobile users in fog computing
Shuaibing Lu, Jie Wu 0001, Yubin Duan, Ning Wang 0018, Juan Fang 0004 |
J. Parallel Distributed Comput. | 5 |
| 2020 | A memory scheduling strategy for eliminating memory access interference in heterogeneous systemabstractAbstract Multiple CPUs and GPUs are integrated on the same chip to share memory, and access requests between cores are interfering with each other. Memory requests from the GPU seriously interfere with the CPU memory access performance. Requests between multiple CPUs are intertwined when accessing memory, and its performance is greatly affected. The difference in access latency between GPU cores increases the average latency of memory accesses. In order to solve the problems encountered in the shared memory of heterogeneous multi-core systems, we propose a step-by-step memory scheduling strategy, which improve the system performance. The step-by-step memory scheduling strategy first creates a new memory request queue based on the request source and isolates the CPU requests from the GPU requests when the memory controller receives the memory request, thereby preventing the GPU request from interfering with the CPU request. Then, for the CPU request queue, a dynamic bank partitioning strategy is implemented, which dynamically maps it to different bank sets according to different memory characteristics of the application, and eliminates memory request interference of multiple CPU applications without affecting bank-level parallelism. Finally, for the GPU request queue, the criticality is introduced to measure the difference of the memory access latency between the cores. Based on the first ready-first come first served strategy, we implemented criticality-aware memory scheduling to balance the locality and criticality of application access. Juan Fang 0004, Mengxuan Wang, Zelin Wei |
J. Supercomput. | 1 |
| 2019 | A Low-Cost and Energy-Efficient NoC Architecture for GPGPUsabstractGPGPU accelerated systems demand high throughput in data communication in order to fully exploit thread-level parallelism. Most of current GPGPU Network-on-Chips (NoCs) employ topology adapted from CPUs, such as mesh and crossbar. However, the trade-off between performance and cost for such networks is sub-optimal, due to the unique traffic pattern of GPUs. In this work, we propose a novel NoC architecture called fused fat tree which modifies the fat tree to match GPU traffic pattern. By separately connecting memory controllers and computing cores to tree roots and leaves, protocol deadlocks can be avoided using just one physical network. However, this modification removes the advantage of path diversity in the original fat tree topology and makes the network vulnerable to hotspot-caused congestion. To solve this problem, we propose to fuse routers with side links to create multiple paths. A load-balancing routing algorithm is also proposed in order to increase network throughput. We also propose a novel preemptive bandwidth allocation scheme to improve resource utilization by taking advantage of request message slacks. Our evaluation results show that our design can improve performance by 46% while achieving 27 % and 25 % area and energy savings on the average. Xianwei Cheng, Yang Zhao 0013, Mohammadreza Robaei, Beilei Jiang, Hui Zhao 0013, Juan Fang 0004 |
ANCS | 6 |
| 2019 | Cost-Efficient Resource Provision for Multiple Mobile Users in Fog ComputingabstractFog computing is an emerging paradigm that brings the computing capabilities close to distributed IoT devices, which provides networking services between end devices and traditional cloud data centers. One important mission is to further reduce the monetary cost of fog resources while meeting the ever-growing demand of multiple users. In this paper, we focus on minimizing the total cost for multiple mobile users to provide an efficient resource provisioning scheme in fog computing. The total cost includes two aspects: the replication cost and the transmission cost. We consider two cases for the resource provision problem by focusing on different cost models. First, one simple case where users can only upload one replication is discussed, and an optimal solution is proposed by converting the original problem into one of bipartite graph matching. Then we consider a more complicated case that each user can upload multiple replications on fog nodes in the resource provisioning. For different transmission cost models, the transmission cost is related to the distance of each pair of fog nodes. This problem is proven to be NP-hard. We first propose a non-adaptive algorithm which is proved to be bounded by 2/3W+1/3OPT. Another 3+ε-approximation algorithm is proposed based on local search, which has better performance with higher complexity. Extensive simulations also prove the efficiency of our schemes. Shuaibing Lu, Jie Wu 0001, Yubin Duan, Ning Wang 0018, Juan Fang 0004 |
ICPADS | 5 |
| 2019 | Miss-aware LLC buffer management strategy based on heterogeneous multi-coreabstractWhen multiple processor (CPU) cores and a GPU integrated together on the same chip share the last-level cache (LLC), the competition for LLC is more serious. CPU and GPU have different memory access characteristics, so that they have differences in the sensitivity of LLC capacity. For many CPU applications, a reduced share of the LLC could lead to significant performance degradation. On the contrary, GPU applications have high number of concurrent threads and they can tolerate access latency. Taking into account the GPU program memory latency tolerance characteristics, we propose an LLC buffer management strategy (buffer-for-GPU, BFG) for heterogeneous multi-core. A buffer is added on the side of LLC to filtrate streaming requests of GPU. Cache-insensitive GPU messages directly access to buffer instead of accessing to LLC, thereby filtering the GPU request and freeing up the LLC space for the CPU application. Then, for the different characteristics of CPU and GPU applications, an improved LRU replacement taking into account the recent access time and access frequency of the cache block is adopted. The cache misses-aware algorithm dynamically selects the improved LRU or LRU algorithm to fit the current operating state by comparing the miss rate of cache in buffer so that the performance of the system will be improved significantly. Juan Fang 0004, Xibei Zhang, Shijian Liu, Zeqing Chang |
J. Supercomput. | 1 |