Xiaoyong Tang

dblp:50/6299 · DBLP profile ↗
← Back
67ranked-venue papers
27as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 17 first-author · 20 since 2021Computer networks · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SPADE: Attention-Guided Split Diffusion for Precise Spatial Control in Interior Layout Image Generation
Lianghao Shen, Qianqian Xing, Ronghui Cao, Xiaoyong Tang, Tan Deng
MMM (2)8
2026 A trustworthy task offloading system for heterogeneous vehicle-edge-cloud collaboration scenarios
Mingfeng Huang, Ronghui Cao, Tan Deng, Xiaoyong Tang
Future Gener. Comput. Syst.4
2026 Dynamic dual hypergraph convolutional neural networks for fine-grained drug-drug interaction prediction
Xiaoyong Tang, Xingyu Du, Hao Li 0025, Tan Deng, Ronghui Cao, Mingfeng Huang
Neurocomputing1
2026 A Dynamic Resource Utilization-Aware Task Scheduling Strategy on Spark Heterogeneous Clusters
abstract
Distributed computing engines, such as Spark, are widely used to process the large volumes of data collected by the Internet of Things (IoT) devices. In IoT systems, computing nodes often exhibit significant heterogeneity. However, most existing task scheduling algorithms neglect the performance differences among system nodes and statically assign tasks based on data locality. When high-performance nodes complete local tasks rapidly, they are subsequently assigned massive non-local tasks, resulting in severe network congestion and performance degradation. To address these issues, we propose a dynamic resource utilization-aware task scheduling strategy (DRUTS), which can efficiently utilize idle bandwidth and high-performance computing resources while also ensuring data locality. Firstly, considering the time-varying characteristics of heterogeneous node performance under various workloads, we use a sliding window to process the task information stream and evaluate the relative performance of nodes in real-time. Then, our proposed strategy adopts weighted random methods to dynamically select taskprovidersandreceiversby combining the cluster network load. Finally, it migrates and pre-executes tasks to high-performance nodes based on data distribution. We evaluate our proposed strategy performance using six typical workloads on two types of real-world heterogeneous clusters. The experimental results clearly demonstrate that our proposed strategy not only improves CPU and network utilization by 15.7% and 13.2%, respectively, but also reduces application execution time by 37.2% compared to existing work.
Xiaoyong Tang, Jiankun Xie, Ronghui Cao, Tan Deng
IEEE Internet Things J.1
2026 Remaining Workload-Aware Dynamic Task Scheduling Algorithm on Spark Heterogeneous Systems
Xiaoyong Tang, Jiankun Xie, Ronghui Cao, Tan Deng
IEEE Trans. Computers1
2025 Neural Network-Enhanced Monte Carlo Tree Search for Adaptive Resource Scheduling in Heterogeneous Spark Environments
Xiaoyong Tang, Ronghui Cao, Tan Deng
ICA3PP (6)1
2025 Fairness-Aware Federated Learning Based on Feature Attention and Contribution Calibration
Hanjing Li, Xiaoyong Tang, Qianqian Xing, Tan Deng, Mingfeng Huang, Ronghui Cao
ICIC (9)4
2025 A Node Load-Aware Horizontal Autoscaling Strategy for FaaS with Shared Resources
abstract
Function as a Service (FaaS) is a popular cloud computing service model that incorporates an auto-scaling mechanism, enabling applications to dynamically adjust computing resources, achieving rapid response to load changes and efficient resource utilization. However, the limited resource allocation mode for function containers can frequently cause function performance degradation before scaling is complete, so some FaaS platforms address this issue by default through a shared-resource mode. But existing constant target load-based autoscalers fail to perceive node-level load under this mode, leading to numerous scaling decisions to nodes that have already reached their load bottlenecks, without bringing actual resource or performance gains. This makes the system underutilized and even degrades its performance. To solve this issue, in this paper, we design a horizontal autoscaler, NDScaler, which efficiently scales functions in shared-resource mode by using the node load-aware scaling strategy, thereby eliminating invalid scaling behaviours and the resulting degradation of function performance. We implement this strategy through the proposed node load-aware and dynamic target load algorithm, which models the scale-up problem as a load transfer problem between nodes and functions and adopts a greedy search strategy to identify the optimal target functions for scale-up. Furthermore, it introduces a dynamic load target to assess the extent of load reduction for functions and accurately scales functions down. We have implemented NDScaler and evaluated it in detail on the OpenFaaS platform. Experimental results show that, compared with existing methods, NDScaler can ensure scaling effectiveness and achieve high-efficiency scaling in both simple single-function scenarios and complex multifunction scenarios, effectively improving function throughput while significantly reducing latency.
Xiaoyong Tang, Sikai Wu, Ronghui Cao, Mingfeng Huang, Tan Deng
ICPADS1
2025 FedAFW:Adaptive Feature-Driven Weighting Based Personalized Federated Learning
abstract
Federated Learning (FL) has gained widespread attention due to its strong privacy protections and collaborative learning capabilities. Recently, Personalized Federated Learning (PFL) has garnered significant attention for its ability to address statistical heterogeneity. Most existing PFL methods either focus on feature extraction, struggling to balance collaborative learning and personalization, or emphasize dynamic weight adjustments, relying on heuristic designs that lead to lower communication efficiency in large-scale federated learning systems. However, these methods fail to effectively integrate these two aspects to achieve both efficient collaborative learning and personalized goals. To address these issues, this paper proposes an Adaptive Feature-Driven Weighting Based Personalized Federated Learning (FedAFW) approach. FedAFW first utilizes local feature representations to guide the generation of global and personalized weights, enhancing the personalization effect. Subsequently, it uses gradient similarity for weight allocation, balancing the relative contributions of the global and personalized models, thus improving overall performance. Experiments on diverse datasets under heterogeneous settings show that FedAFW improves accuracy by up to 5.84%, boosts communication efficiency by 54.6%, and outperforms advanced methods in scalability and stability, demonstrating its robustness in handling statistical heterogeneity.
Ronghui Cao, Xiaoyong Tang, Hanjing Li, Tan Deng, Mingfeng Huang, Qianqian Xin
IJCNN5
2025 TSNet: A Transformer-based Medical Image Segmentation Algorithm for Improving Channel Interaction
abstract
Medical image segmentation is crucial for separating tissue structures and anatomical regions. However, due to significant variations in size, shape, and density of target tissues in medical images, this task faces many challenges. Neural networks are widely used in medical image segmentation due to their powerful feature extraction and pattern recognition capabilities. But traditional Convolutional Neural Networks (CNNs) struggle to capture long-range dependencies, and Transformer models may lack sufficient channel interaction and detail representation. To address the above issues, this paper proposes a novel architecture called TSNet, which innovatively integrates SimAM (Neural Attention Module) and Triplet Attention mechanism. First, Triplet Attention adopts a three-branch structure to effectively encodes channel and spatial information. By reducing information loss and achieving direct correspondence between channels and weights, it significantly enhances the model’s feature extraction and representation capabilities in complex medical image processing. Meanwhile, the parameter-free SimAM module generates adaptive 3D attention weights by optimizing the energy function, further optimizing the interaction and fusion between features. Finally, extensive experiments on real datasets for heart and CT segmentation have shown that the proposed TSNet performs significantly better than the baseline method in terms of Dice Similarity Coefficient (DSC) and the 95th percentile Hausdorff Distance (HD95).
Hujin Peng, Tan Deng, Shiyu Mei, Mingfeng Huang, Ronghui Cao, Xiaoyong Tang
IJCNN7
2025 Remaining Workload-Aware Task Scheduling Strategy in Spark Heterogeneous Environments
abstract
In heterogeneous distributed computing platforms, task execution containers (e.g., Spark executors) often have significant performance differences. However, most task schedulers greedily utilize resources based on the first-release-first-use policy. This leads to load imbalance across heterogeneous executors and poor application performance. To solve the above problem, we first construct a heterogeneous system task execution model. Then, we formalize the load-balancing task scheduling problem in heterogeneous environments as a minimum weighted executor waiting time problem and prove its NP-hardness. Next, a remaining workload-aware task scheduling strategy is proposed to address load imbalance among heterogeneous executors. Additionally, considering the startup overhead difference in executors, we introduce an earliest available executor wait mechanism to further optimize load-balancing. We comprehensively evaluate our proposed approaches using seven typical applications in two real-world heterogeneous environments. The experimental results clearly demonstrate that our approaches not only improve cluster load-balancing but also achieve up to 35.5% performance improvement.
Xiaoyong Tang, Jiankun Xie
IWQoS1
2025 Active-Trust Based Security Service Orchestration Framework for 6G Enabled Massive IoT
abstract
With the support for data-intensive, rate-hungry and delay-sensitive applications, 6G enabled massive IoT is surely becoming the most potential computing paradigm. Along with this trend, the scale of mobile devices and data traffic in the network is increasing explosively, resulting in huge transmission pressure on the backbone network, accompanied by serious security problems. All above call for a secure and high-throughput data communication system for 6G enabled massive IoT. In this paper, an Active-Trust based security Service Orchestration (ATSO) framework is proposed. First, the active-trust evaluation mechanism is introduced at the data acquisition layer, and direct trust is combined with indirect trust to accurately evaluate the trust of data providers. Then, service orchestration mechanism is proposed, which orchestrates data into services through edge devices to implement the service-oriented architecture, and conducts progressive aggregation at routing layer to form more advanced services. Extensive simulation results demonstrate that ATSO effectively improve performance in data security, energy efficiency and delay. Finally, we discuss the potential challenges in promoting the study of ATSO.
Mingfeng Huang, Ronghui Cao, Xiaoyong Tang, Tan Deng
TrustCom3
2025 An online resource-aware leader election algorithm based on Kubernetes load balancing
Xiaoyong Tang, Ronghui Cao
CCF Trans. High Perform. Comput.1
2025 A parallel and pipelined high speed Montgomery modular multiplier for IoT devices
Qianqian Xing, Xiaoyong Tang, Tan Deng, Ronghui Cao, Mingfeng Huang
Comput. Networks5
2025 An Adaptable Pricing-Based Resource Allocation Scheme Considering User Offloading Needs in Edge Computing
abstract
Multiaccess edge computing (MEC) is extensively utilized within the Internet of Things (IoT), wherein end-users pay services to meet the latency demands of their respective tasks. The pricing is impacted not solely by the quantity of data offloaded by the user but also associated with the leased computing and communication resources. Nevertheless, prevailing pricing strategies seldom account for the personalized resource requisites during user offloading. In this article, we present an adaptive pricing-oriented approach for concomitant task offloading and resource allocation, considering hybrid resources, comprising two key components. First, we propose a differential pricing framework for communication and computation resources, where the unit price will be influenced by the proportion of resources rented by users. Subsequently, we design a two-stage Stackelberg game model: 1) employing convex optimization theory to mitigate problem intricacies and 2) employing gradient descent to ascertain the potentially optimal price, thus achieving a balance between minimizing user expenses and maximizing server profitability. Simulation outcomes demonstrate that our approach slashes user costs by 23.3% and enhances average server revenue by 65.6% compared to a flat pricing model with a high-user request rate (five user-initiated requests per 100 ms). This maintains server occupancy within 60% to 80%, thereby alleviating user queuing and refining user Quality of Experience (QoE).
Zhuofan Liao, Xiaoyong Tang, Chaochao Feng
IEEE Internet Things J.3
2025 Context-Aware Proactive Edge Caching for Vehicular Edge Computing Based on Asynchronous Federated Learning
abstract
Edge caching is a promising technique for effectively reducing backhaul pressure and content access latency in the Internet of Vehicles (IoV). The existing content caching solutions still face the following challenges: 1) contents cached on edge servers are outdated quickly as time and user preferences change; 2) the large amount of vehicle data causes huge communication overheads; and 3) limited storage resources of edge servers. Simultaneously considering these issues to reduce transmission latency is a large-scale 0–1 constraint problem, which is NP-hard, and boosting cache hit rates is a key entry point. In this work, we propose a context-aware proactive caching strategy (CPCS) based on asynchronous federated learning (AFL), which works as follows. To improve the accuracy of content popularity prediction, thus improving the cache hit rate, we combine contextual information between different contents and use long and short-term memory networks to analyze the dynamic preferences of vehicle users. After that, vehicles complete the model training and upload via an asynchronous federation learning to complete the popularity prediction. To explore the problem of local models being outdated in AFL, CPCS integrates model compression algorithms, enhancing system efficiency and prediction accuracy. With the prediction results, CPCS gives a content placement algorithm based on the prediction results to approximate the optimal caching scheme. Simulation results show that the CPCS can improve the cache hit rate by 17% at most compared to existing state-of-the-art caching strategies.
Zhuofan Liao, Pang Liu, Xiaoyong Tang
IEEE Internet Things J.4
2025 De-Duplicated Hierarchical Offloading in Vehicular Edge Computing With Task Dependencies
abstract
In vehicular edge computing (VEC), most tasks require high real-time and energy requirements, but the mobility of vehicles and the difficulty of intelligent computing make it hard to meet these requirements. Due to the fact that most VEC tasks can be decomposed into smaller granularity, based on the dependencies between small subtasks, the repetition of tasks can be reduced, thereby improving task completion rates. In this work, we explore the dependencies of subtasks in different applications and design a two-stage multihop clustering de-duplication offloading (MCDO) mechanism. First, MCDO gives a multihop two-layer clustering (MTLC) algorithm to divide clusters based on similarities between different tasks. Based on this, MCDO further designs a de-duplication logical hierarchical offloading (DLHO). DLHO forms a directed acyclic graph (DAG) of de-duplicated subtasks in each cluster and offloads these subtasks in a logical hierarchical manner. Simulation results show that, compared to existing approaches PC5-GO, FedEdge, and MD-TSDQN, MCDO can achieve a minimum improvement of 15.1% in terms of latency and 20.8% in terms of energy consumption.
Zhuofan Liao, Zhenyi Shao, Xiaoyong Tang
IEEE Internet Things J.4
2025 An Adaptive Slicing-Based Task Admission Scheduling Strategy in Multiaccess Edge Computing
abstract
The rise of multiaccess edge computing (MEC) speeds up mobile user services and resolves service delays caused by long-distance transmission to cloud servers. However, in task-intensive scenarios, edge server processing limitations lead to buffer congestion, increasing latency and reducing Quality of Service (QoS). Furthermore, the challenges of edge server task processing are increased by the varying deadline requirements of different tasks, the time variability of task arrivals, and the real-time fluctuations of the network. In this work, we propose an adaptive slicing-based task admission scheduling strategy (ASTA) to address these issues. ASTA consists of an adaptive time slice adjustment algorithm (ASTA-I) and a task admission scheduling algorithm (ASTA-II). ASTA-I dynamically adjusts time slices based on real-time network conditions and task flow. ASTA-II first adjusts task priorities dynamically by considering factors, such as data volume, deadlines, network conditions, and buffer locations. After that, ASTA-II formulates different scheduling strategies based on changes in task priorities. These strategies are formulated to improve the throughput efficiency of edge servers and enhance the average response speed of tasks. Simulation results show that compared with the existing O2A and OTDS in different scenarios, the proposed ASTA can reduce the average number of waiting requests in the edge server buffer by 19.53%–57.73% and 20.42%–50.26%, and accelerates the average response speed of tasks by about 39.76% and 32.41%.
Zhuofan Liao, Yanpu Tang, Xiaoyong Tang, Jiawei Huang 0001
IEEE Internet Things J.3
2025 Resource-Aware Dynamic Scheduling for Tasks With Deadline Constraints on Edge Computing Systems
abstract
The proliferation of various IoT devices has brought about diverse computing requests. Scheduling delay-sensitive tasks to edge nodes closer to data sources can help alleviate core network congestion and improve system quality of service (QoS). However, with the dynamic computing requirements of changing scenarios and the imbalanced performance of limited heterogeneous edge resources, resource competition among multiple tasks has become increasingly fierce. This resource competition leads to inefficient services and performance fluctuations in edge scheduling systems. The key lies in dynamically matching task requirements and limited heterogeneous resources to improve resource utilization efficiency. To overcome this challenge, we propose a resource-aware task grouping scheduling strategy (RATGS) based on our proposed group-based and sharedstate edge scheduling framework, aiming to improve the overall service quality of edge computing systems. We perform extensive evaluation on multiple metrics using realistic workloads and realworld traces. The experimental results demonstrate that RATGS improves the task completion rate by 7.56%∼50.1% before the deadline and improves the efficiency of resource utilization by 17.7%∼94.8% compared with existing baseline strategies. In addition, RATGS performed second best in terms of average completion time.
Wenbiao Cao, Xiaoyong Tang, Tan Deng, Ronghui Cao, Keqin Li 0001
IEEE Trans. Cloud Comput.2
2024 RFR-ABROF: A Multi-Strategy Collaborative Classification Prediction Model Based on Rotation Forest for PM2.5
abstract
Predicting dust pollution is necessary to achieve good air quality and patient recovery. In the era of big data, there are already many machine learning algorithms for predicting the concentration of air pollutant PM2.5. However, these model methods perform poorly when dealing with massive air quality data sets and do not solve the impact of data distribution imbalance. Therefore, a multi-strategy collaborative feature selection method combining random forests and recursive feature elimination with adaptive boosting rotation forest (RFR-ABROF) algorithm is proposed. On this basis, multiclass adaptive synthetic sampling (Multi-ADASYN) strategy is introduced to balance the imbalanced air quality data set. We verified the effectiveness of the model through comparative experiments, which show our proposed model has the higher values of accuracy, precision, recall, f1-score, and ROC curve area under evaluation indicators in the air quality data set of four locations in most cases, and the more considerable the amount of data, the better the model prediction performance.
Xiaoyong Tang, Tan Deng, Ronghui Cao, Zeyuan Tu, XingJiang Hu
CSCWD1
2024 An Adaptive Hoeffding Tree Model Based on Differential Entropy and Relative Entropy for Concept Drift Detection
abstract
The concept drift detection algorithm can timely respond to and adjust the model by monitoring changes in data distribution over time. However, the dynamically adjusted ensemble model may still retain some components with weak adaptability. These components are involved in subsequent training and testing phases, leads to a significant decrease in classification performance. To solve these problems, this paper proposes an Adaptive Hoeffding Tree Model Based on Differential Entropy and Relative Entropy (AHT-DERE) for concept drift detection. It adopts a two-step strategy: a) A differential entropy-based drift detection method, which calculates the information entropy of the two most recently arrived data samples, and quantifies the difference between data distributions by subtracting the entropy values. This measurement serves as the criteria for determining the occurrence of concept drift. b) A relative entropy-based dynamic adjustment method, which utilizes the relative entropy similarity between the fitted and true distributions of the current data samples. This method selects well-adapted components for each round of incremental updates to improve the resilience of the ensemble model to concept drift. Compared to advanced algorithms, experimental results show that in two sets of experiments, the classification performance of AHT-DERE achieved an average improvement of 6.36% and 5.94% on seven publicly available real-world and synthetic datasets, respectively. The maximum improvement reached 13.55% and 10.82%, respectively.
Yongtong Gu, Xiaoyong Tang, Ronghui Cao, Tan Deng
IJCNN5
2024 SecureVeil: A Modular Architecture with Deep Cosine Transformation and Secure Key Fusion for Face Template Protection
abstract
Face template protection has received widespread attention in the field of biometric security. However, most of the recent schemes can not resistant to randomness source exposure, resulting in compromised protected templates, and the verification accuracy needs to be further improved to meet the needs of realistic applications. In this paper, we propose a novel modular architecture called SecureVeil for protecting face templates. SecureVeil protects face templates with a deep cosine transformation network called FlexNet, which performs random orthogonal transforms on face templates using user-specific keys. To protect user-specific keys, SecureVeil employs a secure key fusion construction called SecureFusion, which fuses user-specific keys with face templates and permutation vectors. We evaluate the irreversibility, unlinkability and verification accuracy of SecureVeil on two state-of-the-art face recognition systems, including ArcFace and FaceNet, using three benchmarking datasets, including MOBIO, LFW, and CFP. Experimental results show that the irreversibility of SecureVeil outperforms existing related schemes. Its verification accuracy is superior than all these compared schemes with an average improvement of 7.70%, and improves by 12.45% on FaceNet when using the LFW dataset. Overall, SecureVeil meets the four criteria for face template protection.
Wenzhuo Han, Shun Qin, Xiaoyong Tang, Ronghui Cao, Tan Deng
IJCNN6
2024 Pricing-Based Task Offloading Considering User Energy Consumption in MEC System
abstract
Multi-access Edge Computing (MEC) can find its wide applications in various resource-constrained scenario where the user needs to pay a price to the server for meeting the latency requirements of their own tasks. However, to the best of our knowledge, there is no effective pricing model in MEC taking into account the impact of the user’s own energy consumption on the server’s price scheme. In fact, the transmission energy consumption affects the users’ offloading task decisions, which in turn affects the revenue of the edge servers. To fill in this gap, a novel pricing mechanism for energy-sensitive users is proposed in this paper, specifically: 1) We propose a joint optimization problem of computation offloading decisions and resource pricing that considers user energy consumption. 2) By constructing a two-stage Stackelberg game model, the first stage simplifies the problem by constraint relaxation theory and finds an approximate optimal offloading strategy through the Fast Genetic Algorithm (FGA). The second stage uses the gradient descent method to find the potential optimal prices, thereby achieving a balance between the lowest user cost and the highest server profit. Experimental results show that, compared to traditional algorithms, our approach optimizes the average user cost and user task completion time by 24.65% and 12.37%, respectively.
Zhuofan Liao, Xiaoyong Tang
ISPA3
2024 An Efficient Cooperative Active Caching Strategy in Vehicular Edge Network Based on Asynchronous Federated Learning
abstract
Edge caching is a promising technique for effectively reducing backhaul pressure and content access latency in the Internet of Vehicles (IoV). However, the high mobility of vehicles and dynamic user requests often lead to outdated cached content. Expired cache wastes the limited storage space and transmission in vehicular edge computing. Improving the cache hit rate is an effective approach to address these issues. In this work, we propose a Cooperative Active Caching Strategy (CACS) which works as follows. First, user preferences are analyzed from the historical content data of vehicle users. After that, multiple vehicles cooperate to learn the global model under an asynchronous federated learning framework, and get the content popularity from user preferences and content features. To explore the problem of local models being outdated in asynchronous federated learning, CACS integrates model compression algorithms, enhancing system efficiency and prediction accuracy. Finally, a heuristic cooperative caching content placement algorithm is proposed based on a greedy policy to minimize average access latency. Simulation results show that the CACS can improve the cache hit rate by 15% at most compared to existing state-of-the-art caching strategies.
Pang Liu, Zhuofan Liao, Xiaoyong Tang
ISPA3
2024 A Task Dependency-based Deduplicated Task Offloading Mechanism in Vehicular Edge Computing
abstract
The increasing demand for in-vehicle applications has raised the complexity and computational load, while the in-vehicle tasks exhibit a sensitivity to latency. Previous research has proposed utilizing the idle computational resources of roadside vehicles to alleviate this contradiction. However, the high mobility of vehicles leads to communication interruptions, and the time-varying nature of vehicle density makes resource allocation challenging. In this work, we leverage the dependencies between vehicular computing tasks and design a deduplication offloading mechanism for stable reduction of latency. This mechanism consists of two stages, named the Multi-hop Clustering Deduplication Offloading (MCDO) mechanism. Firstly, a Multi-hop Two Layer Clustering (MTLC) algorithm is designed to divide vehicles based on task dependencies, speed, and position information. Then, a Deduplication Layered Offloading (DLO) algorithm is proposed to identify and remove duplicated tasks within each cluster while maintaining their inter-dependencies. Simulation results demonstrate that MCDO effectively divides vehicle clusters and offloads tasks efficiently under various road conditions. Compared to existing approaches, MCDO significantly enhances system performance, achieving a minimum improvement of 15.1% in terms of latency.
Zhenyi Shao, Zhuofan Liao, Xiaoyong Tang
ISPA3
2024 A Task Admission Scheduling Strategy Based on Adaptive Time Slicing in Mobile Edge Computing
abstract
The rise of Multi-access Edge Computing (MEC) speeds up mobile user services and resolves service delays caused by long-distance transmission to cloud servers. However, in task-intensive scenarios, edge server processing limitations lead to buffer congestion, increasing latency and reducing Quality of Service (QoS). Furthermore, the challenges of edge server task processing are increased by the varying deadline requirements of different tasks, the time variability of task arrivals, and the real-time fluctuations of the network. In this paper, we propose an Adaptive Time Slice Admission Scheduling (ATSAS) strategy to solve these problems. Specifically, ATSAS proposes an Adaptive Time Slice Algorithm (ATSA) and a task Admission Scheduling Algorithm (ASA). ATSA dynamically adjusts the time slice based on real-time network conditions and task flow characteristics. ASA, assisted by ATSA, dynamically adjusts task priorities based on the remaining data amount, deadline requirements, real-time network conditions, and the location of the task buffer. After that, ASA makes different scheduling strategies for tasks. Through the coordination of ATSA and ASA, the proposed strategy ensures the high efficiency and fairness of processing tasks in different realistic scenarios. The simulation results show that, compared with the existing O2A and OTDS in different scenarios, the proposed ATSAS reduces the average number of waiting services in the edge server buffer by 38.63% and 35.34%, and accelerates the average response speed of tasks by about 39.76% and 32.41%.
Yanpu Tang, Zhuofan Liao, Xiaoyong Tang
ISPA3
2024 A SFC Placement Strategy in Low-Visibility Multi-Domains Networks for Low Resource Usage Cost
abstract
The concept of Network Function Virtualization (NFV) enables the realization of services as a Service Function Chain (SFC) that is composed of multiple Virtual Network Functions (VNFs), allowing for flexible deployment across the network. For wireless networks comprised of diverse domains which means be managed by different network providers, each domain has different resource and transmission costs. Due to commercial confidentiality reasons, the internal information of these domains is not interconnected, which increases the difficulty of SFC deployment. How to deploy SFCs in multi-domains networks at the lowest cost is becoming a challenge. This paper proposes a novel strategy for the scenario that information of sub-domain is invisible. The strategy firstly splits the SFC based on the processing dependency of its VNFs to reduce intermediate data transmission. Then it utilizes a Graph Matching Network (GMN) to evaluate the matching degree between sub-domain and SFC fragment. This matching degree acts as a reference for optimizing SFC allocation, thereby avoiding the direct acquisition of internal sub-domain information. Simulation results demonstrate that the proposed strategy not only significantly reduces resource, but also maintains a higher request acceptance ratio.
Qingli Xiong, Zhuofan Liao, Xiaoyong Tang
ISPA3
2024 Entropy Normalization SAC-Based Task Offloading for UAV-Assisted Mobile-Edge Computing
abstract
With the advantages of maneuverability and low cost, Unmanned Aerial Vehicles (UAVs) are widely deployed in mobile edge computing as micro servers to provide computing service. However, tasks usually require a large amount of energy and have strict time constraints, while the battery energy and endurance of UAVs are limited. Therefore, energy consumption and delay have become key issues in such architectures. To address this issue, an Entropy Normalized Soft Actor-Critic (ENSAC) computation offloading algorithm is proposed in this paper, aiming to minimize the weighted sum of task offloading delay and energy consumption. In ENSAC, we formulate the task offloading problem as a Markov Decision Process (MDP). Considering the non-convexity, high-dimensional state space, and continuous action space of this problem, the ENSAC algorithm fully combines deviation strategy and maximum entropy reinforcement learning, and designs a system utility function under entropy normalization as a reward function, thus ensuring fairness in weighted energy consumption and delay. What’s more, ENSAC algorithm also considers UAV trajectory planning, task offloading ratio, and power allocation in the UAV-assisted MEC system. Therefore, compared with previous methods, ENSAC algorithm has stronger stability, better exploration performance, and can handle more complex environments and larger action space. Finally, extensive experiments demonstrate that, in both energy-saving and delay-sensitive scenarios, the ENSAC algorithm can quickly converge to the optimal solution while maintaining stability. Compared with four benchmark algorithms, it reduces the total system cost by 52.73%.
Tan Deng, Ronghui Cao, Yongtong Gu, Jinming Hu, Xiaoyong Tang, Mingfeng Huang, Shixue Li
IEEE Internet Things J.7
2024 An Adaptive Deployment Scheme of Unmanned Aerial Vehicles in Dynamic Vehicle Networking for Complete Offloading
abstract
Unmanned Aerial Vehicles (UAVs), due to their flexible deployment, are used as a three-dimensional space assistant tool for Vehicular Edge Computing (VEC) to cover moving vehicles. However, existing work generally assumes relatively uniform vehicle distribution while the actual road conditions vary over time. The time-varying location of vehicles and road congestion in peak hours pose challenges to vehicular edge computing. First, traffic congestion can lead to imbalanced UAVs load, that is UAVs covering congested areas are overloaded while others remain idle. Second, after the high-speed moving vehicle leaves the service range of the current UAV, it is unable to receive computing results of the original request task, which means task processing failure. In this work, we propose a framework of UAV Clusters Adaptive Deployment (UCAD) to address these issues. By clustering vehicles, UCAD gives a Density-Based Adaptive Region Determination algorithm (DBARD) to determine congested areas and dynamically update them to adapt to dynamic network environments. After that, UCAD presents a Particle Swarm Optimization-based Cluster deployment algorithm (PSOC), deploying UAV clusters within determined areas to provide continuous services for vehicles. Simulation results demonstrate that UCAD can adaptively deploy UAV clusters to assist vehicles based on traffic congestion conditions. Simulation results show that compared to existing works, CONEC, TU2V and IELTS, the proposed UCAD can increase the task success rate by 8.7%, 18.5%, and 23.5%, respectively. UCAD can achieve better performance on UAVs workload balance.
Zhuofan Liao, Chuhao Yuan, Xiaoyong Tang
IEEE Internet Things J.4
2024 Exploring Potential Customized Bus Passengers Across Private Car Trajectory Data
abstract
Customized bus is considered an effective means to alleviate traffic congestion and reduce traffic-related environmental pollution caused by the increasing number of private cars. Exploring potential passenger information as the first stage of customized bus service has become a popular topic. Unlike manual investigation and passenger request methods, current studies utilize data mining methods to actively explore potential passengers from various historical travel data. However, the existing data mining methods only consider the spatiotemporal features of potential passengers and neglect the semantic features related to customized bus services, which play an important role in determining whether passengers are willing to use services. In this paper, we treated the exploration of potential customized bus passengers as a binary classification problem based on private car trajectory data. Then, we propose a novel data mining method, named iTrAdaboost-DTCN, which combines the strengths of deep learning and transfer learning. In detail, it integrates state-of-the-art deep neural networks by constructing a deep trajectory classification network (DTCN), which can automatically extract semantic feature representations to help improve classification accuracy. Due to the lack of city-wide labeled customized bus passenger information in practice, it also integrates instance-based transfer learning through improved TrAdaboost, which solves the learning problem of the target classification domain with limited labeled samples. Experimental results demonstrate that our method can explore potential passengers more effectively than other baseline methods. Furthermore, we apply our method to real-world scenarios and compare three travel characteristics of identified customized and non-customized bus passengers.
Linjiang Zheng, Xiaoyong Tang, Sisi Xiao, Min Zhao 0010, Dihua Sun
IEEE Trans. Intell. Transp. Syst.4
2024 A Cooperative Community-Based Framework for Service Caching and Task Offloading in Multi-Access Edge Computing
abstract
In multi-access edge computing, services are cached from cloud servers to the edge base station providing low-latency services. Horizontal Collaboration (HC) between Base Stations (BSs) can alleviate the problem of resource limitation in edge base stations. However, the initiative of cooperation between base stations and the interests of base stations themselves are not guaranteed. In this paper, we study the Joint Service Caching and Task Offloading decision Problem (JSCTOP). To solve this problem, we propose a Two-Stage Optimization Framework (TSOF) based on cooperative community idea to maximize the benefit of each base station. In the first stage, a Collaborative Community Mechanism (CCM) is designed based on the Dissimilarity of Service Types (DST) to enhance the collaborative capability of BSs. In the second stage, an algorithm for service caching and task decision to update and replace the Service Types (STs) of BSs. We also design an algorithm for service pricing to incentivize BSs. In this algorithm, BSs are able to profit from providing services to users and also profit from assisting other base stations. Extensive simulation results show that TSOF outperforms its counterparts in average delay, total energy, total benefit, and community.
Zhuofan Liao, Guiying Yin, Xiaoyong Tang, Penglu Liu
IEEE Trans. Netw. Serv. Manag.3
2023 Improved Deep Embedded K-Means Clustering with Implicit Orthogonal Space Transformation
abstract
The deep clustering algorithm can learn the latent embedded features of the data through the autoencoder, and cluster the data according to the similarity of the latent features. However, the feature information obtained by the autoencoder may not have a better value for the clustering algorithm and is not suitable for clustering, which greatly reduces the clustering effect. This paper proposes a deep K-means clustering algorithm with implicitly embedded space transformation to answer this question. We implicitly transform the latent feature space into a new type of space that is more friendly to the clustering task, which preserves space invariance. This implicit transformation is done through an orthogonal transformation matrix. The orthogonal transformation matrix is composed of the eigenvectors of the intra-class scattering matrix and the inter-class scattering matrix. In the new space, clusters can be better separated by cluster cohesion and inter-cluster difference. We alternately optimize feature acquisition and clustering to adjust the embedding space and disperse the embedding points, to enrich the clustering information in the latent feature space. Experimental results show that our proposed algorithm can produce better high-quality clusters than many current correlation clustering algorithms on the same experimental dataset.
Xiaoyong Tang, Tan Deng, Ronghui Cao
COMPSAC4
2023 Approx-SMOTE Federated Learning Credit Card Fraud Detection System
Yifei Kou, Dicheng Xiao, Xiaoyong Tang
COMPSAC6
2023 Sequenced Quantization RNN Offloading for Dependency Task in Mobile Edge Computing
Tan Deng, Shixue Li, Xiaoyong Tang, Ronghui Cao, Wenbiao Cao
ICA3PP (2)3
2023 A Grouping-Based Multi-task Scheduling Strategy with Deadline Constraint on Heterogeneous Edge Computing
Xiaoyong Tang, Wenbiao Cao, Tan Deng
ICA3PP (2)1
2023 A Seasonal Decomposition-Based Hybrid-BHPSF Model for Electricity Consumption Forecasting
Xiaoyong Tang, Ronghui Cao
ICA3PP (5)1
2023 A job scheduling algorithm based on parallel workload prediction on computational grid
Xiaoyong Tang, Tan Deng, Zexin Zeng, Haowei Huang, Qiyu Wei, Xiaorong Li
J. Parallel Distributed Comput.1
2023 A Multi-Task BERT-BiLSTM-AM-CRF Strategy for Chinese Named Entity Recognition
Xiaoyong Tang, Chengfeng Long
Neural Process. Lett.1
2023 Service Cost Effective and Reliability Aware Job Scheduling Algorithm on Cloud Computing Systems
abstract
Nowadays, increasing number of services are provided to individuals and organizations through cloud computing systems in apay-as-you-usemodel. This business service paradigm encounters several cloud Quality of Service (QoS) challenges, such as reliability, cost, and response time. The most common mechanism to improve cloud service reliability is a primary/backup (PB) fault-tolerant technique. However, this reliability enhancement technique inevitably results in multiple replications, which lead to high service cost. In recognition of these challenges, we first build a cloud computing systems resources management architecture. Then, we analyze the cloud service execution reliability on the physical resources of a VM and used a CUDA (Compute Unified Device Architecture)-enabled parallel two-dimensional long short-term memory neural network to predict the software faults of a cloud VM. Third, we propose an effective primary/backup cloud service cost calculation approach. To overcome the cloud service response time constraint, we integrate a response time slack factor into this method. Fourth, we formulate the cloud service reliability and cost aware job scheduling problem, which aims at minimizing the total cloud service cost and rejection rate, and improving the system reliability. Fifthly, a heuristic greedy reliability and cost aware job scheduling (RCJS) algorithm is proposed. Finally, a performance evaluation is conducted and the experimental results demonstrate that our proposed RCJS algorithm significantly outperforms optimal redundant VM placement (OPVMP), MIN-MIN algorithms in terms of average service cost and rejection rate. This algorithm also demonstrates good trade-off of reliability when compared to the other two algorithms and is suitable for cloud services with high reliability and low-cost requirements.
Xiaoyong Tang, Zeng Zeng, Bharadwaj Veeravalli
IEEE Trans. Cloud Comput.1
2022 An improved DECPSOHDV-Hop algorithm for node location of WSN in Cyber-Physical-Social-System
Tan Deng, Xiaoyong Tang, Wei Wei 0006, Zeng Zeng
Comput. Commun.2
2022 5G-based smart healthcare system designing and field trial in hospitals
abstract
Abstract With the 5G worldwide deployment, the scale of vertical applications is innovated benefit from 5G technologies including MEC (Multi‐access Edge Computing), network slicing, etc. Especially for healthcare, 5G had been used for COVID‐19 protection and intelligent medical processing. However, limited by the hospital's traditional information infrastructures, those 5G‐based healthcare applications are hard to be deployed and most only for demonstration, also isolated from the existing medical systems. So what is the next generation of smart healthcare information infrastructures is the key issue for the long‐term development of 5G healthcare applications. Even though the standardized 5G MEC framework has been widely used in many vertical scenarios, it is also hard to satisfy hospital‐specific requirements such as hospital‐dedicated deployment, medical data security, and various network connections, etc. This paper proposes a 5G‐based architecture for smart healthcare information infrastructure, a new network element iGW (industry gateway) is defined, and the smart healthcare dedicated cloud platform iMEP (industry multi‐access edge platform) is also introduced here, making it possible to satisfy both the hospital‐specific requirements and the long‐term evolution. Meanwhile, the implementation methodology and the corresponding field test results are presented, which show the significant network performance gain achieved by the proposed new system structure.
Xiaoyong Tang, Jing Chong, Zhengpeng You, Haiying Ren, Yuxiang Shang, Yantao Han
IET Commun.1
2022 A forward and backward private oblivious RAM for storage outsourcing on edge-cloud computing
Zhubin Cai, Xiaoyong Tang, Yuming Xu, Tan Deng
J. Parallel Distributed Comput.3
2022 Reliability-Aware Cost-Efficient Scientific Workflows Scheduling Strategy on Multi-Cloud Systems
abstract
Nowadays, more and more computation-intensive scientific applications with diverse needs are migrating to cloud computing systems. However, the cloud systems alone cannot meet applications’ requirements at all times with the increasing demands from users. Therefore, the multi-cloud systems that can provide scalable storage and computing resources become a good solution. The main challenges for such systems are multiple billing mechanisms, virtual resources heterogeneity, and systems reliability. In response to these challenges, we first build a multi-cloud systems fault-tolerant workflow scheduling framework, which tries to improve the scientific applications execution reliability and reduce their execution cost. Then, we use Weibull distribution to analyze task execution reliability and hazard rate, which is used to duplicate task with high execution hazard rate. Third, we integrate different multi-cloud providers’ billing mechanism into the proposed scheduling framework, and this workflow scheduling problem is mathematically formulated as an optimization problem. Fourth, we define the DAG tasks cost-efficient bottom level, and propose a fault-tolerant cost-efficient workflow scheduling algorithm (FCWS) that minimizes application execution cost, time while ensuring their reliability. Simulation experiments for performance evaluation were conducted based on two real-world applications: Epigenomics and LIGO. The results clearly demonstrate that our proposed FCWS algorithm outperforms existing FR-MOS, CWS in terms of cost and reliability, and FCWS is also better than CWS and inferior to FR-MOS in term of makespan.
Xiaoyong Tang
IEEE Trans. Cloud Comput.1
2022 CC-RRTMG_SW++: Further optimizing a shortwave radiative transfer scheme on GPU
Fei Li 0042, Xiaohui Ji, Jinrong Jiang, Xiaoyong Tang, He Zhang 0005
J. Supercomput.6
2022 Cost-Efficient Workflow Scheduling Algorithm for Applications With Deadline Constraint on Heterogeneous Clouds
abstract
In recent years, more and more large-scale data processing and computing workflow applications run on heterogeneous clouds. Such cloud applications with precedence-constrained tasks are usually deadline-constrained and their scheduling is an essential problem faced by cloud providers. Moreover, minimizing the workflow execution cost based on cloud billing periods is also a complex and challenging problem for clouds. In realizing this, we first model the workflow applications as I/O Data-aware Directed Acyclic Graph (DDAG), according to clouds with global storage systems. Then, we mathematically state this deadline-constrained workflow scheduling problem with the goal of minimum execution financial cost. We also prove that the time complexity of this problem is NP-hard by deducing from a multidimensional multiple-choice knapsack problem. Third, we propose a heuristic cost-efficient task scheduling strategy called CETSS, which includes workflow DDAG model building, task subdeadline initialization, greedy workflow scheduling algorithm, and task adjusting method. The greedy workflow scheduling algorithm mainly consists of dynamical task renting billing period sharing method and unscheduled task subdeadline relax technique. We perform rigorous simulations on some synthetic randomly generated applications and real-world applications, such as Epigenomics, CyberShake, and LIGO. The experimental results clearly demonstrate that our proposed heuristic CETSS outperforms the existing algorithms and can effective save the total workflow execution cost. In particular, CETSS is very suitable for large workflow applications.
Xiaoyong Tang, Wenbiao Cao, Huiya Tang, Tan Deng, Jing Mei, Zeng Zeng
IEEE Trans. Parallel Distributed Syst.1
2020 Multifractal detrended fluctuation analysis parallel optimization strategy based on openMP for image processing
Xiaoyong Tang, Xiaopan Yang, Fan Wu 0016
Neural Comput. Appl.1
2020 Interconnection Network Energy-Aware Workflow Scheduling Algorithm on Heterogeneous Systems
abstract
Heterogeneous systems based on multicore (CPU) and manycore (GPU) processors have been regarded as an important computing infrastructure in recent years. Large-scale computationally intensive scientific workflow applications have recently been deployed on such systems. However, improving the system performance and reducing the energy consumption under user deadline constraints remain challenging problems. In this article, we first investigate the computing node network energy consumption problem of fat-tree interconnection networks for a low communication-to-computation ratio workflow application. We then propose a heuristic list-based network energy-efficient workflow scheduling (NEEWS) algorithm including top-level task computing, task subdeadline initialization, a dynamic adjustment, and an edge data optimization communication method. Extensive simulations were conducted based on randomly generated workflow applications and two real-world scientific applications. The experiment results clearly demonstrate that our proposed workflow scheduling strategy outperforms three other algorithms in terms of energy consumption. In particular, NEEWS is extremely suitable owing to its high parallelism and low communication in large-scale scientific applications.
Xiaoyong Tang, Weiqiang Shi, Fan Wu 0016
IEEE Trans. Ind. Informatics1
2020 Indoor Crowd Density Estimation Through Mobile Smartphone Wi-Fi Probes
abstract
Crowd density estimation is one of the critical issues in social activities. The traditional solution to this problem is to leverage video surveillance to monitor a crowd. However, this is not accurate for crowd density estimation because it is still hard to identify people from background. In the past few years, more and more people use Wi-Fi enabled smartphones. Smartphones can send Wi-Fi request packets periodically, even when they are not connected to access points. This gives another promising solution to the crowd density estimation even for the public environment. In this paper, we first develop a Wi-Fi monitor detection that can capture smartphone passive Wi-Fi signal information including MAC address and received signal strength indicator. Then, we propose a positioning algorithm based on smartphone passive Wi-Fi probe and a dynamic fingerprint management strategy. In real-world public social activities, a person may have zero, one, two, or multiple smartphones with variant Wi-Fi signals. Therefore, we design a method of computing the probability of a user generating one Wi-Fi signal to identify people population. Finally, we propose a crowd density estimation solution based on Wi-Fi probe packets positioning algorithm. Experiments were conducted in an indoor laboratory class and three public social activities, clearly demonstrated that the proposed solution can effectively and accurately estimate crowd density.
Xiaoyong Tang, Bin Xiao 0001, Kenli Li 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Budget-constraint stochastic task scheduling on heterogeneous cloud systems
abstract
Summary In the past few years, more and more business‐to‐consumer and enterprise applications run in the heterogeneous clouds. Such cloud bag‐of‐tasks applications are usually budget constrained, and their scheduling is an essential problem for cloud provider. The problem is even more complex and challenging when the accurate knowledge about task execution time is unknown in advance. Focusing on these challenges, we first build a cloud resource management architecture and stochastic task model, which divides cloud task into two execution parts. Then, we deduce bag‐of‐tasks applications' schedule length (Makespan) and total cost according to heterogeneous clouds' online feedback information of task first part execution. Thirdly, we formulate this stochastic scheduling problem as a linear programming problem. Lastly, we propose a time and cost multiobjective stochastic task scheduling genetic algorithm, in which can find Pareto optimal schedules for stochastic cloud task that meet its budget constraint. The extensive simulation experiments were carried out on a heterogeneous cloud platform with 400 virtual machines, and tasks were derived from Parallel Workloads Archive and the analysis data of real‐world cloud systems. The experimental results show that our proposed stochastic task scheduling genetic algorithm can get shorter schedule length and lower cost with task budget constraints.
Xiaoyong Tang, Zhuojun Fu
Concurr. Comput. Pract. Exp.1
2015 Scheduling Precedence Constrained Stochastic Tasks on Heterogeneous Cluster Systems
abstract
Generally, a parallel application consists of precedence constrained stochastic tasks, where task processing times and intertask communication times are random variables following certain probability distributions. Scheduling such precedence constrained stochastic tasks with communication times on a heterogeneous cluster system with processors of different computing capabilities to minimize a parallel application’s expected completion time is an important but very difficult problem in parallel and distributed computing. In this paper, we present a model of scheduling stochastic parallel applications on heterogeneous cluster systems. We discuss stochastic scheduling attributes and methods to deal with various random variables in scheduling stochastic tasks. We prove that the expected makespan of scheduling stochastic tasks is greater than or equal to the makespan of scheduling deterministic tasks, where all processing times and communication times are replaced by their expected values. To solve the problem of scheduling precedence constrained stochastic tasks efficiently and effectively, we propose a stochastic dynamic level scheduling (SDLS) algorithm, which is based on stochastic bottom levels and stochastic dynamic levels. Our rigorous performance evaluation results clearly demonstrate that the proposed stochastic task scheduling algorithm significantly outperforms existing algorithms in terms of makespan, speedup, and makespan standard deviation.
Kenli Li 0001, Xiaoyong Tang, Bharadwaj Veeravalli, Keqin Li 0001
IEEE Trans. Computers2
2014 Energy-Efficient Stochastic Task Scheduling on Heterogeneous Computing Systems
abstract
In the past few years, with the rapid development of heterogeneous computing systems (HCS), the issue of energy consumption has attracted a great deal of attention. How to reduce energy consumption is currently a critical issue in designing HCS. In response to this challenge, many energy-aware scheduling algorithms have been developed primarily using the dynamic voltage-frequency scaling (DVFS) capability which has been incorporated into recent commodity processors. However, these techniques are unsatisfactory in minimizing both schedule length and energy consumption. Furthermore, most algorithms schedule tasks according to their average-case execution times and do not consider task execution times with probability distributions in the real-world. In realizing this, we study the problem of scheduling a bag-of-tasks (BoT) application, made of a collection of independent stochastic tasks with normal distributions of task execution times, on a heterogeneous platform with deadline and energy consumption budget constraints. We build execution time and energy consumption models for stochastic tasks on a single processor. We derive the expected value and variance of schedule length on HCS by Clark's equations. We formulate our stochastic task scheduling problem as a linear programming problem, in which we maximize the weighted probability of combined schedule length and energy consumption metric under deadline and energy consumption budget constraints. We propose a heuristic energy-aware stochastic task scheduling algorithm called ESTS to solve this problem. Our algorithm can achieve high scheduling performance for BoT applications with low time complexity O(n(M + logn)), where n is the number of tasks and M is the total number of processor frequencies. Our extensive simulations for performance evaluation based on randomly generated stochastic applications and real-world applications clearly demonstrate that our proposed heuristic algorithm can improve the weighted probability that both the deadline and the energy consumption budget constraints can be met, and has the capability of balancing between schedule length and energy consumption.
Kenli Li 0001, Xiaoyong Tang, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.2
2013 Novel clock synchronization algorithm of parametric difference for parallel and distributed simulations
Linjun Fan, Yunxiang Ling, Xiaomin Zhu 0001, Xiaoyong Tang
Comput. Networks5
2012 Energy-Aware Scheduling Algorithm for Task Execution Cycles with Normal Distribution on Heterogeneous Computing Systems
abstract
In the past few years, many energy-aware scheduling algorithms have been developed primarily using the dynamic voltage-frequency scaling (DVFS) capability which has been incorporated into recent commodity processors. However, these techniques are unsatisfied with optimizing both schedule length and energy consumption. Furthermore, most algorithms schedule tasks according to their average case execution time and not consider the task's execution cycles with probability distribution in real-world. In recognition of this, we study the problem of scheduling independent stochastic tasks with normal distribution, deadline and energy consumption budget constraints on a heterogeneous platform. We first formulate this energy-aware stochastic scheduling problem as a linear programming, which maximize the guaranteed confidence probabilities under deadline and energy consumption budget constraints. Then, we propose a heuristic energy-aware stochastic tasks scheduling algorithm (ESTS) to solve this problem, which can achieve high schedule performance for independent tasks with lower complexity. Our extensive simulation performance evaluation study, based on randomly generated stochastic applications and real-world applications, clearly demonstrate that our proposed heuristic algorithm can improve system guaranteed confidence probability and has a good trade-off between schedule length and energy consumption.
Kenli Li 0001, Xiaoyong Tang, Qifeng Yin
ICPP2
2012 A hierarchical reliability-driven scheduling algorithm in grid systems
Xiaoyong Tang, Kenli Li 0001, Meikang Qiu, Edwin H.-M. Sha
J. Parallel Distributed Comput.1
2011 A stochastic scheduling algorithm for precedence constrained tasks on Grid
Xiaoyong Tang, Kenli Li 0001, Guiping Liao, Kui Fang, Fan Wu 0016
Future Gener. Comput. Syst.1
2011 A Novel Security-Driven Scheduling Algorithm for Precedence-Constrained Tasks in Heterogeneous Distributed Systems
abstract
In the recent past, security-sensitive applications, such as electronic transaction processing systems, stock quote update systems, which require high quality of security to guarantee authentication, integrity, and confidentiality of information, have adopted heterogeneous distributed system (HDS) as their platforms. This is primarily due to the fact that single parallel-architecture-based systems may not be sufficient to exploit the available parallelism with the running applications. Most security-aware applications end up in handling dependence tasks, also referred to as Directed Acyclic Graph (DAG), on these HDSs. Unfortunately, most existing algorithms for scheduling such DAGs in HDS fail to fully consider security requirements. In this paper, we systematically design a security-driven scheduling architecture that can dynamically measure the trust level of each node in the system by using differential equations. To do so, we introduce task priority rank to estimate security overhead of such security-critical tasks. Furthermore, we propose a security-driven scheduling algorithm for DAGs which can achieve high quality of security for applications. Our rigorous performance evaluation study results clearly demonstrate that our proposed algorithm outperforms the existing scheduling algorithms in terms of minimizing the makespan, risk probability, and speedup. We also observe that the improvement obtained by our algorithm increases as the security-sensitive data of applications increases.
Xiaoyong Tang, Kenli Li 0001, Zeng Zeng, Bharadwaj Veeravalli
IEEE Trans. Computers1
2010 List scheduling with duplication for heterogeneous computing systems
Xiaoyong Tang, Kenli Li 0001, Guiping Liao, Renfa Li
J. Parallel Distributed Comput.1
2010 Reliability-aware scheduling strategy for heterogeneous distributed computing systems
Xiaoyong Tang, Kenli Li 0001, Renfa Li, Bharadwaj Veeravalli
J. Parallel Distributed Comput.1
2009 Communication contention in APN list scheduling algorithm
Xiaoyong Tang, Kenli Li 0001, David A. Padua
Sci. China Ser. F Inf. Sci.1
2007 An Overview of a Compiler for Mapping Software Binaries to Hardware
abstract
As new applications in embedded communications and control systems push the computational limits of digital signal processing (DSP) functions, there will be an increasing need for software applications to be migrated to hardware in the form of a hardware-software codesign system. In many cases, access to the high-level source code may not be available. It is thus desirable to have a technology to translate the software binaries intended for processors to hardware implementations. This paper provides details on the retargetable FREEDOM compiler. The compiler automatically translates DSP software binaries to register-transfer level (RTL) VHDL and Verilog for implementation on field-programmable gate arrays (FPGAs) as standalone or system-on-chip implementations. We describe the underlying optimizations and some novel algorithms for alias analysis, data dependency analysis, memory optimizations, procedure call recovery, and back-end code scheduling. Experimental results on resource usage and performance are shown for several program binaries intended for the Texas Instruments C 6211 DSP (VLIW) and the ARM 922 T reduced instruction set computer (RISC) processors. Implementation results for four kernels from the Simulink demo library and others from commonly used DSP applications, such as MPEG-4, Viterbi, and JPEG are also discussed. The compiler generated RTL code is mapped to Xilinx Virtex II and Altera Stratix FPGAs. We record overall performance gains of 1.5-26.9 for the hardware implementations of the kernels. Comparisons with the power aware compiler techniques (PACT) high-level synthesis compiler are used to show that software binaries can be used as intermediate representations from any high-level language and generate efficient hardware implementations.
Gaurav Mittal, David Zaretsky, Xiaoyong Tang, Prithviraj Banerjee
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Leakage power optimization with dual-Vth library in high-level synthesis
abstract
In this paper we address the problem of module selection during high-level synthesis. We present a heuristic algorithm for leakage power optimization based on the maximum weight independent set problem. A dual threshold voltage (Vth) technique is used to reduce leakage energy consumption in a data flow graph. Experiments are performed on a data-path dominated test suite of six benchmarks. Our approach achieves an average of 70.9% leakage power reduction, which is very close to the optimal results from an Integer Linear Programming approach.
Xiaoyong Tang, Hai Zhou 0001, Prithviraj Banerjee
DAC1
2004 Automatic translation of software binaries onto FPGAs
abstract
The introduction of advanced FPGA architectures, with built-in DSP support, has given DSP designers a new hardware alternative. By exploiting its inherent parallelism, it is expected that FPGAs can outperform DSP processors. This paper describes the process and considerations for automatically translating binaries targeted for general DSP processors into Register Transfer Level (RTL) VHDL or Verilog code to be mapped onto commercial FPGAs. The Texas Instruments C6000 DSP processor architecture is chosen as the DSP processor platform, and the Xilinx Virtex II as a target FPGA. Various optimizations are discussed, including data dependency analysis, procedure extraction, induction variable analysis, memory optimizations, and scheduling. Experimental results on resource usage and performance are shown for ten software binary benchmarks. Results show performance gains of 3-20X in the FPGA designs over that of the DSP processors in terms of reductions of execution cycles.
Gaurav Mittal, David Zaretsky, Xiaoyong Tang, Prithviraj Banerjee
DAC3
2004 Overview of the FREEDOM Compiler for Mapping DSP Software to FPGAs
abstract
Applications that require digital signal processing (DSP) functions are typically mapped onto general purpose DSP processors. With the introduction of advanced FPGA architectures with built-in DSP support, a new hardware alternative is available for DSP designers. By exploiting its inherent parallelism, it is expected that FPGAs can outperform DSP processors. However, the migration of assembly code to hardware is typically a very arduous process. This paper describes the process and considerations for automatically translating software assembly and binary codes targeted for general DSP processors into register transfer level (RTL) VHDL or Verilog code to be mapped onto commercial FPGAs. The Texas instruments C6000 DSP processor architecture has been used as the DSP processor platform, and the Xilinx Virtex II as the target FPGA. Various optimizations are discussed, including loop unrolling, induction variable analysis, memory and register optimizations, scheduling and resource binding. Experimental results on resource usage and performance are shown for ten software binary benchmarks in the signal processing and image processing domains. Results show performance gains of 3-20x in terms of reductions in execution cycles and 1.3-5x in terms of reductions in execution times for the FPGA designs over that of the DSP processors in terms of reductions in execution cycles.
David Zaretsky, Gaurav Mittal, Xiaoyong Tang, Prithviraj Banerjee
FCCM3
2004 High level area, delay and power estimation for FPGAs
abstract
This paper describes an approach for high-level estimation of area, delay and power for FPGA synthesis. This approach has been integrated within the PACT compiler framework which has an automated design space exploration pass that determines the effects of various compiler optimizations on the synthesized hardware. Such a pass needs early estimation of area, delay and power. Towards this end, we have developed area and delay models for various RTL level operators such as adders, multipliers, and logical operators, which are parameterized with the bit widths of the devices. We have also derived high-level equation based power macro-models which take into account input switching activities, input spatial correlation and input bit width. These models are derived by actual synthesis of the RTL operators using back-end logic synthesis and place-and-route tools. Experimental results show that these area, delay and power models are accurate and efficient.
Tianyi Jiang, Xiaoyong Tang, Prithviraj Banerjee
FPGA2
2004 Macro-models for high level area and power estimation on FPGAs
abstract
As more and more complex applications are implemented on FPGAs, high-level design tools are needed to reduce the design time. A good high-level synthesis tool usually has an automated design space exploration pass to determine the effects of various compiler optimizations on the area and power of the synthesized hardware. Such a pass needs early estimation of area and power. Towards this end, we have developed high-level equation based area and power macro-models for various RTL level operators such as adders, multipliers, and logical operators. The area model is parameterized with the bit width of the device and the power model takes into account input switching activity and input spatial correlation as well as input bit width. These models are derived by actual synthesis of these RTL operators using back-end logic synthesis and place-and-route tools. Compared with the other approaches, our method generated a uniform macro-model for each operator with fewer coefficients and sometimes lower degrees. It is also easier to analyze the power sensitivity to different parameters. Experimental results show that these area and power models are accurate and efficient.
Tianyi Jiang, Xiaoyong Tang, Prithviraj Banerjee
ACM Great Lakes Symposium on VLSI2
2004 Evaluation of scheduling and allocation algorithms while mapping assembly code onto FPGAs
abstract
Migration of software from older general purpose embedded processors onto newer mixed hardware/software Systems-On-Chip (SOC) platforms is becoming an increasingly important topic. Automatic translation of general purpose software binaries and assembly code onto hardware implementations using FPGAs require sophisticated scheduling and allocation algorithms to maximize the resource utilization of such hardware devices. This paper describes the effects of scheduling and chaining of node operations in a CDFG onto an FPGA. The effects of register allocation on scheduled nodes are also discussed. The Texas Instruments C6000 DSP processor architecture was chosen as the DSP processor platform and assembly code, and the Xilinx Virtex II XC2V250 was chosen as the target FPGA. Results are reported on ten benchmarks, which show that scheduling with chaining operations produces the best results on FPGAs, while the addition of register allocation in fact generates poorer designs in terms of area and frequency.
David Zaretsky, Gaurav Mittal, Xiaoyong Tang, Prithviraj Banerjee
ACM Great Lakes Symposium on VLSI3
2002 PACT HDL: a C compiler targeting ASICs and FPGAs with power and performance optimizations
abstract
Chip fabrication technology continues to plunge deeper into sub-micron levels requiring hardware designers to utilize ever-increasing amounts of logic and shorten design time. Toward that end, high-level languages such as C/C++ are becoming popular for hardware description and synthesis in order to more quickly leverage complex algorithms. Similarly, as logic density increases due to technology, power dissipation becomes a progressively more important metric of hardware design. PACT HDL, a C to HDL compiler, merges automated hardware synthesis of high-level algorithms with power and performance optimizations and targets arbitrary hardware architectures, particularly in a System on a Chip (SoC) setting that incorporates reprogrammable and application-specific hardware. PACT HDL is intended for applications well suited to custom hardware implementation such as image and signal processing codes. By making the compiler modular and flexible, optimizations may be executed in any order and at different levels in the compilation process. PACT HDL generates industry standard HDL codes, such as RTL Verilog and VHDL, which may be synthesized and profiled for power using commercial tools. This is the first paper on the PACT compiler project in a series. The compiler framework and introductory optimizations are presented. Later papers will focus on these and other optimizations in detail.
Alex K. Jones, Debabrata Bagchi, Satrajit Pal, Xiaoyong Tang, Alok N. Choudhary, Prithviraj Banerjee
CASES4