Anthony T. Chronopoulos

dblp:69/79 · also Anthony Theodore Chronopoulos · DBLP profile ↗
← Back
89ranked-venue papers
14as first author
23since 2021 · last 2026
0000-0002-0094-1017ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 10 first-author · 7 since 2021Computer networks · 19 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 BPP: Branch pipeline parallelism strategy for accelerating DNN training
Yonggan Cui, Chubo Liu, Anthony T. Chronopoulos, Alexandru Nicolau, Kenli Li 0001
Neurocomputing5
2026 A Resource Reuse Strategy for Large-Scale Matrix Operations in HLS-Based FPGA Design
abstract
Matrix operations (MOPs) are essential for various computational tasks, particularly in deep learning models, which have grown increasingly complex. As these models expand, their demand for computational resources increases significantly, making deployment on resource-limited hardware platforms, such as field-programmable gate arrays, more challenging, especially in balancing resource allocation and computation latency. Reuse-control techniques have been employed to optimize resource allocation by enabling multiple operations to share the same computational units, such as digital signal processors. Yet, this approach presents a tradeoff between resource utilization and latency. In this study, we tackle this challenge by thoroughly analyzing existing reuse-control mechanisms and introducing a novel integer linear programming (ILP)-based strategy. Our experimental results demonstrate that the proposed approach not only improves resource utilization for large-scale MOPs but also significantly reduces latency compared to existing methods. In the best case, on the ResNet model, our ILP-based method achieves up to$2.21\times$lower latency and$2.14\times$lower energy consumption per inference, demonstrating significantly improved performance and energy efficiency. In addition, our work provides a new optimization perspective for hardware design based on high-level synthesis.
Zhihang Lei, Chubo Liu, Baixuan Wu, Anthony T. Chronopoulos, Kenli Li 0001
IEEE Trans. Ind. Informatics5
2026 VM-ORAM: A Novel High-Performance ORAM Architecture for Efficient Data Integrity Verification in Industrial Cloud
abstract
With the rapid surge in industrial data, cloud computing has been integrated into Industrial Internet of Things (IIoT) systems to store, compute, and share massive data. In this process, privacy and data integrity are core concerns. Oblivious RAM (i.e., ORAM) is a technology widely applied to defend against cloud storage access pattern attacks. However, most existing ORAM systems do not consider the integration of data integrity verification technology. Although there are integrity verification systems integrated into conventional Path ORAM and Ring ORAM, they are not suitable for existing new ORAM systems. And the existing data integrity verification ORAM system still has the problem of excessive performance overhead. To address these challenges, this article proposes a novel high-performance data integrity ORAM system, VM-ORAM. Optimizes the ORAM integrity verification process by integrating dynamic scheduling and multipath eviction strategies, thereby minimizing performance loss. The comprehensive analysis and experimental results of this article show that VM-ORAM system not only defends against data tampering attacks, but also maintains high performance of the system.
Chuang Li 0004, Gang Liu 0038, Changyao Tan, Limei Liu, Wenhua Ye, Anthony T. Chronopoulos
IEEE Trans. Ind. Informatics7
2026 Adaptive Block-Wise Mapping With Intra-Block Resource Allocation for Multi-DNN Workloads on Heterogeneous Accelerator Systems
abstract
Deep neural networks (DNNs) dominate workloads on cloud and edge platforms. Meanwhile, the hardware platform towards the heterogeneous system with various accelerators. By mapping layers to their different preferred accelerators, the computation cost of each layer can be reduced. While mapping these layers on the same accelerator can reduce the inter-accelerator communication cost. These two costs are often competing and difficult to optimize simultaneously. Therefore, the core challenge in achieving efficient execution of DNN workloads on heterogeneous systems is: how to map layers to achieve the best trade-off between computation and communication costs. Existing works group layers into blocks and perform blockwise mapping to reduce inter-layer communication within blocks. However, when grouping layers, they typically rely on modelagnostic rules, which fail to hide critical inter-layer communication within blocks for diverse DNNs. Moreover, after block mapping, the lack of intra-block resource allocation further increases computation cost of block. In this paper, we proposeGHCoM, a novel block-wise mapping framework for exploring the effective cost trade-offs.GHCoMemploys an adaptive grouping strategy to guide layer grouping based on the topology of DNNs and dynamically adjust the grouping according to the trade-off target. Furthermore,GHCoMconsiders the fine-grained allocation of computation (i.e., processing elements) and communication (i.e., on-chip bandwidth) resources within each block to mitigate interlayer resource contention. To jointly optimize layer grouping, block-wise mapping and intra-block resource allocation,GHCoMleverages a two-level genetic algorithm (GA) with tailored encodings and operators that capture the interdependence across the entire design space. Experiments across various workloads and system configurations show thatGHCoMconsistently outperforms state-of-the-art baselines, achieving 1.08× to 4.79× speedup in execution latency and reducing energy consumption by 1.83% to 87.71%.
Zhenyu Nie, Haotian Wang 0006, Anthony T. Chronopoulos, Zhuo Tang, Kenli Li 0001, Chubo Liu
IEEE Trans. Parallel Distributed Syst.3
2025 An Input-Aware Sparse Tensor Compiler Empowered by Vectorized Acceleration
abstract
Sparsity is widely prevalent in real-world applications, yet existing compiler optimizations and code generation techniques for sparse computations remain underdeveloped. Sparse matrix-matrix multiplication (SpMM) is a representative operator in sparse computations, whose performance is often limited by the design of sparse formats and the extent of hardware architecture optimization. Most existing solutions achieve highperformance SpMM through two approaches: (1) meticulously designed kernels and specialized sparse formats, which require extensive manual effort, or (2) tensor compilers that support code generation, though these typically offer limited support for sparse patterns, making it challenging to adapt to complex sparsity patterns in practical applications. This paper presents SpMMTC, an input-aware sparse tensor compiler. Given a sparse matrix as input, SpMMTC analyzes its non-zero distribution and generates a vectorized kernel optimized for SpMM on the specific matrix. We evaluated SpMMTC on various workloads. It achieves speedups of 1.21 x to 2.97 x over state-of-the-art methods such as TACO, TVM, and ASpT on different multi-core processors. It also provides a speedup of up to $\mathbf{1. 5 2 x}$ for sparse MobileNetV1 inference on the edge device.
Xianhao He, Haotian Wang 0006, Jiapeng Zhang 0001, Wangdong Yang, Anthony T. Chronopoulos, Kenli Li 0001
DAC5
2025 STREAM: Spatiotemporal Similarity-based Efficient Approximate Median with Tunable Granularity
abstract
The median (MED) is a crucial statistic for measuring the central tendency. However, exact MED computation remains costly, with even state-of-the-art (SOTA) algorithms failing to meet (near) real-time processing demands. While approximate MED algorithm has arisen as a promising candidate, existing approaches ignore the potential opportunity of spatiotemporal similarity within the application and fail to provide applicationspecific trade-offs between execution time and accuracy. Our goal is to design an enhanced approximate MED algorithm STREAM, which is capable of exploiting the spatiotemporal similarity to achieve bucket reuse and establish a tunable-grained bucket mechanism to meet the accuracy of application-specific requirements. Experimental results show that while maintaining nearly identical accuracy, STREAM outperforms the SOTA approximate methods DDSketch (up to $10 \times 4.7 \times$ on average) and KLL (up to $71.2 \times 10.1 \times$ on average).
Fenfang Li, Huizhang Luo, Weichen Liu 0001, Anthony T. Chronopoulos, Kenli Li 0001, Chubo Liu
DAC4
2025 TIPS: A text interaction evaluation metric for learning model interpretation
Zhenyu Nie, Tao Wang 0016, Anthony T. Chronopoulos, Razvan Andonie, Amirhosein Mosavi
Expert Syst. Appl.4
2025 Spectral Efficiency Analysis for Cell-Free Massive MIMO Systems With Low-Resolution ADCs Under Imperfect CSI
Weiyi Ni, Yiling He, Hailin Xiao, Anthony T. Chronopoulos, Petros A. Ioannou
IEEE Internet Things J.4
2025 User Association and Small Base Station Configuration for Energy-Efficiency Maximization in Hybrid-Energy Heterogeneous Cellular Networks
abstract
Dense deployment of small base stations (SBSs) within the coverage of macro base station (MBS) has been spotlighted as a promising solution to conserve grid energy in hybrid-energy heterogeneous cellular networks (HCNs), which caters to the rapidly increasing demand of mobile user (MU). However, MUs in the ultradense cellular network experience handover events more frequently than in conventional networks, which results in increased service interruption time and performance degradation due to blockages. In addition, blindly increasing the number of SBSs not only results in an increased cost for network operators but also brings serious system energy consumption and interference. In this article, we propose a joint user association and SBSs configuration scheme for maximizing energy efficiency (EE) in hybrid-energy HCNs. Specially, an association model with dual connectivity for MUs where they are connected simultaneously with SBSs and MBSs is first proposed to reduce frequent handover, which is also presented to preferentially select SBSs that can provide data transmission for MUs under the user association constraints according to the maximum system EE. And then the ratio between the number of SBSs and the number of MUs for SBSs configuration is analyzed to reduce interference and energy consumption under the tidal effect of HCNs. Furthermore, the EE utility function of joint user association and SBSs configuration is extended, and the Dinkelbach and Lagrangian algorithms are jointly optimized to solve the EE utility function. Finally, numerical simulation results are provided to demonstrate the feasibility of the proposed scheme. It is shown that the proposed scheme outperforms other schemes and can also maximize the EE in hybrid-energy HCNs.
Weiyi Ni, Hailin Xiao, Anthony T. Chronopoulos, Zhongshan Zhang
IEEE Internet Things J.4
2024 Hierarchical Explanations for Text Classification Models: Fast and Effective
abstract
Generating explanations for deep neural networks (DNNs) can make them more trustworthy in real-world applications. For a text classification task, existing methods visualize the contributions of words or word interactions layer by layer in a traversal manner, to assist users in understanding the decision-making of models. However, all these methods only focus on the explanation performance while ignoring inefficiencies in explaining due to the traversal manner. This means that the explanation is not available to users in a timely manner, so users may no longer use it due to the big time cost. To overcome this problem, we propose HETSG, an interaction-based method for explaining text classification models quickly and faithfully, by a simple and effective two-step building strategy. Such a strategy captures the important interaction by first determining the important position and then confirming the direction of interaction, without iterating over all word interactions. We also provide a novel metric to accurately evaluate the performance of each interpretation method. The proposed method is compared with baseline methods (baselines) on six text classification datasets to explain three natural language processing (NLP) models. Experimental results show that our method outperforms all baselines with higher efficiency and it is also competitive in performance.
Zhenyu Nie, Huizhang Luo, Anthony T. Chronopoulos
ICDM5
2024 NGLIC: A Nonaligned-Row Legalization Approach for 3-D Interdie Connection
abstract
3-D placement is an important stage in 3-D physical synthesis. In addition to the need to place the standard cells or Macros inside the die, the placement of interdie connections also needs to be considered. As the density of the interdie connections gradually increases, their placement becomes more critical. However, unlike standard cells, the interdie connection does not need to be aligned into rows, which leads to a larger legalization solution space. Legalization aims at minimizing the total and maximum displacement to maintain the quality of global placement on the premise of no violation of the physical circuit constraint. In this article, we propose a two-stage legalization approach (NGLIC) for nonaligned-row interdie connections. In the initial stage, a single-row height legalization algorithm is used for dense placement. Afterward, the unequal multirow height legalization (UML) is designed for sparse placement in the post-optimization stage. A pruning scheme is adopted to identify and eliminate redundant computations. The performance, effectiveness, and configuration are analyzed empirically based on ICCAD 2022 benchmarks. Compared to the state-of-the-art multirow or single-row height legalization, our approach outperforms well-known algorithms, such as multirow global legalization (MGL) and Abacus by at least 24% averaged total displacement, 3% averaged maximum displacement, and 31% averaged HPWL Growth. Our case study also illustrates the effectiveness of UML and NGLIC. In addition, the effectiveness of pruning is also validated by experiments that show savings of 33% redundant computations.
Yunchuan Qin, Fan Wu 0016, Anthony T. Chronopoulos, Alexandru Nicolau, Kenli Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 RFL-APIA: A Comprehensive Framework for Mitigating Poisoning Attacks and Promoting Model Aggregation in IIoT Federated Learning
abstract
With the development of industrial Internet of Things (IIoT), federated learning (FL) is important for protecting sensitive data from various Internet of Things devices (i.e., clients in FL). Despite FL's privacy benefits, attackers (e.g., untrusted clients) can still compromise the performance of the global model through model poisoning attacks. Unfortunately, two key challenges hinder effective detection and impact the performance of the global model in FL: first, accurately identifying malicious models to defend against attacks, and second, efficiently aggregating local models after detecting malicious clients. To address these challenges, we propose an improved FL system based on fuzzy rules, termed RFL-APIA. Compared to the conventional FL system, we have designed two novel components, federated learning generalized depth detection (FedGDD) and Fedsv-Weighted, to enhance performance and mitigate model poisoning attacks. Specifically, FedGDD introduces variance reduction by examining the relationship between local and global model gradients, thereby mitigating interference in nonindependent and identical distributed settings. It further implements an adaptive penalty factor-based scoring system, leveraging variations in local model updates for precise identification and mitigation of attacks. Based on FedGDD's output, the Fedsv-Weighted mechanism dynamically updates the global model's aggregation weights by considering local models' contributions, thus improving model aggregation. Extensive experiments demonstrate that RFL-APIA effectively prevents model poisoning attacks during training, ensuring model security, and guaranteeing a certain level of accuracy and convergence for the global model.
Chuang Li 0004, Aoli He, Gang Liu 0038, Yanhua Wen, Anthony T. Chronopoulos, Aristotelis Giannakos
IEEE Trans. Ind. Informatics5
2024 A Bilateral Game Approach for Task Outsourcing in Multi-Access Edge Computing
abstract
Multi-access edge computing (MEC) is a promising architecture to provide low-latency applications for future Internet of Things (IoT)-based network systems. Together with the increasing scholarly attention on task offloading, the problem of servers’ resource allocation has been widely studied. The limited computational resources of edge servers (ESs) cannot meet the different demands of terminal entities (TEs). This makes it a challenge to efficiently schedule computational tasks on ESs. In this paper, we consider a MEC resource transaction market with multiple ESs and multiple TEs, which are interdependent and mutually influence each other. This paper aims to investigate the dynamic tasks allocation problem between TEs and ESs and to meet the optimal benefits for both parties in MEC system. However, this many-to-many interaction requires resolving several problems, including task allocation, TEs’ selection on ESs and conflicting interests of both parties. A bilateral game framework is applied to tackle the tasks allocation problem by modeling the problem as two noncooperative games: the supplier and customer side games. The existence and uniqueness of the Nash equilibrium in the aforementioned games are proved. Adistributedtaskoutsourcingalgorithm (DTOA) is designed to determine the equilibrium. Our simulation results have demonstrated the superior performance of DTOA in increasing the ESs’ profit and TEs’ payoffs, as well as flattening the peak and off-peak loads.
Zhao Tong 0001, Dan He 0008, Anthony T. Chronopoulos, Schahram Dustdar
IEEE Trans. Netw. Serv. Manag.5
2023 Latency and Energy-Aware Load Balancing in Cloud Data Centers: A Bargaining Game Based Approach
abstract
With the rapid surge in cloud services, cloud load balancing has become a paramount research issue. The major part of a cloud computing system's operational costs is attributed to energy consumption. Therefore, to provide better QoS, considering the energy minimization factor in load balancing is essential. This paper addresses the latency and energy-aware load balancing problem in a cloud computing system. Specifically, two fundamental performance criteria–response time and energy–for the load balancing problem are considered. To solve this problem, first, the load balancing problem is formulated as an optimization problem. Then it is modeled as a cooperative game so that the solution of the game, called the Nash bargaining solution (NBS), can simultaneously optimize both criteria. The existence and computation of NBS are analyzed theoretically, and an efficient algorithm, called${{\sf L}}$atency and${{\sf E}}$nergy a${{\sf W}}$are load balancIng${{\sf S}}$cheme (${{\sf LEWIS}}$), is proposed to compute the NBS. Further, to assess the efficacy of${{\sf LEWIS}}$, it is compared with three other approaches, i.e.,${\mathsf {Coop\_{RT}}}$,${\mathsf {Coop\_{EN}}}$, and${\mathsf {NCG}}$, on problem instances of various settings. The experimental results show that${{\sf LEWIS}}$not only provides less response time while consuming less energy but also gauntness fairness to the end-users.
Avadh Kishor, Rajdeep Niyogi, Anthony T. Chronopoulos, Albert Y. Zomaya
IEEE Trans. Cloud Comput.3
2023 A Distributed Integrated Feature Selection Scheme for Column Subset Selection
abstract
Most of the existing distributed feature selection schemes neglect how good the subsets are that are mapped to the computational nodes, which causes a waste of time and hardware resources. A distributed integrated feature selection scheme (DIFS) with Subset Quality Evaluation (SQE) is proposed. SQE studies the relevance between the quality of a subset and the number of selected features from this subset, which helps shorten the feature selection time efficiently. We have given the implementation of our scheme for the Column Subset Selection (CSS) problem. We integrate a CSS algorithm in DIFS and information entropy as the SQE metric. We prove that the speedup of DIFS can reach m^3 compared to the centralized algorithm in ideal situations where m is the number of computational nodes, and give a well bounded approximation guarantee of the solution for CSS problem. Extensive experiments on eight data sets are used to verify the performance of scheme. Experiments results demonstrate the effectiveness of SQE and the impressive speedup DIFS can achieve. Although there is a slight increase of the reconstruction error value in some situations. Additional experiments of classification tasks reveal that the performance of DIFS is better than existing state-of-the-art distributed algorithms.
PengCheng Wei, Anthony T. Chronopoulos, Anne C. Elster
IEEE Trans. Knowl. Data Eng.3
2023 Optimal Trading Mechanism Based on Differential Privacy Protection and Stackelberg Game in Big Data Market
abstract
Big data has become a fundamental resource and a commodity in economic activities, thus, it is necessary to build a market model capable of supporting efficient data trading. However, two major challenges remain. First, researches have considered constructing data trading mechanisms, while few of them are based on the method of measuring data value in multiple dimensions. Second, a data market involved an intermediary trading platform (i.e., a third party) which is honest but curious, results may be obtained due the the leakage of private information. In this article, we design TM-OUE, a data trading mechanism based on Optimized Unary Encoding that enables reasonable trading mechanism and protects the privacy of data trading. First of all, we combine qualitative and quantitative methods to measure the value of data in multiple dimensions and formulate data trading model between the data provider and data users. Then, we utilize an Optimized Unary Encoding (OUE) protocol to protect the privacy of the data trading mechanism. Based on the above steps, we develop a two-stage single leader multi-follower Stackelberg game to jointly maximize profits of the data provider and data users. Experimental results demonstrate that TM-OUE can offer appropriate price for data and maximize benefits both for data providers and data users, which guarantees fair data trades while protecting privacy.
Chuang Li 0004, Aoli He, Yanhua Wen, Gang Liu 0038, Anthony T. Chronopoulos
IEEE Trans. Serv. Comput.5
2022 An extended attention mechanism for scene text recognition
Zhenyu Nie, Anthony T. Chronopoulos
Expert Syst. Appl.4
2022 A method for reducing cloud service request peaks based on game theory
Anthony T. Chronopoulos, Jiuchuan Jiang
J. Parallel Distributed Comput.3
2022 A Many-to-Many Demand and Response Hybrid Game Method for Cloud Environments
abstract
In this article, we design a service mechanism for profits optimization between multiple cloud providers and multiple cloud customers (many-to-many). We explore this problem from the perspective of game theory but take a different approach compared with existing cloud resource pricing game methods. First, we regard the relationships among multiple cloud customers as an evolutionary game, and formulate the competitions among the multiple cloud providers as a noncooperative game. Eventually, we form a hybrid game model in which the strategy of each customer and each cloud provider is affected not only by the other side but also by customers or cloud providers other than themselves. Second, based on the hybrid game model, we simulate the bargaining process between cloud providers and customers by controlling supply and demand allocation, and try to ultimately achieve a balanced supply and demand state, i.e., a win-win situation. For each cloud customer and provider, we design a utility function. A customer’s utility involves net profits and the cloud providers’ bidding strategies, and a cloud provider’s utility involves net profits and the cloud customers’ demand strategies. Both sides attempt to maximize their own profits under the influences of each other. We prove that our proposed strategies enable each of the two games to converge to their own equilibrium. Finally, the strategies of cloud customers and providers can be implemented through an iterative proximal algorithm ($\mathcal {IPA}$) and a distributed iterative algorithm ($\mathcal {DIA}$). The experimental results validate our methods and show that the proposed method can benefit both multiple cloud providers and customers.
Gang Liu 0038, Anthony T. Chronopoulos, Chubo Liu, Zhuo Tang
IEEE Trans. Cloud Comput.3
2021 Joint Clustering and Blockchain for Real-Time Information Security Transmission at the Crossroads in C-V2X Networks
abstract
The cellular vehicle-to-everything (C-V2X) networks support diverse kinds of services, such as traffic management, road safety, and sharing data. However, the safety issues cannot be ignored in the process of information transmission. In this article, a joint clustering and blockchain scheme is proposed for real-time information security transmission to prevent some vehicles from sending malicious messages to disrupt the traffic order at the crossroads in C-V2X networks. In this scheme, the dynamic stability of the cluster is maintained by updating the trust value of the vehicle nodes, which can improve the real time and accuracy of the information transmission. The modified Webster algorithm is presented to divert the traffic flow so as to reduce the traffic jams at the crossroads. Meanwhile, the blockchain technology is utilized to establish a vehicle trust management mechanism in C-V2X, which can avoid malicious tampering of vehicle information during information sharing and ensure the safety of vehicle information communication. The simulation results of the Veins simulation platform are provided to demonstrate the effectiveness of the proposed algorithm and verify that the proposed scheme can guarantee the security of real-time information transmission.
Hailin Xiao, Anthony T. Chronopoulos, Zhongshan Zhang
IEEE Internet Things J.4
2021 Guest Editorial Special Issue on Smart IoT System: Opportunities by Linking Cloud, Edge, and AI
Wangdong Yang, Laurence T. Yang, Anthony T. Chronopoulos
IEEE Internet Things J.3
2021 A review on deep learning approaches in healthcare systems: Taxonomies, challenges, and open issues
Shahab B. Band, Mahdis Fathi, Abdollah Dehzangi, Anthony T. Chronopoulos, Hamid Alinejad-Rokny
J. Biomed. Informatics4
2021 Connectivity probability analysis for freeway vehicle scenarios in vehicular networks
Hailin Xiao, Anthony T. Chronopoulos
Wirel. Networks4
2020 Computational intelligence intrusion detection techniques in mobile cloud computing environments: Review, taxonomy, and open research issues
Shahab B. Band, Mahdis Fathi, Anthony T. Chronopoulos, Antonio Montieri, Fabio Palumbo, Antonio Pescapè
J. Inf. Secur. Appl.3
2020 Resource Management for Multi-User-Centric V2X Communication in Dynamic Virtual-Cell-Based Ultra-Dense Networks
abstract
The technology of static user-centric virtual cell (VC) has been designed in the fifth-generation (5G) ultra-dense networks (UDNs) for alleviating both frequent handover and inter-cell interference. However, to provide the user-centric services, the fairness of resource management must be involved when the common vehicular-to-X (V2X) messages are multicast to the vehicle groups. In this paper, a dynamic user-centric virtual cell (DUVC) scheme is proposed for updating adaptively the VC through the mobile tracking of the vehicles. Furthermore, an approximation algorithm is proposed for solving the max-min-fair problem of resource management in V2X communication in order to better support V2X services throughout the service VC. Numerical results are provided for demonstrating that the proposed DUVC scheme is suitable for implementing in the V2X communication. Finally, the proposed algorithm is shown to outperform three other existing algorithms in terms of fairness of resource management.
Hailin Xiao, Anthony T. Chronopoulos, Zhongshan Zhang, Shan Ouyang 0001
IEEE Trans. Commun.3
2020 Game theory-based optimization of distributed idle computing resources in cloud environments
Gang Liu 0038, Guanghua Tan, Kenli Li 0001, Anthony T. Chronopoulos
Theor. Comput. Sci.5
2020 Power Allocation With Energy Efficiency Optimization in Cellular D2D-Based V2X Communication Network
abstract
In vehicle-to-everything (V2X) communication network, cellular device-to-device (D2D) communication can not only improve data rate and spectral utilization but also reduce the traffic load and power consumption. However, cellular D2D-based V2X technology has also a potential deficiency to meet various requirements of V2X communication, particularly in energy efficiency (EE). Power allocation provides an important approach to optimize the EE. In this paper, a new approach for power allocation with EE optimization (EEO) is proposed in cellular D2D-based V2X communication network. The mathematical framework of the new approach is formulated and proved and an algorithm is also proposed. Numerical simulation results are provided to demonstrate the feasibility of the proposed algorithm and the superiority over existing well-known algorithms.
Hailin Xiao, Anthony T. Chronopoulos
IEEE Trans. Intell. Transp. Syst.3
2019 A new malware detection system using a high performance-ELM method
abstract
A vital element of a cyberspace infrastructure is cybersecurity. Many protocols proposed for security issues, which leads to anomalies that affect the related infrastructure of cyberspace. Machine learning (ML) methods used to mitigate anomalies behavior in mobile devices. This paper aims to apply a High-Performance Extreme Learning Machine (HP-ELM) to detect possible anomalies in two malware datasets. Two widely used datasets (the CTU-13 and Malware) are used to test the effectiveness of HP-ELM. Extensive comparisons are carried out in order to validate the effectiveness of the HP-ELM learning method. The experiment results demonstrate that the HP-ELM was the highest accuracy of performance of0.9592 for the top 3 features with one activation function.
Shahab B. Band, Anthony T. Chronopoulos
IDEAS2
2019 Optimising infrastructure as a service provider revenue through customer satisfaction and efficient resource provisioning in cloud computing
abstract
With limited resources, it is quite challenging for cloud providers to meet dynamic and massive customers' demands. Higher utilisation or refusing any service level agreement (SLA) may lead to penalties which play a crucial role in the cloud business. Overutilisation of resources, instead of maximizing the revenue, may lead to a decrease in revenue due to the SLA violations. Various studies have been conducted to investigate these issues; however, there is still room for improvement. In this study, the authors proposed a model to address the resource scalability and SLA violation issues by hiring external resources at low prices. However, in contrast to a federated cloud, the proposed model allows a provider to hire resources from any external provider with flexible terms and price. They designed algorithms to optimise providers' revenue by taking into account different parameters, including resource utilization, customer satisfaction, SLA violation, and prices. Simulation result shows that the proposed model is efficient in handling massive demands, and improves revenue generation and customer satisfaction. Offering joint pricing on customers' choice and outsourcing the extra workload to external resources leads to revenue maximization. Hiring external resources earns external revenue as well as it maximizes the total revenue.
Afzal Badshah, Anwer Ghani, Shahab B. Band, Anthony T. Chronopoulos
IET Commun.4
2019 Joint Clustering and Power Allocation for the Cross Roads Congestion Scenarios in Cooperative Vehicular Networks
abstract
Both clustering and cluster-head vehicles (CHVs) cooperative communication have been employed for reducing traffic congestion to improve road traffic efficiency in cooperative vehicular networks. In this paper, an iterative optimization k-means clustering algorithm with lower complexity than previous algorithms is proposed. It can automatically generate multiple clusters according to the number of vehicles and quickly find the CHVs by avoiding delays caused by complex calculations. Moreover, a new optimization power allocation strategy with bidirectional incremental hybrid decode-amplify-forward protocol focusing on reducing the total power consumption of CHVs is proposed. This strategy can set the signal to noise ratio threshold as the critical point for selecting dynamically the bidirectional incremental amplify-and-forward or decoding and forwarding protocol with a lower outage probability to transmit information. Thus, the proposed power allocation strategy is capable of minimizing the total transmission power while ensuring a lower outage probability than previous approaches. Finally, the numerical results are provided for corroborating the theoretical results and demonstrate the efficiency of the proposed approaches. Note that through the numerical simulations we can find the critical point of the outage probability for the aforementioned protocols under different relay locations. This assists vehicles to select “relays” with the optimal cooperative position for vehicular cooperative communication system.
Hailin Xiao, Anthony T. Chronopoulos, Zhongshan Zhang, Shan Ouyang 0001
IEEE Trans. Intell. Transp. Syst.4
2018 Data Volume Based Data Gathering in WSNs using Mobile Data Collector
abstract
Data collection and transmission are the fundamental operations of WSNs. The performance of WSNs relies upon these essential tasks because data gathering directly affects the efficiency and lifetime of WSNs. This paper presents a data volume based data collection technique using Mobile Data Collector (MDC). In this technique, the MDC uses data volume information to plan visits to the nodes. The MDC visits only those nodes which have generated data while the rests of the nodes are ignored. This scheme is validated with the help of simulations, and the results are compared with existing renowned techniques. The results show that the proposed scheme is energy efficient.
Syed Muhammad Abrar Akber, Imran Ali Khan, Syed Shah Muhammad, Syed Muhammad Mohsin, Iftikhar Ahmed Khan, Shahab B. Band, Anthony T. Chronopoulos
IDEAS7
2018 Computational intelligence approaches for classification of medical data: State-of-the-art, future challenges and research directions
Ali Kalantari, Amirrudin Kamsin, Shahab B. Band, Abdullah Gani, Hamid Alinejad-Rokny, Anthony T. Chronopoulos
Neurocomputing6
2018 Performance Analysis of Multi-Source Multi-Destination Cooperative Vehicular Networks With the Hybrid Decode-Amplify-Forward Cooperative Relaying Protocol
abstract
This paper provides symbol-error-rate (SER) performance analysis and minimum power allocation for multi-source multi-destination cooperative vehicular networks using the hybrid decode-amplify-forward (HDAF) cooperative relaying protocol. Previous studies of power allocation minimize the outage probability subject to a total power constraint. Our approach aims to minimize the power allocation in order to maintain the SER below a specific threshold and thus it achieves lower power consumption. Numerical tests show that HDAF has significantly reduced SER compared with the forward strategies of amplify-and-forward (AF) and decode-and-forward (DF). Furthermore, the power consumption in our proposed approach is much less than that in AF and DF.
Hailin Xiao, Zhongshan Zhang, Anthony T. Chronopoulos
IEEE Trans. Intell. Transp. Syst.3
2017 Introducing ToPe-FFT: An OpenCL-based FFT library targeting GPUs
abstract
Summary In this paper, we present our implementation of the fast Fourier transforms on graphic processing unit (GPU) using OpenCL. This implementation of the FFT (ToPe‐FFT) is based on the Cooley‐Tukey set of algorithms with support for 1D and higher dimensional transforms using different radices. Factorization for mix‐radices enables our code to target FFTs of near arbitrary length. In systems with multiple graphic cards (GPUs), the library automatically balances the FFT computation thus achieving maximum resource utilization and higher speedup. Based on profiling and micro‐benchmarking of ToPe‐FFT, it is observed that the average speedup of our library for different sizes is 48× faster than the single CPU‐based code using FFTW and 3× faster than NVIDIA's GPU‐based cuFFT library.
Bilal Jan, Fiaz Gul Khan, Bartolomeo Montrucchio, Anthony T. Chronopoulos, Shahab B. Band, Abdul Nasir Khan
Concurr. Comput. Pract. Exp.4
2017 An optimized magnetostatic field solver on GPU using open computing language
abstract
Summary Recent graphic processing units (GPUs) have remarkable raw computing power, which can be used for very computationally challenging problems. Like in micromagnetic simulations, where the magnetostatic field computation to analyze the magnetic behavior at very small time and space scale demands a huge computation time. This paper presents a multidimensional FFT‐based parallel implementation of a magnetostatic field computation on GPUs. We have developed a specialized 3D FFT library for magnetostatic field calculation on GPUs. This made it possible to fully exploit the symmetries inherent in the field calculation and other optimizations specific to the GPUs architecture. We have compared our results with the widely used CPU‐based parallel OOMMF program and with an equivalent serial implementation on CPU. The results have shown a speedup of up to 95x and 8.7x for single and 66x and 4.6x for double precision floating point accuracy against equivalent serial implementation and OOMMF, respectively.
Fiaz Gul Khan, Bartolomeo Montrucchio, Bilal Jan, Abdul Nasir Khan, Waqas Jadoon, Shahab B. Band, Anthony T. Chronopoulos, Iftikhar Ahmed Khan
Concurr. Comput. Pract. Exp.7
2017 Load balancing in grid computing: Taxonomy, trends and opportunities
Sumair Khan, Babar Nazir, Iftikhar Ahmed Khan, Shahab B. Band, Anthony T. Chronopoulos
J. Netw. Comput. Appl.5
2017 Non-cooperative power and latency aware load balancing in distributed data centers
Rakesh Tripathi, S. Vignesh, Venkatesh Tamarapalli, Anthony T. Chronopoulos, Hajar Siar
J. Parallel Distributed Comput.4
2015 Data placement using Dewey Encoding in a hierarchical data grid
Amir Masoud Rahmani, Zeinab Fadaie, Anthony T. Chronopoulos
J. Netw. Comput. Appl.3
2014 A Resilient Hierarchical Distributed Loop Self-Scheduling Scheme for Cloud Systems
abstract
In heterogeneous distributed cloud systems, load balance, communication and synchronization overhead must be taken considered. A hierarchical distributed loop self-scheduling scheme is effective and efficient for scientific loop applications. In this paper, we propose a resilient hierarchical distributed loop self-scheduling algorithm suitable for heterogeneous cloud systems. This algorithm is intended to enable the algorithm to continue to work in the event that some virtual machines (VMs) are too slow or cease to respond. We tested our algorithm in a heterogeneous cloud system. The results show that our algorithm can achieve normal operation and good performance.
Yiming Han, Anthony T. Chronopoulos
NCA2
2014 Cost minimization in utility computing systems
abstract
SUMMARY Utility computing is a form of computer service whereby the company providing the service charges the users for using the system resources. In this paper, we present system‐optimal and user‐optimal price‐based job allocation schemes for utility computing systems whose objective is to minimize the cost for the users. The system‐optimal scheme provides an allocation of jobs to the computing resources that minimizes the overall cost for executing all the jobs in the system. The user‐optimal scheme provides an allocation that minimizes the cost for individual users in the system for providing fairness. The system‐optimal scheme is formulated as a constraint minimization problem, and the user‐optimal scheme is formulated as a non‐cooperative game. The prices charged by the computing resource owners for executing the users jobs are obtained using a pricing model based on a non‐cooperative bargaining game theory framework. The performance of the studied job allocation schemes is evaluated using simulations with various system loads and configurations. Copyright © 2012 John Wiley & Sons, Ltd.
Satish Penmatsa, Anthony T. Chronopoulos
Concurr. Comput. Pract. Exp.2
2013 A Hierarchical Distributed Loop Self-Scheduling Scheme for Cloud Systems
abstract
Cloud systems have demonstrated the powerful computation and storage capability in many scientific applications. In this paper, we propose a hierarchical distributed loop self-scheduling scheme to achieve good load balancing by applying weighted self-scheduling scheme on a heterogeneous cloud system. This scheme also considers the distribution of the output data, which can help reduce communication overhead. We evaluated the scheme with two scientific applications: Matrix Multiplication and Quick Sort. The results shows that our schemes achieve better load balancing and better overall performance than standard loop self-scheduling scheme.
Yiming Han, Anthony T. Chronopoulos
NCA2
2012 Two-Dimensional Dynamic Loop Scheduling Schemes for Computer Clusters
abstract
Efficient scheduling of parallel loops in a network of computers can significantly reduce the total execution time of complex scientific applications. In this paper, we compare the performance of two-dimensional dynamic loop scheduling schemes for computer clusters with that of one-dimensional loop scheduling schemes. The loop scheduling schemes are implemented using the Message Passing Interface on a cluster of processors. Experimental results show that the two-dimensional scheduling schemes were found to significantly reduce the total execution time of tasks over the one-dimensional schemes. In addition, the two-dimensional schemes present a more balanced load distribution of the workload among the computers in the cluster.
Anthony T. Chronopoulos, Satish Penmatsa, Naveen Jayakumar, Eric Ogharandukun
NCA1
2012 Deterministic model for Acute Myelogenous Leukemia classification
abstract
Leukemia is a type of cancer that affects the blood and the bone marrow. Manual data analysis is time consuming and not accurate. Attempts to build partial/full automated systems based on segmentation and classification of cells are present in literature, but they are still in prototype stage. Most of the existing automatic systems extract features of the sub-images instead of the complete blood smear. [29]. The main objective of this paper is to a) demonstrate that the classification of peripheral blood smear images containing multiple nuclei can be fully automated, b) to validate the segmented images using hold-out cross validation method. The method has been evaluated using a set of 50 images (with 25 abnormal samples and 25 normal samples) obtained from American Society of Hematology [22]. The computer simulations show that the proposed system robustly segments and classifies Acute Myelogenous Leukemia based on complete microscopic blood images. 93.5% of the cases were correctly classified by the program, suggesting that the method yields good results in terms of classification of leukemia. The developed system can be used as ancillary/backup service to the physician.
Monica Madhukar, Sos S. Agaian, Anthony T. Chronopoulos
SMC3
2012 Towards the optimal synchronization granularity for dynamic scheduling of pipelined computations on heterogeneous computing systems
abstract
SUMMARY Loops are the richest source of parallelism in scientific applications. A large number of loop scheduling schemes have therefore been devised for loops with and without data dependencies (modeled as dependence distance vectors) on heterogeneous clusters. The loops with data dependencies require synchronization via cross‐node communication. Synchronization requires fine‐tuning to overcome the communication overhead and to yield the best possible overall performance. In this paper, a theoretical model is presented to determine the granularity of synchronization that minimizes the parallel execution time of loops with data dependencies when these are parallelized on heterogeneous systems using dynamic self‐scheduling algorithms. New formulas are proposed for estimating the total number of scheduling steps when a threshold for the minimum work assigned to a processor is assumed. The proposed model uses these formulas to determine the synchronization granularity that minimizes the estimated parallel execution time. The accuracy of the proposed model is verified and validated via extensive experiments on a heterogeneous computing system. The results show that the theoretically optimal synchronization granularity, as determined by the proposed model, is very close to the experimentally observed optimal synchronization granularity, with no deviation in the best case, and within 38.4% in the worst case. Copyright © 2012 John Wiley & Sons, Ltd.
Ioannis Riakiotakis, Florina M. Ciorba, Theodore Andronikos, George K. Papakonstantinou, Anthony T. Chronopoulos
Concurr. Comput. Pract. Exp.5
2011 Game-theoretic static load balancing for distributed systems
Satish Penmatsa, Anthony T. Chronopoulos
J. Parallel Distributed Comput.2
2010 Studying the impact of synchronization frequency on scheduling tasks with dependencies in heterogeneous systems
Theodore Andronikos, Florina M. Ciorba, Ioannis Riakiotakis, George K. Papakonstantinou, Anthony T. Chronopoulos
Perform. Evaluation5
2010 A game-theoretic approach to joint rate and power control for uplink CDMA communications
abstract
Next generation wireless systems will be required to support heterogeneous services with different transmission rates that include real time multimedia transmissions, as well as non-real time data transmissions. In order to provide such flexible transmission rates, efficient use of system resources in next generation systems will require control of both data transmission rate and power for mobile terminals. In this paper we formulate the problem of joint transmission rate and power control for the uplink of a single cell CDMA system as a non-cooperative game. We assume that the utility function depends on both transmission rates and powers and show the existence of Nash equilibrium in the non-cooperative joint transmission rate and power control game (NRPG). We include numerical results obtained from simulations that compare the proposed algorithm with a similar one which is also based on game theory and it also updates the transmission rates and powers simultaneously in a single step.
Madhusudhan R. Musku, Anthony T. Chronopoulos, Dimitrie C. Popescu, Anton Stefanescu
IEEE Trans. Commun.2
2009 Comparison of Price-Based Static and Dynamic Job Allocation Schemes for Grid Computing Systems
abstract
Grid computing systems are a cost-effective alternative to traditional high-performance computing systems. However, the computing resources of a grid are usually far apart and connected by Wide Area Networks resulting in considerable communication delays. Hence, efficient allocation of jobs to computing resources for load balancing is essential in these grid systems. In this paper, two price-based dynamic job allocation schemes for computational grids are proposed whose objective is to minimize the execution cost for the grid users' jobs. One scheme tries to provide a system-optimal solution so that the expected price for the execution of all the jobs in the grid system is minimized, while the other tries to provide a job-optimal solution so that all the jobs in the system of the same size will be charged approximately the same expected price independent of the computers allocated for their execution to provide fairness. The performance of the proposed dynamic schemes is compared with static job allocation schemes using simulations.
Satish Penmatsa, Anthony T. Chronopoulos
NCA2
2009 Efficient multi-party digital signature using adaptive secret sharing for low-power devices in wireless networks
abstract
In this paper, we propose an efficient multi-party signature scheme for wireless networks where a given number of signees can jointly sign a document, and it can be verified by any entity who possesses the certified group public key. Our scheme is based on an efficient threshold key generation scheme which is able to defend against both static and adaptive adversaries. Specifically, our key generation method employs the bit commitment technique to achieve efficiency in key generation and share refreshing; our share refreshing method provides proactive protection to long-lasting secret and allows a new signee to join a signing group. We demonstrate that previous known approaches are not efficient in wireless networks, and the proposed multi-party signature scheme is flexible, efficient, and achieves strong security for low-power devices in wireless networks.
Caimu Tang, Dapeng Oliver Wu, Anthony T. Chronopoulos, Cauligi S. Raghavendra
IEEE Trans. Wirel. Commun.3
2008 Cooperative load balancing in distributed systems
abstract
Abstract A serious difficulty in concurrent programming of a distributed system is how to deal with scheduling and load balancing of such a system which may consist of heterogeneous computers. In this paper, we formulate the static load‐balancing problem in single class job distributed systems as a cooperative game among computers. The computers comprising the distributed system are modeled as M/M/1 queueing systems. It is shown that the Nash bargaining solution (NBS) provides an optimal solution (operation point) for the distributed system and it is also a fair solution. We propose a cooperative load‐balancing game and present the structure of NBS. For this game an algorithm for computing NBS is derived. We show that the fairness index is always equal to 1 using NBS, which means that the solution is fair to all jobs. Finally, the performance of our cooperative load‐balancing scheme is compared with that of other existing schemes. Copyright © 2008 John Wiley & Sons, Ltd.
Daniel Grosu, Anthony T. Chronopoulos, Ming-Ying Leung
Concurr. Comput. Pract. Exp.2
2008 Enhancing self-scheduling algorithms via synchronization and weighting
Florina M. Ciorba, Ioannis Riakiotakis, Theodore Andronikos, George K. Papakonstantinou, Anthony T. Chronopoulos
J. Parallel Distributed Comput.5
2007 Studying the impact of synchronization frequency on scheduling tasks with dependencies in heterogeneous systems
Florina M. Ciorba, Ioannis Riakiotakis, George K. Papakonstantinou, Theodore Andronikos, Anthony T. Chronopoulos
PACT5
2007 Multi-dimensional dynamic loop scheduling algorithms
abstract
Distributed computing systems are a viable and less expensive alternative to parallel computers. However, a serious difficulty in concurrent programming of a distributed system is how to deal with scheduling and load balancing of such a system which may consist of heterogeneous computers. Loop scheduling schemes for parallel computers and computer clusters have been proposed in the past. All these schemes are one-dimensional because they partition only the outermost loop of a nested loop construct. In this work, we consider scheduling nested loops with many dimensions. We propose a new methodology which partitions many levels (or dimensions) of nested loops. These new schemes show superior performance over the existing schemes. We implement our new schemes on a network of computers and make performance comparisons with other existing schemes. We expect the new schemes to be particularly useful for multi-core systems because of the fine granularity of the generated tasks.
Anthony T. Chronopoulos, Lionel M. Ni, Satish Penmatsa
CLUSTER1
2007 Optimal synchronization frequency for dynamic pipelined computations on heterogeneous systems
abstract
In this paper we give a theoretical model for determining the synchronization frequency that minimizes the parallel execution time of loops with uniform dependencies dynamically scheduled on heterogeneous systems. Using this model we determine the synchronization frequency that minimizes the estimated parallel time. The accuracy of our method is validated through experiments on a heterogeneous cluster. The results show that the synchronization frequency minimizing the parallel time determined by our method, is very close to the synchronization frequency found experimentally.
Florina M. Ciorba, Ioannis Riakiotakis, Theodore Andronikos, Anthony T. Chronopoulos, George K. Papakonstantinou
CLUSTER4
2007 An optimal scheduling scheme for tiling in distributed systems
abstract
There exist several scheduling schemes for parallelizing loops without dependences for shared and distributed memory systems. However, efficiently parallelizing loops with dependences is a more complicated task. This becomes even more difficult when the loops are executed on a distributed memory cluster where communication and synchronization can be a bottleneck. The problem lies in the processor idle time which occurs during the beginning and final stages of the execution. In this paper we propose a new scheduling scheme that minimizes the processor idle time and thus it enhances load balancing and performance. The new scheme is applied to two-dimensional iteration spaces with dependences. The proposed scheduling scheme follows a tiled wavefront pattern in which the tile size gradually decreases in all dimensions. We have tested the proposed scheme on a dedicated and homogeneous cluster of workstations and we verified that it significantly improves execution times over scheduling using traditional tiling.
Konstantinos Kyriakopoulos, Anthony T. Chronopoulos, Lionel M. Ni
CLUSTER2
2007 Dynamic Multi-User Load Balancing in Distributed Systems
abstract
In this paper, we review two existing static load balancing schemes based on M/M/1 queues. We then use these schemes to propose two dynamic load balancing schemes for multi-user (multi-class) jobs in heterogeneous distributed systems. These two dynamic load balancing schemes differ in their objective. One tries to minimize the expected response time of the entire system while the other tries to minimize the expected response time of the individual users. The performance of the dynamic schemes is compared with that of the static schemes using simulations with various loads and parameters. The results show that, at low communication overheads, the dynamic schemes show superior performance over the static schemes. But as the overheads increase, the dynamic schemes (as expected) yield similar performance to that of the static schemes.
Satish Penmatsa, Anthony T. Chronopoulos
IPDPS2
2007 Implementation of Distributed Loop Scheduling Schemes on the TeraGrid
abstract
Grid computing can be used for high performance computations. However, a serious difficulty in concurrent programming of such heterogeneous systems is how to deal with scheduling and load balancing of such systems which may consist of heterogeneous computers on different sites. Distributed scheduling schemes suitable for parallel loops with independent iterations on heterogeneous computer clusters have been proposed and analyzed in the past. In this article, we implement the previous schemes in MPICH-G2 and MPIg on the TeraGrid. We present performance results for three loop scheduling schemes on single and multi-site TeraGrid clusters.
Satish Penmatsa, Anthony T. Chronopoulos, Nicholas T. Karonis, Brian R. Toonen
IPDPS2
2006 Joint rate and power control using game theory
abstract
Abstract — Efficient use of available resources in next gener-ation wireless systems require control of both data rate and transmitted power for mobile terminals. In this paper the prob-lem of joint transmission rate and power control is approached from the perspective of non-cooperative game theory, and an algorithm for joint rate and power control is presented. A new utility function for mobile terminals is defined, and a detailed analysis of the existence and uniqueness of Nash equilibrium for the non-cooperative joint transmission rate and power control game is presented. The utility function depends on the signal to interference ratio (SIR), and can be adjusted to provide the desired Quality of Service (QoS) requirement. Numerical simulations that compare the proposed algorithm with alternative algorithms developed using game theory are also presented in the paper. I.
Madhusudhan R. Musku, Anthony T. Chronopoulos, Dimitrie C. Popescu
CCNC2
2006 Dynamic multi phase scheduling for heterogeneous clusters
abstract
Distributed computing systems are a viable and less expensive alternative to parallel computers. However, concurrent programming methods in distributed systems have not been studied as extensively as for parallel computers. Some of the main research issues are how to deal with scheduling and load balancing of such a system, which may consist of heterogeneous computers. In the past, a variety of dynamic scheduling schemes suitable for parallel loops (with independent iterations) on heterogeneous computer clusters have been obtained and studied. However, no study of dynamic schemes for loops with iteration dependencies has been reported so far. In this work we study the problem of scheduling loops with iteration dependencies for heterogeneous (dedicated and non-dedicated) clusters. The presence of iteration dependencies incurs an extra degree of difficulty and makes the development of such schemes quite a challenge. We extend three well known dynamic schemes (CSS, TSS and DTSS) by introducing synchronization points at certain intervals so that processors compute in pipelined fashion. Our scheme is called dynamic multi-phase scheduling (DMPS) and we apply it to loops with iteration dependencies. We implemented our new scheme on a network of heterogeneous computers and studied its performance. Through extensive testing on two real-life applications (the heat equation and the Floyd-Steinberg algorithm), we show that the proposed method is efficient for parallelizing nested loops with dependencies on heterogeneous systems.
Florina M. Ciorba, Theodore Andronikos, Ioannis Riakiotakis, Anthony T. Chronopoulos, George K. Papakonstantinou
IPDPS4
2006 Cooperative load balancing for a network of heterogeneous computers
abstract
In this paper, we present a game theoretic approach to solve the static load balancing problem in a distributed system which consists of heterogeneous computers connected by a single channel communication network. We use a cooperative game to model the load balancing problem. Our solution is based on the Nash bargaining solution (NBS) which provides a Pareto optimal solution for the distributed system and is also a fair solution. An algorithm for computing the NBS is derived for the proposed cooperative load balancing game. Our scheme is compared with that of other existing schemes under simulations with various system loads and configurations. We show that the solution of our scheme is near optimal and is superior to the other schemes in terms of fairness.
Satish Penmatsa, Anthony T. Chronopoulos
IPDPS2
2006 Price-based user-optimal job allocation scheme for grid systems
abstract
In this paper, we propose a price-based user-optimal job allocation scheme for grid systems whose nodes are connected by a communication network. The job allocation problem is formulated as a noncooperative game among the users who try to minimize the expected cost of their own jobs. We use the concept of Nash equilibrium as the solution of our noncooperative game and derive a distributed algorithm for computing it. The prices that the grid users has to pay for using the computing resources owned by different resource owners are obtained using a pricing model based on a game theory framework. Finally, our scheme is compared with a system-optimal job allocation scheme under simulations with various system loads and configurations and conclusions are drawn
Satish Penmatsa, Anthony T. Chronopoulos
IPDPS2
2006 An efficient concurrent implementation of a neural network algorithm
abstract
The focus of this study is how we can efficiently implement the neural network backpropagation algorithm on a network of computers (NOC) for concurrent execution. We assume a distributed system with heterogeneous computers and that the neural network is replicated on each computer. We propose an architecture model with efficient pattern allocation that takes into account the speed of processors and overlaps the communication with computation. The training pattern set is distributed among the heterogeneous processors with the mapping being fixed during the learning process. We provide a heuristic pattern allocation algorithm minimizing the execution time of backpropagation learning. The computations are overlapped with communications. Under the condition that each processor has to perform a task directly proportional to its speed, this allocation algorithm has polynomial-time complexity. We have implemented our model on a dedicated network of heterogeneous computers using Sejnowski's NetTalk benchmark for testing. Copyright © 2005 John Wiley & Sons, Ltd.
Razvan Andonie, Anthony T. Chronopoulos, Daniel Grosu, Honorius Gâlmeanu
Concurr. Comput. Pract. Exp.2
2006 Distributed loop-scheduling schemes for heterogeneous computer systems
abstract
Abstract Distributed computing systems are a viable and less expensive alternative to parallel computers. However, a serious difficulty in concurrent programming of a distributed system is how to deal with scheduling and load balancing of such a system which may consist of heterogeneous computers. Some distributed scheduling schemes suitable for parallel loops with independent iterations on heterogeneous computer clusters have been designed in the past. In this work we study self‐scheduling schemes for parallel loops with independent iterations which have been applied to multiprocessor systems in the past. We extend one important scheme of this type to a distributed version suitable for heterogeneous distributed systems. We implement our new scheme on a network of computers and make performance comparisons with other existing schemes. Copyright © 2005 John Wiley & Sons, Ltd.
Anthony T. Chronopoulos, Satish Penmatsa, Siraj Ali
Concurr. Comput. Pract. Exp.1
2005 Joint rate and power control with pricing
abstract
Next generation wireless systems is required to support heterogeneous services with different transmission rates that include real time multimedia transmissions, as well as non-real time data transmissions. In order to provide flexible transmission rates to each terminal, efficient use of system resources requires transmission rate control in addition to power control. In this paper, we present an algorithm for joint transmission rate and power control based on a non-cooperative game theoretic approach. A new utility function that includes pricing is defined for joint transmission rate and power control, and a detailed analysis of the existence and uniqueness of Nash equilibrium for the non-cooperative joint transmission rate and power control game with pricing is presented. Numerical results obtained from simulations that compare the proposed algorithm with alternative algorithms on joint rate and power control are also presented in the paper
Madhusudhan R. Musku, Anthony T. Chronopoulos, Dimitrie C. Popescu
GLOBECOM2
2005 Soft-Timeout Distributed Key Generation for Digital Signature based on Elliptic Curve D-log for Low-Power Devices
abstract
Group based transactions are becoming common via handhelds. Single key based systems may not be able to meet various security requirements. In this paper, we propose a threshold signature scheme based on Pedersen distributed key generation principle which is suitable for handheld devices and ad-hoc networks. Existing distributed key generation protocols use either cryptosystems based on the hardness of discrete logarithm over a finite field or integer factorization. Elliptic curve cryptosystems provide a promising alternative with efficiency which is suitable for low-power devices in terms of memory and processing overhead. In the proposed scheme, the public key from the key generation protocol follows a uniform distribution in the elliptic curve additive group, and the signature can be generated and verified efficiently. We evaluated the proposed key generation protocol and signature scheme using PARI/GP, and the key generation time takes a fraction of a second and the signature signing and verifying can be finished in a few milliseconds on the LINUX Intel PXA 255 processor.
Caimu Tang, Anthony T. Chronopoulos, Cauligi S. Raghavendra
SecureComm2
2005 Noncooperative load balancing in distributed systems
Daniel Grosu, Anthony T. Chronopoulos
J. Parallel Distributed Comput.2
2005 An efficient network-switch scheduling for real-time applications
abstract
Bursts consist of a varying number of asynchronous transfer mode cells corresponding to a datagram. Here, we generalized weighted fair queueing to a burst-based algorithm with preemption. The new algorithm enhances the performance of the switch service for real-time applications, and it preserves the quality of service guarantees. We study this algorithm theoretically and via simulations.
Caimu Tang, Anthony T. Chronopoulos, Ece Yaprak
IEEE Trans. Commun.2
2004 Implementation of Distributed Key Generation Algorithms using Secure Sockets
abstract
Distributed key generation (DKG) protocols are indispensable in the design of any cryptosystem used in communication networks. DKG is needed to generate public/private keys for signatures or more generally for encrypting/decrypting messages. One such DKG (due to Pedersen) has recently been generalized to a provably secure protocol by Gennaro et al. We propose and implement an efficient algorithm to compute the (group generator) parameter g required in the DKG protocol. We also implement the DKG due to Gennaro et al. on a network of computers using secure sockets. We run tests which show the efficiency of the implementation.
Anthony T. Chronopoulos, F. Balbi, D. Veljkovic, N. Kolani
NCA1
2004 Efficient power control for broadcast in wireless communication systems
abstract
Energy efficiency is a measure of performance in wireless networks. Therefore, controlling the transmitter power at a given node increases not only the operating life of the battery but also the overall system capacity by successfully admitting new nodes between a source and a destination. It is essential to find effective means of power control of point-to-point, broadcasting and multicasting scenarios. In past work A. T. Chronopoulos et al. (2002), we presented a new scheme state space-based control design (SSCD) using both the state space and optimal control methodology in discrete-time for power control in wireless systems. Further, we proved the convergence of the overall network with our algorithm using Lyapunov stability analysis. We made a comparison with a well known distributed power control (DPC) scheme. Here we present simulation results and comparisons for point-to-point communication with random node placement. We also combined the schemes SSCD and DPC with a tree based broadcast algorithm BIP to obtain broadcast tree in ad-hoc networks with power control. We show the effectiveness of the new algorithm through simulations.
Anthony T. Chronopoulos, Paul Cotae, S. Ponipireddy
WCNC1
2004 Algorithmic mechanism design for load balancing in distributed systems
abstract
Computational grids are promising next-generation computing platforms for large-scale problems in science and engineering. Grids are large-scale computing systems composed of geographically distributed resources (computers, storage etc.) owned by self interested agents or organizations. These agents may manipulate the resource allocation algorithm in their own benefit, and their selfish behavior may lead to severe performance degradation and poor efficiency. In this paper, we investigate the problem of designing protocols for resource allocation involving selfish agents. Solving this kind of problems is the object of mechanism design theory. Using this theory, we design a truthful mechanism for solving the static load balancing problem in heterogeneous distributed systems. We prove that using the optimal allocation algorithm the output function admits a truthful payment scheme satisfying voluntary participation. We derive a protocol that implements our mechanism and present experiments to show its effectiveness.
Daniel Grosu, Anthony T. Chronopoulos
IEEE Trans. Syst. Man Cybern. Part B2
2003 A Truthful Mechanism for Fair Load Balancing in Distributed Systems
abstract
In this paper we consider the problem of designing load balancing protocols in distributed systems where the participants (e.g. computers, users) are capable of manipulating the load allocation algorithm in their own interest. Using techniques from mechanism design theory we design a mechanism for fair load balancing in heterogeneous distributed systems. We prove that our mechanism is truthful and satisfies the voluntary participation condition. Based on the proposed mechanism we derive a fair load balancing protocol called FAIR-LBM. Finally, we study the effectiveness of our protocol by simulations.
Daniel Grosu, Anthony T. Chronopoulos
NCA2
2003 An efficient 3D grid based scheduling for heterogeneous systems
Anthony T. Chronopoulos, Daniel Grosu, Andrew M. Wissink, Manuel Benche
J. Parallel Distributed Comput.1
2002 Scalable Loop Self-Scheduling Schemes for Heterogeneous Clusters
abstract
Distributed systems (e.g. a LAN of computers) can be used for concurrent processing for some applications. However a serious difficulty in concurrent programming of a distributed system is how to deal with scheduling and load balancing of such a system which may consist of heterogeneous computers. Distributed scheduling schemes suitable for parallel loops with independent iterations on heterogeneous computer clusters have been proposed and analyzed in the past. Here, we implement the previous schemes in CORBA (Orbix). We also present an extension of these schemes implemented in a hierarchical master-slave architecture. We present experimental results and comparisons.
Anthony T. Chronopoulos, Satish Penmatsa
CLUSTER1
2002 Algorithmic Mechanism Design for Load Balancing in Distributed Systems
abstract
Computational grids are large scale computing systems composed of geographically distributed resources (computers, storage etc.) owned by self interested agents or organizations. These agents may manipulate the resource allocation algorithm for their own benefit and their selfish behavior may lead to severe performance degradation and poor efficiency. In this paper we investigate the problem of designing protocols for resource allocation involving selfish agents. Solving this kind of problem is the object of mechanism design theory. Using this theory we design a truthful mechanism for solving the static load balancing problem in heterogeneous distributed systems. We prove that by using the optimal allocation algorithm the output function admits a truthful payment scheme satisfying voluntary participation. We derive a protocol that implements our mechanism and present experiments to show its effectiveness.
Daniel Grosu, Anthony T. Chronopoulos
CLUSTER2
2002 Distributed power control in wireless communication systems
abstract
Energy efficiency is a measure of performance in wireless networks. Therefore, controlling the transmitter power at a given node increases not only battery operating life, but also overall system capacity by successfully admitting new links. It is essential to find effective means of power control in point-to-point, broadcasting and multicasting scenarios. Wireless networking presents formidable challenges and we consider the problem of unicast or point-to-point (peer-to-peer) communication in wireless networks in the presence of other nodes. We study the feasibility of admitting new links in an wireless network operating area while maintaining quality of service (QoS), in terms of signal-to-interference ratio (SIR), for each link. SIR is maintained by adjusting the transmitter power levels at each source for a given link. Distributed power control (DPC) is a natural choice for this purpose because, unlike centralized power control, DPC should be able to adjust the power levels of each transmitted signal using local measurements, so that in a reasonable time, all nodes/links maintain the desired SIR. We present a suite of DPC schemes using both state space and optimal control methodology in discrete-time. Further, we prove the convergence of the overall network with our algorithm using Lyapunov stability analysis in comparison with a well known DPC scheme (see Bambos, N. et al., IEEE ACM Trans. on Networking, p.583-97, 2000). We present simulation results and comparisons for point to point communications in an overlapping scenario.
Sarangapani Jagannathan, Anthony T. Chronopoulos, S. Ponipireddy
ICCCN2
2001 A Class of Loop Self-Scheduling for Heterogeneous Clusters
abstract
Distributed Computing Systems are a viable and less expensive alternative to parallel computers. However, a serious difficulty in concurrent programming of a distributed system is how to deal with scheduling and load balancing of such a system which may consist of heterogeneous computers. Distributed scheduling schemes suitable for parallel loops with independent iterations on heterogeneous computer clusters have been designed in the past. In this work we consider a class of Self-Scheduling schemes for parallel loops with independent iterations which have been applied to multiprocessor systems. We extend this type of schemes to heterogeneous distributed systems. We present tests that the distributed versions of these schemes maintain load balanced execution on heterogeneous systems.
Anthony T. Chronopoulos, Manuel Benche, Daniel Grosu, Razvan Andonie
CLUSTER1
2001 A Parallel Krylov-Type Method for Nonsymmetric Linear Systems
Anthony T. Chronopoulos, Andrey B. Kucherov
HiPC1
2001 Static Load Balancing for CFD Simulations on a Network of Workstations
abstract
In distributed simulations, the delivered performance of networks of heterogeneous computers degrades severely if the computations are not load balanced. In this work we consider the distributed simulation of TURNS (Transonic Unsteady Rotor Navier Stokes), a 3-D space CFD code. We propose a load balancing heuristic for simulations on networks of heterogeneous workstations. Our algorithm takes into account the CPU speed and memory capacity of the workstations. Test run comparisons with the equal task allocation algorithm demonstrated significant efficiency gains.
Anthony T. Chronopoulos, Daniel Grosu, Manuel Benche, Andrew M. Wissink
NCA1
2000 A new-generation parallel computer and its performance evaluation
Sotirios G. Ziavras, Haim Grebel, Anthony T. Chronopoulos, Florent Marcelli
Future Gener. Comput. Syst.3
1999 Dynamic buffer allocation in an ATM switch
Ece Yaprak, Anthony T. Chronopoulos, Kleanthis Psarris
Comput. Networks2
1998 A cell burst scheduling for ATM networking. I. Theory
abstract
Fair queueing is a useful queueing discipline for packet switching systems. It was developed in last decade and was aimed at the general packet switching systems with varying packet length. However it is not suitable for use in the ATM networking, because the ATM cell length is very small and fixed, and so the scheduling scheme on a per cell basis is not practical. We introduce the burst and quality unit concepts in the scheduling algorithm and we make some significant modification on the fair queueing and adapt it to ATM networking to meet QoS requirements. Under the work-conserving assumption, we show that the burst based non-preemptive and preemptive algorithms provide throughput and fairness guarantees.
Caimu Tang, Anthony T. Chronopoulos, Ece Yaprak
ISCC2
1998 A cell burst scheduling for ATM networking. II. Implementation
abstract
For pt.I see ibid. p.445-61. The scheduling scheme of a switch affects the delay, throughput and fairness of a network and thus has a great impact on the quality of service (QoS). In part I, we present a theoretical analysis of a burst scheduling for ATM switches and proved QoS guarantees on throughput and fairness of the applications. Here, we use simulation to demonstrate the superiority of the burst based weighted fair queueing over the non-burst version. Our simulation study is based on backbone and access subnetworks, which are common in the real world.
Caimu Tang, Anthony T. Chronopoulos, Ece Yaprak
ISCC2
1997 Parallel Solution of a Traffic Flow Simulation Problem
Anthony T. Chronopoulos
Parallel Comput.1
1996 Parallel Iterative S-Step Methods for Unsymmetric Linear Systems
Anthony T. Chronopoulos, Charles D. Swanson
Parallel Comput.1
1992 Orthogonal s-step methods for nonsymmetric linear systems of equations
abstract
Conjugate Gradient-like methods such as Orthomin(k) have been developed to obtain a good numerical approximation to the solution of Ax = f when the matrix A is large, sparse, and nonsymmetric. In s-step variations of these iterative methods, s consecutive steps of the one-step methods are performed simultaneously. The number of inner products required is reduced and the resulting algorithms are more suitable for parallel computations. However, lack of orthogonality between the s direction vectors at each iteration leads to instability unless s is small (s≤5). In this project, the effect of orthogonalizing the s direction vectors at each iteration was studied. The ATA-Orthogonal s-Step Orthomin(k) and p-Orthogonal s-Step Orthomin(k) algorithms were developed and shown to be stable for large values of s (up to s=20). The performance of these algorithms on a multiple processor CRAY Y-MP8 computer was analyzed.
Charles D. Swanson, Anthony T. Chronopoulos
ICS2
1991 An Efficient Arnoldi Method Implemented on Parallel Computers
Sun Kyung Kim, Anthony T. Chronopoulos
ICPP (3)2
1991 Towards efficient parallel implementation of the CG method applied to a class of block tridiagonal linear systems
abstract
An efficient implementation of Conjugate Gradient (CG) methods on vector and parallel machines is presented.The two different architecture models considered are the shared memory machines with memory hierarchy and the message passing private memory machines.For a parametrized vector architecture similu to CRAY-2 we present (theoretically) an implementation of the s-step CG used to solve an elliptic partiaJ differential-equation fast as that of the standard parallel computem we show of s-step CG can be up to mance of the standatd CG.problem is twice as CG.For Hypcrcube that the performance 2s times the perfor-
Anthony T. Chronopoulos
SC1
1991 A class of Lanczos-like algorithms implemented on parallel computers
Sun Kyung Kim, Anthony T. Chronopoulos
Parallel Comput.2
1989 On the efficient implementation of preconditioned s-step conjugate gradient methods on multiprocessors with memory hierarchy
Anthony T. Chronopoulos, Charles William Gear
Parallel Comput.1