VLDB 2026 Research / reviewers in the wild / expert
Iraklis Anagnostopoulos
dblp:12/7914
· DBLP profile ↗
37ranked-venue papers
3as first author
21since 2021 · last 2025
0000-0003-0985-3045ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 3 first-author · 18 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RankMap: Priority-Aware Multi-DNN Manager for Heterogeneous Embedded DevicesabstractModern edge data centers simultaneously handle multiple Deep Neural Networks (DNNs), leading to significant challenges in workload management. Thus, current management systems must leverage the architectural heterogeneity of new embedded systems to efficiently handle multi-DNN workloads. This paper introduces RankMap, a priority-aware manager specifically designed for multi-DNN tasks on heterogeneous embedded devices. RankMap addresses the extensive solution space of multi-DNN mapping through stochastic space exploration combined with a performance estimator. Experimental results show that RankMap achieves x3.6 higher average throughput compared to existing methods, while preventing DNN starvation under heavy workloads and improving the prioritization of specified DNNs by x 57.5. Andreas Karatzas, Dimitrios Stamoulis, Iraklis Anagnostopoulos |
DATE | 3 |
| 2025 | Late Breaking Results: Leveraging Approximate Computing for Carbon-Aware DNN AcceleratorsabstractThe rapid growth of Machine Learning (ML) has increased demand for DNN hardware accelerators, but their embodied carbon footprint poses significant environmental challenges. This paper leverages approximate computing to design sustainable accelerators by minimizing the Carbon Delay Product (CDP). Using gate-level pruning and precision scaling, we generate area-aware approximate multipliers and optimize the accelerator design with a genetic algorithm. Results demonstrate reduced embodied carbon while meeting performance and accuracy requirements. Aikaterini Maria Panteleaki, Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Iraklis Anagnostopoulos |
DATE | 5 |
| 2025 | Less is More: Optimizing Function Calling for LLM Execution on Edge DevicesabstractThe advanced function-calling capabilities of foundation models open up new possibilities for deploying agents to perform complex API tasks. However, managing large amounts of data and interacting with numerous APIs makes function calling hardware-intensive and costly, especially on edge devices. Current Large Language Models (LLMs) struggle with function calling at the edge because they cannot handle complex inputs or manage multiple tools effectively. This results in low task-completion accuracy, increased delays, and higher power consumption. In this work, we introduce Less-is-More, a novel fine-tuning-free function-calling scheme for dynamic tool selection. Our approach is based on the key insight that selectively reducing the number of tools available to LLMs significantly improves their function-calling performance, execution time, and power efficiency on edge devices. Experimental results with state-of-the-art LLMs on edge hardware show agentic success rate improvements, with execution time reduced by up to 70% and power consumption by up to 40%. Varatheepan Paramanayakam, Andreas Karatzas, Iraklis Anagnostopoulos, Dimitrios Stamoulis |
DATE | 3 |
| 2025 | Sponge Attacks on Sensing AI: Energy-Latency Vulnerabilities and Defense via Model PruningabstractRecent studies have shown that sponge attacks can significantly increase the energy consumption and inference latency of deep neural networks (DNNs). However, prior work has focused primarily on computer vision and natural language processing tasks, overlooking the growing use of lightweight AI models in sensing-based applications on resource-constrained devices, such as those in Internet of Things (IoT) environments. These attacks pose serious threats of energy depletion and latency degradation in systems where limited battery capacity and real-time responsiveness are critical for reliable operation. This paper makes two key contributions. First, we present the first systematic exploration of energy-latency sponge attacks targeting sensing-based AI models. Using wearable sensing-based AI as a case study, we demonstrate that sponge attacks can substantially degrade performance by increasing energy consumption, leading to faster battery drain, and by prolonging inference latency. Second, to mitigate such attacks, we investigate model pruning, a widely adopted compression technique for resource-constrained AI, as a potential defense. Our experiments show that pruning-induced sparsity significantly improves model resilience against sponge poisoning. We also quantify the trade-offs between model efficiency and attack resilience, offering insights into the security implications of model compression in sensing-based AI systems deployed in IoT environments. Syed Mhamudul Hasan, Hussein Zangoti, Iraklis Anagnostopoulos, Abdur Rahman Bin Shahid |
GLOBECOM | 3 |
| 2025 | Leveraging Image Difficulty for Run-Time Adaptive DNN Inference on Embedded DevicesabstractDeep Neural Networks (DNNs) impose great challenges on resource-constraint embedded devices since they employ billions of computational operations. To satisfy these computational demands such devices utilize hardware accelerators, which can lead to increased power consumption whatsoever. To that end, the compression of DNNs to lower precision has been proposed, in order to achieve savings in energy consumption at the cost of some accuracy loss during inference. However, DNNs do not behave similarly under lower-precision execution and the accuracy degradation can be severe. In this work, we utilize the notion of image difficulty and explore how we can change DNN precision during inference to achieve gains in energy consumption without big drops in accuracy. We evaluate our work on the ImageNet dataset and show how the proposed framework achieves energy savings at run-time. Vasileios Pentsos, Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos |
ISCAS | 4 |
| 2025 | Approximate Multiplier Mapping for Unfairness Mitigation in Energy-Efficient DNNsabstractEmbedded devices struggle with the heavy computational demands of extensive neural network models, a problem partially addressed by integrating accelerators with numerous multiply-accumulate units. However, this solution increases energy consumption. While using approximate circuits in accelerators can lower energy usage, it compromises accuracy and raises concerns about maintaining fair inference across diverse populations, particularly in the medical field. This work leverages approximate multipliers in deep neural networks to retain fairness and reduce energy consumption while keeping the inference accuracy within strict thresholds. Ourania Spantidi, Georgios Zervakis 0001, Jörg Henkel, Iraklis Anagnostopoulos |
ISCAS | 4 |
| 2025 | Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge ServersabstractEdge computing systems struggle to efficiently manage multiple concurrent deep neural network (DNN) workloads while meeting strict latency requirements, minimizing power consumption, and maintaining environmental sustainability. This paper introduces Ecomap, a sustainability-driven framework that dynamically adjusts the maximum power threshold of edge devices based on real-time carbon intensity. Ecomap incorporates the innovative use of mixed-quality models, allowing it to dynamically replace computationally heavy DNNs with lighter alternatives when latency constraints are violated, ensuring service responsiveness with minimal accuracy loss. Additionally, it employs a transformer-based estimator to guide efficient workload mappings. Experimental results using NVIDIA Jetson AGX Xavier demonstrate that Ecomap reduces carbon emissions by an average of 30% and achieves a 25% lower carbon delay product (CDP) compared to state-of-the-art methods, while maintaining comparable or better latency and power efficiency. Varatheepan Paramanayakam, Andreas Karatzas, Dimitrios Stamoulis, Iraklis Anagnostopoulos |
IEEE Trans. Computers | 4 |
| 2024 | Exploration of TPU Architectures for the Optimized Transformer in Drainage Crossing DetectionabstractUnderstanding hydrologic connectivity within landscapes is crucial for managing environmental challenges. Despite advancements in high-resolution Digital Elevation Models (DEMs) derived from Light Detection and Ranging (LiDAR) technology, accurately delineating hydrologic connectivity remains challenging due to disruptions caused by virtual flow barriers, such as roads and bridges. This study addresses this issue by enhancing the detection performance and reducing the latency of Transformer models for image detection of drainage crossings. We retrained a Detection Transformer (DETR) with a specialized recipe to improve culvert detection performance. Owing to the high susceptibility of LiDAR-based DEMs to measurement noise and varying data modalities, we conducted extensive data preprocessing to ensure DETR compatibility with the culvert dataset. Ablation studies on input size indicate that the model performs optimally with 800×800 pixel inputs, demonstrating its adaptability to new data modalities. Additionally, we employed Tensor Processing Units (TPUs) to decrease the model’s latency. We developed a novel strategy to optimize TPU architecture, utilizing genetic algorithms to expedite the discovery of optimal TPU configurations for detection deployment. Our model surpasses the performance of previous models on the same task. This work not only addresses the computational complexities of deploying advanced object detection in environmental contexts but also significantly contributes to the precise and efficient monitoring of hydrologic connectivity. Amirhossein Nazeri, Denys W. Godwin, Aikaterini Maria Panteleaki, Iraklis Anagnostopoulos, Michael Edidem, Ruopu Li, Tong Shu |
IEEE Big Data | 4 |
| 2024 | MapFormer: Attention-based multi-DNN manager for throughout & power co-optimization on embedded devicesabstractIn the context of modern services that use multiple Deep Neural Networks (DNNs), managing workloads on embedded devices presents unique challenges. These devices often incorporate diverse architectures, necessitating advanced management solutions to efficiently deploy multi-DNN workloads. Traditionally, the focus has been on improving throughput, while power optimization has received less attention. This paper presents MapFormer, a new manager that uses attention-based mechanisms to enhance both throughput and power efficiency. MapFormer intelligently assigns multi-DNN workloads to different computing components of embedded systems---CPU, GPU, and DLA---and adjusts operational frequencies to optimize power use. Experimental results show that MapFormer significantly improves average throughput under set power budgets by 90.8%, offering a promising approach for managing complex workloads on heterogeneous embedded systems. Andreas Karatzas, Iraklis Anagnostopoulos |
ICCAD | 2 |
| 2023 | OmniBoost: Boosting Throughput of Heterogeneous Embedded Devices under Multi-DNN WorkloadabstractModern Deep Neural Networks (DNNs) exhibit profound efficiency and accuracy properties. This has introduced application workloads that comprise of multiple DNN applications, raising new challenges regarding workload distribution. Equipped with a diverse set of accelerators, newer embedded system present architectural heterogeneity, which current run-time controllers are unable to fully utilize. To enable high throughput in multi-DNN workloads, such a controller is ought to explore hundreds of thousands of possible solutions to exploit the underlying heterogeneity. In this paper, we propose OmniBoost, a lightweight and extensible multi-DNN manager for heterogeneous embedded devices. We leverage stochastic space exploration and we combine it with a highly accurate performance estimator to observe a ×4.6 average throughput boost compared to other state-of-the-art methods. The evaluation was performed on the HiKey970 development board. Our code is publicly available at https://github.com/AndreasKaratzas/omniboost-v1. Andreas Karatzas, Iraklis Anagnostopoulos |
DAC | 2 |
| 2023 | Automated Energy-Efficient DNN Compression under Fine-Grain Accuracy ConstraintsabstractDeep Neural Networks (DNNs) are utilized in a variety of domains, and their computation intensity is stressing embedded devices that comprise limited power budgets. DNN compression has been employed to achieve gains in energy consumption on embedded devices at the cost of accuracy loss. Compression-induced accuracy degradation is addressed through fine-tuning or retraining, which can not always be feasible. Additionally, state-of-art approaches compress DNNs with respect to the average accuracy achieved during inference, which can be a misleading evaluation metric. In this work, we explore more fine-grain properties of DNN inference accuracy, and generate energy-efficient DNNs using signal temporal logic and falsification jointly through pruning and quantization. We offer the ability to control at run-time the quality of the DNN inference, and propose an automated framework that can generate compressed DNNs that satisfy tight fine-grain accuracy requirements. The conducted evaluation on the ImageNet dataset has shown over 30% in energy consumption gains when compared to baseline DNNs. Ourania Spantidi, Iraklis Anagnostopoulos |
DATE | 2 |
| 2023 | The Perfect Match: Selecting Approximate Multipliers for Energy-Efficient Neural Network InferenceabstractReconfigurable approximate multipliers have been proposed as a way to improve the energy efficiency of neural network inference. However, selecting the optimal combination of approximate modes is a challenging problem due to the tradeoff between energy savings and accuracy loss. In this paper, we propose a methodology for selecting the best triad of approximate multipliers to form a reconfigurable approximate multiplier that can satisfy a maximum accuracy drop threshold and achieve the highest possible energy savings. We use formal methods to produce a Pareto-front of solutions that satisfy the accuracy constraint and maximize energy savings. Experimental results show that our methodology can achieve significant gains in energy with negligible drops in accuracy when compared to the baseline. Ourania Spantidi, Iraklis Anagnostopoulos |
HPSR | 2 |
| 2022 | Fair Scheduling Through Collaborative Filtering on Multicore SystemsabstractModern applications are being increasingly demanding in terms of computing capabilities, and high performance is required at all times. Chip multiprocessors (CMPs) comprise multiple cores and have been widely employed to address this demand. However, the cores of a CMP share several components of the memory hierarchy for which concurrent executing applications compete to access at run-time. This contention can lead to severe performance loss and has a different impact on each application, resulting in potential starvation for selected applications. Thus, there is a need for a scheduling policy to efficiently address this contention-induced unfairness. In this work, we utilize matrix reconstruction techniques to enhance scheduling decisions at run-time, ensuring the fair and efficient execution of any given application workload. Our evaluation shows that our proposed scheduling policy can achieve up to 25.8% gains in fairness when compared to the Linux completely fair scheduler, and up to 6.1% when compared to another state-of-the-art approach, without inflicting performance degradation. Ourania Spantidi, Theodoros Marinakis, Iraklis Anagnostopoulos |
ISCAS | 3 |
| 2022 | A Pressure-Aware Policy for Contention Minimization on Multicore SystemsabstractModern Chip Multiprocessors (CMPs) are integrating an increasing amount of cores to address the continually growing demand for high-application performance. The cores of a CMP share several components of the memory hierarchy, such as Last-Level Cache (LLC) and main memory. This allows for considerable gains in multithreaded applications while also helping to maintain architectural simplicity. However, sharing resources can also result in performance bottleneck due to contention among concurrently executing applications. In this work, we formulate a fine-grained application characterization methodology that leverages Performance Monitoring Counters (PMCs) and Cache Monitoring Technology (CMT) in Intel processors. We utilize this characterization methodology to develop two contention-aware scheduling policies, one static and one dynamic , that co-schedule applications based on their resource-interference profiles. Our approach focuses on minimizing contention on both the main-memory bandwidth and the LLC by monitoring the pressure that each application inflicts on these resources. We achieve performance benefits for diverse workloads, outperforming Linux and three state-of-the-art contention-aware schedulers in terms of system throughput and fairness for both single and multithreaded workloads. Compared with Linux, our policy achieves up to 16% greater throughput for single-threaded and up to 40% greater throughput for multithreaded applications. Additionally, the policies increase fairness by up to 65% for single-threaded and up to 130% for multithreaded ones. Shivam Kundan, Theodoros Marinakis, Iraklis Anagnostopoulos, Dimitrios Kagaris |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Thermal-Aware Design for Approximate DNN AcceleratorsabstractRecent breakthroughs in Neural Networks (NNs) have made DNN accelerators ubiquitous and led to an ever-increasing quest on adopting them from Cloud to edge computing. However, state-of-the-art DNN accelerators pack immense computational power in a relatively confined area, inducing significant on-chip power densities that lead to intolerable thermal bottlenecks. Existing state of the art focuses on using approximate multipliers only to trade-off efficiency with inference accuracy. In this work, we present a thermal-aware approximate DNN accelerator design in which we additionally trade-off approximation with temperature effects towards designing DNN accelerators that satisfy tight temperature constraints. Using commercial multi-physics tool flows for heat simulations, we demonstrate how our thermal-aware approximate design reduces the temperature from 139$^{\circ }$C, in an accurate circuit, down to 79$^{\circ }$C. This enables DNN accelerators to fulfill tight thermal constraints, while still maximizing the performance and reducing the energy by around 75% with a negligible accuracy loss of merely 0.44% on average for a wide range of NN models. Furthermore, using physics-based transistor aging models, we demonstrate how reductions in voltage and temperature obtained by our approximate design considerably improve the circuit’s reliability. Our approximate design exhibits around 40% less aging-induced degradation compared to the baseline design. Georgios Zervakis 0001, Iraklis Anagnostopoulos, Sami Salamin, Ourania Spantidi, Isai Roman-Ballesteros, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 2 |
| 2022 | Energy-Efficient DNN Inference on Approximate Accelerators Through Formal Property ExplorationabstractDeep neural networks (DNNs) are being heavily utilized in modern applications, putting energy-constraint devices to the test. To bypass high energy consumption issues, approximate computing has been employed in DNN accelerators to balance out the accuracy-energy reduction trade-off. However, the approximation-induced accuracy loss can be very high and drastically degrade the performance of the DNN. Therefore, there is a need for a fine-grain mechanism that would assign specific DNN operations to approximation to maintain acceptable DNN accuracy, while achieving low energy consumption. We present an automated framework for weight-to-approximation mapping through formal property exploration for approximate DNN accelerators. At the MAC unit level, our experimental evaluation surpassed already energy-efficient mappings by more than$\times 2$in terms of energy gains, while supporting a fine-grain control over the introduced approximation. Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Control Variate Approximation for DNN AcceleratorsabstractIn this work, we introduce a control variate approximation technique for low error approximate Deep Neural Network (DNN) accelerators. The control variate technique is used in Monte Carlo methods to achieve variance reduction. Our approach significantly decreases the induced error due to approximate multiplications in DNN inference, without requiring time-exhaustive retraining compared to state-of-the-art. Leveraging our control variate method, we use highly approximated multipliers to generate power-optimized DNN accelerators. Our experimental evaluation on six DNNs, for Cifar-10 and Cifar100 datasets, demonstrates that, compared to the accurate design, our control variate approximation achieves same performance and 24% power reduction for a merely 0.16% accuracy loss. Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel |
DAC | 3 |
| 2021 | Reliability-Aware Quantization for Anti-Aging NPUs
Sami Salamin, Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Jörg Henkel, Hussam Amrouch |
DATE | 4 |
| 2021 | Efficient Resource Management of Clustered Multi-Processor Systems Through Formal Property ExplorationabstractModern embedded systems have adopted the clustered Chip Multi-Processor (CMP) paradigm in conjunction with dynamic frequency scaling techniques to improve application performance and power consumption. Nonetheless, modern applications are becoming more aggressive in terms of computational power. At the same time, the integration of multiple cores in the same cluster has resulted in significant increase of power consumption creating thermal hotspots. Conventional design approaches consider fixed power and temperature constraints, which are mostly extracted experimentally leading many times to pessimistic run-time decisions and performance losses. In this paper, we present a unified framework for efficient resource management of clustered CMPs by enabling formal property exploration and integrating robustness analysis. Specifically, we bridge the gap between run-time decisions and design-time exploration by using Parametric Signal Temporal Logic (PSTL) for mining the values of system constraints. Then, we utilize the extracted values to enhance the decisions of the run-time resource manager. Results on the Odroid-XU3 show that the proposed methodology offers more coarse- and fine-grain optimizations. Ourania Spantidi, Iraklis Anagnostopoulos, Georgios Fainekos |
DATE | 2 |
| 2021 | Positive/Negative Approximate Multipliers for DNN AcceleratorsabstractRecent Deep Neural Networks (DNNs) manage to deliver superhuman accuracy levels on many AI tasks. DNN accelerators are becoming integral components of modern systems-on-chips. DNNs perform millions of arithmetic operations per inference and DNN accelerators integrate thousands of multiply-accumulate units leading to increased energy requirements. To lower the energy consumption of DNN accelerators, approximate computing principles are employed. However, complex DNNs can be increasingly sensitive to approximation. In this work, we present a dynamically configurable approximate multiplier that supports three operation modes, i.e., exact, positive error, and negative error. In addition, we propose a filter-oriented approximation method to map the weights to the appropriate modes of the approximate multiplier. Our mapping algorithm balances the positive with the negative errors due to the approximate multiplications, aiming at maximizing the energy reduction while minimizing the overall convolution error. We evaluate our approach on multiple DNNs and datasets against state-of-the-art approaches, where our method achieves 18.33% energy gains on average across 7 NNs on 4 different datasets for a maximum accuracy drop of only 1%. Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel |
ICCAD | 3 |
| 2021 | Priority-Aware Scheduling under Shared-Resource Contention on Chip Multicore ProcessorsabstractIn this paper, we present a priority-aware scheduling methodology for concurrent application execution on chip multi-core processors. Our methodology improves the performance of up to 4 high-priority applications while also preventing resource starvation for low-priority ones by means of progress-aware scheduling. We compare our results with Linux's completely fair scheduler and two state-of-the-art progress aware schedulers. Experimental results on an Intel Xeon Gold 6130 CPU demonstrate an average increase in high-priority application performance of 36.4% over Linux while also maintaining high throughput for low-priority applications. In addition, we show that our methodology achieves high-priority application performance comparable to a state-of-the-art hardware cache partitioning method (Intel's POCAT). Our method achieves average performance within 14.5% to 0.2% of POCAT without the need for hardware support. Shivam Kundan, Iraklis Anagnostopoulos |
ISCAS | 2 |
| 2020 | A Machine Learning Approach for Improving Power Efficiency on Clustered Multi-Processor SystemabstractModern embedded systems have adopted the clustered Chip Multi-Processor (CMP) paradigm in conjunction with dynamic frequency scaling techniques for improving application performance and power consumption. However, applications suffer from performance saturation due to resource contention. Thus, after a certain point, any frequency increase results only in power overheads without any performance gains. In this work, we present a run-time manager that focuses on power efficiency improvement for clustered CMPs. Specifically, it monitors the activity of concurrently executing applications and utilizes neural networks to select an appropriate frequency that keeps performance high, while reducing power consumption. Experimental results on the Odroid-XU3 board show that the proposed methodology improves power efficiency (MIPS/Watt) by 23% compared to Linux's performance governor with a negligible performance drop of only 3%. Shivam Kundan, Iraklis Anagnostopoulos |
ISCAS | 2 |
| 2020 | ARIAN: A Scalable Method for Adding aRbItrAry Numbers on Modern ProcessorsabstractHigh precision calculations that exceed the register's width on a computing system require arbitrary arithmetic. An application example where arbitrary long numbers are widely used is cryptography because longer numbers offer higher encryption security. Modern systems typically employ up to 64-bit registers, way less than what an arbitrary number requires, while conventional algorithms do not exploit hardware characteristics as well. In this paper, we propose ARIAN, a new scalable method to add arbitrary long numbers which utilizes logical operations rather than arithmetic to perform calculations. We also extended our algorithm (AR-AVX) to utilize AVX (Advanced Vector eXtensions) instructions to exploit parallelization and further increase calculation speed up. Experimental results show that the proposed methodology achieves a speed up of more than 120× on average, comparing to current implementations, while with the addition of AVX we achieve a 300× speed up on average. Konstantinos Poulos, Iraklis Anagnostopoulos, Themistoklis Haniotakis |
ISCAS | 2 |
| 2020 | NPU Thermal ManagementabstractNeural processing units (NPUs) are becoming an integral part in all modern computing systems due to their substantial role in accelerating neural networks (NNs). The significant improvements in cost-energy-performance stem from the massive array of multiply accumulate (MAC) units that remarkably boosts the throughput of NN inference. In this work, we are the first to investigate the thermal challenges that NPUs bring, revealing how MAC arrays, which form the heart of any NPU, impose serious thermal bottlenecks to on-chip systems due to their excessive power densities. For the first time, we explore: 1) the effectiveness of precision scaling and frequency scaling (FS) in temperature reductions and 2) how advanced on-chip cooling using superlattice thin-film thermoelectric (TE) open doors for new tradeoffs between temperature, throughput, cooling cost, and inference accuracy in NPU chips. Our work unveils that hybrid thermal management, which composes different means to reduce the NPU temperature, is a key. To achieve that, we propose and implement PFS-TE technique that couples precision and FS together with superlattice TE cooling for effective NPU thermal management. Using commercial signoff tools, we obtain accurate power and timing analysis of MAC arrays after a full-chip design is performed based on 14-nm Intel FinFET technology. Then, multiphysics simulations using finite-element methods are carried out for accurate heat simulations in the presence and absence of on-chip cooling. Afterward, comprehensive design-space exploration is presented to demonstrate the Pareto frontier and the existing tradeoffs between temperature reductions, power overheads due to cooling, throughput, and inference accuracy. Using a wide range of NNs trained for image classification, experimental results demonstrate that our novel NPU thermal management increases the inference efficiency (TOPS/Joule) by 1.33×, 1.87×, and 2× under different temperature constraints; 105 °C, 85 °C, and 70 °C, respectively, while the average accuracy drops merely from 89.0% to 85.5%. Hussam Amrouch, Georgios Zervakis 0001, Sami Salamin, Hammam Kattan, Iraklis Anagnostopoulos, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | A Message-Passing Microcoded Synchronization for Distributed Shared Memory ArchitecturesabstractImplementation of concurrent data structures in architectures that provide limited synchronization primitives is a critical challenge. Typical lock-based implementations suffer from well-known problems such as poor scalability and unfairness. In this paper, we propose a client-server based synchronization model that can be applied in data structures with low level of parallelism for distributed shared memory many-core systems that support also message-passing communication. Additionally, we utilize a programmable hardware accelerator with appropriate application interfaces to overcome the performance-flexibility dilemma. Experimental results show that the proposed work performs 20$\times$ faster than the single lock model with 88$\times$ less idle cycles and 7$\times$ less power consumption. Zois-Gerasimos Tasoulas, Iraklis Anagnostopoulos, Lazaros Papadopoulos, Dimitrios Soudris |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | SPA: Simple pool architecture for application resource allocation in many-core systemsabstractThe technology push by Moore's law brings a paradigm shift in the adaption of many core systems which replace high frequency superscalar processors with many simpler ones. On the software side, in order to utilize the available computational power, applications are following the high performance parallel/multi-threading model. Thus, many-core systems raise the challenges of resource allocation and fragmentation making necessary efficient run-time resource management techniques. In this paper, we propose SPA, a Simple Pool Architecture for managing resource allocation in many-core systems. The proposed framework follows a distributed approach in which cores are organized into clusters and multiple clusters form a pool. Clusters are created based on system's characteristics and the allocation of cores is performed in a distributed manner so as to take advantage of spatial features, shared resources and reduce scattering of cores. Experimental results show that SPA produces on average 15% better application response time while waiting time is reduced by 45% on average compared to other state-of-art methodologies. Jayasimha Sai Koduri, Iraklis Anagnostopoulos |
DATE | 2 |
| 2018 | Throughput optimization and resource allocation on GPUs under multi-application executionabstractPlatform heterogeneity prevails as a solution to the throughput and computational challenges imposed by parallel applications and technology scaling. Specifically, Graphics Processing Units (GPUs) are based on the Single Instruction Multiple Thread (SIMT) paradigm and they can offer tremendous speedup for parallel applications. However, GPUs were designed to execute a single application at a time. In case of simultaneous multi-application execution, due to the GPUs' massive multi-threading paradigm, applications compete against each other using destructively the shared resources (caches and memory controllers) resulting in significant throughput degradation. In this paper, a methodology for minimizing interference in shared resources and provide efficient concurrent execution of multiple applications on GPUs is presented. Particularly, the proposed methodology (i) performs application classification; (ii) analyzes the per-class interference; (iii) finds the best matching between classes; and (iv) employs an efficient resource allocation. Experimental results showed that the proposed approach increases the throughput of the system for two concurrent applications by an average of 36% compared to the default execution and 10% compared to an exahustive profile-based optimization technique. Srinivasa Reddy Punyala, Theodoros Marinakis, Arash Komaee, Iraklis Anagnostopoulos |
DATE | 4 |
| 2018 | Weather-based road condition estimation in the era of Internet-of-Vehicles (IoV)abstractModern high-end vehicles are equipped with advanced embedded computing elements and wireless connectivity capabilities. All these new features can provide new services, introducing the smart car era. This evolution has created a large network that includes cars as preliminary entities, extending the Internet of Things (IoT) to what is often referred as Internet-of-Vehicles (IoV). Thus, smart cars rise the challenge of efficient utilization of the available computational, sensing and connection potential. This paper presents a systematic methodology for combining different information sources in order to estimate the status of the car and perform specific actions. As a use case, the proposed approach considers the estimation of the road condition utilizing weather information. Specifically, it informs the driver about a recommended (dynamic) speed limit that the car should adapt to. Ioannis Galanis, Priyaa Gurunathan, Dona Burkard, Iraklis Anagnostopoulos |
ISCAS | 4 |
| 2018 | Fog Computing and Efficient Resource Management in the era of Internet-of-Video Things (IoVT)abstractInternet-of-Things (IoT) consists of interconnected devices with sensing, monitoring and processing functionalities that work in a cooperative way to offer services. Smart buildings, self-driving cars, house monitoring and management, city electricity and pollution monitoring are some examples where IoT systems have been already deployed. Amongst different kinds of devices in IoT, cameras have a key role, since they can capture rich and resourceful content. The number of embedded cameras is rising rapidly establishing the term Internet-of-Video Things (IoVT). However, since multiple IoT devices share the same gateway, the data that is produced from high definition cameras struggle the network and the available computational resources, resulting in Quality-of-Service degradation regarding visual content. In this work, we propose a methodology that tries to balance the content generation rate of cameras in an IoT environment. Specifically, the targeted use case is face recognition for video surveillance under local storage, network utilization and computational constraints while achieving the highest possible accuracy. Sai Saketh Nandan Perala, Ioannis Galanis, Iraklis Anagnostopoulos |
ISCAS | 3 |
| 2018 | A Hierarchical Distributed Runtime Resource Management Scheme for NoC-Based Many-CoresabstractAs technology constantly strengthens its presence in all aspects of human life, computing systems integrate a high number of processing cores, whereas applications become more complex and greedy for computational resources. Inevitably, this high increase in processing elements combined with the unpredictable resource requirements of executed applications at design time impose new design constraints to resource management of many-core systems, turning the distributed functionality into a necessity. In this work, we present a distributed runtime resource management framework for many-core systems utilizing a network-on-chip (NoC) infrastructure. Specifically, we couple the concept of distributed management with parallel applications by assigning different roles to the available computing resources. The presented design is based on the idea of local controllers and managers, whereas an on-chip intercommunication scheme ensures decision distribution. The evaluation of the proposed framework was performed on an Intel Single-Chip Cloud Computer, an actual NoC-based, many-core system. Experimental results show that the proposed scheme manages to allocate resources efficiently at runtime, leading to gains of up to 30% in application execution latency compared to relevant state-of-the-art distributed resource management frameworks. Vasileios Tsoutsouras, Iraklis Anagnostopoulos, Dimosthenis Masouros, Dimitrios Soudris |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | Performance-Aware Resource Management of Multi-Threaded Applications on Many-Core SystemsabstractModern computing systems employ a large number of processing elements leaving behind traditional design approaches and architectures. On the software side, this evolution in system architecture has driven rapid changes on the field of application development too, by increasing the usage of highly parallel/multi-threading and demanding applications. Thus, many-core systems raise the challenge of efficient resource management especially in cases where changes occur at run-time with a rapid pace. In this paper, a performance-aware resource management scheme for many-core architectures is presented. Particular, the developed framework takes as input parallel applications and performs an application profiling. Based on that profile information, a thread to core mapping algorithm finds (i) the appropriate number of threads that this application will have in order to maximize the utilization of the system; and (ii) the best mapping for maximizing the performance of the application. Experimental results showed that our mapping framework produces on average 23% and 18% better application turnaround time compared to another state-of-art run-time manager. Daniel Olsen, Iraklis Anagnostopoulos |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Application resource management for exploitation of non-volatile memory in many-core systemsabstractCurrent trends in many-core system design show an increasing desire to integrate several cores and accelerators on a single chip in order to support better parallel applications. Moreover, the increasing number of parallel applications and tasks result in heavy memory utilization. Non-Volatile Memories (NVMs) prevail as a replacement of DRAM due to low leakage power and higher density. However, NVMs support limited number of write operations and consequently suffer from short lifetime issues. Additionally, non-uniform distribution of write operations on NVMs has worsen this situation. Most of the state-of-the-art approaches improve NVM lifetime by re-designing the system at the architecture level. In this paper, we propose an application mapping framework for parallel applications exploiting and improving the lifetime of NVMs in distributed shared memory systems. Results show an improvement of ×3.2 in lifetime and a reduction of 60% in write variation. Setareh Behroozi, Iraklis Anagnostopoulos |
ISCAS | 2 |
| 2017 | A multi-agent based system for run-time distributed resource managementabstractModern embedded systems tend to employ a plethora of inter-connected components leaving behind complex superscalar/centralized approaches. This evolution in system architecture has driven rapid changes on the field of application development too, by increasing the usage of complex and demanding applications. Thus, multi-agent systems raise the challenge of efficient resource management especially in cases where changes occur at run-time with a rapid pace. This paper couples the concept of multi-agent systems with run-time resource management techniques in order to develop a distributed framework for run-time management of multiagent systems. The proposed framework is based on local runtime services and distributed agents in order to provide self-management functions while keeping system requirements. The motivation for this work is coming from the automotive industry and specifically from the Infotainment domain. Ioannis Galanis, Daniel Olsen, Iraklis Anagnostopoulos |
ISCAS | 3 |
| 2017 | An efficient and fair scheduling policy for multiprocessor platformsabstractScheduling is a decision-making process that deals with the assignment of resources to tasks over given periods, aiming to optimize one or more objectives. Responsible for efficient distribution of the CPU time among the processes, scheduler has become an essential part of computer systems. While applications run on neighboring cores of a many-core system, they compete with each other for the shared resources (cache, memory etc.). This contention can result in great performance degradation for the applications that are concurrently executed. For this reason, treating the cores of a many-core systems as isolated and independent units is a very optimistic abstraction and can cause great problems to the objectives a scheduler tries to optimize. This paper presents a scheduler that focuses on improving the system's fairness by deciding the group of applications that will be executed together based on the progress they have performed. Results shows that the proposed scheduler achieves on average 86% fairness improvement compared to two state-of-art schedulers. Theodoros Marinakis, Alexandros-Herodotos Haritatos, Konstantinos Nikas, Georgios I. Goumas, Iraklis Anagnostopoulos |
ISCAS | 5 |
| 2013 | Distributed run-time resource management for malleable applications on many-core platformsabstractTodays prevalent solutions for modern embedded systems and general computing employ many processing units connected by an on-chip network leaving behind complex superscalar architectures In this paper, we couple the concept of distributed computing with parallel applications and present a workload-aware distributed run-time framework for malleable applications on many-core platforms. The presented framework is responsible for serving in a distributed way and at run-time, the needs of malleable applications, maximizing resource utilization avoiding dominating effects and taking into account the type of processors supporting platform heterogeneity, while having a small overhead in overall inter-core communication. Our framework has been implemented as part of a C simulator and additionally as a run-time service on the Single-Chip Cloud Computer (SCC), an experimental processor created by Intel Labs, and we compared it against a state-of-art run-time resource manager. Experimental results showed that our framework has on average 70% less messages, 64% smaller message size and 20% application speed-up gain. Iraklis Anagnostopoulos, Vasileios Tsoutsouras, Alexandros Bartzas, Dimitrios Soudris |
DAC | 1 |
| 2013 | Power-aware dynamic memory management on many-core platforms utilizing DVFSabstractToday multicore platforms are already prevalent solutions for modern embedded systems. In the future, embedded platforms will have an even more increased processor core count, composing many-core platforms. In addition, applications are becoming more complex and dynamic and try to efficiently utilize the amount of available resources on the embedded platforms. Efficient memory utilization is a key challenge for application developers, especially since memory is a scarce resource and often becomes the system's bottleneck. To cope with this dynamism and achieve better memory footprint utilization (low memory fragmentation) application developers resort to the usage of dynamic memory (heap) management techniques, by allocating and deallocating data at runtime. Moreover, overall power consumption is another key challenge that needs to be taken into consideration. Towards this, designers employ the usage of Dynamic Voltage and Frequency Scaling (DVFS) mechanisms, adapting to the application's computational demands at runtime. In this article, we propose the combination of dynamic memory management techniques with DVFS ones. This is performed by integrating, within the memory manager, runtime monitoring mechanisms that steer the DVFS mechanisms to adjust clock frequency and voltage supply based on heap performance. The proposed approach has been evaluated on a distributed shared-memory many-core platform composed of multiple LEON3 processors interconnected by a Network-on-Chip infrastructure, supporting DVFS. Experimental results show that by using the proposed method for monitoring and applying DVFS mechanisms the power consumption concerning dynamic memory management was reduced by approximately 37%. In addition we present the trade-offs the proposed approach. Last, by combining the developed method with heap fragmentation-aware dynamic memory managers, we achieve low heap fragmentation values combined with low power consumption. Iraklis Anagnostopoulos, Jean-Michel Chabloz, Ioannis Koutras, Alexandros Bartzas, Ahmed Hemani, Dimitrios Soudris |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | A divide and conquer based distributed run-time mapping methodology for many-core platformsabstractReal-time applications are raising the challenge of unpredictability. This is an extremely difficult problem in the context of modern, dynamic, multiprocessor platforms which, while providing potentially high performance, make the task of timing prediction extremely difficult. In this paper, we present a flexible distributed run-time application mapping framework for both homogeneous and heterogeneous multi-core platforms that adapts to application's needs and application's execution restrictions. The novel idea of this article is the application of autonomic management paradigms in a decentralized manner inspired by Divide-and-Conquer (D&C) method. We have tested our approach in a Leon-based Network-on-Chip platform using both synthetic and real application workload. Experimental results showed that our mapping framework produces on average 21% and 10% better on-chip communication cost for homogeneous and heterogeneous platform respectively. Iraklis Anagnostopoulos, Alexandros Bartzas, Georgios Kathareios, Dimitrios Soudris |
DATE | 1 |