Lev Mukhanov

dblp:135/5385 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-0119-4359ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 LLM-VeriOpt: Verification-Guided Reinforcement Learning for LLM-Based Compiler Optimization
abstract
Large Language Models (LLMs) for compiler optimization have recently emerged as a frontier research direction, with many studies demonstrating their potential to automate and improve low-level code transformations. While various techniques have been proposed to enhance LLMs’ ability to optimize LLVM IR or assembly code, ensuring the semantic equivalence of transformed instructions remains a fundamental prerequisite for safe and effective performance improvement. At the same time, code generated by LLMs is often so far away from being correct that it is very difficult to work out how to improve them to proceed in generating optimizations using the output.In this work, we present LLM-VeriOpt, a novel reinforcement-learning methodology that incorporates feedback from a formal verifier, Alive2, to guide the training of a small-scale model, Qwen-3B. This facilitates the use of Guided Reinforcement via Group Relative Policy Optimization (GRPO), using semantic-equivalence signals from the Alive2 formal verification tool as part of the reward function. This allows the model to self-correct based on observing and subsequently learning to give correctness feedback during training, giving high code coverage by successfully transforming large amounts of code, while also optimizing it significantly.We demonstrate our technique by designing an LLM-based peephole optimizer over LLVM-IR. Our method significantly improves the correctness of IR optimizations versus the base LLM Qwen-3B applied with just a prompt and no fine-tuning — achieving a 5.4× improvement in code successfully modified. The resulting model produces verifiably correct output 90% of the time, comfortably outperforming larger state-of-the-art LLMs, including Meta’s LLM Compiler. This yields speedups of 2.3× over O0-optimized code, comparable to the handwritten LLVM-instcombine pass, and producing emergent optimizations that outperform it in 20% of cases.
Xiangxin Fang, Jiaqin Kang, Rodrigo Rocha, Sam Ainsworth 0001, Lev Mukhanov
CGO5
2025 A Deep Technical Review of nZDC Fault Tolerance
abstract
Faults within CPU circuits, which generate incorrect results and thus silent data corruption, have become endemic at scale. The only generic techniques to detect one-time or intermittent soft errors, such as particle strikes or voltage spikes, require redundant execution, where copies of each instruction in a program are executed twice and compared. The only software solution for this task that is open source and available for use today is nZDC, which aims to achieve “near-zero silent data corruption” through control- and data-flow redundancy. However, when we tried to apply this to large-scale workloads, we found it suffered a wide set of false positives, negatives, compiler bugs and run-time crashes, which meant it was impossible to benchmark against. This document details the wide set of fixes and workarounds we had to put in place to make nZDC work across full suites. We provide many new insights as to the edge cases that make such instruction duplication tricky under complex ISAs such as AArch64 and their similarly complex ABIs. Evaluation across SPECint 2006 and PARSEC with our extensions takes us from no workloads executing to all bar four, with 2× and 1.6× geomean overhead respectively relative to execution with no fault tolerance.
Minli Julie Liao, Sam Ainsworth 0001, Lev Mukhanov, Timothy M. Jones 0001
CC3
2025 Parallaft: Runtime-Based CPU Fault Tolerance via Heterogeneous Parallelism
abstract
The increasing vulnerability of microprocessors due to frequent silicon faults greatly exacerbates the risks of silent data corruption. Existing software-based schemes to detect these suffer from high power and performance overhead. State-of-the-art hardware fault-tolerance techniques exploit processor heterogeneity to minimize power, performance, and area overhead, but have not seen deployment in production due to their complexity. This paper shows for the first time that the same insights of heterogeneous parallelism can be repurposed without any hardware support. We present Parallaft, a parallel software-based error detection technique taking the insights of state-of-the-art hardware techniques and repurposing them with tools more suited to the hardware of today, such as copy-on-write checkpointing, dirty-page tracking and performance-counter synchronization. This allows error checking to be offloaded to little cores of an Apple M2 heterogeneous processor, achieving half energy cost while maintaining performance comparable to the homogeneous duplication mechanism RAFT.
Boyue Zhang 0003, Sam Ainsworth 0001, Lev Mukhanov, Timothy M. Jones 0001
CGO3
2025 ParaVerser: Harnessing Heterogeneous Parallelism for Affordable Fault Detection in Data Centers
abstract
Data-Center operators have awoken to the fact that silent data corruption resulting from defective silicon compute units is endemic at scale. Software scanners have been deployed to mitigate the issue, but either have low coverage or take months, leaving long windows of incorrect behaviour. By contrast, the redundancy mechanisms used in automotive double the required power and area, so cannot be practically deployed in server-space.We present ParaVerser, a high-coverage, low-overhead solution to hardware-level error detection in servers. Through minor architectural modifications, we enable conventional cores in heterogeneous server-grade processors to act as checker cores, thus exploiting heterogeneity, frequency scaling and the inherent parallelism in repeat runs to provide energy-efficient error checking. By dynamically coupling big.LITTLE-style out-of-order superscalar cores with in-order ones, we reduce energy overheads relative to a typical lockstep system by 70% with identical guarantees, at only 4.3% performance degradation, and 1064B per-core area overhead.
Minli Julie Liao, Sam Ainsworth 0001, Lev Mukhanov, Adrián Barredo, Markos Kynigos, Timothy M. Jones 0001
DSN3
2025 Unlocking the power of 4G/5G mobile networks: An empirical dive into quality and energy efficiency in YouTube Edge services
abstract
The advancements in 5G mobile networks and Edge computing offer great potential for services like augmented reality and Cloud gaming, thanks to their low latency and high bandwidth capabilities. However, the practical limitations of achieving optimal latency on real applications remain uncertain. This paper investigates the latency, bandwidth, and energy consumption of 5G Networks and leverages YouTube Edge service as the practical use case. We analyze how latency, bandwidth, and energy consumption differ between 4G LTE and 5G networks and how the location of YouTube Edge servers impacts these metrics. Surprisingly, our observations show that the 5G ecosystems have average latency hikes of up to 2 × , demonstrating that they are far from achieving their proclaimed promises. Our research study reveals over 10 significant observations and implications, indicating that the primary constraints on 4G/LTE and 5G capabilities are the ecosystem and energy efficiency of mobile devices’ downstream data. Moreover, our study demonstrates that to unlock the potential of 5G and its applications fully, it is crucial to prioritize efforts to improve the 5G ecosystem and introduce better methods and techniques to enhance energy efficiency.
Peixuan Song, Ahmed M. Abdelmoniem, Lev Mukhanov
Comput. Networks4
2024 Triangel: A High-Performance, Accurate, Timely On-Chip Temporal Prefetcher
abstract
Temporal prefetching, where correlated pairs of addresses are logged and replayed on repeat accesses, has recently become viable in commercial designs. Arm’s latest processors include Correlating Miss Chaining prefetchers, which store such patterns in a partition of the on-chip cache. However, the state-of-the-art on-chip temporal prefetcher in the literature, Triage, features some design inconsistencies and inaccuracies that pose challenges for practical implementation. We first examine and design fixes for these inconsistencies to produce an implementable baseline. We then introduce Triangel, a prefetcher that extends Triage with novel sampling-based methodologies to allow it to be aggressive and timely when the prefetcher is able to handle observed long-term patterns, and to avoid inaccurate prefetches when less able to do so. Triangel gives a 26.4% speedup compared to a baseline system with a conventional stride prefetcher alone, compared with 9.3% for Triage at degree 1 and 14.2% at degree 4. At the same time Triangel only increases memory traffic by 10% relative to baseline, versus 28.5% for Triage.
Sam Ainsworth 0001, Lev Mukhanov
ISCA2
2023 ARETE: Accurate Error Assessment via Machine Learning-Guided Dynamic-Timing Analysis
abstract
Nanometer circuits are increasingly prone to timing errors, escalating the need forfault injectionframeworks to accurately evaluate their impact on applications. In this paper, we propose ARETE, a novel cross-layer, fault-injection framework that combines dynamic-binary instrumentation with machine learning-guided dynamic-timing analysis. ARETE enables accurate fault-injection into any application by estimating the location of the injecting errors via dynamic-timing analysis. To accelerate fault-injection, we develop a novel, data-aware, machine learning-based mechanism that dynamically pre-selects the error-prone instructions and limits the application of the costly dynamic-timing analysis only to them. To evaluate ARETE's accuracy, our fully automated toolflow is configured to support fault-injection based on detailed post-layout gate-level simulations as well as via existing workload-agnostic error models. Our results for various workloads, including an autonomous-driving library, show that the location and time of injected errors performed by ARETE, is 89.9% consistent with fault-injection based on full gate-level simulation. On average, ARETE executes 84.6× faster than gate-level simulation and at a cost of 3.4% loss in the program output quality estimation. When compared to the existing statistical fault-injection tools that are based on workload-agnostic error models, ARETE improves the accuracy of fault-injection rate and output quality estimation by 143.9% and 40.4% on average, respectively.
Ioannis Tsiokanos, Styliani Tompazi, Giorgis Georgakoudis, Lev Mukhanov, Georgios Karakonstantis
IEEE Trans. Computers4
2022 Instruction-aware Learning-based Timing Error Models through Significance-driven Approximations
abstract
The adoption of aggressively down-scaled voltages along with worsening process variations, render nanometer devices prone to timing errors that threaten system functionality. The increased vulnerability of nanometer circuits to these errors attracted recent efforts in the development of timing error prediction models using machine learning (ML) methods. However, the majority of such models may be inaccurate, since they either neglect important microarchitecture properties and workload-dependent parameters, affecting timing error manifestation, or are constrained to limited operating areas. In this paper, we propose microarchitecture- and workload-aware ML models for timing error prediction that jointly consider various instruction types as well as all in-flight instructions in a pipeline. Our proposed models are able to predict the exact time and location (i.e., cycle, instruction and bit position) of timing errors with over 98% accuracy across multiple, critical operating regions. To circumvent the increased model complexity due to the considered features, we apply for the first time significance-driven approximations. Evaluation results for various workloads and voltage reduction levels show that our significance-driven precision scaling improves the models’ inference time up to 4.66×, with less than 4% accuracy loss. Finally, we use the proposed model to accurately and realistically inject timing errors during the evaluation of application resiliency. When compared to prior timing error evaluation frameworks that rely on workload-agnostic models, our framework improves the output quality estimation up to 82.6%.
Styliani Tompazi, Ioannis Tsiokanos, Jesús Martínez del Rincón, Lev Mukhanov, Georgios Karakonstantis
ICCD4
2022 On the Evaluation of the Total-Cost-of-Ownership Trade-Offs in Edge vs Cloud Deployments: A Wireless-Denial-of-Service Case Study
abstract
We are witnessing an explosive growth in the number of Internet-connected devices and the emergence of several new classes of Internet of Things (IoT) applications that require rapid processing of an abundance of data. To overcome the resulting need for more network bandwidth and low network latency, a new paradigm has emerged that promotes the offering of Cloud services at the Edge, closer to users. However, the Edge is a highly constrained environment with limited power budget for servers per Edge installation which, in turn, limits the number of Internet-connected devices, such as sensors, that an installation can service. Consequently, the limited number of sensors leads to a reduction in the area coverage provided by them and puts in question the effectiveness for deploying IoT applications at the Edge. In this paper, we investigate the benefits of running an emerging security focused IoT application, (jamming detection), at the Edge vs. the Cloud by developing a Total Cost of Ownership (TCO) model, which considers the application's requirements as well as the Edge's constraints. For the first time, we build such a model based on realistic performance and energy-efficiency measurements obtained from commodity 64-bit ARM based micro-servers that are excellent candidates for supporting Cloud services at the Edge. Such servers represent the type of devices that can provide the right balance between power and performance, without requiring any complicate cooling and power supply infrastructure, which will not be available at the de-centralized deployments. Aiming at improving the energy efficiency, we exploit the pessimistic design margins adopted conventionally in such devices and investigate their operation under lower than nominal supply voltage and memory refresh-rate. Our results show that the jamming detection application deployed at an Edge environment is superior to a Cloud based solution by up to 2.13 times in terms of TCO. Moreover, when servers operate below nominal conditions, we can achieve up to 9 percent power savings which enables in several situations 100 percent gains in the TCO/area-coverage metric, i.e double area can be served with the same TCO.
Panagiota Nikolaou, Yiannakis Sazeides, Alejandro Lampropulos, Denis Guilhot, Andrea Bartoli, George Papadimitriou 0001, Athanasios Chatzidimitriou, Dimitris Gizopoulos, Konstantinos Tovletoglou, Lev Mukhanov, Georgios Karakonstantis
IEEE Trans. Sustain. Comput.10
2021 Revealing DRAM Operating GuardBands Through Workload-Aware Error Predictive Modeling
abstract
Improving the energy efficiency of DRAMs becomes very challenging due to the growing demand for storage capacity and failures induced by the manufacturing process. To protect against failures, vendors adopt conservative margins in the refresh period and supply voltage. Previously, it was shown that these margins are too pessimistic and will become impractical due to high-power costs, especially in future DRAM technologies. In this article, we present a new technique for automatic scaling the DRAM refresh period under reduced supply voltage that minimizes the probability of failures. The main idea behind the proposed approach is that DRAM error behavior is workload-dependent and can be predicted based on particular program inherent features. We use a Machine Learning (ML) method to build a workload-aware DRAM error behavior model based on the program features which we extract from real workloads during our DRAM error characterization campaign. With such a model, we identify the marginal value of the DRAM refresh period under relaxed voltage for each DRAM module of a server that enable us to reduce the DRAM power. We implement a temperature-driven OS governor which automatically sets the module-specific marginal DRAM parameters discovered by the ML model. Our governor reduces the DRAM power by 24 percent on average while minimizing the probability of failures. Unlike previous studies, our technique: i) does not require intrusive changes to hardware; ii) is implemented on a real server; iii) uses a mechanism that prevents any abnormal DRAM error behavior; and iv) can be easily deployed in data centers.
Lev Mukhanov, Konstantinos Tovletoglou, Hans Vandierendonck, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IEEE Trans. Computers1
2020 HaRMony: Heterogeneous-Reliability Memory and QoS-Aware Energy Management on Virtualized Servers
abstract
The explosive growth of data increases the storage needs, especially within servers, making DRAM responsible for more than 40% of the total system power. Such a reality has made researchers focus on energy saving schemes that relax the pessimistic DRAM circuit parameters at the cost of potential faults. In an effort to limit the resultant risk of critical data disruption, new methods were introduced that split DRAM into domains with varying reliability and power. The benefits of such schemes may have been showcased on simulators but have neither been implemented on real systems with a complete software stack, nor have been combined with any energy-reliability OS management policies. In this paper, we are the first to implement and evaluate HaRMony, a heterogeneous-reliability memory framework, in conjunction with QoS-aware energy management policies on a server with a complete virtualization stack. HaRMony overcomes the practical restrictions stemming from default hardware specifications, which were neglected in prior works, by introducing a software-based memory interleaving scheme. Furthermore, we expose the capabilities of HaRMony to the QEMU-KVM hypervisor through two unique policies. The first policy enables the hypervisor to seek the most power efficient DRAM circuit parameters based on the server availability requested by the user. The second policy enables users to exploit the inherent application error-resiliency by allowing them to limit the error protection mechanisms and allocate data structures on variably-reliable memory domains. Our evaluation shows that HaRMony reduces the performance overhead incurred due to disabling hardware interleaving from 29.3% down to 1.1% and leads to 17.7% DRAM energy savings and 8.6% total system energy savings on average in case of native execution of 28 benchmarks on an ARMv8-based server. Finally, we demonstrate that our QoS-aware scaling governor integrated with QEMU-KVM can dynamically scale the DRAM parameters, while reducing the system energy by 8.4% and meeting the targeted QoS even under extreme temperatures.
Konstantinos Tovletoglou, Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
ASPLOS2
2020 Increasing the Profit of Cloud Providers through DRAM Operation at Reduced Margins
abstract
Energy reduction is a key objective in cloud computing, and DRAM memories are responsible for an important amount of the energy consumption of data center nodes. Vendors adopt very conservative margins for DRAM operating parameters, such as the refresh rate and supply voltage, to guarantee correct operation even under the worst process variation and operating conditions. In this paper, we investigate the exploitation of DRAM margins to improve the energy efficiency of data center nodes, without triggering penalties due to service level agreement (SLA) violations. We introduce a model that captures the most important aspects of job management and system configuration. We also introduce RM-DRAM, a scheduling and node configuration policy that exploits the extended margins of DRAMs to reduce the operator's cost, considering the tradeoff between the cost of energy consumption and potential SLA violations. RM-DRAM also employs cost-aware (rather than threshold-based) VM consolidation. We extract the parameters used in the simulation (particularly power consumption and error rates) by characterizing a commercial ARM-based server. We perform simulations to evaluate the effectiveness of our approach, showing that significant gains, up to 34.84% and 29.53% in energy and cost, respectively, can be achieved compared with a state-of-the-art policy.
Christos Kalogirou, Christos D. Antonopoulos, Nikolaos Bellas, Spyros Lalis, Lev Mukhanov, Georgios Karakonstantis
CCGRID5
2020 DEFCON: Generating and Detecting Failure-prone Instruction Sequences via Stochastic Search
abstract
The increased variability and adopted low supply voltages render nanometer devices prone to timing failures, which threaten the functionality of digital circuits. Recent schemes focused on developing instruction-aware failure prediction models and adapting voltage/frequency to avoid errors while saving energy. However, such schemes may be inaccurate when applied to pipelined cores since they consider only the currently executed instruction and the preceding one, thereby neglecting the impact of all the concurrently executing instructions on failure occurrence. In this paper, we first demonstrate that the order and type of instructions in sequences with a length equal to the pipeline depth affect significantly the failure rate. To overcome the practically impossible evaluation of the impact of all possible sequences on failures, we present DEFCON, a fully automated framework that stochastically searches for the most failure-prone instruction sequences (ISQs). DEFCON generates such sequences by integrating a properly formulated genetic algorithm with accurate post-layout dynamic timing analysis, considering the data-dependent path sensitization and instruction execution history. The generated micro-architecture aware ISQs are then used by DEFCON to estimate the failure vulnerability of any application. To evaluate the efficacy of the proposed framework, we implement a pipelined floating-point unit and perform dynamic timing analysis based on input data that we extract from a variety of applications consisting of up-to 43.5M ISQs. Our results show that DEFCON reveals quickly ISQs that maximize the output quality loss and correctly detects 99.7% of the actual faulty ISQs in different applications under various levels of variation-induced delay increase. Finally, DEFCON enable us to identify failure-prone ISQs early at the design cycle, and save 26.8% of energy on average when combined with a clock stretching mechanism.
Ioannis Tsiokanos, Lev Mukhanov, Giorgis Georgakoudis, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
DATE2
2020 DStress: Automatic Synthesis of DRAM Reliability Stress Viruses using Genetic Algorithms
abstract
Failures become inevitable in DRAM devices, which is a major obstacle for scaling down the density of cells in future DRAM technologies. These failures can be detected by specific DRAM tests that implement the data and memory access patterns having a strong impact on DRAM reliability. However, the design of such tests is very challenging, especially for testing DRAM devices in operation, due to an extremely large number of possible cell-to-cell interference effects and combinations of patterns inducing these effects.In this paper, we present a new framework for the synthesis of DRAM reliability stress viruses, DStress. This framework automatically searches for the data and memory access patterns that induce the worst-case DRAM error behavior regardless the internal DRAM design. The search engine of our framework is based on Genetic Algorithms (GA) and a programming tool that we use to specify the patterns examined by GA. To evaluate the effect of program viruses on DRAM reliability, we integrate DStress with an experimental server where 72 DRAM chips can operate under various operating parameters and temperatures.We present the results of our 7-month experimental study on the search of DRAM reliability stress viruses. We show that DStress finds the worst-case data pattern virus and the worst-case memory access virus with probabilities of 1 - 4 × 10-7and 0.95, respectively. We demonstrate that the discovered patterns induce by at least 45% more errors than the traditional data pattern micro-benchmarks used in previous studies. We show that DStress enables us to detect the marginal DRAM operating parameters reducing the DRAM power by 17.7 % on average without compromising reliability. Overall, our framework facilitates the exploration of new data patterns and memory access scenarios increasing the probability of DRAM errors, which is essential for improving the state-of-the-art DRAM testing mechanisms.
Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
MICRO1
2019 Low-Power Variation-Aware Cores based on Dynamic Data-Dependent Bitwidth Truncation
abstract
Increasing variability of transistor parameters in nanoscale era renders modern circuits prone to timing failures. To address such failures, designers adopt pessimistic timing/voltage guardbands, which are estimated under rare worst-case conditions, thus leading to power and performance overheads. Recent approximation schemes based on precision reduction may help to limit the incurred overheads, but the precision is reduced statically in all operations. This results in unnecessary quality loss, since these schemes neglect the fact that only few long latency paths (LLPs) may be prone to failures, and such paths may be activated rarely. In this paper, we propose a variation-aware framework that minimizes any quality loss by dynamically truncating the bitwidth only for operands triggering the LLPs. This is achieved by predicting at runtime the excitation of the LLPs based on the processed operands. The applied truncation, which we implement by setting a number of least-significant bits to a constant value of zero, can effectively reduce the delay of the excited LLPs, providing sufficient timing slack to avoid failures without using conservative guardbands. To facilitate the adoption of such a scheme within pipelined cores and limit the incurred overheads, we also shape the path distribution appropriately for isolating the LLPs in a single pipeline stage. Additionally, to evaluate the efficacy of our framework, we perform post-layout dynamic timing analysis based on real operands that we extract from a variety of applications. When applied to the implementation of an IEEE-754 compatible double precision floating-point unit (FPU) in a 45nm technology, our approach eliminates timing failures under 8% delay variations with no performance loss. Our design comes at a cost of up-to 4.48% power and 0.34% area overheads, while the occasional operand truncation incurs minimal quality-loss in terms of relative error, up-to 4.1 · 10-6. Finally, when compared to an FPU with pessimistic margins, our techniqutechniquee can save up-to 44.3% power.
Ioannis Tsiokanos, Lev Mukhanov, Georgios Karakonstantis
DATE2
2018 An energy-efficient and error-resilient server ecosystem exceeding conservative scaling limits
abstract
The explosive growth of Internet-connected devices will soon result in a flood of generated data, which will increase the demand for network bandwidth as well as compute power to process the generated data. Consequently, there is a need for more energy efficient servers to empower traditional centralized Cloud data-centers as well as emerging decentralized data-centers at the Edges of the Cloud. In this paper, we present our approach, which aims at developing a new class of micro-servers - the UniServer - that exceed the conservative energy and performance scaling boundaries by introducing novel mechanisms at all layers of the design stack. The main idea lies on the realization of the intrinsic hardware heterogeneity and the development of mechanisms that will automatically expose the unique varying capabilities of each hardware. Low overhead schemes are employed to monitor and predict the hardware behavior and report it to the system software. The system software including a virtualization and resource management layer is responsible for optimizing the system operation in terms of energy or performance, while guaranteeing non-disruptive operation under the extended operating points. Our characterization results on a 64-bit ARMv8 micro-server in 28nm process reveal large voltage margins in terms of Vmin variation among the 8 cores of the CPU chip, among three different sigma chips, and among different benchmarks with the potential to obtain up-to 38.8% energy savings. Similarly, DRAM characterizations show that refresh rate and voltage can be relaxed by 35x and 5%, respectively, leading to 23.2% power savings on average.
Georgios Karakonstantis, Konstantinos Tovletoglou, Lev Mukhanov, Hans Vandierendonck, Dimitrios S. Nikolopoulos, Peter Lawthers, Panos K. Koutsovasilis, Manolis Maroudas, Christos D. Antonopoulos, Christos Kalogirou, Nikolaos Bellas, Spyros Lalis, Srikumar Venugopal, Arnau Prat-Pérez, Alejandro Lampropulos, Marios Kleanthous, Andreas Diavastos, Zacharias Hadjilambrou, Panagiota Nikolaou, Yiannakis Sazeides, Pedro Trancoso, George Papadimitriou 0001, Manolis Kaliorakis, Athanasios Chatzidimitriou, Dimitris Gizopoulos, Shidhartha Das
DATE3
2018 DRAM Characterization under Relaxed Refresh Period Considering System Level Effects within a Commodity Server
abstract
Today's rapid generation of data and the increased need for higher memory capacity has triggered a lot of studies on aggressive scaling of refresh period, which is currently set according to rare worst case conditions. Such studies analysed in detail the data-dependent circuit level factors and indicated the need for online DRAM characterization due to the variable cell retention time. They have done so by executing few test data patterns on FPGAs under controlled temperatures by using thermal testbeds, which however cannot be available in the field. Moreover, the existing studies were not able to reveal any system level effects, which may be excited under the execution of workloads on real systems and directly or indirectly affect DRAM reliability. In this paper, we develop an experimental framework based on a state-of-the-art 64-bit ARM based server with Linux OS, in which we enabled the DRAM characterization under relaxed refresh period by executing conventional test data patterns as well as popular HPC and Cloud workloads. Our results indicate that common test patterns are ineffective in identifying error-prone locations at low DRAM temperatures. Furthermore, we reveal that there is a strong correlation between the SOC utilization and DRAM reliability. By exploiting such findings, we developed a benchmark, which can indirectly stress the DRAM temperature and thus used for characterization in the field without needing any complicated thermal equipment. Our study shows that the refresh period can be relaxed by 35 times on such a commodity system with all errors being corrected by the available error correcting codes, resulting in 11.5% power savings on average.
Lev Mukhanov, Konstantinos Tovletoglou, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IOLTS1
2018 Minimization of Timing Failures in Pipelined Designs via Path Shaping and Operand Truncation
abstract
The continuous scaling of transistor sizes and the increased static and dynamic parametric variations render nanometer circuits more prone to timing failures. To protect circuits from such failures, typically designers adopt pessimistic timing margins, which are estimated statically under rare worst-case conditions. In this paper, we aim at minimizing the timing failures, while avoiding such pessimistic margins by proposing an approach that initially minimizes the number of long latency paths within each processor pipeline stage and con- straints them in as few stages as possible. Such an approach, not only reduces the timing failures, but also limits the potential error prone locations to only few pipeline registers/stages. To further reduce these failures, we exploit the path excitation dependence on data patterns and we truncate the bit-width of the operands in the few remaining long latency paths by setting a number of least significant bits to a constant value zero. Such truncation may incur quality loss, but this can be controlled by carefully selecting the number of truncated bits and will be in any case less than the catastrophic loss that may be incurred under random timing failures. Additionally, our framework performs post- place and route dynamic timing analysis based on real operands that are extracted from a variety of applications, helping to estimate the dynamic timing failures, while considering the data dependent path excitation. When applied to an IEEE-754 compatible double precision Floating Point Unit (FPU), the proposed approach reduces the timing failures by 104.5× on average compared to a reference FPU design under an assumed 8.1% variation-induced worst-case path delay increase in a 45 nm process. Finally, results show that path shaping alone introduces an insignificant 0.25% area and 5.7% power overhead with no performance cost. The combination of path shaping with aggressive operand bit- width truncation leads to up-to 44.7% on average power savings due to the substantially reduced switching activity at a minimal quality loss.
Ioannis Tsiokanos, Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IOLTS2
2018 Variation-Aware Pipelined Cores through Path Shaping and Dynamic Cycle Adjustment: Case Study on a Floating-Point Unit
abstract
In this paper, we propose a framework for minimizing variation-induced timing failures in pipelined designs, while limiting any overhead incurred by conventional guardband based schemes. Our approach initially limits the long latency paths (LLPs) and isolates them in as few pipeline stages as possible by shaping the path distribution. Such a strategy, facilitates the adoption of a special unit that predicts the excitation of the isolated LLPs and dynamically allows an extra cycle for the completion of only these error-prone paths. Moreover, our framework performs post-layout dynamic timing analysis based on real operands that we extract from a variety of applications. This allows us to estimate the bit error rates under potential delay variations, while considering the dynamic data dependent path excitation. When applied to the implementation of an IEEE-754 compatible double precision floating-point unit (FPU) in a 45nm process technology, the path shaping helps to reduce the bit error rates on average by 2.71 x compared to the reference design under 8% delay variations. The integrated LLPs prediction unit and the dynamic cycle adjustment avoid such failures and any quality loss at a cost of up-to 0.61% throughput and 0.3% area overheads, while saving 37.95% power on average compared to an FPU with pessimistic margins.
Ioannis Tsiokanos, Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
ISLPED2
2017 ALEA: A Fine-Grained Energy Profiling Tool
abstract
Energy efficiency is becoming increasingly important, yet few developers understand how source code changes affect the energy and power consumption of their programs. To enable them to achieve energy savings, we must associate energy consumption with software structures, especially at the fine-grained level of functions and loops. Most research in the field relies on direct power/energy measurements taken from on-board sensors or performance counters. However, this coarse granularity does not directly provide the needed fine-grained measurements. This article presents ALEA, a novel fine-grained energy profiling tool based on probabilistic analysis for fine-grained energy accounting. ALEA overcomes the limitations of coarse-grained power-sensing instruments to associate energy information effectively with source code at a fine-grained level. We demonstrate and validate that ALEA can perform accurate energy profiling at various granularity levels on two different architectures: Intel Sandy Bridge and ARM big.LITTLE. ALEA achieves a worst-case error of only 2% for coarse-grained code structures and 6% for fine-grained ones, with less than 1% runtime overhead. Our use cases demonstrate that ALEA supports energy optimizations, with energy savings of up to 2.87 times for a latency-critical option pricing workload under a given power budget.
Lev Mukhanov, Pavlos Petoumenos, Zheng Wang 0001, Konstantinos Parasyris, Dimitrios S. Nikolopoulos, Bronis R. de Supinski, Hugh Leather
ACM Trans. Archit. Code Optim.1
2015 ALEA: Fine-Grain Energy Profiling with Basic Block Sampling
abstract
Energy efficiency is an essential requirement for all contemporary computing systems. We thus need tools to measure the energy consumption of computing systems and to understand how workloads affect it. Significant recent research effort has targeted direct power measurements on production computing systems using on-board sensors or external instruments. These direct methods have in turn guided studies of software techniques to reduce energy consumption via workload allocation and scaling. Unfortunately, direct energymeasurementsarehamperedbythelowpowersampling frequency of power sensors. The coarse granularity of power sensing limits our understanding of how power is allocated in systems and our ability to optimize energy efficiency via workload allocation. We present ALEA, a tool to measure power and energy consumption at the granularity of basic blocks, using a probabilistic approach. ALEA provides fine-grained energy profiling via statistical sampling, which overcomes the limitations of power sensing instruments. Compared to state-of-the-art energy measurement tools, ALEA provides finer granularity without sacrificing accuracy. ALEA achieves low overhead energy measurements with mean error rates between 1.4% and 3.5% in 14 sequential and parallel benchmarks tested on both Intel and ARM platforms. The sampling method caps execution time overhead at approximately 1%. ALEA is thus suitable for online energy monitoring and optimization. Finally, ALEA is a user-space tool with a portable, machine-independent sampling method. We demonstrate three use cases of ALEA, where we reduce the energy consumption of a k-means computational kernel by 37%, an ocean modeling code by 33%, and a ray tracing code by 6% compared to high-performance execution baselines, by varying the power optimization strategy between basic blocks.
Lev Mukhanov, Dimitrios S. Nikolopoulos, Bronis R. de Supinski
PACT1
2015 Power Capping: What Works, What Does Not
abstract
Peak power consumption is the first order design constraint of data centers. Though peak power consumption is rarely, if ever, observed, the entire data center facility must prepare for it, leading to inefficient usage of its resources. The most prominent way for addressing this issue is to limit the power consumption of the data center IT facility far below its theoretical peak value. Many approaches have been proposed to achieve that, based on the same small set of enforcement mechanisms, but there has been no corresponding work on systematically examining the advantages and disadvantages of each such mechanism. In the absence of such a study, it is unclear what is the optimal mechanism for a given computing environment, which can lead to unnecessarily poor performance if an inappropriate scheme is used. This paper fills this gap by comparing for the first time five widely used power capping mechanisms under the same hardware/software setting. We also explore possible alternative power capping mechanisms beyond what has been previously proposed and evaluate them under the same setup. We systematically analyze the strengths and weaknesses of each mechanism, in terms of energy efficiency, overhead, and predictable behavior. We show how these mechanisms can be combined in order to implement an optimal power capping mechanism which reduces the slowdown compared to the most widely used mechanism by up to 88%. Our results provide interesting insights regarding the different trade-offs of power capping techniques, which will be useful for designing and implementing highly efficient power capping in the future.
Pavlos Petoumenos, Lev Mukhanov, Zheng Wang 0001, Hugh Leather, Dimitrios S. Nikolopoulos
ICPADS2
2013 Dynamic memory access monitoring based on tagged memory
abstract
Software vulnerabilities become one of the top threats to world security in the coming decade. The most of such vulnerabilities are based on memory leaks and memory corruption. Many memory access monitoring tools exist, but most of them suffer from high overhead what makes it impossible to use such tools in the “real world” software projects. The goal of this research is to develop and investigate a memory access monitoring tool that detects memory leaks and corruption using tagged memory without reduce of an original application performance.
Mikhail A. Gorelov, Lev Mukhanov
PACT2