Ayse K. Coskun

dblp:79/3945 · also Ayse Kivilcim Coskun · DBLP profile ↗
← Back
111ranked-venue papers
18as first author
37since 2021 · last 2026
0000-0002-6554-088XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 97 · 18 first-author · 28 since 2021Software engineering, systems software and programming languages · 23 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCARLET: A Scalable OPCM-Based Accelerator for Transformer Inference with Tiled Crossbars
abstract
While transformer-based large language models (LLMs) have achieved state-of-the-art performance on a wide range of natural language processing tasks, their massive computational demands, especially during inference, pose a significant challenge. Photonic accelerators offer a promising solution, but existing designs struggle with the precision, dynamism, and storage requirements of modern LLMs. This paper introduces SCARLET, a hybrid photonic architecture that addresses these limitations through two key components. First, we design a high-density optical phase-change memory (OPCM) crossbar for static matrix multiplications, achieving 5.6× higher bit density and 86.43% lower energy compared to previous OPCM crossbar designs. Second, we introduce an approximate photonic floating-point multiplier to handle dynamic matrix multiplications and quantization steps by approximating floating-point computations with weighted integer sums, thus, eliminating the need for frequent memory reprogramming. Our evaluation on models with up to 13 billion parameters demonstrates significant performance improvements, including up to 17.15× and 8.45× lower latency during prefill and generation phases, respectively.
Sina Karimi, Guowei Yang 0005, Carlos A. Ríos Ocampo, Ajay Joshi, Ayse K. Coskun
DATE5
2026 Privacy-Preserving Data Center Demand Response Using Multi-party Computation
abstract
The rapid growth of AI has significantly increased data center energy demand, placing increasing pressure on power grids. Demand Response (DR) programs utilize the flexibility of power consumers, such as data centers, to help balance the supply and demand in power grids. In collaborative data center DR participation, multiple data centers share information with an external coordinator for improved power dispatch and quality of service for their workloads. However, disclosing sensitive information such as power usage and workload performance to an untrusted third party raises significant privacy concerns. To address this, we use Multi-Party Computation (MPC) to perform secure power dispatch without revealing sensitive inputs. Standard MPC has heavy communication overhead, which challenges real-time DR requirements. To meet real-time DR requirements, we optimize our system by tailoring fixed-point bitwidths and substituting expensive divisions with polynomial approximations and Newton-Raphson iterations. Evaluated using the MP-SPDZ library, our optimizations yield up to 33 $$\times $$ fewer communication rounds and a 45 $$\times $$ speedup, delivering sub-second latency for up to 16 data centers while matching the power dispatch accuracy of the baseline.
Seyda Nur Güzelhan, Fatih Acun, Can Hankendi, Ayse K. Coskun, Ajay Joshi
Euro-Par (1)4
2026 PrvTel: Lightweight Models for Private and Accurate Telemetry Data Retention
Fuheng Zhao, Eric S. Wang, Ayse K. Coskun, Divyakant Agrawal, Amr El Abbadi, Zaoxing Liu
NSDI4
2026 CarbonMeter: A Lightweight Framework for Achieving Sustainable and Cost-Efficient Data Centers
abstract
As both direct and embodied carbon emissions of data centers is expected to increase in coming years, estimating data center carbon footprints has become essential to understand and address the sustainability challenges. Due to the lack of publicly available data regarding data center operations and infrastructure, there is a need for tools that can provide carbon emission estimations with a minimal amount of data. To address this challenge, we present our open-source tool,CarbonMeter, which allows users to quickly generate carbon footprint estimates for a given data center, using only data center area and data center power capacity information.CarbonMeterallows users to easily adjust parameters that affect carbon emissions and provides a breakdown of operational and embodied footprint. Furthermore,CarbonMeterenables swift evaluation of workload migration strategies to optimize electricity costs, carbon emissions and renewable curtailment by using real-time signals in the Nordic region. We evaluate the benefits of workload migration decisions byCarbonMeteron multiple case studies encompassing data centers located in Norway, Denmark, and Germany. Our results show 6% to 21% electricity cost reductions, while reducing the carbon emissions up to 49%.
Can Hankendi, Ayse K. Coskun, Benjamin K. Sovacool
IEEE Trans. Computers2
2025 Fast Machine Learning Based Prediction for Temperature Simulation Using Compact Models
abstract
As transistor densities increase, managing thermal challenges in 3D IC designs becomes more complex. Traditional methods like finite element methods and compact thermal models (CTMs) are computationally expensive, while existing machine learning (ML) models require large datasets and a long training time. To address these challenges with the ML models, we introduce a novel ML framework that integrates with CTMs to accelerate steady-state thermal simulations without needing large datasets. Our approach achieves up to 70 × speedup over state-of-the-art simulators, enabling real-time, high-resolution thermal simulations for 2D and 3D IC designs.11This research was partially funded by the NSF CCF 2131127 grant
Mohammadamin Hajikhodaverdian, Sherief Reda, Ayse K. Coskun
DATE3
2025 Lessons Learned from Anomaly Detection in Chameleon Cloud
abstract
Cloud computing has become integral to modern technology infrastructure, supporting a wide range of services from e-commerce to AI applications. Chameleon is a large-scale, configurable testbed designed to enable edge-to-cloud research through full bare-metal provisioning, virtualization, and diverse hardware resources, which is built on a leading open source cloud platform OpenStack. However, monitoring Chameleon’s heterogeneous infrastructure is challenging, particularly across Open-Stack services and hardware components. Traditional threshold-based alerting methods struggle to keep up with the scale and complexity of such environments. In this work, we present an anomaly detection framework for OpenStack services in the Chameleon Cloud. We curate and publish the first dataset of resource usage metrics collected from OpenStack control plane services. We evaluate four state-of-the-art unsupervised multivariate time series models, namely TranAD, Prodigy, USAD, and OmniAnomaly, on this dataset and share key insights from deploying them. Our findings indicate that for our use case, while all models achieve high F1 scores, training with three days of healthy data effectively balances training cost and detection accuracy.
S. M. Qasim, Can Hankendi, Kate Keahey, Gianluca Stringhini, Ayse K. Coskun
IC2E6
2025 Tutorial: The Energy Cost of Privacy and Security
abstract
Security and privacy are key enablers (and often also a legal requirement) for a number of applications, including smart grid and smart cities, health care, data analytics, and personalized services. Because of this, research in the domain of security and privacy-preserving techniques is progressing at high pace. However, if on the one side the research community devoted large attention to the study of more efficient algorithms and the design of more efficient architectures implementing them, on the other, the energy cost and the energy implications of the use of these technologies have not yet been explored in the needed depth, with the majority of literature focusing on block ciphers. This tutorial exposes the community to the main current research results and best practices in this research area, and aims to foster the exchange of ideas between all the involved stakeholders. The tutorial presents the background and latest achievements in the field of energy assessment and reduction for security and privacy-preserving primitives. This tutorial covers the needed background on security algorithms, discusses their energy consumption, and presents, by means of relevant examples, how to design and implement security primitives that achieve a limited energy footprint. In particular, the focus is on two families of security primitives: block ciphers and privacy-preserving primitives. The tutorial introduces the basic concepts and the main algorithms belonging to these families, discusses recent advances in the domain, and presents in detail the energy consumption of these technologies and in their applications such as machine learning. Further, the tutorial will show optimizations that have been proposed to minimize the energy footprint of security primitives, with a particular focus on block ciphers, discussing also the design of lightweight and low-energy cryptographic algorithms. The tutorial concludes discussing open problems, limitations, and possible research directions. The tutorial is divided into three sections and will begin with a talk providing a detailed introduction of the needed concepts, to allow attendees not familiar with the topic to be able to successfully follow the whole tutorial. More in details, the sections are: • "Introduction to Security Primitives and Privacy Preserving Technologies". This talk will introduce the audience to the security primitives and the relevant privacy preserving technologies and protocols, [1] that will be analyzed in the rest of the tutorial. • "Energy Assessment of Security Primitives". This talk summarizes current research in energy assessment [2] of security primitives, reporting the method used to assess them and presenting literature result on the energy consumption of security primitives. The talk will conclude presenting open problems and future research directions. • "Energy Efficient Design and Implementation of Security Primitives". This talk reviews the strategies that have been applied to security primitives to reduce their energy consumption and presents the algorithms that have been designed, since the beginning, to achieve a limited energy footprint [3] , [4] . The talk will conclude presenting open problems and future research directions.
Ayse K. Coskun, Paolo Palmieri 0001, Francesco Regazzoni 0001
ISLPED1
2025 Job Grouping Based Intelligent Resource Prediction Framework
Beste Oztop, Benjamin Schwaller, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun
JSSPP7
2025 Distributed Economic Dispatch in Power Networks Incorporating Data Center Flexibility
abstract
We consider Data Centers (DCs) as flexible loads that can alter their power consumption to alleviate congestion in the electric power network. We model DCs using a queuing-theoretic view and we form a Quality of Service (QoS)-based cost function that signifies how well a DC can carry out its workload given an amount of active servers. We integrate DCs in a centralized economic dispatch problem that determines, apart from power generation, DC workload shifting and server utilization, while respecting transmission line constraints. We further present a tractable decentralized formulation obtained via Lagrangian decomposition, which we solve using a dual gradient ascent algorithm. Experimental results on a standard power network explore the system-wide benefits of DC flexibility in “coupled” data and power networks, emphasizing on the trade-offs between the DC location, QoS, and efficiency.
Athanasios Tsiligkaridis, Panagiotis Andrianesis 0001, Ayse K. Coskun, Michael C. Caramanis, Ioannis Paschalidis
IEEE Trans. Sustain. Comput.3
2024 Data Center Demand Response for Sustainable Computing: Myth or Opportunity?
abstract
In our computing-driven era, the escalating power consumption of modern data centers, currently constituting approximately 3% of global energy use, is a burgeoning concern. With the anticipated surge in usage accompanying the widespread adoption of AI technologies, addressing this issue becomes imperative. This paper discusses a potential solution: integrating data centers into grid programs such as “demand response” (DR). This strategy not only optimizes power usage without requiring new fossil-fuel infrastructure but also facilitates more ambitious renewable deployment by adding demand flexibility to the grid. However, the unique scale, operational knobs and constraints, and future projections of data centers present distinct opportunities and urgent challenges for implementing DR. This paper delves into the myths and opportunities inherent in this perspective on improving data center sustainability. While obstacles including creating the requisite software infrastructure, establishing institutional trust, and addressing privacy concerns remain, the landscape is evolving to meet the challenges. Noteworthy achievements have emerged in the development of intelligent solutions that can be swiftly implemented in data centers to accelerate the adoption of DR. These multifaceted solutions encompass dynamic power capping, load scheduling, load forecasting, market bidding, and collaborative optimization. We offer insights into this promising step towards making sustainable computing a reality.
Ayse K. Coskun, Fatih Acun, Quentin Clark, Can Hankendi, Daniel C. Wilson
DATE1
2024 Unleashing Performance Insights with Online Probabilistic Tracing
abstract
Distributed tracing has become a fundamental tool for diagnosing performance issues in the cloud by recording causally ordered, end-to-end workflows of request executions. However, tracing workloads in production can introduce significant overheads due to the extensive instrumentation needed for identifying performance variations. This paper addresses the trade-off between the cost of tracing and the utility of the “spans” within that trace through Astraea, an online probabilistic distributed tracing system. Astraea is based on our technique that combines online Bayesian learning and multi-armed bandit frameworks. This formulation enables Astraea to effectively steer tracing towards the useful instrumentation needed for accurate performance diagnosis. Astraea localizes performance variations using only 20-35% of available instrumentation, markedly reducing tracing overhead, storage, compute costs, and trace analysis time.
M. Toslali, S. Qasim, F. A. Oliveira, Gianluca Stringhini, Ayse K. Coskun
IC2E8
2024 PraxiPaaS: A Decomposable Machine Learning System for Efficient Container Package Discovery
abstract
Due to the increasing complexity of cloud architectures, automatically tracking and inspecting container packages in Platform-as-a-Service (PaaS) clusters are challenging tasks. This introspection capability, however, is critical to identify vulnerable packages and compile an accurate Software Bill of Materials (SBOM). Motivated by introspection frameworks focusing on virtual machine (VM) settings and ML methods for software discovery, we design PraxiPaaS as a framework to inspect PaaS container images with a highly scalable ML inference pipeline by scanning file changes during package installations. Our ML pipeline includes a structured collection of word2vec encoders and a corresponding structured ML model to achieve short incremental training time for incorporating additional packages while maintaining a high F1-score in generating the SBOM. Our evaluation shows that our structured ML pipeline provides an exponential drop in incremental training time from 2.8 hours to $8.6 \mathbf{s}$ with 32 CPU cores, while maintaining an F1-score of 0.82, compared to the traditional monolithic model design. We deploy a prototype of PraxiPaaS in the New England Research Cloud (NERC) OpenShift cluster and evaluate the inference time comparing structured versus monolithic model design.
Zongshun Zhang, Lisa Korver, Anthony Byrne, Gianluca Stringhini, Abraham Matta, Ayse K. Coskun
IC2E8
2024 SOPHIE: A Scalable Recurrent Ising Machine Using Optically Addressed Phase Change Memory
abstract
Ising problems are nondeterministic-polynomial-hard (NP-hard) problems prevalent in various domains, such as statistical physics, circuit design, and machine learning. They pose significant challenges for traditional algorithms and architectures. Researchers have recently developed nature-inspired Ising machines to tackle these optimization problems efficiently. Many optimization problems can be mapped to the Ising model, and physical laws will drive the Ising machine towards the solution. However, existing Ising machines suffer from scalability issues, i.e., performance drops when problem sizes exceed their physical capacity. In this paper, we propose SOPHIE, a Scalable Optical PHase-change memory (OPCM) based Ising Engine. SOPHIE integrates architectural, algorithmic, and device optimizations to address scalability challenges in Ising machines. We architect SOPHIE using 2.5D integration, where we integrate a controller chiplet, a DRAM chiplet, laser sources, and multiple OPCM chiplets. SOPHIE utilizes OPCMs to perform matrix-vector multiplications efficiently. Our symmetric tile mapping at the architecture level reduces approximately half of the OPCM array area, enhancing the scalability of SOPHIE. We use algorithmic optimizations to efficiently handle large problems that cannot fit within hardware constraints. Specifically, we adopt a symmetric local update technique and a stochastic global synchronization strategy. These two algorithmic approaches decompose large problems into isolated tiles, reduce computation requirements, and minimize communication in SOPHIE. We apply device-level optimizations to adopt the modified algorithm. These device-level optimizations include employing bi-directional OPCM arrays and dual-precision analog-to-digital converters. SOPHIE is 3 x faster than the state-of-the-art photonic Ising machines on small graphs and 125x faster than the FPGA-based designs on large problems. SOPHIE alleviates the hardware capacity constraints, offering a scalable and efficient alternative for solving Ising problems.
Guowei Yang 0005, Sina Karimi, Carlos A. Ríos Ocampo, Ayse K. Coskun, Ajay Joshi
MICRO4
2024 LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
abstract
Large Language Models (LLMs) have been suggested for use in automated vulnerability repair, but benchmarks showing they can consistently identify security-related bugs are lacking. We thus develop SecLLMHolmes, a fully automated evaluation framework that performs the most detailed investigation to date on whether LLMs can reliably identify and reason about security-related bugs. We construct a set of 228 code scenarios and analyze eight of the most capable LLMs across eight different investigative dimensions using our framework. Our evaluation shows LLMs provide non-deterministic responses, incorrect and unfaithful reasoning, and perform poorly in real-world scenarios. Most importantly, our findings reveal significant non-robustness in even the most advanced models like ‘PaLM2’ and ‘GPT-4’: by merely changing function or variable names, or by the addition of library functions in the source code, these models can yield incorrect answers in 26% and 17% of cases, respectively. These findings demonstrate that further LLM advances are needed before LLMs can be used as general purpose security assistants.
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond A. Pearce, Ayse K. Coskun, Gianluca Stringhini
SP5
2024 Runtime Performance Anomaly Diagnosis in Production HPC Systems Using Active Learning
abstract
With the increasing scale and complexity of High-Performance Computing (HPC) systems, performance variations in applications caused by anomalies have become significant bottlenecks in system health and operational efficiency. As we move towards exascale systems, these variations become more prominent due to the increased sharing of resources. Such variations lead to lower energy efficiency and higher operational costs. To mitigate these problems, one must quickly and accurately diagnose the root cause of the anomalies at scale. One way to evaluate system health and identify the underlying causes is by manually examining certain performance metrics in telemetry data or using rule-based methods. Due to the daily size of telemetry data reaching terabytes and the fact that the numeric telemetry data contains thousands of metrics, manual analysis of telemetry to diagnose problems becomes challenging. Given these limitations, Machine Learning (ML)-based approaches have been gaining popularity as they have been shown to be effective and practical in diagnosing previously encountered performance anomalies. One primary challenge for supervised ML models is that they require a significant amount of labeled samples during training. However, obtaining many labels for anomalies is extremely difficult and costly, considering anomalies occur infrequently and real-world numeric system telemetry data is hard to label since it contains thousands of metrics. This paper proposes a novel active learning-based framework that diagnoses performance anomalies (i.e., identifying the type of an anomaly) in HPC systems at runtime using significantly fewer labeled samples compared to state-of-the-art ML-based approaches. We show that the proposed framework achieves the same F1-score compared to a supervised approach using much fewer labeled samples (i.e., 16x fewer samples for achieving a 0.78 F1-score, 11x fewer samples for achieving a 0.82 F1-score), even when there are previously unseen applications and application inputs in the test dataset.
Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun
IEEE Trans. Parallel Distributed Syst.9
2023 Temperature-Aware Sizing of Multi-Chip Module Accelerators for Multi-DNN Workloads
abstract
This paper demonstrates the need for temperature awareness in sizing accelerators to target multi-DNN workloads. To that end, we build TESA, a TEmperature-aware methodology that Sizes and places Accelerators to balance both the cost and power of a multi-chip module (MCM), including DRAM power for multi-deep neural network workloads. TESA tunes the accelerator chiplet size and inter-chiplet spacing to generate a temperature-aware MCM layout, subject to user-defined latency, area, power, and thermal constraints. Using TESA for both 2D and 3D systolic array-based chiplets, we demonstrate up to 44% MCM cost savings and 63% DRAM power savings, respectively, over a temperature-unaware baseline at iso-frequency and iso-interposer area. We also demonstrate a need for TESA to obtain feasible MCM configurations for multi-DNN workloads such as augmented/virtual reality (AR/VR).
Prachi Shukla, Derrick Aguren, Thomas Burd, Ayse K. Coskun, John Kalamatianos
DATE4
2023 MicroFaaS on OpenFaaS: An Embedded Platform for Running Cloud Functions
abstract
Function-as-a-Service (FaaS) platforms present a new cloud computing paradigm by enabling serverless deployment and execution of functions. MicroFaaS, a cost-effective and energy-efficient datacenter architecture, replaces x86-based rack servers with ARM-based single-board computers (SBCs). This paper focuses on enhancing MicroFaaS by incorporating OpenFaaS, a popular secure function building and deployment framework. Through extensive experimentation, MicroFaaS on OpenFaaS showcases improved energy efficiency, ease-of-use, and scalability over traditional cloud systems.
Abin B. George, Anthony Byrne, Ayse K. Coskun
IC2E3
2023 Poster Paper: Efficient Navigation of Cloud Performance with 'nuffTrace
abstract
Distributed tracing has become an essential tool to navigate performance of complex, distributed cloud-native applications, providing a comprehensive view of a request from end-to-end. However, the sheer amount of data generated by distributed tracing can be overwhelming, making it difficult to store, process, and extract meaningful insights. This paper presents the vision of ’nuffTrace that embodies a novel tree-based probabilistic data structure that summarizes trace data in a compact form without storing all of the data, enabling developers to analyze cloud application performance with high accuracy and efficiency.
S. Qasim, Mert Toslali, Quentin Clark, Srinivasan Parthasarathy 0001, Fábio Oliveira, Gianluca Stringhini, Ayse K. Coskun
IC2E8
2023 Enabling Privacy-preserving Multidimensional Network Telemetry with Autoencoders
abstract
Network telemetry systems are essential for monitoring network traffic and informing management decisions. However, increasing privacy concerns make user data access and analysis challenging for operators. We introduce PrvTel, a privacy-preserving telemetry system that uses an AutoEncoder model to encode user traffic data and preserve telemetry query ability with differential privacy guarantees. PrvTel features a lightweight model to be stored, facilitates quick training, and executes queries with minimal delay.
Gianluca Stringhini, Ayse K. Coskun, Zaoxing Liu
IC2E4
2023 Processing-in-Memory Using Optically-Addressed Phase Change Memory
abstract
Today's Deep Neural Network (DNN) inference systems contain hundreds of billions of parameters, resulting in significant latency and energy overheads during inference due to frequent data transfers between compute and memory units. Processing-in-Memory (PiM) has emerged as a viable solution to tackle this problem by avoiding the expensive data movement. PiM approaches based on electrical devices suffer from throughput and energy efficiency issues. In contrast, Optically-addressed Phase Change Memory (OPCM) operates with light and achieves much higher throughput and energy efficiency compared to its electrical counterparts. This paper introduces a system-level design that takes the OPCM programming overhead into consideration, and identifies that the programming cost dominates the DNN inference on OPCM-based PiM architectures. We explore the design space of this system and identify the most energy-efficient OPCM array size and batch size. We propose a novel thresholding and reordering technique on the weight blocks to further reduce the programming overhead. Combining these optimizations, our approach achieves up to 65.2 × higher throughput than existing photonic accelerators for practical DNN workloads.
Guowei Yang 0005, Cansu Demirkiran, Zeynep Ece Kizilates, Carlos A. Ríos Ocampo, Ayse K. Coskun, Ajay Joshi
ISLPED5
2023 Prodigy: Towards Unsupervised Anomaly Detection in Production HPC Systems
abstract
Performance variations caused by anomalies in modern High Performance Computing (HPC) systems lead to decreased efficiency, impaired application performance, and increased operational costs. While machine learning (ML)-based frameworks for automated anomaly detection (often based on time series telemetry data) are gaining popularity in the literature, practical deployment challenges are often overlooked. Some ML-based frameworks require extensive customization, while others need a rich set of labeled samples, none of which are feasible for a production HPC system.
Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Manuel Egele, Ayse K. Coskun
SC9
2023 TREAD-M3D: Temperature-Aware DNN Accelerators for Monolithic 3-D Mobile Systems
abstract
Monolithic 3-D (MONO3 D) integration provides performance and power efficiency benefits over 2-D circuits and, thus, is a potent technology for the design of deep neural network (DNN) accelerators with enhanced energy efficiency. However, high IC temperatures are major challenges for the design of MONO3 D systems. To this end, this article focuses on designing temperature-aware MONO3 D DNN accelerators. We propose a new automated method, called TREAD- M3 D, that provides a near-optimal MONO3 D DNN accelerator architecture in terms of systolic array size, SRAM organization, partition across 3-D layers, and operating frequency, for a given DNN, optimization goal, and temperature constraint. TREAD- M3 D incorporates circuit- and architecture-level models to evaluate the power and performance characteristics of different partitions. Our method reveals valuable insights and enables tradeoff analysis for achieving high energy efficiency in MONO3 D systolic arrays. In comparison to recent works that adopt a fixed partition choice to design MONO3 D DNN systems, TREAD- M3 D yields up to 22% higher energy efficiency. Using TREAD- M3 D, we further demonstrate that temperature unawareness not only leads to infeasible configurations due to temperature violations but also over-estimates energy-delay-product benefits by up to 24%.
Prachi Shukla, Vasilis F. Pavlidis, Emre Salman, Ayse K. Coskun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 ALBADross: Active Learning Based Anomaly Diagnosis for Production HPC Systems
abstract
Diagnosing causes of performance variations in High-Performance Computing (HPC) systems is a daunting chal-lenge due to the systems' scale and complexity. Variations in application performance result in premature job termination, lower energy efficiency, or wasted computing resources. One potential solution is manual root-cause analysis based on system telemetry data. However, this approach has become an increasingly time-consuming procedure as the process relies on human expertise and the size of telemetry data is voluminous. Recent research employs supervised machine learning (ML) models to diagnose previously encountered performance anomalies in compute nodes automatically. However, these models generally necessitate vast amounts of labeled samples that represent anomalous and healthy states of an application during training. The demand for labeled samples is constraining because gathering labeled samples is difficult and costly, especially considering anomalies that occur infrequently. This paper proposes a novel active learning-based framework that diagnoses previously encountered performance anomalies in HPC systems using significantly fewer labeled samples compared to state-of-the-art ML-based frameworks. Our framework combines an active learning-based query strategy and a supervised classifier to minimize the number of labeled samples required to achieve a target performance score. We evaluate our framework on a production HPC system and a testbed HPC cluster using real and proxy applications. We show that our framework, ALBADross, achieves a 0.95 Fl-score using 28x fewer labeled samples compared to a supervised approach with equal Fl-score, even when there are previously unseen applications and application inputs in the test dataset.
Burak Aksar, Efe Sencan, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Brian Kulis, Ayse K. Coskun
CLUSTER8
2022 MicroFaaS: Energy-efficient Serverless on Bare-metal Single-board Computers
abstract
Serverless function-as-a-service (FaaS) platforms offer a radically-new paradigm for cloud software development, yet the hardware infrastructure underlying these platforms is based on a decades-old design pattern. The rise of FaaS presents an opportunity to reimagine cloud infrastructure to be more energy-efficient, cost-effective, reliable, and secure. In this paper, we show how replacing handfuls of x86-based rack servers with hundreds of ARM-based single-board computers could lead to a virtualization-free, energy-proportional cloud that achieves this vision. We call our systematically-designed implementation MicroFaaS, and we conduct a thorough evaluation and cost analysis comparing MicroFaaS to a throughput-matched FaaS platform implemented in the style of conventional virtualization-based cloud systems. Our results show a 5.6x increase in energy efficiency and 34.2% decrease in total cost of ownership compared to our baseline.
Anthony Byrne, Yanni Pang, Allen Zou, Shripad Nadgowda, Ayse K. Coskun
DATE5
2022 Architecting Optically Controlled Phase Change Memory
abstract
Phase Change Memory (PCM) is an attractive candidate for main memory, as it offers non-volatility and zero leakage power while providing higher cell densities, longer data retention time, and higher capacity scaling compared to DRAM. In PCM, data is stored in the crystalline or amorphous state of the phase change material. The typical electrically controlled PCM (EPCM), however, suffers from longer write latency and higher write energy compared to DRAM and limited multi-level cell (MLC) capacities. These challenges limit the performance of data-intensive applications running on computing systems with EPCMs. Recently, researchers demonstrated optically controlled PCM (OPCM) cells with support for 5 bits / cell in contrast to 2 bits / cell in EPCM. These OPCM cells can be accessed directly with optical signals that are multiplexed in high-bandwidth-density silicon-photonic links. The higher MLC capacity in OPCM and the direct cell access using optical signals enable an increased read/write throughput and lower energy per access than EPCM. However, due to the direct cell access using optical signals, OPCM systems cannot be designed using conventional memory architecture. We need a complete redesign of the memory architecture that is tailored to the properties of OPCM technology. This article presents the design of a unified network and main memory system called COSMOS that combines OPCM and silicon-photonic links to achieve high memory throughput. COSMOS is composed of a hierarchical multi-banked OPCM array with novel read and write access protocols. COSMOS uses an Electrical-Optical-Electrical (E-O-E) control unit to map standard DRAM read/write commands (sent in electrical domain) from the memory controller on to optical signals that access the OPCM cells. Our evaluation of a 2.5D-integrated system containing a processor and COSMOS demonstrates 2.14 × average speedup across graph and HPC workloads compared to an EPCM system. COSMOS consumes 3.8× lower read energy-per-bit and 5.97× lower write energy-per-bit compared to EPCM. COSMOS is the first non-volatile memory that provides comparable performance and energy consumption as DDR5 in addition to increased bit density, higher area efficiency, and improved scalability.
Aditya Narayan, Yvain Thonnart, Pascal Vivet, Ayse K. Coskun, Ajay Joshi
ACM Trans. Archit. Code Optim.4
2022 PACT: An Extensible Parallel Thermal Simulator for Emerging Integration and Cooling Technologies
abstract
Thermal analysis is an essential step that enables co-design of the computing system (i.e., integrated circuits and computer architectures) with the cooling system (e.g., heat sink). Existing thermal simulation tools are limited by several major challenges that prevent them from providing fast solutions to large problem sizes that are necessary to conduct standard-cell level thermal analysis or to evaluate new technologies or large chips. To overcome these challenges, we introduce a SPICE-based parallel compact thermal simulator (PACT) that achieves fast and accurate, standard cell to architecture-level, steady-state, and transient parallel thermal simulations. PACT utilizes the advantages of multicore processing (OpenMPI) and includes several solvers to speed up both steady-state and transient simulations. PACT can be easily extended to model a variety of emerging integration and cooling technologies by simply modifying the thermal netlist. In addition, PACT can also be used with popular architecture-level performance and power simulators. In comparison to a state-of-the-art finite-element method (FEM)-based simulator (COMSOL), PACT has a maximum error of 2.77% and 3.28% for steady-state and transient thermal simulations, respectively. Compared to a popular compact thermal simulator, HotSpot, PACT demonstrates a speedup of up to$1.83\times $and$186\times $for steady-state and transient simulations, respectively. We also show the applicability and extensibility of PACT through modeling emerging integration and cooling technologies, such as monolithic 3-D integrated circuits and liquid cooling via microchannels, and full-system simulation integration on a 2.5-D system with silicon-photonic network-on-chips (PNoCs).
Prachi Shukla, Sofiane Chetoui, Sean S. Nemtzow, Sherief Reda, Ayse K. Coskun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 Praxi: Cloud Software Discovery That Learns From Practice
abstract
With today’s rapidly-evolving cloud landscape embracing continuous integration and delivery, users of cloud systems must monitor software running on theircontainersandvirtual machines (VMs)to ensure compliance, security, and efficiency. Traditional solutions to this problem rely on manually-createdrulesthat identify software installations and modifications, but these require expert authors and are often unmaintainable. Recently, automated techniques for software discovery have emerged. Some techniques use examples of software to train machine learning models to predict which software has been installed on a system. Others leverage the knowledge of packaging practices to aid in discovery without requiring any pre-training, but these practice-based methods cannot provide precise-enough information to perform discovery by themselves. This article introduces Praxi, a new software discovery method that builds upon the strengths of prior approaches by combining the accuracy of learning-based methods with the efficiency of practice-based methods. In tests using samples collected on real-world cloud systems, Praxi correctly classifies installations at least 97.6 percent of the time, while running 14.8 times faster and using 87 percent less disk space than a similar learning-based method. Using a diverse software dataset, this article quantitatively compares Praxi to systematic rule-, learning-, and practice-based methods, and discusses the best uses for each.
Anthony Byrne, Emre Ates, Ata Turk, Vladimir Pchelin, Sastry S. Duri, Shripad Nadgowda, Canturk Isci, Ayse K. Coskun
IEEE Trans. Cloud Comput.8
2022 HPC Data Center Participation in Demand Response: An Adaptive Policy With QoS Assurance
abstract
Demand response programs help stabilize the electricity grid by providing monetary stimulus to consumers if they regulate their power consumption following market requirements. Regulation service, a market that requires participants to regulate power by following a signal updated every few seconds, is particularly beneficial to HPC data centers since data centers are capable of increasing/decreasing power consumption owing to the flexibility in running workloads and the availability of power control mechanisms. While prior works have explored how data centers can provide regulation service reserves, Quality-of-Service (QoS) provisioning for the jobs running at the data centers has not been considered. In this work, we propose an Adaptive policy with QoS Assurance that enables data centers to participate in regulation service programs with assurance on job QoS. Our policy regulates data center power through job scheduling and server power capping. QoS assurance is achieved by applying a queueing-theoretic result to our job scheduling strategy. We evaluate our policy by experiments on a real cluster. Our results demonstrate that the proposed policy reduces electricity costs by 25-56% while providing QoS assurance. On the other hand, the baseline policies cannot meet QoS constraints in 9 of the 14 workload traces tested.
Yijia Zhang 0002, Daniel C. Wilson, Ioannis Paschalidis, Ayse K. Coskun
IEEE Trans. Sustain. Comput.4
2022 High Bandwidth Thermal Covert Channel in 3-D-Integrated Multicore Processors
abstract
Exploiting thermal coupling among the cores of a processor to secretly communicate sensitive information is a serious threat in mobile, desktop, and server platforms. Existing works on temperature-based covert communication typically rely on controlling the execution of high-power CPU stressing programs to transmit confidential information. Such covert channels with high-power programs are typically easier to detect as they cause significant rise in temperature. In this work, we demonstrate that by leveraging vertical integration, it is sufficient to execute typical SPLASH-2 benchmark applications to transfer 200 bits per second (bps) of secret data via thermal covert channels. The strong vertical thermal coupling among the cores of a 3-D multicore processor increases the rates of covert communication by$3.4\times $compared to covert communication in conventional 2-D integrated circuits (ICs). Furthermore, we show that the bandwidth of this thermal communication in 3-D ICs is more resilient to thermal interference caused by applications running in other cores. This reduced interference significantly increases the danger posed by such attacks. We also investigate the effect of reducing intertier overlap between colluded cores and show that the covert channel bandwidth is reduced by up to 62% with no overlap.
Krithika Dhananjay, Vasilis F. Pavlidis, Ayse K. Coskun, Emre Salman
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Temperature-Aware Optimization of Monolithic 3D Deep Neural Network Accelerators
abstract
We propose an automated method to facilitate the design of energy-efficient Mono3D DNN accelerators with safe on-chip temperatures for mobile systems. We introduce an optimizer to investigate the effect of different aspect ratios and footprint specifications of the chip, and select energy-efficient accelerators under user-specified thermal and performance constraints. We also demonstrate that using our optimizer, we can reduce energy consumption by 1.6x and area by 2x with a maximum of 9.5% increase in latency compared to a Mono3D DNN accelerator optimized only for performance.
Prachi Shukla, Sean S. Nemtzow, Vasilis F. Pavlidis, Emre Salman, Ayse K. Coskun
ASP-DAC5
2021 Iter8: Online Experimentation in the Cloud
abstract
Online experimentation is an agile software development practice that plays an essential role in enabling rapid innovation. Existing solutions for online experimentation in Web and mobile applications are unsuitable for cloud applications. There is a need for rethinking online experimentation in the cloud to advance the state-of-the-art by considering the unique challenges posed by cloud environments.
Mert Toslali, Srinivasan Parthasarathy 0001, Fábio Oliveira, Hai Huang 0002, Ayse K. Coskun
SoCC5
2021 Automating instrumentation choices for performance problems in distributed applications with VAIF
abstract
Developers use logs to diagnose performance problems in distributed applications. However, it is difficult to know a priori where logs are needed and what information in them is needed to help diagnose problems that may occur in the future. We present the Variance-driven Automated Instrumentation Framework (VAIF), which runs alongside distributed applications. In response to newly-observed performance problems, VAIF automatically searches the space of possible instrumentation choices to enable the logs needed to help diagnose them. To work, VAIF combines distributed tracing (an enhanced form of logging) with insights about how response-time variance can be decomposed on the critical-path portions of requests' traces. We evaluate VAIF by using it to localize performance problems in OpenStack and HDFS. We show that VAIF can localize problems related to slow code paths, resource contention, and problematic third-party code while enabling only 3-34% of the total tracing instrumentation.
Mert Toslali, Emre Ates, Alex Ellis, Darby Huye, Samantha Puterman, Ayse K. Coskun, Raja R. Sambasivan
SoCC8
2021 A Data Center Demand Response Policy for Real-World Workload Scenarios in HPC
abstract
Demand response programs offer an opportunity for large power consumers to save on electricity costs by modulating their power consumption in response to demand changes in the electricity grid. Multiple types of such programs exist; for example, regulation service programs enable a consumer to bid for a sustainable amount of power draw over a time period, along with a reserve amount they are able to provide at request of the electricity service provider. Data centers offer unique capabilities to participate in these programs since they have significant capacity to modify their power consumption through workload scheduling and CPU power limiting. This paper proposes a novel power management policy and a bidding policy that enable data centers to participate in regulation service programs under real-world constraints. The power management policy schedules computing jobs and applies server power-capping under both the constraints of power programs and the constraints of job Quality-of-Service (QoS). Simulations with workload traces from a real data center show that the proposed policies enable data centers to meet both the requirement of regulation service programs and the QoS requirement of jobs. We demonstrate that, by applying our policies, data centers can save their electricity costs by 10% while abiding by all the QoS constraints in a real-world scenario.
Yijia Zhang 0002, Daniel C. Wilson, Ioannis Paschalidis, Ayse K. Coskun
DATE4
2021 E2EWatch: An End-to-End Anomaly Diagnosis Framework for Production HPC Systems
Burak Aksar, Benjamin Schwaller, Omar Aaziz, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun
Euro-Par7
2021 Introducing Application Awareness Into a Unified Power Management Stack
abstract
Effective power management in a data center is critical to ensure that power delivery constraints are met while maximizing the performance of users' workloads. Power limiting is needed in order to respond to greater-than-expected power demand. HPC sites have generally tackled this by adopting one of two approaches: (1) a system-level power management approach that is aware of the facility or site-level power requirements, but is agnostic to the application demands; OR (2) a job-level power management solution that is aware of the application design patterns and requirements, but is agnostic to the site-level power constraints. Simultaneously incorporating solutions from both domains often leads to conflicts in power management mechanisms. This, in turn, affects system stability and leads to irreproducibility of performance. To avoid this irreproducibility, HPC sites have to choose between one of the two approaches, thereby leading to missed opportunities for efficiency gains.This paper demonstrates the need for the HPC community to collaborate towards seamless integration of system-aware and application-aware power management approaches. This is achieved by proposing a new dynamic policy that inherits the benefits of both approaches from tight integration of a resource manager and a performance-aware job runtime environment. An empirical comparison of this integrated management approach against state-of-the-art solutions exposes the benefits of investing in end-to-end solutions to optimize for system-wide performance or efficiency objectives. With our proposed system-application integrated policy, we observed up to 7% reduction in system time dedicated to jobs and up to 11% savings in compute energy, compared to a baseline that is agnostic to system power and application design constraints.
Daniel C. Wilson, Siddhartha Jana, Aniruddha Marathe, Stephanie Brink, Christopher Cantalupo, Diana R. Guttman, Brad Geltz, Lowren H. Lawson, Asma Al-Rawi, Ali Mohammad, Fuat Keceli, Federico Ardanaz, Jonathan Eastep, Ayse K. Coskun
IPDPS14
2021 PROWAVES: Proactive Runtime Wavelength Selection for Energy-Efficient Photonic NoCs
abstract
2.5-D manycore systems running parallel applications are severely bottlenecked by network-on-chip (NoC) latencies and bandwidth. Traditionally, NoCs are composed of electrical links that exhibit constrained bandwidth, increased energy consumption at high-speed communication, and long latencies. Photonic NoCs (PNoCs) have been shown to provide high bandwidth at low latencies and negligible data-dependent power. However, the power overheads of lasers, thermal tuning, and electrical-optical conversion present major challenges against wide-scale adoption of PNoCs. A primary factor that impacts PNoC power is the number of activated laser wavelengths in the system. Applications’ dynamic bandwidth needs provide the opportunity to selectively deactivate laser wavelengths when there is a lower bandwidth demand to alleviate high PNoC power concerns. This article analyzes dynamic PNoC activity of applications at runtime so as to select laser wavelengths depending on an application’s bandwidth requirements. The article then proposesPROWAVES, a proactive runtime wavelength selection policy that forecasts the bandwidth needs and activates the minimum laser wavelengths for each application phase. We develop a cross-layer simulation framework to model the system performance, PNoC power and transient thermal distribution in a manycore system with PNoCs. We comparePROWAVESwith prior system-level policies and our simulation results on a 2.5-D system demonstrate thatPROWAVESprovides 18% and 33% power savings with only 1% and 5% loss in performance, respectively, compared to activating all laser wavelengths in the system.
Aditya Narayan, Yvain Thonnart, Pascal Vivet, Ayse K. Coskun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 ECOGreen: Electricity Cost Optimization for Green Datacenters in Emerging Power Markets
abstract
Modern datacenters need to tackle efficiently the increasing demand for computing resources while minimizing energy usage and monetary costs. Power market operators have recently introduced emerging demand-response programs, in which electricity consumers regulate their power usage following provider requests to reduce monetary costs. Among different programs, regulation service (RS) reserves are particularly promising for datacenters due to the high credit gain possibilities and datacenters' flexibility in regulating their power consumption. Therefore, it is essential to develop bidding strategies for datacenters to participate in emerging power markets together with power management policies that are aware of power market requirements at runtime. In this paper we propose ECOGreen, a holistic strategy to jointly optimize the datacenter RS problem and virtual machine (VM) allocation that satisfies the hour-ahead power market constraints in the presence of electrical energy storage (EES) and renewable energy. We first find the best power and reserve bidding values as well as the number of active servers in a fast analytical way that works well in practice. Then, we present an online adaptive policy that modulates datacenter power consumption by controlling VMs CPU resource limits and efficiently utilizing demand-side EES and renewable power, while guaranteeing quality-of-service (QoS) constraints. Our results demonstrate that ECOGreen can provide 76 percent of the datacenter power consumption on average as reserves to the market, due to largely operating on renewable sources and EES. This translates into ECOGreen saving up to 71 percent electricity costs when compared to other state-of-the-art datacenter electricity cost minimization techniques that participate in the power market.
Ali Pahlevan, Marina Zapater, Ayse K. Coskun, David Atienza 0001
IEEE Trans. Sustain. Comput.3
2020 Quantifying the impact of network congestion on application performance and network metrics
abstract
In modern high-performance computing (HPC) systems, network congestion is an important factor that contributes to performance degradation. However, how network congestion impacts application performance is not fully understood. As Aries network, a recent HPC network architecture featuring a dragonfly topology, is equipped with network counters measuring packet transmission statistics on each router, these network metrics can potentially be utilized to understand network performance. In this work, by experiments on a large HPC system, we quantify the impact of network congestion on various applications' performance in terms of execution time, and we correlate application performance with network metrics. Our results demonstrate diverse impacts of network congestion: while applications with intensive MPI operations (such as HACC and MILC) suffer from more than 40% extension in their execution times under network congestion, applications with less intensive MPI operations (such as Graph500 and HPCG) are mostly not affected. We also demonstrate that a stall-to-flit ratio metric derived from Aries network counters is positively correlated with performance degradation and, thus, this metric can serve as an indicator of network congestion in HPC systems.
Yijia Zhang 0002, Taylor L. Groves, Brandon Cook 0001, Nicholas J. Wright, Ayse K. Coskun
CLUSTER5
2020 System-level Evaluation of Chip-Scale Silicon Photonic Networks for Emerging Data-Intensive Applications
abstract
Emerging data-driven applications such as graph processing applications are characterized by their excessive memory footprint and abundant parallelism, resulting in high memory bandwidth demand. As the scale of datasets for applications is reaching orders of TBs, performance limitation due to bandwidth demands is a major concern. Traditional on-chip electrical networks fail to meet such high bandwidth demands due to increased energy-per-bit or physical limitations with pin counts. Silicon photonic networks have emerged as a promising alternative to electrical interconnects, owing to their high bandwidth density and low energy-per-bit communication with negligible data-dependent power. Wide-scale adoption of silicon photonics at chip level, however, is hampered by their high sensitivity to process and thermal variations, high laser power due to losses along the network, and power consumption of the electrical-optical conversion. Device-level technological innovations to mitigate these issues are promising, yet they do not consider the system-level implications of the applications running on manycore systems with photonic networks. This work aims to bridge the gap between the system-level attributes of applications with the underlying architectural and device-level characteristics of silicon photonic networks to achieve energy-efficient computing. We particularly focus on graph applications, which involve unstructured yet abundant parallel memory accesses that stress the on-chip communication networks, and develop a cross-layer framework to evaluate 2.5D systems with silicon photonic networks. We demonstrate 38% power savings through system-level management using wavelength selection policies with only 1% loss in system performance and further evaluate architectural design choices on 2.5D systems with photonic networks.
Aditya Narayan, Yvain Thonnart, Pascal Vivet, Ajay Joshi, Ayse K. Coskun
DATE5
2020 POPSTAR: a Robust Modular Optical NoC Architecture for Chiplet-based 3D Integrated Systems
abstract
Silicon photonics technology is now gaining maturity with increasing levels of design complexity from devices to large photonic integrated circuits. Close integration of control electronics with 3D assembly of photonics and CMOS opens the way to high-performance computing architectures partitioned in chiplets connected by optical NoC on silicon photonic interposers. In this paper, we give an overview of our works on optical links and NoC for manycore systems, from low-level control of photonic devices to high-level system optimization of the optical communications. We detail the POPSTAR optical NoC topology and architecture (Processors On Photonic Silicon interposer Terascale ARchitecture) with electro-optical interface chiplets, the corresponding nested spiral topology for single-writer multiple- reader links and the associated control electronics, in charge of high-speed drivers, thermal stabilization and handling of the protocol stack, from data integrity to flow-control, routing and arbitration of the optical communications. The strengths and opportunities for this architecture will be discussed, with a shift in system & implementation constraints with respect to previous optical NoC proposals, and new challenges to be addressed.
Yvain Thonnart, Stéphane Bernabé, Jean Charbonnier, Christian Bernard, David Coriat, César Fuguet Tortolero, Pierre Tissier, Benoît Charbonnier, Stephane Malhouitre, Damien Saint-Patrice, Myriam Assous, Aditya Narayan, Ayse K. Coskun, Denis Dutoit, Pascal Vivet
DATE13
2020 A Learning-Based Thermal Simulation Framework for Emerging Two-Phase Cooling Technologies
abstract
Future high-performance chips will require new cooling technologies that can extract heat efficiently. Two-phase cooling is a promising processor cooling solution owing to its high heat transfer rate and potential benefits in cooling power. Two-phase cooling mechanisms, including microchannel-based two-phase cooling or two-phase vapor chambers (VCs), are typically modeled by computing the temperature-dependent heat transfer coefficient (HTC) of the evaporator or coolant using an iterative simulation framework. Precomputed HTC correlations are specific to a given cooling system design and cannot be applied to even the same cooling technology with different cooling parameters (such as different geometries). Another challenge is that HTC correlations are typically calculated with computational fluid dynamics (CFD) tools, which induce long design and simulation times. This paper introduces a learning-based temperature-dependent HTC simulation framework that is used to model a two-phase cooling solution with a wide range of cooling design parameters. In particular, the proposed framework includes a compact thermal model (CTM) of two-phase VCs with hybrid wick evaporators (of nanoporous membrane and microchannels). We build a new simulation tool to integrate the proposed simulation framework and CTM. We validate the proposed simulation framework as well as the new CTM through comparisons against a CFD model. Our simulation framework and CTM achieve a speedup of 21 × with an average error of 0.98° C (and a maximum error of 2.59° C). We design an optimization flow for hybrid wicks to select the most beneficial hybrid wick geometries. Our flow is capable of finding a geometry- coolant combination that results in a lower (or similar) maximum chip temperature compared to that of the best coolant-geometry pair selected by grid search, while providing a speedup of 9.4 x.
Geoffrey Vaartstra, Prachi Shukla, Zhengmao Lu, Evelyn Wang, Sherief Reda, Ayse K. Coskun
DATE7
2020 Cross-Layer Co-Optimization of Network Design and Chiplet Placement in 2.5-D Systems
abstract
2.5-D integration technology is gaining attention and popularity in manycore computing system design. 2.5-D systems integrate homogeneous or heterogeneous chiplets in a flexible and cost-effective way. The design choices of 2.5-D systems impact overall system performance, manufacturing cost, and thermal feasibility. This article proposes a cross-layer co-optimization methodology for 2.5-D systems. We jointly optimize the network topology and chiplet placement across logical, physical, and circuit layers to improve system performance, reduce manufacturing cost, and lower operating temperature, while ensuring thermal safety and routability. We also propose a novel gas-station link, which enables pipelined interchiplet links in passive interposers. Our cross-layer methodology achieves better performance-cost tradeoffs of 2.5-D systems and yields better solutions in optimizing interchiplet network and 2.5-D system designs than prior methods. Compared to single-chip systems, 2.5-D systems designed using our new approach achieve 88% higher performance at the same manufacturing cost, or 29% lower cost with the same performance. Compared to the closest state-of-the-art, our new approach achieves 40%-68% (49% on average) iso-cost performance improvement and 30%-38% (32% on average) iso-performance cost reduction.
Ayse K. Coskun, Furkan Eris, Ajay Joshi, Andrew B. Kahng, Yenai Ma, Aditya Narayan, Vaishnav Srinivas
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 LoCool: Fighting Hot Spots Locally for Improving System Energy Efficiency
abstract
Elevated on-chip temperatures significantly degrade performance, energy-efficiency, and lifetime of processors. The cooling system for a chip is typically designed to remove the worst-case heat generated per unit area. Cooling demand, however, spatially and temporally varies across a chip as hot spots occur on different locations with different intensities. Thus, designing a homogeneous cooling system for a chip can be inefficient. Recently, hybrid cooling strategies, such as integrating thermoelectric coolers (TECs) with microchannel liquid cooling, have been explored for hot spot mitigation. The efficiency of such a cooling system strongly depends on the operating point of each cooling method, as well as the locations and intensities of the hot spots. To this end, we first devise a compact thermal modeling method for the design and evaluation of hybrid cooling systems in a fast and accurate way. The proposed model provides up to four orders of magnitude speedup in simulation time compared to COMSOL multiphysics simulations with less than 2.9 °C average temperature error. Leveraging our fast model, we develop LoCool, a hybrid cooling optimization method, which jointly determines the most energy-efficient cooling settings for a given chip power distribution and temperature constraint. LoCool determines the liquid flow rate and the input current for each TEC depending on the cooling requirements for individual hot spots as well as for the background heat. Experimental evaluation shows up to 40% cooling energy savings compared to designing homogeneous cooling systems under the same thermal constraints.
Fulya Kaplan, Mostafa Said, Sherief Reda, Ayse K. Coskun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 An automated, cross-layer instrumentation framework for diagnosing performance problems in distributed applications
abstract
Diagnosing performance problems in distributed applications is extremely challenging. A significant reason is that it is hard to know where to place instrumentation a priori to help diagnose problems that may occur in the future. We present the vision of an automated instrumentation framework, Pythia, that runs alongside deployed distributed applications. In response to a newly-observed performance problem, Pythia searches the space of possible instrumentation choices to enable the instrumentation needed to help diagnose it. Our vision for Pythia builds on workflow-centric tracing, which records the order and timing of how requests are processed within and among a distributed application's nodes (i.e., records their workflows). It uses the key insight that localizing the sources high performance variation within the workflows of requests that are expected to perform similarly gives insight into where additional instrumentation is needed.
Emre Ates, Lily Sturmann, Mert Toslali, Orran Krieger, Richard Megginson, Ayse K. Coskun, Raja R. Sambasivan
SoCC6
2019 Towards Practical Record and Replay for Mobile Applications
abstract
The ability to repeat the execution of a program is a fundamental requirement in evaluating computer systems and apps. Reproducing executions of mobile apps has proven difficult under real-life scenarios due to different sources of external inputs and interactive nature of the apps. We present a new practical record/replay framework for Android, RandR, which handles multiple sources of input and provides cross-device replay capabilities through a dynamic instrumentation approach. We demonstrate the feasibility of RandR by recording and replaying a set of real-world apps.
Onur Sahin, Assel Aliyeva, Hariharan Mathavan, Ayse K. Coskun, Manuel Egele
DAC4
2019 WAVES: Wavelength Selection for Power-Efficient 2.5D-Integrated Photonic NoCs
abstract
Photonic Network-on-Chips (PNoCs) offer promising benefits over Electrical Network-on-Chips (ENoCs) in many-core systems owing to their lower latencies, higher bandwidth, and lower energy-per-bit communication with negligible data-dependent power. These benefits, however, are limited by a number of challenges. Microring resonators (MRRs) that are used for photonic communication have high sensitivity to process variations and on-chip thermal variations, giving rise to possible resonant wavelength mismatches. State-of-the-art microheaters, which are used to tune the resonant wavelength of MRRs, have poor efficiency resulting in high thermal tuning power. In addition, laser power and high static power consumption of drivers, serializers, comparators, and arbitration logic partially negate the benefits of the sub-pJ operating regime that can be obtained with PNoCs. To reduce PNoC power consumption, this paper introduces WAVES, a wavelength selection technique to identify and activate the minimum number of laser wavelengths needed, depending on an application's bandwidth requirement. Our results on a simulated 2.5D manycore system with PNoC demonstrate an average of 23% (resp. 38%) reduction in PNoC power with only <;1% (resp. <;5%) loss in system performance.
Aditya Narayan, Yvain Thonnart, Pascal Vivet, César Fuguet Tortolero, Ayse K. Coskun
DATE5
2019 An Overview of Thermal Challenges and Opportunities for Monolithic 3D ICs
abstract
Monolithic 3D (Mono3D) is a three-dimensional integration technology that can overcome some of the fundamental limitations faced by traditional, two-dimensional scaling. This paper analyzes the unique thermal characteristics of Mono3D ICs by simulating a two-tier flip-chip Mono3D IC and highlights the primary differences in comparison to a similarly-sized flip-chip TSV-based 3D IC. Specifically, we perform architectural-level thermal simulations for both technologies and demonstrate that vertical thermal coupling is stronger in Mono3D ICs, leading to lower upper tier temperatures. We also investigate the significance of lateral versus vertical flow of heat in Mono3D ICs. We simulate different hot spot scenarios in a two-tier Mono3D IC and show that although the lateral heat flow is limited as compared to TSV-based 3D ICs, ignoring this mechanism can cause nonnegligible error (~4°C) in temperature estimation, particularly for layers farther from the heat sink. In addition, we show that with increasing interconnect utilization (due to the contribution of Joule heating to overall temperature), the on-chip temperatures and the significance of lateral heat flow within the two-tier Mono3D IC also increase. Finally, we discuss potential opportunities in Mono3D ICs to enhance their thermal integrity.
Prachi Shukla, Ayse K. Coskun, Vasilis F. Pavlidis, Emre Salman
ACM Great Lakes Symposium on VLSI2
2019 HPAS: An HPC Performance Anomaly Suite for Reproducing Performance Variations
abstract
Modern high performance computing (HPC) systems, including supercomputers, routinely suffer from substantial performance variations. The same application with the same input can have more than 100% performance variation, and such variations cause reduced efficiency and wasted resources. There have been recent studies on performance variability and on designing automated methods for diagnosing "anomalies" that cause performance variability. These studies either observe data collected from HPC systems, or they rely on synthetic reproduction of performance variability scenarios. However, there is no standardized way of creating performance variability inducing synthetic anomalies; so, researchers rely on designing ad-hoc methods for reproducing performance variability.
Emre Ates, Yijia Zhang 0002, Burak Aksar, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun
ICPP7
2019 Modeling and Optimization of Chip Cooling with Two-Phase Vapor Chambers
abstract
Ultra-high power densities that are expected in future processors cannot be efficiently mitigated by conventional cooling solutions. Using two-phase vapor chambers (VCs) with micropillar wick evaporators is an emerging cooling technique that can effectively remove high heat fluxes through the evaporation process of a coolant. Two-phase VCs with micropillar wicks offer high cooling efficiency by leveraging a capillary-driven flow, where the coolant is passively driven by the wicking structure that eliminates the need for an external pump. Thermal models for such emerging cooling technologies are essential to evaluate their impact on future processors. Existing thermal models for two-phase VCs use computational fluid dynamics (CFD) modules, which incur long design and simulation times. This paper presents a fast and accurate compact thermal model for two-phase VCs with micropillar wicks. Our model achieves a maximum error of 1.25°C with a speedup of 214x in comparison to a CFD model. Using our proposed thermal model, we build an optimization flow that selects the best cooling solution and its cooling parameters to minimize the cooling power under a temperature constraint for a given processor and power profile. We then demonstrate our optimization flow on different chip sizes and hot spot distributions to choose the optimal cooling technique among VCs, microchannel-based two-phase cooling, liquid cooling via microchannels, and a hybrid cooling technique with thermoelectric coolers and liquid cooling with microchannels.
Geoffrey Vaartstra, Prachi Shukla, Sherief Reda, Evelyn Wang, Ayse K. Coskun
ISLPED6
2019 RANDR: Record and Replay for Android Applications via Targeted Runtime Instrumentation
abstract
The ability to repeat the execution of a program is a fundamental requirement in many areas of computing from computer system evaluation to software engineering. Reproducing executions of mobile apps, in particular, has proven difficult under real-life scenarios due to multiple sources of external inputs and interactive nature of the apps. Previous works that provide record/replay functionality for mobile apps are restricted to particular input sources (e.g., touchscreen events) and present deployment challenges due to intrusive modifications to the underlying software stack. Moreover, due to their reliance on record and replay of device specific events, the recorded executions cannot be reliably reproduced across different platforms. In this paper, we present a new practical approach, RandR, for record and replay of Android applications. RandR captures and replays multiple sources of input (i.e., UI and network) without requiring source code (OS or app), administrative device privileges, or any special platform support. RandR achieves these qualities by instrumenting a select set of methods at runtime within an application's own sandbox. In addition, to enable portability of recorded executions across different platforms for replay, RandR contextualizes UI events as interactions with particular UI components (e.g., a button) as opposed to relying on platform specific features (e.g., screen coordinates). We demonstrate RandR's accurate cross-platform record and replay capabilities using over 30 real-world Android apps across a variety of platforms including emulators as well as commercial off-the-shelf mobile devices deployed in real life.
Onur Sahin, Assel Aliyeva, Hariharan Mathavan, Ayse K. Coskun, Manuel Egele
ASE4
2019 CAPE: A cross-layer framework for accurate microprocessor power estimation
Monir Zaman, Mustafa M. Shihab, Ayse K. Coskun, Yiorgos Makris
Integr.3
2019 Maestro: Autonomous QoS Management for Mobile Applications Under Thermal Constraints
abstract
Power densities of modern mobile system-on-a-chip designs can quickly exceed the thermal design limits during typical application use such as gaming or Web browsing. Resulting high temperatures lead to frequent thermal throttling and significant loss in quality-of-service (QoS) delivered to users. Thus, a joint consideration of thermal constraints and QoS requirements is essential to maximize the overall user experience. Prior techniques either rely on users to determine the best tradeoff point between QoS and temperature, or greedily utilize the thermal headroom to maximize performance, causing QoS to drop below user tolerable levels over extended durations of use. This paper introduces the MAESTRO framework to automatically manage QoS at runtime depending on application characteristics and thermal constraints. MAESTRO builds on the observation that increased temperatures can be tolerated for applications with bursty compute patterns due to idle periods between activities, while causing large QoS degradations for long-running applications with continuous computations. MAESTRO: 1) detects such continuous computations that are susceptible to throttling; 2) proactively finds a QoS level to balance user experience and temperature; and 3) performs closed-loop DVFS and thermally efficient thread mapping to meet the target QoS on a heterogeneous multicore CPU. Such application-adaptive control of QoS-temperature tradeoffs allows MAESTRO to sustain a target QoS level within a user tolerable range for longer durations without sacrificing the performance of latency-sensitive bursty computations. Evaluations on a real system prototype validates MAESTRO's ability to accurately detect potential throttlinginduced QoS degradations and demonstrates 41% to 6.7× longer durations of sustained QoS compared to state-of-the-art for a set of mobile applications.
Onur Sahin, Lothar Thiele, Ayse K. Coskun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning
abstract
As the size and complexity of high performance computing (HPC) systems grow in line with advancements in hardware and software technology, HPC systems increasingly suffer from performance variations due to shared resource contention as well as software- and hardware-related problems. Such performance variations can lead to failures and inefficiencies, which impact the cost and resilience of HPC systems. To minimize the impact of performance variations, one must quickly and accurately detect and diagnose the anomalies that cause the variations and take mitigating actions. However, it is difficult to identify anomalies based on the voluminous, high-dimensional, and noisy data collected by system monitoring infrastructures. This paper presents a novel machine learning based framework to automatically diagnose performance anomalies at runtime. Our framework leverages historical resource usage data to extract signatures of previously-observed anomalies. We first convert collected time series data into easy-to-compute statistical features. We then identify the features that are required to detect anomalies, and extract the signatures of these anomalies. At runtime, we use these signatures to diagnose anomalies with negligible overhead. We evaluate our framework using experiments on a real-world HPC supercomputer and demonstrate that our approach successfully identifies 98 percent of injected anomalies and consistently outperforms existing anomaly diagnosis techniques.
Ozan Tuncer, Emre Ates, Yijia Zhang 0002, Ata Turk, Jim M. Brandt, Vitus J. Leung, Manuel Egele, Ayse K. Coskun
IEEE Trans. Parallel Distributed Syst.8
2018 Taxonomist: Application Detection Through Rich Monitoring Data
Emre Ates, Ozan Tuncer, Ata Turk, Vitus J. Leung, Jim M. Brandt, Manuel Egele, Ayse K. Coskun
Euro-Par7
2018 A cross-layer methodology for design and optimization of networks in 2.5D systems
abstract
2.5D integration technology is gaining popularity in the design of homogeneous and heterogeneous many-core computing systems. 2.5D network design, both inter- and intra-chiplet, impacts overall system performance as well as its manufacturing cost and thermal feasibility. This paper introduces a cross-layer methodology for designing networks in 2.5D systems. We optimize the network design and chiplet placement jointly across logical, physical, and circuit layers to achieve an energy-efficient network, while maximizing system performance, minimizing manufacturing cost, and adhering to thermal constraints. In the logical layer, our co-optimization considers eight different network topologies. In the physical layer, we consider routing, microbump assignment, and microbump pitch constraints to account for the extra costs associated with microbump utilization in the inter-chiplet communication. In the circuit layer, we consider both passive and active links with five different link types, including a gas station link design. Using our cross-layer methodology results in more accurate determination of (superior) inter-chiplet network and 2.5D system designs compared to prior methods. Compared to 2D systems, our approach achieves 29% better performance with the same manufacturing cost, or 25% lower cost with the same performance.
Ayse K. Coskun, Furkan Eris, Ajay Joshi, Andrew B. Kahng, Yenai Ma, Vaishnav Srinivas
ICCAD1
2018 MOCA: Memory Object Classification and Allocation in Heterogeneous Memory Systems
abstract
In the era of abundant-data computing, main memory's latency and power significantly impact overall system performance and power. Today's computing systems are typically composed of homogeneous memory modules, which are optimized to provide either low latency, high bandwidth, or low power. Such memory modules do not cater to a wide range of applications with diverse memory access behavior. Thus, heterogeneous memory systems, which include several memory modules with distinct performance and power characteristics, are becoming promising alternatives. In such a system, allocating applications to their best-fitting memory modules improves system performance and energy efficiency. However, such an approach still leaves the full potential of heterogeneous memory systems under-utilized because not only applications, but also the memory objects within that application differ in their memory access behavior. This paper proposes a novel page allocation approach to utilize heterogeneous memory systems at the memory object level. We design a memory object classification and allocation framework (MOCA) to characterize memory objects and then allocate them to their best-fit memory module to improve performance and energy efficiency. We experiment with heterogeneous memory systems that are composed of a Reduced Latency DRAM (RLDRAM) for latency-sensitive objects, a 2.5D-stacked High Bandwidth Memory (HBM) for bandwidth-sensitive objects, and a Low Power DDR (LPDDR) for non-memory-intensive objects. The MOCA framework includes detailed application profiling, a classification mechanism, and an allocation policy to place memory objects. Compared to a system with homogeneous memory modules, we demonstrate that heterogeneous memory systems with MOCA improve memory system energy efficiency by up to 63%. Compared to a heterogeneous memory system with only application-level page allocation, MOCA achieves a 26% memory performance and a 33% energy efficiency improvement for multi-program workloads.
Aditya Narayan, Tiansheng Zhang, Shaizeen Aga, Satish Narayanasamy, Ayse K. Coskun
IPDPS5
2018 Level-Spread: A New Job Allocation Policy for Dragonfly Networks
abstract
The dragonfly network topology has attracted attention in recent years owing to its high radix and constant diameter. However, the influence of job allocation on communication time in dragonfly networks is not fully understood. Recent studies have shown that random allocation is better at balancing the network traffic, while compact allocation is better at harnessing the locality in dragonfly groups. Based on these observations, this paper introduces a novel allocation policy called Level-Spread for dragonfly networks. This policy spreads jobs within the smallest network level that a given job can fit in at the time of its allocation. In this way, it simultaneously harnesses node adjacency and balances link congestion. To evaluate the performance of Level-Spread, we run packet-level network simulations using a diverse set of application communication patterns, job sizes, and communication intensities. We also explore the impact of network properties such as the number of groups, number of routers per group, machine utilization level, and global link bandwidth. Level-Spread reduces the communication overhead by 16% on average (and up to 71%) compared to the state-of-the-art allocation policies.
Yijia Zhang 0002, Ozan Tuncer, Fulya Kaplan, Katzalin Olcoz, Vitus J. Leung, Ayse K. Coskun
IPDPS6
2018 Design Optimization of 3D Multi-Processor System-on-Chip with Integrated Flow Cell Arrays
abstract
Integrated flow cell array (FCA) is an emerging technology, targeting the cooling and power delivery challenges of modern 2D/3D Multi-Processor Systems-on-Chip (MPSoCs). In FCA, electrolytic solutions are pumped through microchannels etched in the silicon of the chips, removing heat from the system, while, at the same time, generating power on-chip. In this work, we explore the impact of FCA system design on various 3D architectures and propose a methodology to optimize a 3D MPSoC with integrated FCA to run a given workload in the most energy-efficient way. Our results show that an optimized configuration can save up to 50% energy with respect to sub-optimal 3D MPSoC configurations.
Artem Aleksandrovich Andreev, Fulya Kaplan, Marina Zapater, Ayse K. Coskun, David Atienza 0001
ISLPED4
2018 Proteus: Detecting Android Emulators from Instruction-Level Profiles
Onur Sahin, Ayse K. Coskun, Manuel Egele
RAID2
2017 User-profile-based analytics for detecting cloud security breaches
abstract
While the growth of cloud-based technologies has benefited the society tremendously, it has also increased the surface area for cyber attacks. Given that cloud services are prevalent today, it is critical to devise systems that detect intrusions. One form of security breach in the cloud is when cyber-criminals compromise Virtual Machines (VMs) of unwitting users and, then, utilize user resources to run time-consuming, malicious, or illegal applications for their own benefit. This work proposes a method to detect unusual resource usage trends and alert the user and the administrator in real time. We experiment with three categories of methods: simple statistical techniques, unsupervised classification, and regression. So far, our approach successfully detects anomalous resource usage when experimenting with typical trends synthesized from published real-world web server logs and cluster traces. We observe the best results with unsupervised classification, which gives an average F1-score of 0.83 for web server logs and 0.95 for the cluster traces.
Trishita Tiwari, Ata Turk, Alina Oprea, Katzalin Olcoz, Ayse K. Coskun
IEEE BigData5
2017 Unveiling the Interplay Between Global Link Arrangements and Network Management Algorithms on Dragonfly Networks
abstract
Network messaging delay historically constitutes a large portion of the wall-clock time for High Performance Computing (HPC) applications, as these applications run on many nodes and involve intensive communication among their tasks. Dragonfly network topology has emerged as a promising solution for building exascale HPC systems owing to its low network diameter and large bisection bandwidth. Dragonfly includes local links that form groups and global links that connect these groups via high bandwidth optical links. Many aspects of the dragonfly network design are yet to be explored, such as the performance impact of the connectivity of the global links, i.e., global link arrangements, the bandwidth of the local and global links, or the job allocation algorithm. This paper first introduces a packet-level simulation framework to model the performance of HPC applications in detail. The proposed framework is able to simulate known MPI (message passing interface) routines as well as applications with custom-defined communication patterns for a given job placement algorithm and network topology. Using this simulation framework, we investigate the coupling between global link bandwidth and arrangements, communication pattern and intensity, job allocation and task mapping algorithms, and routing mechanisms in dragonfly topologies. We demonstrate that by choosing the right combination of system settings and workload allocation algorithms, communication overhead can be decreased by up to 44%. We also show that circulant arrangement provides up to 15% higher bisection bandwidth compared to the other arrangements, but for realistic workloads, the performance impact of link arrangements is less than 3%.
Fulya Kaplan, Ozan Tuncer, Vitus J. Leung, Karl S. Hemmert, Ayse K. Coskun
CCGrid5
2017 Adaptive Tuning of Photonic Devices in a Photonic NoC Through Dynamic Workload Allocation
abstract
Photonic network-on-chip (PNoC) is a promising candidate to replace traditional electrical NoC in manycore systems that require substantial bandwidths. The photonic links in the PNoC comprise laser sources, optical ring resonators, passive waveguides, and photodetectors. Reliable link operation requires laser sources and ring resonators to have matching optical frequencies. However, inherent thermal sensitivity of photonic devices and manufacturing process variations can lead to a frequency mismatch. To avoid this mismatch, micro-heaters are used for thermal trimming and tuning, which can dissipate a significant amount of power. This paper proposes a novel FreqAlign workload allocation policy, accompanying an adaptive frequency tuning (AFT) policy, that is capable of reducing thermal tuning power of PNoC. FreqAlign uses thread allocation and thread migration to control temperature for matching the optical frequencies of ring resonators in each photonic link. The AFT policy reduces the remaining optical frequency difference among ring resonators and corresponding on-chip laser sources by hardware tuning methods. We use a full modeling stack of a PNoC that includes a performance simulator, a power simulator, and a thermal simulator with a temperature-dependent laser source power model to design and evaluate our proposed policies. Our experimental results demonstrate that FreqAlign reduces the resonant frequency gradient between ring resonators by 50%-60% when compared to existing workload allocation policies. Coupled with AFT, FreqAlign reduces localized thermal tuning power by 19.28 W on average, and is capable of saving up to 34.57 W when running realistic loads in a 256-core system without any performance degradation.
José L. Abellán, Ayse K. Coskun, Anjun Gu, Warren Jin 0002, Ajay Joshi, Andrew B. Kahng, Jonathan Klamkin, Cristian Morales, John Recchio, Vaishnav Srinivas, Tiansheng Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Scale & Cap: Scaling-Aware Resource Management for Consolidated Multi-threaded Applications
abstract
As the number of cores per server node increases, designing multi-threaded applications has become essential to efficiently utilize the available hardware parallelism. Many application domains have started to adopt multi-threaded programming; thus, efficient management of multi-threaded applications has become a significant research problem. Efficient execution of multi-threaded workloads on cloud environments, where applications are often consolidated by means of virtualization, relies on understanding the multi-threaded specific characteristics of the applications. Furthermore, energy cost and power delivery limitations require data center server nodes to work under power caps, which bring additional challenges to runtime management of consolidated multi-threaded applications. This article proposes a dynamic resource allocation technique for consolidated multi-threaded applications for power-constrained environments. Our technique takes into account application characteristics specific to multi-threaded applications, such as power and performance scaling, to make resource distribution decisions at runtime to improve the overall performance, while accurately tracking dynamic power caps. We implement and evaluate our technique on state-of-the-art servers and show that the proposed technique improves the application performance by up to 21% under power caps compared to a default resource manager.
Can Hankendi, Ayse K. Coskun
ACM Trans. Design Autom. Electr. Syst.2
2016 DeltaSherlock: Identifying changes in the cloud
abstract
To track security and compliance requirements and perform problem diagnosis, administrators of cloud computing systems need to monitor significant system changes occurring on the set of cloud instances under their supervision. Considering the large number of instances (virtual machines, containers) possibly operating under multiple configurations, this is a difficult-to-track process. Standard solutions to this problem rely on manually-created rules to identify changes. These techniques suffer from a limited scope, rely on domain expertise, and are time-consuming and error-prone. Recently, more streamlined approaches that automatically determine the type of individual system changes have been proposed, but these techniques assume that system states right before and after each individual change can be captured, a rather difficult requirement to enforce in real world usage. This paper proposes DeltaSherlock, a practical system change discovery framework that can capture system states on-demand and detect multiple system changes between them. We evaluate DeltaSherlock over 25,000 system changes caused by software installations collected from virtual machines (VMs) deployed over a commercial cloud. DeltaSherlock can accurately identify multiple software installations with 96.8% accuracy when supplied with a non-overlapping record of system changes and with 77.8% accuracy when supplied with random irregular observations possibly containing overlapping or incomplete changes.
Ata Turk, Hao Chen 0024, Anthony Byrne, John Knollmeyer, Sastry S. Duri, Canturk Isci, Ayse K. Coskun
IEEE BigData7
2016 Cross-layer floorplan optimization for silicon photonic NoCs in many-core systems
Ayse K. Coskun, Anjun Gu, Warren Jin 0002, Ajay Joshi, Andrew B. Kahng, Jonathan Klamkin, Yenai Ma, John Recchio, Vaishnav Srinivas, Tiansheng Zhang
DATE1
2016 QScale: thermally-efficient QoS management on heterogeneous mobile platforms
abstract
Single-ISA heterogeneous mobile processors integrate low-power and power-hungry CPU cores together to combine energy efficiency with high performance. While running computationally demanding applications, current power management and scheduling techniques greedily maximize quality-of-service (QoS) within thermal constraints using power-hungry cores. We show that such an approach delivers short bursts of high QoS, but also causes severe QoS loss over time due to thermal throttling. To provide mobile users with sustainable QoS over extended durations, this paper proposes QScale. QScale is a novel thermally-efficient QoS management framework for mobile devices with heterogeneous multi-core CPUs. QScale leverages two novel observations to provide thermally-efficient QoS: (1) threads of a mobile application exhibit significant heterogeneity, which can be exploited during scheduling; (2) thermal efficiency of core allocation decisions is significantly altered by thermal interactions across system-on-a-chip (SoC) components and application characteristics. QScale coordinates closed-loop frequency control with thermally-efficient scheduling to deliver the desired QoS with minimal exhaustion of processor thermal headroom. Our experiments on a state-of-the-art heterogeneous mobile platform show that QScale meets target QoS levels while minimizing heating, achieving up to 8x longer durations of sustainable QoS.
Onur Sahin, Ayse K. Coskun
ICCAD2
2016 Communication and cooling aware job allocation in data centers for communication-intensive workloads
Eduard Llamosí, Fulya Kaplan, Chulian Zhang, Jiayi Sheng, Martin C. Herbordt, Gunar Schirner, Ayse K. Coskun
J. Parallel Distributed Comput.8
2015 Just Enough is More: Achieving Sustainable Performance in Mobile Devices under Thermal Limitations
abstract
With the integration of high-performance multicore processors and multiple accelerators into modern mobile system-on-chips (SoCs), power densities have grown substantially. As a result, thermal management policies, which ensure operation at thermally safe conditions, became essential components of state-of-the-art mobile systems. Traditional thermal throttling approaches aim at maximum utilization of the available thermal headroom to minimize the performance loss and maximize user performance. This paper demonstrates that, in a mobile platform, such greedy techniques can lead to significant degradation in the quality-of-service (QoS) levels as the duration of device activity increases, leading to inconsistent user experience over time. We demonstrate that incorporating user/application QoS requirements into mobile power management to provide “just enough” performance (instead of always maximizing performance) allows for more efficient usage of the thermal headroom, which translates to substantially extended durations of sustainable performance. We propose a closed-loop QoS control policy, including an efficient dynamic voltage and frequency scaling (DVFS) state scheduling technique, to minimize the thermal impact for extending the sustainability of desired QoS levels. Experiments on a modern smartphone show that the proposed technique provides up to 74% longer sustainable performance while meeting the target QoS demands for a variety of real-life applications.
Onur Sahin, Paul Thomas Varghese, Ayse K. Coskun
ICCAD3
2015 PaCMap: Topology Mapping of Unstructured Communication Patterns onto Non-contiguous Allocations
abstract
In high performance computing (HPC), applications usually have many parallel tasks running on multiple machine nodes. As these tasks intensively communicate with each other, the communication overhead has a significant impact on an application's execution time. This overhead is determined by the application's communication pattern as well as the network distances between communicating tasks. By mapping the tasks to the available machine nodes in a communication-aware manner, the network distances and the execution times can be significantly reduced.
Ozan Tuncer, Vitus J. Leung, Ayse K. Coskun
ICS3
2015 Adaptive sprinting: How to get the most out of Phase Change based passive cooling
abstract
CMOS scaling trends lead to elevated on-chip temperatures, which substantially limit the performance of today's processors. To improve thermal efficiency, Phase Change Materials (PCMs) have recently been used as passive cooling solutions. PCMs store large amount of heat at near-constant temperature during phase change, allowing strategies such as computational sprinting. While existing sprinting methods allow short performance boosts, there is significant unexplored potential in improving performance on systems with PCM-enhanced cooling. To this end, this paper proposes a novel runtime management policy driven by observations that are not captured by prior techniques: (i) PCM melts non-uniformly due to spatially heterogeneous on-chip heat distribution; (ii) power consumption during sprinting is highly application dependent and assuming a fixed sprinting power leads to lower thermal efficiency; (iii) if we monitor the remaining PCM energy at various locations, we can utilize the PCM heat storage capability much more efficiently. The proposed Adaptive Sprinting policy exploits these observations to extend sprinting duration for increased performance gains. Our policy monitors the remaining PCM energy corresponding to each core at runtime, and using this information, it decides on the number, the location and the voltage-frequency (V/f) setting of the sprinting cores. Experimental evaluation including a detailed phase change thermal model demonstrates 29% performance improvement, 22% energy savings, and 43% energy delay product (EDP) reduction on average, compared to prior strategies.
Fulya Kaplan, Ayse K. Coskun
ISLPED2
2015 Dynamic Cache Pooling in 3D Multicore Processors
abstract
Resource pooling, where multiple architectural components are shared among cores, is a promising technique for improving system energy efficiency and reducing total chip area. 3D stacked multicore processors enable efficient pooling of cache resources owing to the short interconnect latency between vertically stacked layers. This article first introduces a 3D multicore architecture that provides poolable cache resources. We then propose a runtime management policy to improve energy efficiency in 3D systems by utilizing the flexible heterogeneity of cache resources. Our policy dynamically allocates jobs to cores on the 3D system while partitioning cache resources based on cache hungriness of the jobs. We investigate the impact of the proposed cache resource pooling architecture and management policy in 3D systems, both with and without on-chip DRAM. We evaluate the performance, energy efficiency, and thermal behavior for a wide range of workloads running on 3D systems. Experimental results demonstrate that the proposed architecture and policy reduce system energy-delay product (EDP) and energy-delay-area product (EDAP) by 18.8% and 36.1% on average, respectively, in comparison to 3D processors with static cache sizes.
Tiansheng Zhang, Ayse K. Coskun
ACM J. Emerg. Technol. Comput. Syst.3
2015 Leakage-Aware Cooling Management for Improving Server Energy Efficiency
abstract
The computational and cooling power demands of enterprise servers are increasing at an unsustainable rate. Understanding the relationship between computational power, temperature, leakage, and cooling power is crucial to enable energy-efficient operation at the server and data center levels. This paper develops empirical models to estimate the contributions of static and dynamic power consumption in enterprise servers for a wide range of workloads, and analyzes the interactions between temperature, leakage, and cooling power for various workload allocation policies. We propose a cooling management policy that minimizes the server energy consumption by setting the optimum fan speed during runtime. Our experimental results on a presently shipping enterprise server demonstrate that including leakage awareness in workload and cooling management provides additional energy savings without any impact on performance.
Marina Zapater, Ozan Tuncer, José Luis Ayala, José Manuel Moya, Kalyan Vaidyanathan, Kenny C. Gross, Ayse K. Coskun
IEEE Trans. Parallel Distributed Syst.7
2014 The data center as a grid load stabilizer
abstract
To accommodate the increasing presence of volatile and intermittent renewable energy sources in power generation, independent system operators (ISO) offer opportunities for demand side regulation service (RS) so as to stabilize the grid load. These power market features allow the demand side to earn monetary credits by modulating its power consumption dynamically following an RS signal broadcast by ISO. This paper studies the capacities and benefits of a major potential demand side, the data center, to provide RS. We propose a dynamic control policy that modulates the data center power consumption in response to ISO requests by leveraging server power capping techniques and various server power states. Results demonstrate that using our policy, data centers can provide fast reserves in quantities that are substantial proportions (around 50%) of their average energy consumption, with no major deterioration in quality of service (QoS). By doing so, data centers decrease their energy costs around 50%, while providing the ISOs and the society in general with cost effective demand side reserves that render massive renewable generation adoption affordable.
Hao Chen 0024, Michael C. Caramanis, Ayse K. Coskun
ASP-DAC3
2014 Detecting and identifying system changes in the cloud via discovery by example
abstract
Discovering and identifying system changes caused by events such as software installation and updates, configuration changes, and security patches are important functionalities for change management, security, compliance and problem diagnosis in emerging cloud platforms. Currently, most discovery tools use manually written rules, which require specific knowledge of software and systems. Approaches based on manually written rules are often fragile and require constant maintenance in this era of continuous integration. In this paper, we propose a novel “discovery by example” approach to autonomously search for and identify system changes. Our approach learns characteristic features of system changes automatically, without requiring any explicit rule definitions or specific knowledge of the underlying software or systems. In this approach, given a system change, our method searches a repository that contains previous stored system changes and returns those that are similar to it. We further explore the use of various forms of “fingerprints” to represent system changes efficiently and faithfully in a compact manner. We propose and evaluate two types of fingerprints: the “basename fingerprint” and the “1-D histogram fingerprint”. We show that both fingerprints exhibit different efficiency and accuracy trade-offs, and they can be effectively employed in different use cases. We evaluate the performance of our approach with both techniques and further present an application of it in system real-time streaming monitoring.
Hao Chen 0024, Sastry S. Duri, Vasanth Bala, Nilton Bila, Canturk Isci, Ayse K. Coskun
IEEE BigData6
2014 Thermal management of manycore systems with silicon-photonic networks
abstract
Silicon-photonic network-on-chips (NoCs) provide high bandwidth density; therefore, they are promising candidates to replace electrical NoCs in manycore systems. The silicon-photonic NoCs, however, are sensitive to the temperature gradients that typically occur on the chip, and hence, require proactive thermal management. This paper first provides a design space exploration of silicon-photonic networks in manycore systems and quantifies the performance impact of the temperature gradients for various network bandwidths. The paper then introduces a novel job allocation technique that minimizes the temperature gradients among the ring modulators/filters to improve the application performance. Experimental results for a single-chip 256-core system demonstrate that our policy is able to maintain the maximum network bandwidth. Compared to existing workload allocation policies, the proposed policy improves system performance by up to 26.1% when running a single application and 18.3% for multi-program scenarios.
Tiansheng Zhang, José L. Abellán, Ajay Joshi, Ayse K. Coskun
DATE4
2014 Modeling and analysis of Phase Change Materials for efficient thermal management
abstract
Direct placement of Phase Change Materials (PCMs) on the chip has been recently explored as a passive temperature management solution. PCMs provide the ability to store large amounts of heat at a close-to-constant temperature during the phase change (solid to liquid and vice versa). This latent heat capacity can be used to provide higher performance while reducing hot spots. Detailed modeling of the phase change behavior is essential for the design and evaluation of systems with PCM. This paper proposes an accurate phase change model that is integrated into the commonly used thermal simulation tool, HotSpot. It also provides validation of the proposed model by carrying out computational fluid dynamics (CFD) simulations using COMSOL Multiphysics®. This paper also explores the impact of PCM properties on the thermal profile of a processor, and demonstrates that PCM material choices can affect peak temperatures by up to 20.1°C. Experimental results show that dynamic policy decisions change dramatically when using the proposed detailed phase change model, as prior simpler PCM models can substantially over/under-estimate temperature and PCM melting duration. The proposed model helps design more effective dynamic management policies and enables realistic evaluation of systems with PCM.
Fulya Kaplan, Charlie De Vivero, Samuel Howes, Manish Arora, Houman Homayoun, Wayne P. Burleson, Dean M. Tullsen, Ayse K. Coskun
ICCD8
2014 CoolBudget: Data center power budgeting with workload and cooling asymmetry awareness
abstract
Power over-subscription challenges and emerging cost management strategies motivate designing efficient data center power capping techniques. During capping, provisioned power must be budgeted among the computational and cooling units. This work presents a data center power budgeting policy that simultaneously improves the quality-of-service (QoS) and power efficiency by considering the workload- and cooling-induced asymmetries among the servers. Proposed policy finds the most efficient data center temperature and the power distribution among servers while guaranteeing reliable temperature levels for the server internal components. Experiments based on real servers demonstrate 21% increase in throughput compared to existing techniques.
Ozan Tuncer, Kalyan Vaidyanathan, Kenny C. Gross, Ayse K. Coskun
ICCD4
2014 Sharing and placement of on-chip laser sources in silicon-photonic NoCs
abstract
Silicon-photonic links are projected to replace the electrical links for global on-chip communications in future manycore systems. The use of off-chip laser sources to drive these silicon-photonic links can lead to higher link losses, thermal mismatch between laser source and on-chip photonic devices, and packaging challenges. Therefore, on-chip laser sources are being evaluated as candidates to drive the on-chip photonic links. In this paper, we first explore the power, efficiency and temperature tradeoffs associated with an on-chip laser source. Using a 3D stacked system that integrates a manycore chip with the optical devices and laser sources, we explore the design space for laser source sharing (among waveguides) and placement to minimize laser power by simultaneously considering the network bandwidth requirements, thermal constraints, and physical layout constraints. As part of this exploration we consider Clos and crossbar logical topologies, U-shaped and W-shaped physical layouts, and various sharing/placement strategies: locally-placed dedicated laser sources for waveguides, locally-placed shared laser sources, and shared laser sources placed remotely along the chip edges. Our analysis shows that logical topology, physical layout, and photonic device losses strongly drive the laser source sharing and placement choices to minimize laser power.
Chao Chen 0003, Tiansheng Zhang, Pietro Contu, Jonathan Klamkin, Ayse K. Coskun, Ajay Joshi
NOCS5
2013 Leakage and temperature aware server control for improving energy efficiency in data centers
abstract
Reducing the energy consumption for computation and cooling in servers is a major challenge considering the data center energy costs today. To ensure energy-efficient operation of servers in data centers, the relationship among computational power, temperature, leakage, and cooling power needs to be analyzed. By means of an innovative setup that enables monitoring and controlling the computing and cooling power consumption separately on a commercial enterprise server, this paper studies temperature-leakage-energy tradeoffs, obtaining an empirical model for the leakage component. Using this model, we design a controller that continuously seeks and settles at the optimal fan speed to minimize the energy consumption for a given workload. We run a customized dynamic load-synthesis tool to stress the system. Our proposed cooling controller achieves up to 9% energy savings and 30W reduction in peak power in comparison to the default cooling control scheme.
Marina Zapater, José Luis Ayala, José Manuel Moya, Kalyan Vaidyanathan, Kenny C. Gross, Ayse K. Coskun
DATE6
2013 3D-MMC: a modular 3D multi-core architecture with efficient resource pooling
abstract
This paper demonstrates a fully functional hardware and software design for a 3D stacked multi-core system for the first time. Our 3D system is a low-power 3D Modular Multi-Core (3D-MMC) architecture built by vertically stacking identical layers. Each layer consists of cores, private and shared memory units, and communication infrastructures. The system uses shared memory communication and Through-Silicon-Vias (TSVs) to transfer data across layers. A serialization scheme is employed for inter-layer communication to minimize the overall number of TSVs. The proposed architecture has been implemented in HDL and verified on a test chip targeting an operating frequency of 400MHz with a vertical bandwidth of 3.2Gbps. The paper first evaluates the performance, power and temperature characteristics of the architecture using a set of software applications we have designed. We demonstrate quantitatively that the proposed modular 3D design improves upon the cost and performance bottlenecks of traditional 2D multi-core design. In addition, a novel resource pooling approach is introduced to efficiently manage the shared memory of the 3D stacked system. Our approach reduces the application execution time significantly compared to 2D and 3D systems with conventional memory sharing.
Tiansheng Zhang, Alessandro Cevrero, Giulia Beanato, Panagiotis Athanasopoulos, Ayse K. Coskun, Yusuf Leblebici
DATE5
2013 Dynamic server power capping for enabling data center participation in power markets
abstract
Today's US power markets offer new opportunities for the energy consumers to reduce their energy costs by first promising an average consumption rate for the next hour and then by following a regulation signal broadcast by the independent system operators (ISOs), who need to match supply and demand in real time in presence of volatile and intermittent renewable energy generation. This paper leverages the power regulation capabilities of the servers so as to enable the data centers to participate in these emerging power markets. As the data center energy consumption continues to grow, proposed participation in the power markets has the promise to achieve significant monetary savings. The paper first solves a data center regulation service (RS) optimization problem to determine the optimal average power consumption and regulation quantity that minimize the energy cost. We then propose a dynamic server power capping technique to modulate the real-time power consumption in response to ISO requests while maintaining the desired quality-of-service (QoS). Experiments on a real-life server demonstrate that our technique can reduce the energy cost by 29% on average compared to using a fixed power cap.
Hao Chen 0024, Can Hankendi, Michael C. Caramanis, Ayse K. Coskun
ICCAD4
2013 vCap: Adaptive power capping for virtualized servers
abstract
Power capping on server nodes has become an essential feature in data centers for controlling energy costs and peak power consumption. More than half of the server nodes are virtualized in today's data centers; thus, providing a practical power capping technique for consolidated virtual environments is a significant research problem. This paper proposes a power capping technique, vCap, which makes resource allocation decisions to maximize the Quality-of-Service (QoS) while meeting the power constraints in virtualized servers that run multi-threaded applications. For a given set of applications, vCap first decides which applications to co-schedule based on application scalability and then optimizes the QoS in an application-aware manner for each VM by adaptively adjusting the CPU resources. Experiments on real-life multi-core servers show that vCap provides 12% higher energy efficiency in comparison to the state-of-the-art power capping techniques, while adhering to the power cap 92% of the time within a 2W error margin.
Can Hankendi, Sherief Reda, Ayse K. Coskun
ISLPED3
2013 Dynamic cache pooling for improving energy efficiency in 3D stacked multicore processors
abstract
Resource pooling, where multiple architectural components are shared among multiple cores, is a promising technique for improving the system energy efficiency and reducing the total chip area. 3D stacked multicore processors enable efficient pooling of cache resources owing to the short interconnect latency between vertically stacked layers. This paper introduces a 3D multicore architecture that provides poolable cache resources. We propose a runtime policy that improves energy efficiency in 3D stacked processors by providing flexible heterogeneity of the cache resources. Our policy dynamically allocates jobs to cores on the 3D stacked system in a way that pairs applications with contrasting cache use, while also partitioning the cache resources based on the cache hungriness of the applications. Experimental results demonstrate that the proposed policy improves system energy-delay product (EDP) and energy-delay-area product (EDAP) by up to 39.2% and 57.2%, respectively, compared to 3D processors with static cache sizes.
Tiansheng Zhang, Ayse K. Coskun
VLSI-SoC3
2013 GreenCool: An Energy-Efficient Liquid Cooling Design Technique for 3-D MPSoCs Via Channel Width Modulation
abstract
Liquid cooling using interlayer microchannels has appeared as a viable and scalable packaging technology for 3-D multiprocessor system-on-chips (MPSoCs). Microchannel-based liquid cooling, however, can substantially increase the on-chip thermal gradients, which are undesirable for reliability, performance, and cooling efficiency. In this paper, we present GreenCool, an optimal design methodology for liquid-cooled 3-D MPSoCs. GreenCool simultaneously minimizes the cooling energy for a given system while maintaining thermal gradients and peak temperatures under safe limits. This is accomplished by tuning the heat transfer characteristics of the microchannels using channel width modulation. Channel width modulation is compatible with the current process technologies and incurs minimal additional fabrication costs. Through an extensive set of experiments, we show that channel width modulation is capable of complementing and enhancing the benefits of temperature-aware floorplanning. We also experiment with a 16-core 3-D system with stacked dynamic random-access memory, for which GreenCool improves energy efficiency by up to 53% with respect to no channel modulation.
Mohamed M. Sabry, Arvind Sridhar, Ayse K. Coskun, David Atienza 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2012 Optimizing energy efficiency of 3-D multicore systems with stacked DRAM under power and thermal constraints
abstract
3D multicore systems with stacked DRAM have the potential to boost system performance significantly; however, this performance increase may cause 3D systems to exceed the power budget or create thermal hot spots. This paper introduces a framework to model on-chip DRAM accesses and analyzes performance, power, and temperature tradeoffs of 3D systems. We propose a runtime optimization policy to maximize performance while maintaining power and thermal constraints. Our policy dynamically monitors workload behavior and selects among low-power and turbo operating modes accordingly. Experiments with multithreaded workloads demonstrate up to 49% energy efficiency improvements compared to existing thermal management policies.
Katsutoshi Kawakami, Ayse K. Coskun
DAC3
2012 Quantifying the impact of frequency scaling on the energy efficiency of the single-chip cloud computer
abstract
Dynamic frequency and voltage scaling (DVFS) techniques have been widely used for meeting energy constraints. Single-chip many-core systems bring new challenges owing to the large number of operating points and the shift to message passing interface (MPI) from shared memory communication. DVFS, however, has been mostly studied on single-chip systems with one or few cores, without considering the impact of the communication among cores. This paper evaluates the impact of frequency scaling on the performance and power of many-core systems with MPI. We conduct experiments on the Single-Chip Cloud Computer (SCC), an experimental many-core processor developed by Intel. The paper first introduces the run-time monitoring infrastructure and the application suite we have designed for an in-depth evaluation of the SCC. We provide an extensive analysis quantifying the effects of frequency perturbations on performance and energy efficiency. Experimental results show that run-time communication patterns lead to significant differences in power/performance tradeoffs in many-core systems with MPI.
Andrea Bartolini, MohammadSadegh Sadri, John-Nicholas Furst, Ayse K. Coskun, Luca Benini
DATE4
2012 Reducing the energy cost of computing through efficient co-scheduling of parallel workloads
abstract
Future computing clusters will prevalently run parallel workloads to take advantage of the increasing number of cores on chips. In tandem, there is a growing need to reduce energy consumption of computing. One promising method for improving energy efficiency is co-scheduling applications on compute nodes. Efficient consolidation for parallel workloads is a challenging task as a number of factors, such as scalability, inter-thread communication patterns, or memory access frequency of the applications affect the energy/performance tradeoffs. This paper evaluates the impact of co-scheduling parallel workloads on the energy consumed per useful work done on real-life servers. Based on this analysis, we propose a novel multi-level technique that selects the best policy to co-schedule multiple workloads on a multi-core processor. Our measurements demonstrate that the proposed multi-level co-scheduling method improves the overall energy per work savings of the multi-core system up to 22% compared to state-of-the-art techniques.
Can Hankendi, Ayse K. Coskun
DATE2
2012 Analysis and runtime management of 3D systems with stacked DRAM for boosting energy efficiency
abstract
3D stacked systems with on-chip DRAM provide high speed and wide bandwidth for accessing main memory, overcoming the limitations of slow off-chip buses. Power densities and temperatures on the chip, however, increase following the performance improvement. The complex interplay between performance, energy, and temperature on 3D systems with on-chip DRAM can only be addressed using a comprehensive evaluation framework. This paper first presents such a framework for 3D multicore systems capable of running architecture-level performance simulations along with energy and thermal evaluations, including a detailed analysis of the DRAM layers. Experimental results on 16-core 3D systems running parallel applications demonstrate up to 88.5% improvement in energy delay product compared to equivalent 2D systems. We also present a memory management policy that targets applications with spatial variations in DRAM accesses and performs temperature-aware mapping of memory accesses to DRAM banks.
Ayse K. Coskun
DATE2
2012 Topology-aware reliability optimization for multiprocessor systems
Fulya Kaplan, Ming-yu Hsieh, Ayse K. Coskun
VLSI-SoC4
2012 Introduction to the special section on adaptive power management for energy and temperature-aware computing systems
abstract
No abstract available.
Ayse K. Coskun, Yung-Hsiang Lu, Qinru Qiu
ACM Trans. Design Autom. Electr. Syst.1
2011 Run-time energy management of manycore systems through reconfigurable interconnects
abstract
The active on-chip network channel width has a direct impact on the cache and memory access latency in manycore processors. A good choice of channel width improves the application performance and energy efficiency. In manycore systems, where workload patterns change significantly over time, setting the network channel width statically for the average or worst-case traffic gives sub-optimal energy efficiency. This paper proposes a novel, low-cost method to reconfigure the network channel width at run time to maximize energy efficiency of applications. We analyze the effect of channel width choices for two commonly used cache hierarchies, private and distributed L2 caches, on manycore systems with a bus or crossbar architecture running parallel workloads. The proposed reconfiguration policy predicts the energy-delay product (EDP) for the currently running application at various channel widths and chooses the best fitting width to minimize EDP. The experimental results show that in systems with private and distributed L2 caches our policy reduces EDP by 49.3% and 23.9%, and 65.5% and 20.6% on average with bus and crossbar, respectively, in comparison to statically setting the channel width.
Chao Chen 0003, Ayse K. Coskun, Ajay Joshi
ACM Great Lakes Symposium on VLSI3
2011 Identifying the optimal energy-efficient operating points of parallel workloads
abstract
As the number of cores per processor grows, there is a strong incentive to develop parallel workloads to take advantage of the hardware parallelism. In comparison to single-threaded applications, parallel workloads are more complex to characterize due to thread interactions and resource stalls. This paper presents an accurate and scalable method for determining the optimal system operating points (i.e., number of threads and DVFS settings) at runtime for parallel workloads under a set of objective functions and constraints that optimize for energy efficiency in multi-core processors. Using an extensive training data set gathered for a wide range of parallel workloads on a commercial multi-core system, we construct multinomial logistic regression (MLR) models that estimate the optimal system settings as a function of workload characteristics. We use L1-regularization to automatically determine the relevant workload metrics for energy optimization. At runtime, our technique determines the optimal number of threads and the DVFS setting with negligible overhead. Our experiments demonstrate that our method outperforms prior techniques with up to 51% improved decision accuracy. This translates to up to 10.6% average improvement in energy-performance operation, with a maximum improvement of 30.9%. Our technique also demonstrates superior scalability as the number of potential system operating points increases.
Ryan Cochran, Can Hankendi, Ayse K. Coskun, Sherief Reda
ICCAD3
2011 Thermal analysis and active cooling management for 3D MPSoCs
abstract
3D stacked architectures reduce communication delay in multiprocessor system-on-chips (MPSoCs) and allowing more functionality per unit area. However, vertical integration of layers exacerbates the reliability and thermal problems, and cooling is a limiting factor in multi-tier systems. Liquid cooling is a highly efficient solution to overcome the accelerated thermal problems in 3D architectures. However, liquid cooling brings new challenges in modeling and run-time management. This paper proposes a design-time/run-time thermal management policy for 3D MPSoCs with inter-tier liquid cooling. First, we perform a design-time analysis to estimate the thermal impact of liquid cooling and dynamic voltage frequency scaling (DVFS) on 3D MPSoCs. Based on this analysis, we define a set of management rules for run-time thermal management. We utilize these rules to control and adjust the liquid flow rate in order to match the cooling demand for preventing energy wastage of over-cooling, while maintaining a stable thermal profile in the 3D MPSoCs. Experimental results on multi-tier 3D MPSoCs show that proposed design-time/run-time management policy prevents the system to exceed the given threshold temperature while reducing cooling energy by 50% on average and system-level energy by 18% on average in comparison to using a static worst-case flow rate setting.
Mohamed M. Sabry, David Atienza 0001, Ayse K. Coskun
ISCAS3
2011 Pack & Cap: adaptive DVFS and thread packing under power caps
abstract
The ability to cap peak power consumption is a desirable feature in modern data centers for energy budgeting, cost management, and efficient power delivery. Dynamic voltage and frequency scaling (DVFS) is a traditional control knob in the tradeoff between server power and performance. Multi-core processors and the parallel applications that take advantage of them introduce new possibilities for control, wherein workload threads are packed onto a variable number of cores and idle cores enter low-power sleep states. This paper proposes Pack & Cap, a control technique designed to make optimal DVFS and thread packing control decisions in order to maximize performance within a power budget. In order to capture the workload dependence of the performance-power Pareto frontier, a multinomial logistic regression (MLR) classifier is built using a large volume of performance counter, temperature, and power characterization data. When queried during runtime, the classifier is capable of accurately selecting the optimal operating point. We implement and validate this method on a real quad-core system running the PARSEC parallel benchmark suite. When varying the power budget during runtime, Pack & Cap meets power constraints 82% of the time even in the absence of a power measuring device. The addition of thread packing to DVFS as a control knob increases the range of feasible power constraints by an average of 21% when compared to DVFS alone and reduces workload energy consumption by an average of 51.6% compared to existing control techniques that achieve the same power range.
Ryan Cochran, Can Hankendi, Ayse K. Coskun, Sherief Reda
MICRO3
2011 Energy-Efficient Multiobjective Thermal Control for Liquid-Cooled 3-D Stacked Architectures
abstract
3-D stacked systems reduce communication delay in multiprocessor system-on-chips (MPSoCs) and enable heterogeneous integration of cores, memories, sensors, and RF devices. However, vertical integration of layers exacerbates temperature-induced problems such as reliability degradation. Liquid cooling is a highly efficient solution to overcome the accelerated thermal problems in 3-D architectures; however, it brings new challenges in modeling and run-time management for such 3-D MPSoCs with multitier liquid cooling. This paper proposes a novel design-time/run-time thermal management strategy. The design-time phase involves a rigorous thermal impact analysis of various thermal control variables. We then utilize this analysis to design a run-time fuzzy controller for improving energy efficiency in 3-D MPSoCs through liquid cooling management and dynamic voltage and frequency scaling (DVFS). The fuzzy controller adjusts the liquid flow rate dynamically to match the cooling demand of the chip for preventing overcooling and for maintaining a stable thermal profile. The DVFS decisions increase chip-level energy savings and help balance the temperature across the system. Our controller is used in conjunction with temperature-aware load balancing and dynamic power management strategies. Experimental results on 2-tier and 4-tier 3-D MPSoCs show that our strategy prevents the system from exceeding the given threshold temperature. At the same time, we reduce cooling energy by up to 63% and system-level energy by up to 21% in comparison to statically setting a flow rate setting to handle worst-case temperatures.
Mohamed M. Sabry, Ayse K. Coskun, David Atienza 0001, Tajana Rosing, Thomas Brunschwiler
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Hybrid dynamic energy and thermal management in heterogeneous embedded multiprocessor SoCs
abstract
Heterogeneous multiprocessor system-on-chips (MPSoCs) which consist of cores with various power and performance characteristics can customize their configuration to achieve higher performance per Watt. On the other hand, inherent imbalance in power densities across MPSoCs leads to non-uniform temperature distributions, which affect performance and reliability adversely. In addition, managing temperature might result in conflicting decisions with achieving higher energy efficiency. In this work, we propose a joint thermal and energy management technique specifically designed for heterogeneous MPSoCs. Our technique identifies the performance demands of the current workload. By utilizing job scheduling and voltage/frequency scaling dynamically, we meet the desired performance while minimizing the energy consumption and the thermal imbalance. In comparison to performance-aware policies such as load balancing, our technique simultaneously reduces the thermal hot spots, temperature gradients, and energy consumption significantly.
Shervin Sharifi, Ayse K. Coskun, Tajana Rosing
ASP-DAC2
2010 DynAHeal: Dynamic energy efficient task assignment for wireless healthcare systems
abstract
Energy consumption is a critical parameter in wireless healthcare systems which consist of battery operated devices such as sensors and local aggregators. The system battery lifetime depends on the allocation of processing, sensing, and communication tasks to devices of the system. In this paper, we optimize the battery life of a wireless healthcare system by efficiently assigning tasks to the available resources. There are several dynamically changing characteristics in the system, such as task parameters (processing complexity, arrival rate, and output data), each device's available battery capacity, varying wireless channel conditions, and network load. Our dynamic task assignment algorithm, ¿DynAHeal¿ adapts to such changing conditions, and improves the battery life. Our experiments show that the task assignment given by DynAHeal improves the overall system lifetime under varying dynamic conditions on an average 60% relative to sending all the data for processing to the base station, and 35% with respect to an optimal static design time assignment.
Priti Aghera, Dilip Krishnaswamy, Diana Fang, Ayse K. Coskun, Tajana Rosing
DATE4
2010 Energy-efficient variable-flow liquid cooling in 3D stacked architectures
abstract
Liquid cooling has emerged as a promising solution for addressing the elevated temperatures in 3D stacked architectures. In this work, we first propose a framework for detailed thermal modeling of the microchannels embedded between the tiers of the 3D system. In multicore systems, workload varies at runtime, and the system is generally not fully utilized. Thus, it is not energy-efficient to adjust the coolant flow rate based on the worst-case conditions, as this would cause an excess in pump power. For energy-efficient cooling, we propose a novel controller to adjust the liquid flow rate to meet the desired temperature and to minimize pump energy consumption. Our technique also includes a job scheduler, which balances the temperature across the system to maximize cooling efficiency and to improve reliability. Our method guarantees operating below the target temperature while reducing the cooling energy by up to 30%, and the overall energy by up to 12% in comparison to using the highest coolant flow rate.
Ayse K. Coskun, David Atienza 0001, Tajana Rosing, Thomas Brunschwiler, Bruno Michel
DATE1
2010 Fuzzy control for enforcing energy efficiency in high-performance 3D systems
abstract
3D stacked circuits reduce communication delay in multicore system-on-chips (SoCs) and enable heterogeneous integration of cores, memories, sensors, and RF devices. However, vertical integration of layers exacerbates the reliability and thermal problems, and cooling is a limiting factor in multi-tier systems. Liquid cooling is a highly efficient solution to overcome the accelerated thermal problems in 3D architectures; however, liquid cooling brings new challenges in modeling and runtime management. This paper proposes a novel controller for improving energy efficiency and reliability in 3D systems through liquid cooling management and dynamic voltage frequency scaling (DVFS). The proposed fuzzy controller adjusts the liquid flow rate at runtime to match the cooling demand for preventing energy wastage of over-cooling and for maintaining a stable thermal profile. The DVFS decisions provide chip-level energy savings and help balancing the temperature across the system. Experimental results on 8- and 16-core multicore SoCs show that the controller prevents the system to exceed the given threshold temperature while reducing cooling energy by up to 50% and system-level energy by up to 21% in comparison to using a static worst-case flow rate setting.
Mohamed M. Sabry, Ayse K. Coskun, David Atienza 0001
ICCAD2
2009 Dynamic thermal management in 3D multicore architectures
abstract
Technology scaling has caused the feature sizes to shrink continuously, whereas interconnects, unlike transistors, have not followed the same trend. Designing 3D stack architectures is a recently proposed approach to overcome the power consumption and delay problems associated with the interconnects by reducing the length of the wires going across the chip. However, 3D integration introduces serious thermal challenges due to the high power density resulting from placing computational units on top of each other. In this work, we first investigate how the existing thermal management, power management and job scheduling policies affect the thermal behavior in 3D chips. We then propose a dynamic thermally-aware job scheduling technique for 3D systems to reduce the thermal problems at very low performance cost. Our approach can also be integrated with power management policies to reduce energy consumption while avoiding the thermal hot spots and large temperature variations.
Ayse K. Coskun, José Luis Ayala, David Atienza 0001, Tajana Rosing, Yusuf Leblebici
DATE1
2009 Temperature- and Cost-Aware Design of 3D Multiprocessor Architectures
abstract
3D stacked architectures provide significant benefits in performance, footprint and yield. However, vertical stacking increases the thermal resistances, and exacerbates temperature-induced problems that affect system reliability, performance, leakage power and cooling cost. In addition, the overhead due to through-silicon-vias (TSVs) and scribe lines contribute to the overall area, affecting wafer utilization and yield. As any of the aforementioned parameters can limit the 3D stacking process of a multiprocessor SoC (MPSoC), in this work we investigate the tradeoffs between cost and temperature profile across various technology nodes. We study how the manufacturing costs change when the number of layers, defect density, number of cores, and power consumption vary. For each design point, we also compute the steady state temperature profile, where we utilize temperature-aware floorplan optimization to eliminate the adverse effects of inefficient floorplan decisions on temperature. Our results provide guidelines for temperature-aware floorplanning in 3D MPSoCs. For each technology node, we point out the desirable design points from both cost and temperature standpoints. For example, for building a many-core SoC with 64 cores at 32 nm, stacking 2 layers provides a desirable design point. On the other hand, at 45 nm technology, stacking 3 layers keeps temperatures at an acceptable range while reducing the cost by an additional 17% in comparison to 2 layers.
Ayse K. Coskun, Andrew B. Kahng, Tajana Rosing
DSD1
2009 Thermal Modeling and Management of Liquid-Cooled 3D Stacked Architectures
Ayse K. Coskun, José Luis Ayala, David Atienza 0001, Tajana Rosing
VLSI-SoC1
2009 Utilizing Predictors for Efficient Thermal Management in Multiprocessor SoCs
abstract
Conventional thermal management techniques are reactive, as they take action after temperature reaches a threshold. Such approaches do not always minimize and balance the temperature, and they control temperature at a noticeable performance cost. This paper investigates how to use predictors for forecasting temperature and workload dynamics, and proposes proactive thermal management techniques for multiprocessor system-on-chips. The predictors we study include autoregressive moving average modeling and lookup tables. We evaluate several reactive and predictive techniques on an UltraSPARC T1 processor and an architecture-level simulator. Proactive methods achieve significantly better thermal profiles and performance in comparison to reactive policies.
Ayse K. Coskun, Tajana Rosing, Kenny C. Gross
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Temperature-aware MPSoC scheduling for reducing hot spots and gradients
abstract
Thermal hot spots and temperature gradients on the die need to be minimized to manufacture reliable systems while meeting energy and performance constraints. In this work, we solve the task scheduling problem for multiprocessor system-on-chips (MPSoCs) using Integer Linear Programming (ILP). The goal of our optimization is minimizing the hot spots and balancing the temperature distribution on the die for a known set of tasks. Under the given assumptions about task characteristics, the solution is optimal. We compare our technique against optimal scheduling methods for energy minimization, energy balancing, and hot spot minimization, and show that our technique achieves significantly better thermal profiles. We also extend our technique to handle workload variations at runtime.
Ayse K. Coskun, Tajana Rosing, Keith Whisnant, Kenny C. Gross
ASP-DAC1
2008 Temperature management in multiprocessor SoCs using online learning
abstract
In deep submicron circuits, thermal hot spots and high temperature gradients increase the cooling costs, and degrade reliability and performance. In this paper, we propose a low-cost temperature management strategy for multicore systems to reduce the adverse effects of hot spots and temperature variations. Our technique utilizes online learning to select the best policy for the current workload characteristics among a given set of expert policies. We achieve 20% and 60% average decrease in the frequency of hot spots and thermal cycles respectively in comparison to the best performing expert, and reduce the spatial gradients to below 5%.
Ayse K. Coskun, Tajana Rosing, Kenny C. Gross
DAC1
2008 Proactive temperature balancing for low cost thermal management in MPSoCs
abstract
Designing thermal management strategies that reduce the impact of hot spots and on-die temperature variations at low performance cost is a very significant challenge for multiprocessor system-on-chips (MPSoCs). In this work, we present a proactive MPSoC thermal management approach, which predicts the future temperature and adjusts the job allocation on the MPSoC to minimize the impact of thermal hot spots and temperature variations without degrading performance. In addition, we implement and compare several reactive and proactive management strategies, and demonstrate that our proactive temperature-aware MPSoC job allocation technique is able to dramatically reduce the adverse effects of temperature at very low performance cost. We show experimental results using a simulator as well as an implementation on an UltraSPARC T1 system.
Ayse K. Coskun, Tajana Rosing, Kenny C. Gross
ICCAD1
2008 Proactive temperature management in MPSoCs
abstract
Preventing thermal hot spots and large temperature variations on the die is critical for addressing the challenges in system reliability, performance, cooling cost and leakage power. Reactive thermal management methods, which take action after temperature reaches a given threshold, maintain the temperature below a critical level at the cost of performance, and do not address the temperature variations. In this work, we propose a proactive thermal management approach, which estimates the future temperature using regression, and allocates workload on a multicore system to reduce and balance the temperature to avoid temperature induced problems. Our technique reduces the hot spots and temperature variations significantly in comparison to reactive strategies.
Ayse K. Coskun, Tajana Rosing, Kenny C. Gross
ISLPED1
2008 Static and Dynamic Temperature-Aware Scheduling for Multiprocessor SoCs
abstract
Thermal hot spots and high temperature gradients degrade reliability and performance, and increase cooling costs and leakage power. In this paper, we explore the benefits of temperature-aware task scheduling for multiprocessor system-on-a-chip (MPSoC). We evaluate our techniques using workload characteristics collected from a real system by Sun's Continuous System Telemetry. We first solve the task scheduling problem statically using integer linear programming (ILP). The ILP solution is guaranteed to be optimal for the given assumptions for tasks. We formulate ILPs for minimizing energy, balancing energy, and reducing hot spots, and provide an extensive comparison of their thermal behavior against our technique. Our static solution can reduce the frequency of hot spots by 35%, spatial gradients by 85%, and thermal cycles by 61% in comparison to the ILP for minimizing energy. We then design dynamic scheduling policies at the OS-level with negligible performance overhead. Our adaptive dynamic policy reduces the frequency of high-magnitude thermal cycles and spatial gradients by around 50% and 90%, respectively, in comparison to state-of-the-art schedulers. Reactive thermal management strategies, such as thread migration, can be combined with our scheduling policy to further reduce hot spots, temperature variations, and the associated performance cost.
Ayse K. Coskun, Tajana Rosing, Keith Whisnant, Kenny C. Gross
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Temperature aware task scheduling in MPSoCs
Ayse K. Coskun, Tajana Rosing, Keith Whisnant
DATE1
2007 Transient fault prediction based on anomalies in processor events
Satish Narayanasamy, Ayse K. Coskun, Brad Calder
DATE2
2006 A simulation methodology for reliability analysis in multi-core SoCs
abstract
Reliability has become a significant challenge for system design in new process technologies. Higher integration levels dramatically increase power densities, which leads to higher temperature and adverse effects on reliability. In this paper, we introduce a simulation methodology to analyze reliability of multi-core SoCs. The proposed simulator is the first to provide system-on-chip level fine-grained reliability analysis. We use our simulation methodology to study the reliability effects of design choices such as thermal packaging and placement, as well as runtime events such as power management policies and workload distributions.
Ayse K. Coskun, Tajana Rosing, Yusuf Leblebici, Giovanni De Micheli
ACM Great Lakes Symposium on VLSI1