Dejan S. Milojicic

dblp:53/726 · DBLP profile ↗
← Back
68ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-9830-8588ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 28 · 5 first-author · 9 since 2021Systems, architecture and hardware · 24 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Computer networks · 5Human-computer interaction and ubiquitous computing · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Variability-Guided Performance Optimization
abstract
The past few decades have seen software and hardware growing more heterogeneous and layered in abstractions. This trend produced many benefits for hiding complexity and increasing efficiency and modularity. But it also makes reasoning about performance and identifying its underlying factors more challenging because of the presence of performance variability. Moreover, performance variability can prevent synchronous applications from scaling and server applications from meeting service-level agreements.
Eitan Frachtenberg, Viyom Mittal, Mohammed Baydoun, Aditya Dhakal, Izzat El Hajj, Dejan S. Milojicic
ICPE6
2026 Are We There Yet? Predicting if Executing Applications are Near Completion
Mohammad Sonji, Mohammed Baydoun, Safaa Diab, Amir Nassereldine, Pedro Bruel, Aditya Dhakal, Rolando P. Hong Enriquez, Gourav Rattihalli, Diman Zad Tootaghaj, Gallig Renaud, Barbara M. Chapman, Fatima K. Abu Salem, Eitan Frachtenberg, Dejan S. Milojicic, Izzat El Hajj
ICPE14
2026 StreamDedup: Distributed In-line Deduplication for Disaggregated Storage
abstract
Efficient data reduction techniques, including deduplication and compression, are essential in storage systems, affecting performance and longevity. Existing data deduplication approaches often focus on intra-SSD deduplication, missing opportunities for cross-node deduplication, or have scalability issues when aiming for low latency and high-throughput data reduction on large-scale, distributed SSD arrays. We propose StreamDedup, a distributed stream accelerator implementing a transparent layer of deduplication as a network-attached, middle-tier service between the compute and storage tiers. StreamDedup manages all aspects of data deduplication and compression and can be seamlessly integrated into existing systems. It is RDMA-enabled and highly scalable, enhancing data processing capacities for large-scale storage systems. Our prototype, deployed on FPGAs, demonstrates that StreamDedup achieves a throughput of 12.7 GB/s on a single node, matching the network bandwidth of disaggregated storage, with a latency of less than 50 µs. Across 10 nodes, StreamDedup shows an almost linear increase in throughput with less than 60 µs of latency.
Jiayong Li, Jonas Dann, Zhenhao He, Gustavo Alonso, Sai Rahul Chalamalasetti, Dejan S. Milojicic, Lance Evans, Alex Veprinsky, Runbin Shi
ACM Trans. Reconfigurable Technol. Syst.6
2025 The Technology Megatrends Predictions Retrospective and Outlook
abstract
The world has never been as unpredictable as today. Technological milestones alone are being announced almost weekly, such as large language models DeepSeek, followed by Qwen2.5, and Grok 3. Socio-political, economic, and ecological development trump technology advancements and form a complex net of interdependencies. Wars are being won and lost using technology, as are the largest wins and losses on the market, and unfortunately, impact in nature. With such a speed of development and drastic impact on all facets of our lives, the ability to predict trends, especially megatrends, is becoming essential for business, technology, politics, and daily life.In computer society, we have been making annual technology predictions for over 15 years and documenting them for 10 years (see Fig. 1). In conjunction, with the IEEE Future Directions Committee, we have been exploring technology megatrends for more than two years, issuing biennial reports. Both efforts have been releasing reports and acted in a coherent, mutually supporting role even if using different approaches (online vs inperson), scope (trends vs. megatrends, technology-only vs taking into account other megatrends), and cadence (annual vs biennial). Our reports have made a tangible impact on businesses, and education, and they are widely cited and used. In this paper, we provide a short retrospective and we outline our next steps.
Dejan S. Milojicic, Phillip A. Laplante
COMPSAC1
2025 Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
abstract
In recent years, Large Language Models (LLM) such as ChatGPT, Copilot, and Gemini have been widely adopted in different areas. As the use of LLMs continues to grow, many efforts have focused on reducing the massive training overheads of these models. But it is the environmental impact of handling user requests to LLMs that is increasingly becoming a concern. Recent studies estimate that the costs of operating LLMs in their inference phase can exceed training costs by 25× per year. As LLMs are queried incessantly, the cumulative carbon footprint for the operational phase has been shown to far exceed the footprint during the training phase. Further, estimates indicate that 500 ml of fresh water is expended for every 20–50 requests to LLMs during inference. To address these important sustainability issues with LLMs, we propose a novel framework called SLIT to co-optimize LLM quality of service (time-to-first token), carbon emissions, water usage, and energy costs. The framework utilizes a machine learning (ML) based metaheuristic to enhance the sustainability of LLM hosting across geo-distributed cloud datacenters. Such a framework will become increasingly vital as LLMs proliferate.
Hayden Moore, Sirui Qi, Ninad Hogade, Dejan S. Milojicic, Cullen E. Bash, Sudeep Pasricha
ACM Great Lakes Symposium on VLSI4
2025 HeteroBench: Multi-kernel Benchmarks for Heterogeneous Systems
Hongzheng Tian, Alok Mishra 0002, Rolando P. Hong Enriquez, Dejan S. Milojicic, Eitan Frachtenberg, Sitao Huang
ICPE5
2024 2024 IEEE World Congress on Services
abstract
A warm welcome to the 2024 IEEE World Congress on Services (SERVICES). With Professor Zhi Jin and Professor Michael Sheng serving as the Congress General Chairs, I trust everyone will have a rewarding experience participating in the IEEE Computer Society's flagship annual event in services computing, whether attending on-site or remotely.
Elisa Bertino, Carl K. Chang, Rong Chang 0001, Peter Chen, Ernesto Damiani, Abdelsalam Helal, Dennis Gannon, Frank Leymann, Hong Mei 0001, Dejan S. Milojicic, Stephen S. Yau
CLOUD10
2024 Message from Rong N. Chang, Steering Committee Chair
abstract
A warm welcome to the 2024 IEEE World Congress on Services (SERVICES). With Professor Zhi Jin and Professor Michael Sheng serving as the Congress General Chairs, I trust everyone will have a rewarding experience participating in the IEEE Computer Society's flagship annual event in services computing, whether attending on-site or remotely.
Elisa Bertino, Carl K. Chang, Rong Chang 0001, Peter Chen, Ernesto Damiani, Abdelsalam Helal, Dennis Gannon, Frank Leymann, Hong Mei 0001, Dejan S. Milojicic, Stephen S. Yau
SSE10
2024 Opportunistic Energy-Aware Scheduling for Container Orchestration Platforms Using Graph Neural Networks
abstract
Reducing the energy consumption of data centers is critical to meeting international climate goals and lowering operation costs. Container orchestration platforms can help counteract this trend by optimally placing applications across the infrastructure to increase resource utilization and reduce energy consumption. But platforms in use today are still energy-agnostic and do not offer any insights into energy consumption. In this paper, we present a monitoring framework and a new modeling approach for resource usage in data centers. The model captures heterogeneous hardware and software and acts as input for a Graph Neural Network (GNN) to predict power consumption. Based on this model, we derive a set of container scheduling algorithms that opportunistically schedule applications based on the estimated energy impact of incoming containers. Our results show that the GNN-based prediction model is very accurate and achieves an average RMSE (Root Mean Square Error) of 7.5%. We have implemented a custom scheduler to demonstrate the benefits of using our prediction, and our scheduler can decrease energy consumption on average by 6.2% without any code changes for the application and without increasing workload completion time compared to the default Kubernetes scheduler.
Philipp Raith, Gourav Rattihalli, Aditya Dhakal, Sai Rahul Chalamalasetti, Dejan S. Milojicic, Eitan Frachtenberg, Stefan Nastic, Schahram Dustdar
CCGrid5
2024 Quantum optimization algorithms: Energetic implications
abstract
Summary Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. In this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least to more energy efficient than their classical counterparts; why feedback‐based quantum optimization is energy‐inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of 20. Finally, we use the NchooseK high‐level programming model to run optimization problems on both gate‐based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate‐based counterparts.
Rolando P. Hong Enriquez, Rosa M. Badia, Barbara M. Chapman, Kirk Bresniker, Scott Pakin, Alok Mishra 0002, Pedro Bruel, Aditya Dhakal, Gourav Rattihalli, Ninad Hogade, Eitan Frachtenberg, Dejan S. Milojicic
Concurr. Comput. Pract. Exp.12
2023 Fine-Grained Heterogeneous Execution Framework with Energy Aware Scheduling
abstract
The growing convergence of high-performance, data analytics, and machine-learning applications is increasingly pushing computing systems toward heterogeneous processors and specialized hardware accelerators. Hardware heterogeneity, in turn, leads to finer-grained workflows. State-of-the-art server-less computing resource managers do not currently provide efficient scheduling of such fine-grained tasks on systems with heterogeneous CPUs and specialized hardware accelerators (e.g., GPUs and FPGAs). Working with fine-grained tasks presents an opportunity for more efficient energy use via new scheduling models. Our proposed scheduler enables technologies like Nvidia's Multi-Process Service (MPS) to pack multiple fine-grained tasks on GPUs efficiently. Its advantages include better co-location of jobs and better sharing of hardware resources such as GPUs that were not previously possible on container orchestration systems. We propose a Kubernetes-native energy-aware scheduler that integrates with our heterogeneous framework. Combining fine-grained resource scheduling on heterogeneous hardware and energy-aware scheduling results in up to 17.6% improvement in makespan, up to 20.16% reduction in energy consumption for CPU workloads, and up to 58.15% improvement in makespan, and up to 28.92% reduction in energy consumption for GPU workloads.
Gourav Rattihalli, Ninad Hogade, Aditya Dhakal, Eitan Frachtenberg, Rolando P. Hong Enriquez, Pedro Bruel, Alok Mishra 0002, Dejan S. Milojicic
CLOUD8
2023 The Art of Prediction
abstract
Predictions have always attracted interest, because seeing the future could be very useful, powerful, and also fun. Those who can predict ahead of others have a strategic advantage. Predictions are essential in business, the military [1], and healthcare. As of recently, with COVID and with recent wars, predictions became critical to humankind's survival: for predicting when to open or close countries, predicting supply chains, vaccines, etc. Predictions are hard because they depend on many factors. There is a social aspect to making any prediction that pertains to humans hard. There is an ecological aspect that is extremely complex on its own. Technology predictions may be simplest, but technology also depends on business success, i.e. economics. For these reasons, technology predictions are a combination of art, science, and business. Teams in the IEEE Computer Society and Future Directions, led by the author of this paper, have conducted technology predictions for more than a decade. We predicted technologies for the following year, eight years out, we predicted individual technologies and megatrends comprising a multitude of technologies. We evaluated our predictions to gain insights into how successful we were and to improve our processes. Our predictions have gained a lot of interest in the community, resulting in regular press releases growing in target audiences. The most recent one has a target audience larger than 200M. We started special issues of IEEE Computer on Technology Predictions (the 4thone will appear in July 2023) and a quarterly column on Predictions in the same magazine (the 8thcolumn will also appear in July 2023). We have regularly organized panels and given keynotes on technology predictions. In this paper, we summarize our experience and “predict” our future work.
Dejan S. Milojicic
SSE1
2023 Predicting the Performance-Cost Trade-off of Applications Across Multiple Systems
abstract
In modern computing environments, users may have multiple systems accessible to them such as local clusters, private clouds, or public clouds. This abundance of choices makes it difficult for users to select the system and configuration for running an application that best meet their performance and cost objectives. To assist such users, we propose a prediction tool that predicts the full performance-cost trade-off space of an application across multiple systems. Our tool runs and profiles a submitted application on a small number of configurations from some of the systems, and uses that information to predict the application's performance on all configurations in all systems. The prediction models are trained offline with data collected from running a large number of applications on a wide variety of configurations. Notable aspects of our tool include: providing different scopes of prediction with varying online profiling requirements, automating the selection of the small number of configurations and systems used for online profiling, performing online profiling using partial runs thereby make predictions for applications without running them to completion, employing a classifier to distinguish applications that scale well from those that scale poorly, and predicting the sensitivity of applications to interference from other users. We evaluate our tool using 69 data analytics and scientific computing benchmarks executing on three different single-node CPU systems with 8–9 configurations each and show that it can achieve low prediction error with modest profiling overhead.
Amir Nassereldine, Safaa Diab, Mohammed Baydoun, Kenneth Leach, Maxim Alt, Dejan S. Milojicic, Izzat El Hajj
CCGrid6
2023 MOSAIC: A Multi-Objective Optimization Framework for Sustainable Datacenter Management
abstract
In recent years, cloud service providers have been building and hosting datacenters across multiple geographical locations to provide robust services. However, the geographical distribution of datacenters introduces growing pressure to both local and global environments, particularly when it comes to water usage and carbon emissions. Unfortunately, efforts to reduce the environmental impact of such datacenters often lead to an increase in the cost of datacenter operations. To co-optimize the energy cost, carbon emissions, and water footprint of datacenter operation from a global perspective, we propose a novel framework for multi-objective sustainable datacenter management (MOSAIC) that integrates adaptive local search with a collaborative decomposition-based evolutionary algorithm to intelligently manage geographical workload distribution and datacenter operations. Our framework sustainably allocates workloads to datacenters while taking into account multiple geography- and time-based factors including renewable energy sources, variable energy costs, power usage efficiency, carbon factors, and water intensity in energy. Our experimental results show that, compared to the best-known prior work frameworks, MOSAIC can achieve 27.45× speedup and 1.53 × improvement in Pareto Hypervolume while reducing the carbon footprint by up to 1.33 ×, water footprint by up to 3.09 ×, and energy costs by up to 1.40 ×. In the simultaneous three-objective co-optimization scenario, MOSAIC achieves a cumulative improvement across all objectives (carbon, water, cost) of up to 4.61 × compared to the state-of-the-arts.
Sirui Qi, Dejan S. Milojicic, Cullen E. Bash, Sudeep Pasricha
HiPC2
2023 Kernel-as-a-Service: A Serverless Programming Model for Heterogeneous Hardware Accelerators
abstract
With the slowing of Moore's law and decline of Dennard scaling, computing systems increasingly rely on specialized hardware accelerators in addition to general-purpose compute units. Increased hardware heterogeneity necessitates disaggregating applications into workflows of fine-grained tasks that run on a diverse set of CPUs and accelerators. Current accelerator delivery models cannot support such applications efficiently, as (1) the overhead of managing accelerators erases performance benefits for fine-grained tasks; (2) exclusive accelerator use per task leads to underutilization; and (3) specialization increases complexity for developers.
Tobias Pfandzelter, Aditya Dhakal, Eitan Frachtenberg, Sai Rahul Chalamalasetti, Darel Emmot, Ninad Hogade, Rolando P. Hong Enriquez, Gourav Rattihalli, David Bermbach, Dejan S. Milojicic
Middleware10
2022 Farview: Disaggregated Memory with Operator Off-loading for Database Engines
Dario Korolija, Dimitrios Koutsoukos, Kimberly Keeton, Konstantin Taranov, Dejan S. Milojicic, Gustavo Alonso
CIDR5
2022 Sharing non-cache-coherent memory with bounded incoherence
abstract
Summary Cache coherence in modern computer architectures enables easier programming by sharing data across multiple processors. Unfortunately, it can also limit scalability due to cache coherency traffic initiated by competing memory accesses. Rack‐scale systems introduce shared memory across a whole rack, but without inter‐node cache coherence. This poses memory management and concurrency control challenges for applications that must explicitly manage cache‐lines. To fully utilize rack‐scale systems for low‐latency and scalable computation, applications need to maintain cached memory accesses in spite of non‐coherency. This paper introduces Bounded Incoherence, a memory consistency model that enables cached access to shared data‐structures in non‐cache‐coherency memory. It ensures that updates to memory on one node are visible within at most a bounded amount of time on all other nodes. We evaluate this memory model on modified PowerGraph graph processing framework, and boost its performance by 30% with eight sockets by enabling cached‐access to data‐structures.
Yuxin Ren 0001, Gabriel Parmer, Dejan S. Milojicic
Concurr. Comput. Pract. Exp.3
2021 Mixed Precision Quantization for ReRAM-based DNN Inference Accelerators
abstract
ReRAM-based accelerators have shown great potential for accelerating DNN inference because ReRAM crossbars can perform analog matrix-vector multiplication operations with low latency and energy consumption. However, these crossbars require the use of ADCs which constitute a significant fraction of the cost of MVM operations. The overhead of ADCs can be mitigated via partial sum quantization. However, prior quantization flows for DNN inference accelerators do not consider partial sum quantization which is not highly relevant to traditional digital architectures. To address this issue, we propose a mixed precision quantization scheme for ReRAM-based DNN inference accelerators where weight quantization, input quantization, and partial sum quantization are jointly applied for each DNN layer. We also propose an automated quantization flow powered by deep reinforcement learning to search for the best quantization configuration in the large design space. Our evaluation shows that the proposed mixed precision quantization scheme and quantization flow reduce inference latency and energy consumption by up to 3.89x and 4.84x, respectively, while only losing 1.18% in DNN inference accuracy.
Sitao Huang, Aayush Ankit, Plínio Silveira, Rodrigo Antunes, Sai Rahul Chalamalasetti, Izzat El Hajj, Dong Eun Kim, Glaucimar Aguiar, Pedro Bruel, Sergey Serebryakov, Can Li 0024, Paolo Faraboschi, John Paul Strachan, Deming Chen, Kaushik Roy 0001, Wen-Mei W. Hwu, Dejan S. Milojicic
ASP-DAC18
2021 Resource Sharing and Security Implications on Machine Learning Inference Accelerators
abstract
Due to the increasing adoption of Machine Learning (ML) and in particular Deep Learning (DL), many specialized energy efficient accelerators are being proposed by academia and industry. A number of these accelerators are designed to run a single application at a time in exclusive access mode. This approach gives applications maximum performance but reduces resource efficiency, resulting in increased costs over time. Sharing the device among multiple jobs increases resource utilization and amplifies return on investment. This study is driven by a broad investigation of various spatial resource sharing strategies in machine learning hardware accelerators and performance evaluation in a novel memristor-based accelerator called PUMA [1]. Two methods of spatial sharing are discussed: Model Packing and Logical Allocation. Simulations showed that both methods can be implemented on the PUMA accelerator and have advantages in terms of increased resource utilization. The former spatial sharing strategy achieves higher level of parallelism, fitting more models per device (7 models on 11 tiles), but has higher interference overhead (up to 49%), still being in most cases better than the overhead found for GPUs. The latter spatial sharing strategy achieves better isolation with almost no interference overhead (<1%) with the cost of leaving resources unused (same 7 models consumed 16 tiles). Finally, we discuss security implications of resource sharing for ML and other concerns, presenting a novel ML model integrity check and model bias verification.
Plínio Silveira, César A. F. De Rose, Avelino Francisco Zorzo, Miguel G. Xavier, Dejan S. Milojicic, Sai Rahul Chalamalasetti, Sergey Serebryakov
COMPSAC5
2021 Future of HPC: Diversifying Heterogeneity
abstract
After the end of Dennard scaling and with the imminent end of Moore's Law, it has become challenging to continue scaling HPC systems within a given power envelope. This is exacerbated most in large systems, such as high end supercomputers. To alleviate this problem, general purpose is no longer sufficient, and HPC systems and components are being augmented with special-purpose hardware. By definition, because of the narrow applicability of specialization, broad supercomputing adoption requires using different heterogeneous components, each optimized for a specific application domain. In this paper, we discuss the impact of the introduced heterogeneity of specialization across the HPC stack: interconnects including memory models, accelerators including power and cooling, use cases and applications including AI, and delivery models, such as traditional, as-a-Service, and federated. We believe that a stack that supports diversification across hardware and software is required to continue scaling performance and maintaining energy efficiency.
Dejan S. Milojicic, Paolo Faraboschi, Nicolas Dubé, Duncan Roweth
DATE1
2020 PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-Efficient ReRAM
abstract
The wide adoption of deep neural networks has been accompanied by ever-increasing energy and performance demands due to the expensive nature of training them. Numerous special-purpose architectures have been proposed to accelerate training: both digital and hybrid digital-analog using resistive RAM (ReRAM) crossbars. ReRAM-based accelerators have demonstrated the effectiveness of ReRAM crossbars at performing matrix-vector multiplication operations that are prevalent in training. However, they still suffer from inefficiency due to the use of serial reads and writes for performing the weight gradient and update step. A few works have demonstrated the possibility of performing outer products in crossbars, which can be used to realize the weight gradient and update step without the use of serial reads and writes. However, these works have been limited to low precision operations which are not sufficient for typical training workloads. Moreover, they have been confined to a limited set of training algorithms for fully-connected layers only. To address these limitations, we propose a bit-slicing technique for enhancing the precision of ReRAM-based outer products, which is substantially different from bit-slicing for matrix-vector multiplication only. We incorporate this technique into a crossbar architecture with three variants catered to different training algorithms. To evaluate our design on different types of layers in neural networks (fully-connected, convolutional, etc.) and training algorithms, we develop PANTHER, an ISA-programmable training accelerator with compiler support. Our design can also be integrated into other accelerators in the literature to enhance their efficiency. Our evaluation shows that PANTHER achieves up to 8.02×, 54.21×, and 103× energy reductions as well as 7.16×, 4.02×, and 16× execution time reductions compared to digital accelerators, ReRAM-based accelerators, and GPUs, respectively.
Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Sapan Agarwal, Matthew J. Marinella, Martin Foltin, John Paul Strachan, Dejan S. Milojicic, Wen-Mei W. Hwu, Kaushik Roy 0001
IEEE Trans. Computers8
2019 PUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning Inference
abstract
Memristor crossbars are circuits capable of performing analog matrix-vector multiplications, overcoming the fundamental energy efficiency limitations of digital logic. They have been shown to be effective in special-purpose accelerators for a limited set of neural network applications. We present the Programmable Ultra-efficient Memristor-based Accelerator (PUMA) which enhances memristor crossbars with general purpose execution units to enable the acceleration of a wide variety of Machine Learning (ML) inference workloads. PUMA's microarchitecture techniques exposed through a specialized Instruction Set Architecture (ISA) retain the efficiency of in-memory computing and analog circuitry, without compromising programmability. We also present the PUMA compiler which translates high-level code to PUMA ISA. The compiler partitions the computational graph and optimizes instruction scheduling and register allocation to generate code for large and complex workloads to run on thousands of spatial cores. We have developed a detailed architecture simulator that incorporates the functionality, timing, and power models of PUMA's components to evaluate performance and energy consumption. A PUMA accelerator running at 1 GHz can reach area and power efficiency of 577 GOPS/s/mm 2 and 837~GOPS/s/W, respectively. Our evaluation of diverse ML applications from image recognition, machine translation, and language modelling (5M-800M synapses) shows that PUMA achieves up to 2,446× energy and 66× latency improvement for inference compared to state-of-the-art GPUs. Compared to an application-specific memristor-based accelerator, PUMA incurs small energy overheads at similar inference latency and added programmability.
Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Geoffrey Ndu, Martin Foltin, R. Stanley Williams, Paolo Faraboschi, Wen-Mei W. Hwu, John Paul Strachan, Kaushik Roy 0001, Dejan S. Milojicic
ASPLOS11
2019 Analysis and Modeling of Collaborative Execution Strategies for Heterogeneous CPU-FPGA Architectures
abstract
Heterogeneous CPU-FPGA systems are evolving towards tighter integration between CPUs and FPGAs for improved performance and energy efficiency. At the same time, programmability is also improving with High Level Synthesis tools (e.g., OpenCL Software Development Kits), which allow programmers to express their designs with high-level programming languages, and avoid time-consuming and error-prone register-transfer level (RTL) programming. In the traditional loosely-coupled accelerator mode, FPGAs work as offload accelerators, where an entire kernel runs on the FPGA while the CPU thread waits for the result. However, tighter integration of the CPUs and the FPGAs enables the possibility of fine-grained collaborative execution, i.e., having both devices working concurrently on the same workload. Such collaborative execution makes better use of the overall system resources by employing both CPU threads and FPGA concurrency, thereby achieving higher performance. In this paper, we explore the potential of collaborative execution between CPUs and FPGAs using OpenCL High Level Synthesis. First, we compare various collaborative techniques (namely, data partitioning and task partitioning), and evaluate the tradeoffs between them. We observe that choosing the most suitable partitioning strategy can improve performance by up to 2x. Second, we study the impact of a common optimization technique, kernel duplication, in a collaborative CPU-FPGA context. We show that the general trend is that kernel duplication improves performance until the memory bandwidth saturates. Third, we provide new insights that application developers can use when designing CPU-FPGA collaborative applications to choose between different partitioning strategies. We find that different partitioning strategies pose different tradeoffs (e.g., task partitioning enables more kernel duplication, while data partitioning has lower communication overhead and better load balance), but they generally outperform execution on conventional CPU-FPGA systems where no collaborative execution strategies are used. Therefore, we advocate even more integration in future heterogeneous CPU-FPGA systems (e.g., OpenCL 2.0 features, such as fine-grained shared virtual memory).
Sitao Huang, Li-Wen Chang, Izzat El Hajj, Simon Garcia de Gonzalo, Juan Gómez-Luna, Sai Rahul Chalamalasetti, Mohamed El-Hadedy 0001, Dejan S. Milojicic, Onur Mutlu, Deming Chen, Wen-Mei W. Hwu
ICPE8
2019 Memory-Side Protection With a Capability Enforcement Co-Processor
abstract
Byte-addressable nonvolatile memory (NVM) blends the concepts of storage and memory and can radically improve data-centric applications, from in-memory databases to graph processing. By enabling large-capacity devices to be shared across multiple computing elements, fabric-attached NVM changes the nature of rack-scale systems and enables short-latency direct memory access while retaining data persistence properties and simplifying the software stack. An adequate protection scheme is paramount when addressing shared and persistent memory, but mechanisms that rely on virtual memory paging suffer from the tension between performance (pushing toward large pages) and protection granularity (pushing toward small pages). To address this tension, capabilities are worth revisiting as a more powerful protection mechanism, but the long time needed to introduce new CPU features hampers the adoption of schemes that rely on instruction-set architecture support. This article proposes the Capability Enforcement Co-Processor (CEP), a programmable memory controller that implements fine-grain protection through the capability model without requiring instruction-set support in the application CPU. CEP decouples capabilities from the application CPU instruction-set architecture, shortens time to adoption, and can rapidly evolve to embrace new persistent memory technologies, from NVDIMMs to native NVM devices, either locally connected or fabric attached in rack-scale configurations. CEP exposes an application interface based on memory handles that get internally converted to extended-pointer capabilities. This article presents a proof of concept implementation of a distributed object store (Redis) with CEP. It also demonstrates a capability-enhanced file system (FUSE) implementation using CEP. Our proof of concept shows that CEP provides fine-grain protection while enabling direct memory access from application clients to the NVM, and that by doing so opens up important performance optimization opportunities (up to 4× reduction in latency in comparison to software-based security enforcement) without compromising security. Finally, we also sketch how a future hybrid model could improve the initial implementation by delegating some CEP functionality to a CHERI-enabled processor.
Leonid Azriel, Lukas Humbel, Reto Achermann, Alex Richardson 0001, Moritz Hoffmann 0001, Avi Mendelson, Timothy Roscoe, Robert N. M. Watson, Paolo Faraboschi, Dejan S. Milojicic
ACM Trans. Archit. Code Optim.10
2018 Computing In-Memory, Revisited
abstract
The Von Neumann's architecture has been the dominant computing paradigm ever since its inception in the mid-forties. It revolves around the concept of a "stored program" in memory, and a central processing unit that executes the program. As an alternative, Processing-In-Memory (PIM) ideas have been around for at least two decades, however with very limited adoption. Today, three trends are creating a compelling motivation to take a second look. Novel devices such as memristor blur the boundary between memory and compute, effectively providing both in the same element. Power efficiency has become very important, both in the datacenter and at the edge. Machine learning applications driven by a data-flow model have become ubiquitous. In this paper, we sketch our Computing-In-Memory (CIM) vision, and its substantial performance and power improvement potential. Compared to PIM models, CIM more clearly separates computing from memory. We then discuss the programming model, which we consider the biggest challenge. We close by describing how CIM impacts non-functional characteristics, such as reliability, scale, and configurability.
Dejan S. Milojicic, Kirk Bresniker, Gary Campbell, Paolo Faraboschi, John Paul Strachan, Stan Williams
ICDCS1
2017 Separating Translation from Protection in Address Spaces with Dynamic Remapping
abstract
It is time to reconsider memory protection. The emergence of large non-volatile main memories, scalable interconnects, and rack-scale computers running large numbers of small "micro services" creates significant challenges for memory protection based solely on MMU mechanisms. Central to this is a tension between protection and translation: optimizing for translation performance often comes with a cost in protection flexibility.
Reto Achermann, Chris I. Dalton, Paolo Faraboschi, Moritz Hoffmann 0001, Dejan S. Milojicic, Geoffrey Ndu, Alex Richardson 0001, Timothy Roscoe, Adrian L. Shaw, Robert N. M. Watson
HotOS5
2017 Adaptive scheduling of parallel jobs in spark streaming
abstract
Streaming data analytics has become increasingly vital in many applications such as dynamic content delivery (e.g., advertisements), Twitter sentiment analysis, and security event processing (e.g., intrusion detection systems, and spam filters). Emerging stream processing systems, such as Spark Streaming, treat the continuous stream as a series of micro-batches of data and continuously process these micro-batch jobs. Such micro-batch based stream processing provides several advantages over traditional stream processing systems, which process streaming data one record at a time, including fast recovery from failures, better load balancing and scalability. However, efficient scheduling of micro-batch jobs to achieve high throughput and low latency is very challenging due to the complex data dependency and dynamism inherent in streaming workloads. In this paper, we propose A-scheduler, an adaptive scheduling approach that dynamically schedules parallel micro-batch jobs in Spark Streaming and automatically adjusts scheduling parameters to improve performance and resource efficiency. Specifically, A-scheduler dynamically schedules multiple jobs concurrently using different policies based on their data dependencies and automatically adjusts the level of job parallelism and resource shares among jobs based on workload properties. We implemented A-scheduler and evaluated it with a real-time security event processing workload. Our experimental results show that A-scheduler can reduce end-to-end latency by 42% and improve workload throughput and energy efficiency by 21% and 13%, respectively, compared to the default Spark Streaming scheduler.
Dazhao Cheng, Yuan Chen 0001, Xiaobo Zhou 0002, Daniel Gmach, Dejan S. Milojicic
INFOCOM5
2017 InterSCSimulator: Large-Scale Traffic Simulation in Smart Cities Using Erlang
Eduardo Felipe Zambom Santana, Nelson Lago, Fabio Kon, Dejan S. Milojicic
MABS4
2017 SAVI objects: sharing and virtuality incorporated
abstract
Direct sharing and storing of memory objects allows high-performance and low-overhead collaboration between parallel processes or application workflows with loosely coupled programs. However, sharing of objects is hindered by the inability to use subtype polymorphism which is common in object-oriented programming languages. That is because implementations of subtype polymorphism in modern compilers rely on using virtual tables stored at process-specific locations, which makes objects unusable in processes other than the creating process. In this paper, we present SAVI Objects, objects with Sharing and Virtuality Incorporated. SAVI Objects support subtype polymorphism but can still be shared across processes and stored in persistent data structures. We propose two different techniques to implement SAVI Objects and evaluate the tradeoffs between them. The first technique is virtual table duplication which adheres to the virtual-table-based implementation of subtype polymorphism, but duplicates virtual tables for shared objects to fixed memory addresses associated with each shared memory region. The second technique is hashing-based dynamic dispatch which re-implements subtype polymorphism using hashing-based look-ups to a global virtual table. Our results show that SAVI Objects enable direct sharing and storing of memory objects that use subtype polymorphism by adding modest overhead costs to object construction and dynamic dispatch time. SAVI Objects thus enable faster inter-process communication, improving the overall performance of production applications that share polymorphic objects.
Izzat El Hajj, Thomas B. Jablin, Dejan S. Milojicic, Wen-Mei W. Hwu
Proc. ACM Program. Lang.3
2017 Concurrent Log-Structured Memory for Many-Core Key-Value Stores
abstract
Key-value stores are an important tool in managing and accessing large in-memory data sets. As many applications benefit from having as much of their working state fit into main memory, an important design of the memory management of modern key-value stores is the use of log-structured approaches, enabling efficient use of the memory capacity, by compacting objects to avoid fragmented states. However, with the emergence of thousand-core and peta-byte memory platforms (DRAM or future storage-class memories) log-structured designs struggle to scale, preventing parallel applications from exploiting the full capabilities of the hardware: careful coordination is required for background activities (compacting and organizing memory) to remain asynchronous with respect to the use of the interface, and for insertion operations to avoid contending for centralized resources such as the log head and memory pools. In this work, we present the design of a log-structured key-value store called Nibble that incorporates a multi-head log for supporting concurrent writes, a novel distributed epoch mechanism for scalable memory reclamation, and an optimistic concurrency index. We implement Nibble in the Rust language in ca. 4000 lines of code, and evaluate it across a variety of data-serving workloads on a 240-core cache-coherent server. Our measurements show Nibble scales linearly in uniform YCSB workloads, matching competitive non-log-structured key-value stores for write- dominated traces at 50 million operations per second on 1 TiB-sized working sets. Our memory analysis shows Nibble is efficient, requiring less than 10% additional capacity, whereas memory use by non-log-structured key-value store designs may be as high as 2x.
Alex Merritt, Ada Gavrilovska, Yuan Chen 0001, Dejan S. Milojicic
Proc. VLDB Endow.4
2016 Enabling Elastic Stream Processing in Shared Clusters
abstract
Distributed data stream processing has become an increasingly popular computational framework due to many emerging applications which require real-time processing of data such as dynamic content delivery and security event analysis. These distributed data stream processing applications are often run on shared, multi-tenant clusters as companies try to consolidate from dedicated clusters for each application (batch and streaming) to a single cluster using a global cluster manager such as Hadoop YARN. In shared cluster environments, guaranteeing the quality of service constraints for throughput and response time for both stream processing applications and batch applications is a significant challenge. Stream processing applications often face an elastic demand where the input rate can vary drastically. The typical solution to solve workload elasticity is to guarantee enough resources to the application, but this solution is not possible when resources are being shared among multiple applications. In this paper, we present an approach for supporting elastic scaling of distributed data stream processing applications and efficiently scheduling and coordinating stream processing with batch processing in shared clusters. Our solution consists of a congestion detection monitor which detects bottlenecks in the streaming system and a global state manager that performs non-disruptive, stateful scaling of streaming applications. We implemented our solution using Storm, a popular stream processing framework, and tested our implementation on a Hadoop YARN cluster using a real-time security event processing workload. Our experimental results show that our solution improves stream processing application throughput by 49% over default Storm while decreasing average request response times by 58%.
Jack Li 0001, Calton Pu, Yuan Chen 0001, Daniel Gmach, Dejan S. Milojicic
CLOUD5
2016 SpaceJMP: Programming with Multiple Virtual Address Spaces
abstract
Memory-centric computing demands careful organization of the virtual address space, but traditional methods for doing so are inflexible and inefficient. If an application wishes to address larger physical memory than virtual address bits allow, if it wishes to maintain pointer-based data structures beyond process lifetimes, or if it wishes to share large amounts of memory across simultaneously executing processes, legacy interfaces for managing the address space are cumbersome and often incur excessive overheads. We propose a new operating system design that promotes virtual address spaces to first-class citizens, enabling process threads to attach to, detach from, and switch between multiple virtual address spaces. Our work enables data-centric applications to utilize vast physical memory beyond the virtual range, represent persistent pointer-rich data structures without special pointer representations, and share large amounts of memory between processes efficiently.
Izzat El Hajj, Alex Merritt, Gerd Zellweger, Dejan S. Milojicic, Reto Achermann, Paolo Faraboschi, Wen-Mei W. Hwu, Timothy Roscoe, Karsten Schwan
ASPLOS4
2016 KLAP: Kernel launch aggregation and promotion for optimizing dynamic parallelism
abstract
Dynamic parallelism on GPUs simplifies the programming of many classes of applications that generate paral-lelizable work not known prior to execution. However, modern GPUs architectures do not support dynamic parallelism efficiently due to the high kernel launch overhead, limited number of simultaneous kernels, and limited depth of dynamic calls a device can support. In this paper, we propose Kernel Launch Aggregation and Promotion (KLAP), a set of compiler techniques that improve the performance of kernels which use dynamic parallelism. Kernel launch aggregation fuses kernels launched by threads in the same warp, block, or kernel into a single aggregated kernel, thereby reducing the total number of kernels spawned and increasing the amount of work per kernel to improve occupancy. Kernel launch promotion enables early launch of child kernels to extract more parallelism between parents and children, and to aggregate kernel launches across generations mitigating the problem of limited depth. We implement our techniques in a real compiler and show that kernel launch aggregation obtains a geometric mean speedup of 6.58x over regular dynamic parallelism. We also show that kernel launch promotion enables cases that were not originally possible, improving throughput by a geometric mean of 30.44 x.
Izzat El Hajj, Juan Gómez-Luna, Cheng Li 0014, Li-Wen Chang, Dejan S. Milojicic, Wen-Mei W. Hwu
MICRO5
2016 Evaluating and Improving the Performance and Scheduling of HPC Applications in Cloud
abstract
Cloud computing is emerging as a promising alternative to supercomputers for some high-performance computing (HPC) applications. With cloud as an additional deployment option, HPC users and providers are faced with the challenges of dealing with highly heterogeneous resources, where the variability spans across a wide range of processor configurations, interconnects, virtualization environments, and pricing models. In this paper, we take a holistic viewpoint to answer the question-why and whoshould choose cloud for HPC, for what applications, and how should cloud be used for HPC? To this end, we perform comprehensive performance and cost evaluation and analysis of running a set of HPC applications on a range of platforms, varying from supercomputers to clouds. Further, we improve performance of HPC applications in cloud by optimizing HPC applications' characteristics for cloud and cloud virtualization mechanisms for HPC. Finally, we present novel heuristics for online application-aware job scheduling in multi-platform environments. Experimental results and simulations using CloudSim show that current clouds cannot substitute supercomputers but can effectively complement them. Significant improvement in average turnaround time (up to 2X)and throughput (up to 6X) can be attained using our intelligent application-aware dynamic scheduling heuristics compared tosingle-platform or application-agnostic scheduling.
Abhishek Gupta 0002, Paolo Faraboschi, Filippo Gioachin, Laxmikant V. Kalé, Richard Kaufmann, Bu-Sung Lee, Verdi March, Dejan S. Milojicic, Chun Hui Suen
IEEE Trans. Cloud Comput.8
2015 Beyond Processor-centric Operating Systems
Paolo Faraboschi, Kimberly Keeton, Tim Marsland, Dejan S. Milojicic
HotOS4
2015 Not Your Parents' Physical Address Space
Simon Gerber, Gerd Zellweger, Reto Achermann, Kornilios Kourtis, Timothy Roscoe, Dejan S. Milojicic
HotOS6
2015 Improving Preemptive Scheduling with Application-Transparent Checkpointing in Shared Clusters
abstract
Modern data center clusters are shifting from dedicated single framework clusters to shared clusters. In such shared environments, cluster schedulers typically utilize preemption by simply killing jobs in order to achieve resource priority and fairness during peak utilization. This can cause significant resource waste and delay job response time.
Jack Li 0001, Calton Pu, Yuan Chen 0001, Vanish Talwar, Dejan S. Milojicic
Middleware5
2015 Bringing Test-Driven Development to web service choreographies
Felipe M. Besson, Paulo Moura, Fabio Kon, Dejan S. Milojicic
J. Syst. Softw.4
2013 Improving HPC Application Performance in Cloud through Dynamic Load Balancing
abstract
Driven by the benefits of elasticity and pay-as-you-go model, cloud computing is emerging as an attractive alternative and addition to in-house clusters and supercomputers for some High Performance Computing (HPC) applications. However, poor interconnect performance, heterogeneous and dynamic environment, and interference by other virtual machines (VMs) are some bottlenecks for efficient HPC in cloud. For tightly-coupled iterative applications, one slow processor slows down the entire application, resulting in poor CPU utilization. In this paper, we present a dynamic load balancer for tightly-coupled iterative HPC applications in cloud. It infers the static hardware heterogeneity in virtualized environments, and also adapts to the dynamic heterogeneity caused by the interference arising due to multi-tenancy. Through continuous live monitoring, instrumentation, and periodic refinement of task distribution to VMs, our load balancer adapts to the dynamic variations in cloud resources. Through experimental evaluation on a private cloud with 64 VMs using benchmarks and a real science application, we demonstrate performance benefits up to 45%. Finally, we analyze the effect of load balancing frequency, problem size, and computational granularity (problem decomposition) on the performance and scalability of our techniques.
Abhishek Gupta 0002, Osman Sarood, Laxmikant V. Kalé, Dejan S. Milojicic
CCGRID4
2013 The Who, What, Why, and How of High Performance Computing in the Cloud
abstract
Cloud computing is emerging as an alternative to supercomputers for some of the high-performance computing (HPC) applications that do not require a fully dedicated machine. With cloud as an additional deployment option, HPC users are faced with the challenges of dealing with highly heterogeneous resources, where the variability spans across a wide range of processor configurations, interconnections, virtualization environments, and pricing rates and models. In this paper, we take a holistic viewpoint to answer the question - why and who should choose cloud for HPC, for what applications, and how should cloud be used for HPC? To this end, we perform a comprehensive performance evaluation and analysis of a set of benchmarks and complex HPC applications on a range of platforms, varying from supercomputers to clouds. Further, we demonstrate HPC performance improvements in cloud using alternative lightweight virtualization mechanisms - thin VMs and OS-level containers, and hyper visor- and application-level CPU affinity. Next, we analyze the economic aspects and business models for HPC in clouds. We believe that is an important area that has not been sufficiently addressed by past research. Overall results indicate that current public clouds are cost-effective only at small scale for the chosen HPC applications, when considered in isolation, but can complement supercomputers using business models such as cloud burst and application-aware mapping.
Abhishek Gupta 0002, Laxmikant V. Kalé, Filippo Gioachin, Verdi March, Chun Hui Suen, Bu-Sung Lee, Paolo Faraboschi, Richard Kaufmann, Dejan S. Milojicic
CloudCom (1)9
2013 HPC-Aware VM Placement in Infrastructure Clouds
abstract
Cloud offerings are increasingly serving workloads with a large variability in terms of compute, storage and networking resources. Computing requirements (all the way to High Performance Computing or HPC), criticality, communication intensity, memory requirements, and scale can vary widely. Virtual Machine (VM) placement and consolidation for effective utilization of a common pool of resources for efficient execution of such diverse class of applications in the cloud is challenging, resulting in higher cost and missed Service Level Agreements (SLAs). For HPC, current cloud providers either offer dedicated cloud with dedicated nodes, losing out on consolidation benefits of virtualization, or use HPC-agnostic cloud scheduling resulting in poor HPC performance. In this work, we address application-aware allocation of n VM instances (comprising a single job request) to physical hosts from a single pool. We design and implement an HPC-aware scheduler on top of Open Stack Compute (Nova) and also incorporate it in a simulator (Cloud Sim). Through various optimizations, specifically topology- and hardware-awareness, cross-VM interference accounting and application-aware consolidation, we demonstrate enhanced VM placements which achieve up to 45% improvement in HPC performance and/or 32% increase in job throughput while limiting the effect of jitter (or noise) to 8%.
Abhishek Gupta 0002, Laxmikant V. Kalé, Dejan S. Milojicic, Paolo Faraboschi, Susanne M. Balle
IC2E3
2013 A Change Impact Analysis Approach for Workflow Repository Management
abstract
Large and complex workflow repositories include a series of interdependent workflows. In this scenario, it becomes hard to estimate the effort required to accomplish changes to workflows. Furthermore, ad-hoc changes may induce side and ripple effects, which ultimately hamper the reliability of the repository. In this paper, we introduce a static dependency-centric change impact analysis approach for workflow repository management. The approach relies on metrics and visualizations that makes it easy and quick to estimate change impact. We implemented the approach, incorporated it into HP Operations Orchestration (HP OO), and conducted an exploratory study in which we thoroughly analyzed the workflow repository of 8 HP OO customers. Besides being able to characterize and compare the repositories against each other, we found that while the out-of-the-box repository provided by HP OO has 10 flows with high change impact, 5 customer repositories had higher values that ranged from 11 (+10%) to 35 (+250%).
Gustavo Ansaldi Oliva, Marco Aurélio Gerosa, Dejan S. Milojicic, Virginia Smith
ICWS3
2013 Optimizing Checkpoints Using NVM as Virtual Memory
abstract
Rapid checkpointing will remain key functionality for next generation high end machines. This paper explores the use of node-local nonvolatile memories (NVM) such as phase-change memory, to provide frequent, low overhead checkpoints. By adapting existing multi-level checkpoint techniques, we devise new methods, termed NVM-checkpoints, that efficiently store checkpoints on both local and remote node NVM. The checkpoint frequencies are guided by failure models that capture the expected accessibility of such data after failure. To lower overheads, NVM-checkpoints reduce the NVM and interconnect bandwidth used with a novel pre-copy mechanism, which incrementally moves checkpoint data from DRAM to NVM before a local checkpoint is started. This reduces local checkpoint cost by limiting the instantaneous data volume moved at checkpoint time, thereby freeing bandwidth for use by applications. In fact, the pre-copy method can reduce peak interconnect usage up to 46%. Since our approach treats NVM as memory rather than as 'Ramdisk', pre-copying can be generalized to directly move data to remote NVMs. This results in 40% faster application execution times compared to asynchronous approaches not using pre-copying.
Sudarsun Kannan, Ada Gavrilovska, Karsten Schwan, Dejan S. Milojicic
IPDPS4
2013 A systematic literature review of service choreography adaptation
Leonardo A. F. Leite, Gustavo Ansaldi Oliva, Guilherme M. Nogueira, Marco Aurélio Gerosa, Fabio Kon, Dejan S. Milojicic
Serv. Oriented Comput. Appl.6
2013 Guest Editors' Introduction: Special Issue on Cloud Computing
abstract
The articles in this special section focus on the topic of cloud computing, technologies, applications, and new areas of technological innovation.
Vojislav B. Misic, Rajkumar Buyya, Dejan S. Milojicic, Yong Cui 0001
IEEE Trans. Parallel Distributed Syst.3
2012 Quantifying Manageability of Cloud Platforms
abstract
Cloud computing has changed the way companies are managing their IT infrastructure and application development and/or deployment. While there are a number of Cloud platforms available, it is a challenge to identify a cloud platform that best suits the business needs. As a consequence Cloud service providers may end up selecting a platform that is either too low- or too high-level to manage, or does not offer the paradigms needed. In this paper, we introduce the Cloud manageability metrics and an approach to quantify the manageability of cloud platforms using the introduced metrics. We then use the metrics and approach to compare manageability of different cloud IaaS and PaaS platforms for specific use case scenarios. The values for the metrics are derived by executing the use cases in each platform using test environments. Based on the results of the comparison, we recommend a cloud platform that is best suited for different organizations needs with respect to cloud management. We expect that the proposed approach will help organizations to evaluate various cloud platforms and choose the platform that best matches their needs.
Madhavi Maiya, Sai Dasari, Sandhya Shivaprasad, Dejan S. Milojicic
IEEE CLOUD5
2012 Exploring the performance and mapping of HPC applications to platforms in the cloud
abstract
This paper presents a scheme to optimize the mapping of HPC applications to a set of hybrid dedicated and cloud resources. First, we characterize application performance on dedicated clusters and cloud to obtain application signatures. Then, we propose an algorithm to match these signatures to resources such that performance is maximized and cost is minimized. Finally, we show simulation results revealing that in a concrete scenario our proposed scheme reduces the cost by 60% at only 10-15% performance penalty vs. a non optimized configuration. We also find that the execution overhead in cloud can be minimized to a negligible level using thin hypervisors or OS-level containers.
Abhishek Gupta 0002, Laxmikant V. Kalé, Dejan S. Milojicic, Paolo Faraboschi, Richard Kaufmann, Verdi March, Filippo Gioachin, Chun Hui Suen, Bu-Sung Lee
HPDC3
2009 Reiki: Serviceability Architecture and Approach for Reduction and Management of Product Service Incidents
abstract
There is a significant number of IT failures per year because parts fail, products are used in ways they were not designed for, and humans make errors in using products. These failures result in incidents that product vendors service as a part of the warranty or contracts. Incidents incur significant costs for servicing them, including call centers, parts, and field engineers. Some of the major problems include lack of coherent incident information, leading to inaccurate service diagnosis and inability to forecast failures. At the same time, technology has evolved. Hardware is generally more reliable, failures are moving from hardware to firmware, software, and applications. The scale effect limits human operator engagement, prevents centralized approaches, and expands automation. Traditional ways of handling incidents are not appropriate any more. In this paper we present a set of tools and approaches that enable unified serviceability with self-healing, automated learning, and an analysis engine. Unified serviceability with self-healing results in clean incident data and it reduces criticality of incidents into deferred maintenance. Automated learning produces empirically proven actionable knowledge enabling cost reduction of automated incident resolution. Using clean data and actionable knowledge, the analysis engine helps predict failures and determine trends, resulting in preventive maintenance. Collectively, preventive and deferred maintenance and automated incident service significantly reduce the costs. This way we have aligned incidents cost with the technology trends.
Chris Connelly, Brian Cox, Tim Forell, Dejan S. Milojicic, Alan Nemeth, Peter Piet, Suhas Shivanna
ICWS5
2009 A systematic and practical approach to generating policies from service level objectives
abstract
In order to manage a service to meet the agreed upon SLA, it is important to design a service of the required capacity and to monitor the service thereafter for violations at runtime. This objective can be achieved by translating SLOs specified in the SLA into lower-level policies that can then be used for design and enforcement purposes. Such design and operational policies are often constraints on thresholds of lower level metrics. In this paper, we propose a systematic and practical approach that combines fine-grained performance modeling with regression analysis to translate service level objectives into design and operational policies for multi-tier applications. We demonstrate that our approach can handle both request-based and session-based workloads and deal with workload changes in terms of both request volume and transaction mix. We validate our approach using both the RUBiS e-commerce benchmark and a trace-driven simulation of a business-critical enterprise application. These results show the effectiveness of our approach.
Yuan Chen 0001, Subu Iyer, Dejan S. Milojicic, Akhil Sahai
Integrated Network Management3
2009 Modeling remote desktop systems in utility environments with application to QoS management
abstract
A remote desktop utility system is an emerging client/server networked model for enterprise desktops. In this model, a shared pool of consolidated compute and storage servers host users' desktop applications and data respectively. End-users are allocated resources for a desktop session from the shared pool on-demand, and they interact with their applications over the network using remote display technologies. Understanding the detailed behavior of applications in these remote desktop utilities is crucial for more effective QoS management. However, there are challenges due to hard-to-predict workloads, complexity, and scale. In this paper, we present a detailed modeling of a remote desktop system through case study of an Office application - email. The characterization provides insights into workload and user model, the effect of remote display technology, and implications of shared infrastructure. We then apply these learnings and modeling results for improved QoS resource management decisions - achieving over 90% improvement compared to state of the art allocation mechanisms. We also present discussion on generalizing a methodology for a broader applicability of model-driven resource management.
Vanish Talwar, Klara Nahrstedt, Dejan S. Milojicic
Integrated Network Management3
2008 Automatically Determining Compatibility of Evolving Services
abstract
A major advantage of Service-Oriented Architectures (SOA) is composition and coordination of loosely coupled services. Because the development lifecycles of services and clients are decoupled, multiple service versions have to be maintained to continue supporting older clients. Typically versions are managed within the SOA by updating service descriptions using conventions on version numbers and namespaces. In all cases, the compatibility among services description must be evaluated, which can be hard, error-prone and costly if performed manually, particularly for complex descriptions. In this paper, we describe a method to automatically determine when two service descriptions are backward compatible. We then describe a case study to illustrate how we leveraged version compatibility information in a SOA environment and present initial performance overheads of doing so. By automatically exploring compatibility information, a) service developers can assess the impact of proposed changes; b) proper versioning requirements can be put in client implementations guaranteeing that incompatibilities will not occur during run-time; and c) messages exchanged in the SOA can be validated to ensure that only expected messages or compatible ones are exchanged.
Karin Becker, Andre Lopes, Dejan S. Milojicic, Jim Pruyne, Sharad Singhal
ICWS3
2008 Moara: Flexible and Scalable Group-Based Querying System
Steven Y. Ko, Praveen Yalagandula, Indranil Gupta, Vanish Talwar, Dejan S. Milojicic, Subu Iyer
Middleware5
2008 Improving distributed service management using Service Modeling Language (SML)
abstract
Automatic service and application deployment and management is becoming possible through the use of service and infrastructure discovery and policy systems. But using the infrastructure optimally requires intimate knowledge of the hardware and the interaction of its components in order to make optimal allocation of shared resources. This paper proposes an architecture where the hardware infrastructure not only makes operational parameters available (disk size, network bandwidth) but also presents to the service management components, relationships and constraints between the hardware components. We present an implementation which uses the Service Modeling Language, SML, to communicate this information and show how this architecture saves service management from knowing intimate knowledge of the hardware. This enhances optimal service deployment and management in a heterogeneous hardware environment and is a step toward autonomic computing.
Robert Adams 0003, Ricardo Rivaldo, Guilherme Germoglio, Flavio Santos, Yuan Chen 0001, Dejan S. Milojicic
NOMS6
2007 Automated Availability Management Driven by Business Policies
abstract
Policy-driven service management helps reduce IT management cost and it keeps the service management aligned with business objectives. While most of the previous research focuses on performance/resource managements, little has been researched in the area of availability management driven by business policies. This is a critically important task in enterprise IT, because a single failure in enterprise IT could cause huge business loss. It is still unclear how we can automate the availability management in a highly dynamic and complex system according to business level objectives for performance and risk attitude/preference. As a consequence, users can not manage the availability/performance ratio to match their risk tolerance. In this paper, we propose a policy-driven approach to automate run-time availability management in IT systems, according to high level availability and performance objectives. We further apply von Neumann-Morgenstern utility theory to deal with users' risk attitude policies, an important class of business policies specific for availability management, and the associated preference structures. Based on the proposed approach, we implement an automated decision engine for availability management. The initial evaluation of the solution illustrates the significance of the policy-driven approach and it demonstrates its applicability for availability management in complex IT environments. This way IT users can customize their availability to the risks tolerable by business objectives.
Zhongtang Cai, Yuan Chen 0001, Vibhore Kumar, Dejan S. Milojicic, Karsten Schwan
Integrated Network Management4
2007 SML Model-based Management
abstract
Models enable a clear separation between domain knowledge and application-specific details. In the system management arena, there are multiple implementations of model-based system management solutions, but until now, there was no industry-wide agreement on a common language or paradigm to enable interoperability. A new standard has been proposed to describe IT Services, the service modeling language ("SML"). While SML enables interoperability, it still poses challenges in terms of scalable model stores, model validation, and use. This paper discusses an architecture for SML model management and validation and describe and evaluate a prototype implementation with a use case on system management of PlanetLab.
Ricardo Rivaldo, Guilherme Germoglio, Flavio Santos, Yuan Chen 0001, Dejan S. Milojicic, Robert Adams 0003
Integrated Network Management5
2006 UsingWeb Services for Configuration and Deployment according to the CDDLM Standard
abstract
As Web services and service oriented architectures are adopted, it is increasingly important to have standard and interoperable means to deploy and configure Web services. Within the Global Grid Forum, HP, NEC, and Softricity have been developing a standard for configuration description, deployment, and lifecycle management (CDDLM). In order to prove its feasibility, reference implementations are being developed. This paper describes an independent reference implementation of CDDLM and the experience in using Web Services for deployment in a standardized manner. Our main contributions are: the lessons learned in implementing this WS-based standard and an architecture for implementing CDDLM
Ayla Débora Dantas de Souza Rebouças, Guilherme Germoglio, Flavio Santos, Marcelo Iury S. Oliveira, Walfredo Cirne, Francisco Vilar Brasileiro, Dejan S. Milojicic, Sandro Rafaeli, Katia Barbosa Saikoski
ICWS7
2006 Specification-Enhanced Policies for Automated Management of Changes in IT Systems
Chetan Shiva Shankar, Vanish Talwar, Subu Iyer, Yuan Chen 0001, Dejan S. Milojicic, Roy H. Campbell
LISA5
2005 Comparison of Approaches to Service Deployment
abstract
IT today is driven by the trend of increasing scale and complexity. Utility and Grid computing models, PlanetLab, and traditional data centers, are reaching the scale of thousands of computers. Installed software consists of dozens of interdependent applications and services. As the complexity and scale of these systems continues to grow, it becomes increasingly difficult to administer and manage them. At the same time, the service deployment technologies are still based on scripts and configuration files with minimal ability to express dependencies, to document and to verify configurations. This results in hard-to-use and erroneous system configurations. Language- and model-based tools, such as SmartFrog and Radia, are proposed for addressing these deployment challenges, but it is unclear whether they are beneficial over traditional solutions. In this paper, we quantitatively compare manual, script-, language-, and model-based deployment solutions as a function of scale, complexity, and susceptibility to change. We also qualitatively compare them in terms of expressiveness and barrier to first use. We demonstrate that script-based solutions are well matched for large scale deployments, language-based for services of large complexity, and model-based for dynamic changes to the design. Finally, we offer a table summarizing rules of thumb regarding which solution to use in which case, subject to deployment needs.
Vanish Talwar, Qinyi Wu, Calton Pu, Wenchang Yan, Gueyoung Jung, Dejan S. Milojicic
ICDCS6
2005 Quality of Manageability of Web Services
abstract
Web Services has become a predominant paradigm in delivering services to users. A number of standards and solutions have paved the way to reliable and secure way of Web Services execution. One area that still has not sufficiently matured is management of Web Services. There are currently a couple of standards that are being defined, such as Web Services Distributed Management and Web Services Management. However, this is a complex problem, encompassing dependencies on the underlying resources, dealing with distributed services, their availability, scalability, security, etc. Of particular concern is the automation of management and cost of management.
Dejan S. Milojicic
ICWS1
2005 Dealing with Scale and Adaptation of Global Web Services Management
abstract
Service oriented architectures (SOA) are becoming the prevalent approach for realizing modern services and systems. SOA offers superior support for autonomy (decoupling) and heterogeneity compared to previous generation middleware systems, resulting in more scalable and adaptive solutions. However, SOA have not adequately addressed management, while traditional management solutions do not sufficiently scale to address the needs of (global) Web services. We propose scalable management based on models and industry standards. We discuss a use case for global service management, we present its design, implementation and preliminary evaluation. We retain all the benefits of SOA while also enabling global scale manageability. Our approach provides manageability that is comprehensible for administrators yet automated enough for integration into autonomous systems.
William Vambenepe, Carol Thompson, Vanish Talwar, Sandro Rafaeli, Bryan S. Murray, Dejan S. Milojicic, Subu Iyer, Keith I. Farkas, Martin F. Arlitt
ICWS6
2004 ContentCascade Incremental Content Exchange between Public Displays and Personal Devices
abstract
Public exhibits or displays are widely used to present information in public and semipublic places. Users may pass by several of these and observe a significant amount of interesting content, much of which may be forgotten. Unfortunately, there is no widely adopted way for users to acquire related information in locale for their subsequent perusal. We present ContentCascade, a mechanism that allows users to implicitly download various levels of detail of summary information about the content available at the public displays to their personal devices. It also naturally and informally enables users to trade off quality of summary information with presence time. This way, ContentCascade can potentially enhance the overall user experience with public displays as users can recall the summary information later.
Himanshu Raj, Rich Gossweiler, Dejan S. Milojicic
MobiQuitous3
2004 Susceptibility of Commodity Systems and Software to Memory Soft Errors
abstract
It is widely understood that most system downtime is accounted for by programming errors and administration time. However, a growing body of work has indicated an increasing cause of downtime may stem from transient errors in computer system hardware due to external factors, such as cosmic rays. This work indicates that moving to denser semiconductor technologies at lower voltages has the potential to increase these transient errors. In this paper, we investigate the susceptibility of commodity operating systems and applications on commodity PC processors to these soft-errors and we introduce ideas regarding the improved recovery from these transient errors in software. Our results indicate that, for the Linux kernel and a Java virtual machine running sample workloads, many errors are not activated, mostly due to overwriting. In addition, given current and upcoming microprocessor support, our results indicate that those errors activated, which would normally lead to system reboot, need not be fatal to the system if software knowledge is used for simple software recovery. Together, they indicate the benefits of simple memory soft error recovery handling in commodity processors and software.
Alan Messer, Philippe Bernadat, Guangrui Fu, DeQing Chen, Zoran Dimitrijevic, David Jeun Fung Lie, Durga Mannaru, Alma Riska, Dejan S. Milojicic
IEEE Trans. Computers9
2003 Adaptive Offloading Inference for Delivering Applications in Pervasive Computing Environments
abstract
Pervasive computing allows a user to access an application on heterogeneous devices continuously and consistently. However it is challenging to deliver complex applications on resource-constrained mobile devices, such as cellular telephones and PDA. Different approaches, such as application-based or system-based adaptations, have been proposed to address the problem. However existing solutions often require degrading application fidelity. We believe that this problem can be overcome by dynamically partitioning the application and offloading part of the application execution to a powerful nearby surrogate. This will enable pervasive application delivery to be realized without significant fidelity degradation or expensive application rewriting. Because pervasive computing environments are highly dynamic, the runtime offloading system needs to adapt to both application execution patterns and resource fluctuations. Using the fuzzy control model, we have developed an offloading inference engine to adaptively solve two key decision-making problems during runtime offloading: (1) timely triggering of adaptive offloading, and (2) intelligent selection of an application partitioning policy. Extensive trace-driven evaluations show the effectiveness of the offloading inference engine.
Xiaohui Gu, Klara Nahrstedt, Alan Messer, Ira Greenberg 0002, Dejan S. Milojicic
PerCom5
2002 Towards a Distributed Platform for Resource-Constrained Devices
abstract
Many visions of the future predict a world with pervasive computing, where computing services and resources permeate the environment. In these visions, people will want to execute a service on any available device without worrying about whether the service has been tailored for the device. We believe that it will be difficult to create services that can execute well on the wide variety of devices that are being developed because of problems with diversity and resource constraints. We believe that these problems can be greatly reduced by using an ad-hoc distributed platform to transparently off-load portions of a service from a resource-constrained device to a nearby server. We implemented a preliminary prototype and emulator to study this approach. Our experiments show the beneficial use of nearby resources to relieve both memory and processing constraints, when it is appropriate to do so. We believe that this approach will reduce the burden on developers by masking more device details.
Alan Messer, Ira Greenberg 0002, Philippe Bernadat, Dejan S. Milojicic, DeQing Chen, Thomas J. Giuli, Xiaohui Gu
ICDCS4
2002 Case Studies in Security and Resource Management for Mobile Object Systems
Dejan S. Milojicic, Gul A. Agha, Philippe Bernadat, Deepika Chauhan, Shai Guday, Nadeem Jamali, Dan Lambright, Franco Travostino
Auton. Agents Multi Agent Syst.1
2002 Editorial: Mobile agent systems
Danny B. Lange, Dejan S. Milojicic
Softw. Pract. Exp.2
1998 MASIF: The OMG Mobile Agent System Interoperability Facility
Dejan S. Milojicic, Markus Breugst, Ingo Busse, John Campbell 0002, Stefan Covaci, Barry Friedman, Kazuya Kosaka, Danny B. Lange, Kouichi Ono, Mitsuru Oshima, Cynthia Tham, Sankar Virdhagriswaran, Jim White
Pers. Ubiquitous Comput.1
1998 Extended Memory Management (XMM): Lessons Learned
abstract
This paper describes the lessons learned from the development and use of the eXtended Memory Management (XMM) subsystem in the Mach microkernel. XMM provides distributed memory management for single system image Unix on systems ranging from a few nodes to over 1000 nodes. XMM enables complete cross-node transparency at the Mach virtual memory interface in support of distributed file systems and distributed process execution, including making distributed shared memory available to applications. Our experience with XMM has revealed severe problems in scalability, performance, and complexity, leading us to conclude that significant portions of XMM's architecture and design should be viewed as failed experiments. Among the lessons we have learned from our experience are that copy on reference should be used instead of copy on write for memory management in distributed systems, and that virtual copy optimizations are inappropriate for distributed coherent virtual memory. © 1998 John Wiley & Sons, Ltd.
Dejan S. Milojicic, Randall W. Dean, Michelle Dominijanni, Alan Langerman, Steven J. Sears
Softw. Pract. Exp.1