Tamar Eilam

dblp:16/1242 · DBLP profile ↗
← Back
27ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0002-0912-0776ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 3 since 2021Computer networks · 5 · 4 first-author · 1 since 2021Theory of computation · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
abstract
As AI workloads drive increases in datacenter power consumption, accurate GPU power estimation is critical for proactive power management. However, existing power models face a scalability bottleneck not in the modeling techniques themselves, but in obtaining the hardware utilization inputs they require. Conventional approaches rely on either costly simulation or hardware profiling, which makes them impractical when rapid predictions are required. This work presents EnergAIzer, which addresses this scalability bottleneck by developing a lightweight solution to predict utilization inputs, reducing the estimation walltime from hours to seconds. Our key insight is that kernels in AI workloads commonly employ optimizations that create structured patterns, which analytically determine memory traffic and execution timeline. We construct a performance model using these patterns as an analytical scaffold for empirical data fitting, which also naturally exposes module-level utilization. This predicted utilization is then fed into our power model to estimate dynamic power consumption. EnergAIzer achieves $8 \%$ power errors on NVIDIA Ampere GPUs, competitive with traditional power models with elaborate cycle-level simulation or hardware profiling. We demonstrate EnergAIzer’s exploration capabilities for frequency scaling and architectural configurations, including forecasting the power of NVIDIA H100 with just $7 \%$ error. In summary, EnergAIzer provides fast and accurate power prediction for AI workloads, paving the way for power-aware design explorations.
Kyungmi Lee, Zhiye Song, Xin Zhang 0025, Tamar Eilam, Anantha P. Chandrakasan
ISPASS5
2024 Process-Based Efficient Power Level Exporter
abstract
In this paper, we present the Kepler framework, designed to address the critical need for precise power and energy measurement in on-prem cloud-native, containerized environments, with a specific focus on processes, containers, and Kubernetes pods. The framework aims to support other tools in making informed decisions regarding provisioning, scheduling, and energy-optimization in cloud environments. Our approach involves leveraging the Kepler framework to create power models using Hardware Counters (HC), and real- time system power metrics from hardware sensors like x86 Running Average Power Limit (RAPL). Unlike previous methods that create and validate power models using aggregated system metrics, we propose a versatile process-level power model trained with per-process metrics. Those metrics are collected via a series of experiments in a controlled environment, measuring the incremental power consumption of processes under different scenarios. The collected data is then utilized to create a power model to be used in a shared cloud environment, and to validate the created power models using different set of input metrics. Our results show a significant improvement in the model accuracy compared to prior works, when incorporating per- process metrics and real-time system power metrics into the power estimation process. For instance, using the simplest power model, which is based on CPU utilization ratio, resulted in a Sum of Squared Error (SSE) of 75. In contrast, a power model created using aggregated system metrics, as the related works, had an SSE of 175 without real-time power metrics, and 5.6 with our proposed model refinement by normalizing the model results with the real-time system power metrics. On the other hand, training the power model with per-process metrics from controlled experiments yielded an SSE as low as 1.68 using real- time system power metrics, representing a 70% improvement in model accuracy compared to using aggregated system metrics, and an SSE 8.7 without power metrics, representing a 95% improvement in model accuracy. Furthermore, the results show that Kepler has a notable lower overhead by utilizing extended Berkeley Packet Filter (eBPF) for HC collection than alternative methods.
Marcelo Amaral, Tatsuhiro Chiba, Rina Nakazawa, Sunyanan Choochotkaew, Tamar Eilam
CLOUD7
2024 Harnessing the Power of Specialization for Sustainable Computing
abstract
Computing is critical to address some of the most pressing needs of humanity today, including climate change mitigation and adaptation. However, it is also the source of a significant and steadily increasing carbon toll, attributed in part to the exponential growth in energy-demanding workloads, such as artificial intelligence (AI). Due to the demise of Dennard scaling, we can no longer count on exponentially-improve energy efficiency of general-purpose processors. Therefore, today's operational efficiency gains rely on specialized hardware.
Tamar Eilam
ASPLOS (3)1
2024 FedCore: Straggler-Free Federated Learning with Distributed Coresets
abstract
Federated learning (FL) is a machine learning paradigm that allows multiple clients to collaboratively train a shared model while keeping their data on-premise. However, the straggler issue, due to slow clients, often hinders the efficiency and scalability of FL. This paper presents FedCore, an algorithm that innovatively tackles the straggler problem via the decentralized selection of coresets, representative subsets of a dataset. Contrary to existing centralized coreset methods, FedCore creates coresets directly on each client in a distributed manner, ensuring privacy preservation in FL. FedCore translates the coreset optimization problem into a more tractable k-medoids clustering problem and operates distributedly on each client. Theoretical analysis confirms FedCore's convergence, and practical evaluations demonstrate an 8x reduction in FL training time, without compromising model accuracy. Our extensive evaluations also show that FedCore generalizes well to existing FL frameworks11Code: https://github.com/hongpeng-guo/PedCore.
Hongpeng Guo, Haotian Gu, Bo Chen 0025, Tamar Eilam, Deming Chen, Klara Nahrstedt
ICC6
2024 Best-Effort Power Model Serving for Energy Quantification of Cloud Instances
abstract
Quantifying energy consumption is a fundamental element of green computing. Power models trained by resource utilization allow quantifying the energy number and enable energy-efficient resource management systems without raising the concerns of complexity, cost, and security. However, energy consuming behavior on different machines varies by several factors. In this paper, we address the challenges of power modeling for cloud instances where information about these factors is obscured or unseen in the training set, and propose a best-effort method to train and serve a power model as precise as possible by leveraging a large, industry-standard power database. The proposed method prioritizes the modeling precision, and offers similarity and uncertainty indicators to elucidate the confidence level when serving an unseen instance. The results have demonstrated feasibility and precision of the proposed method against comparable approaches.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Marcelo Amaral, Rina Nakazawa, Scott Trent, UmaMaheswari Devi, Tamar Eilam
MASCOTS8
2023 Kepler: A Framework to Calculate the Energy Consumption of Containerized Applications
abstract
Energy accounting is crucial in data centers for optimizing power provisioning, capping, and tuning. This paper introduces the Kepler framework, which estimates power consumption at the process, container, and Kubernetes pod levels. Kepler offers a set of power models applicable to various architectures and metrics. In this study, we propose a generic power model that utilizes hardware counters (HC) and realtime system power metrics (e.g., running average power limit (RAPL)) as independent variables in a regression model. Unlike previous approaches that rely on aggregate power consumption, our methodology measures individual process power consumption to train the power model. We provide step-by-step instructions to measure process power consumption in a controlled environment, considering the activation constant and load-dependent dynamic power consumption in different executions. By following the Greenhouse Gas (GHG) Protocol, our approach ensures the fair distribution of constant power among the user's processes. The results demonstrate significantly improved accuracy with a mean squared error (MSE) as low as 0.010 for the proposed method, compared with an MSE of 0.16 for a simple ratio approach and 0.92 when training the model using aggregated workload power.
Marcelo Amaral, Tatsuhiro Chiba, Rina Nakazawa, Sunyanan Choochotkaew, Tamar Eilam
CLOUD7
2023 Advancing Cloud Sustainability: A Versatile Framework for Container Power Model Training
abstract
Estimating power consumption in modern Cloud is important to account for the power consumed by each container. The challenge is that multiple customers are sharing the same hardware platform, where physical information is mostly obscured. In addition, there is the overhead in power consumption that the Cloud control plane induces. This paper addresses these challenges and introduces a pipeline framework for container power model training on the basis of available performance counters and other metrics. The proposed model utilizes machine learning techniques to predict the power consumed by the control plane and associated processes when running together with the user containers, and uses it for isolating the dynamic power consumed by the user-inducing workload. Applying the proposed power model does not require online power measurements, nor does it need machine information, or information on other tenants sharing the same machine. The results of cross-workload, cross-platform experiments demonstrated the higher accuracy of the model when predicting power consumption of unseen containers on unknown platforms, including on virtual machines.
Sunyanan Choochotkaew, Chen Wang 0039, Tatsuhiro Chiba, Marcelo Amaral, Tamar Eilam
MASCOTS7
2022 BoFL: bayesian optimized local training pace control for energy efficient federated learning
abstract
Federated learning (FL) is a machine learning paradigm that enables a cluster of decentralized edge devices to collaboratively train a shared machine learning model without exposing users' raw data. However, the intensive model training computation is energy-demanding and poses severe challenges to end devices' battery life. In this paper, we present BoFL, a training pace controller deployed on the edge devices that actuates the hardware operational frequencies over multiple configurations to achieve energy-efficient federated learning. BoFL operates in an explore-then-exploit manner within limited rounds of FL tasks. BoFL explores the large hardware frequency space strategically with a tailor-designed Bayesian optimization algorithm. BoFL first finds a set of good operational configurations within few task training rounds, and then exploits these configurations in the remaining rounds to achieve minimized energy consumption for model training. Experiments on multiple real-world edge devices with different FL tasks suggest that BoFL can reduce energy consumption of model training by around 26%, and achieve near-optimal energy efficiency.
Hongpeng Guo, Haotian Gu, Zhe Yang 0010, Nandhini Chandramoorthy, Tamar Eilam, Deming Chen, Klara Nahrstedt
Middleware7
2013 Testing Idempotence for Infrastructure as Code
Waldemar Hummer, Florian Rosenberg, Fábio Oliveira, Tamar Eilam
Middleware4
2012 Semantically-Rich Composition of Virtual Images
abstract
Virtualization promises to reduce data centers' total cost of ownership by enabling the creation of a small set of standardized building blocks to be shared and used many times indifferent software stacks. However, without proper methodology and tools, an organization can easily end up with a large number of one-off virtual images, adversely affecting the cost. We propose an approach, tool, and algorithms for constructing high-quality, semantically-rich image building blocks that are easy to share, compose, and reuse. In our approach, domain experts codify knowledge of a particular software product (or a combination thereof) in a platform- and cloud-agnostic software bundle. Image builders easily construct virtual images by composing a set of standardized bundles. Semantic-based validation guarantees a valid and complete image design. Moreover, we propose algorithms to automate image design by searching for an optimal set of building blocks taking into account multiple metrics such as cost, size, and expected build duration.
Fábio Oliveira, Tamar Eilam, Michael H. Kalantar, Florian Rosenberg
IEEE CLOUD2
2011 Pattern-based composite application deployment
abstract
The deployment of composite applications and services in distributed or compute cloud environments is still a challenging task that is a key source of operational cost and risk. Current approaches to composite deployments can be categorized as workflow based and model based. In the workflow based approach, deployers create end-to-end workflows to automate application deployment, while in the model based approach architects design detailed “desired state” models and validate they meet all requirements and constraints. Today, there is no formally understood relationship and mapping between “desired state” models and deployment workflows, posing a challenge and limitation on architects and deployers. In this paper we propose a new model based approach to bridge the gap between deployment models and workflows. Our approach supports separation of concerns where basic automation building blocks (such as scripts and workflows) can be developed independently of the resource model and with no knowledge of it. Therefore, the method enables deployers to continue to leverage useful libraries of automation building blocks, while enjoying the benefits of a sound resource model, used for validation and constraint satisfaction. We propose algorithms and implementation to generate end-to-end workflows for input “desired state” resource models, based on given libraries of automation building blocks. Our approach has been incorporated into IBM's leading deployment modeling platform [15] and is in active use by customers in a large range of applications.
Tamar Eilam, Michael Elder, Alexander V. Konstantinou, Ed C. Snible
Integrated Network Management1
2008 Automatic Realization of SOA Deployment Patterns in Distributed Environments
William C. Arnold, Tamar Eilam, Michael H. Kalantar, Alexander V. Konstantinou, Alexander Totok
ICSOC2
2007 Pattern Based SOA Deployment
William C. Arnold, Tamar Eilam, Michael H. Kalantar, Alexander V. Konstantinou, Alexander Totok
ICSOC2
2007 Average stretch analysis of compact routing schemes
Tamar Eilam, Cyril Gavoille, David Peleg
Discret. Appl. Math.1
2006 Model Driven Provisioning: Bridging the Gap Between Declarative Object Models and Procedural Provisioning Tools
Kaoutar El Maghraoui, Alok Meghranjani, Tamar Eilam, Michael H. Kalantar, Alexander V. Konstantinou
Middleware3
2005 Reducing the complexity of application deployment in large data centers
abstract
The deployment and configuration of distributed applications is a human intensive and highly complex process that poses significant challenges to data center operators. The process involves many cross-cutting concerns such as connectivity, performance, and security requirements, as well as resource availability, policies and best practices. These interdependencies represent a significant source of complexity, cost, and risk in data center management. In this paper we address this problem using a new approach that leverages concepts from the model-driven architecture research domain. We describe a prototype application deployment automation system based on model transformation techniques. We show how model transformation techniques can replace the manual process of writing and adapting scripts and workflows, reduce the deployment complexity, guarantee configuration integrity and consistency, and allow for a separation of concerns.
Tamar Eilam, Michael H. Kalantar, Alexander V. Konstantinou, Giovanni Pacifici
Integrated Network Management1
2002 Neptune: A Dynamic Resource Allocation and Planning System for a Cluster Computing Utility
abstract
We present Neptune - the resource director of Océano, a policy driven fabric management system that dynamically reconfigures resources in a computing utility cluster. Neptune implements an on-line control mechanism subject to policy-based performance and resource configuration objectives. Neptune reassigns servers and bandwidth among a set of service domains, based on pre-defined policy, in response to workload changes. It builds and executes a reconfiguration plan through a planning framework, breaking reconfiguration objectives into individual tasks delegated to set of lower level resource managers. We describe an example decision policy algorithm that we implemented and demonstrated in an 80 server multi-domain computing utility.
Donald P. Pazel, Tamar Eilam, Liana L. Fong, Michael H. Kalantar, Karen Appleby, Germán S. Goldszmidt
CCGRID2
2002 Lightpath arrangement in survivable rings to minimize the switching cost
abstract
This paper studies the design of low-cost survivable wavelength-division-multiplexing (WDM) networks. To achieve survivability, lightpaths are arranged as a set of rings. Arrangement in rings is also necessary to support SONET/SDH protection schemes such as 4FBLSR above the optical layer. This is expected to be the most common architecture in regional (metro) networks. We assume that we are given a set of lightpaths in an arbitrary network topology and aim at finding a partition of the lightpaths to rings adding a minimum number of lightpaths to the original set. The cost measure that we consider (number of lightpaths) reflects the switching cost of the entire network. In the case of a SONET/SDH higher layer, the number of lightpaths is equal to the number of add-drop multiplexers (ADMs) (since two subsequent lightpaths in a ring can share an ADM at the common node). We prove some negative results on the tractability and approximability of the problem and provide an approximation algorithm with a worst case approximation ratio of 8/5. We study some special cases in which the performance of the algorithm is improved. A similar problem was introduced, motivated, and studied by Liu, Li, Wan and Frieder (see Proc. INFOCOM 2000, p.1020-1025, 2000) Gerstel, Lin and Sasaki, (see Proc. IEEE INFOCOM '98, p. 94-101, 1998)(where it was termed minimum ADM problem). However, these two works focused on a ring topology while we generalize the problem to an arbitrary network topology.
Tamar Eilam, Shlomo Moran, Shmuel Zaks
IEEE J. Sel. Areas Commun.1
2002 The complexity of the characterization of networks supporting shortest-path interval routing
Tamar Eilam, Shlomo Moran, Shmuel Zaks
Theor. Comput. Sci.1
2000 Approximation Algorithms for Survivable Optical Networks
Tamar Eilam, Shlomo Moran, Shmuel Zaks
DISC1
2000 Transparently Obtaining Scalability for Java Applications on a Cluster
Yariv Aridor, Michael Factor, Avi Teperman, Tamar Eilam, Assaf Schuster
J. Parallel Distributed Comput.4
2000 On the totalk-diameter of connection networks
Yefim Dinitz, Tamar Eilam, Shlomo Moran, Shmuel Zaks
Theor. Comput. Sci.2
1999 Lower bounds for linear interval routing
abstract
Linear interval routing is a space-efficient routing method for point-to-point communication networks. It is a restricted variant of interval routing where the routing range associated with every link is represented by an interval with no wraparound. A common way to measure the efficiency of such routing methods is in terms of the maximal length of a path a message traverses. For interval routing, the upper bound and lower bound on this quantity are 2D and 2D − 3, respectively, where D is the diameter of the network. We prove a lower bound of Ω(D2) on the length of a path a message traverses under linear interval routing. We further extend the result by showing a connection between the efficiency of linear interval routing and the total2-diameter (defined in Section 4) of the network, and by presenting a family of graphs for which this lower bound is tight. © 1999 John Wiley & Sons, Inc. Networks 34: 37–46, 1999
Tamar Eilam, Shlomo Moran, Shmuel Zaks
Networks1
1998 Compact Routing Schemes with Low Stretch Factor (Extended Abstract)
abstract
This paper presents a routing strategy called Pivot Interval Routing (PIR), which allows inessage routing on every weighted n-node network along paths whose stretch (namely, the ratio between their length and the distance between their endpoints) is at most five, and whose average stretch is at inost three, with routing tables of size O(n3/" log3/' n) bits in total.A similar routing strategy for unweighted networks which guarantees the same bounds on the stretch factor and in addition a bound of r1.501 on the route lengths, where D is the dianreter of the network, is also presented.Moreover, it is shown that the PIR strategy can be implemented so that the generated scheme is in the forin of an interval routing scheme (IRS), using at most 2dm intervals per link in the first case and 3,/m in the second case.As a result, the scheines are siinpler than previous ones and they imply that paths of messages are loop-free.Finally, it is showu that there is no loop-free routing strategy guaranteeing a inemory bound of J5i bits per router for all networks, regardless of the route lengths.
Tamar Eilam, Cyril Gavoille, David Peleg
PODC1
1997 A Complete Characterization of the Path Layout Construction Problem for ATM Networks with Given Hop Count and Load (Extended Abstract)
Tamar Eilam, Michele Flammini, Shmuel Zaks
ICALP1
1997 The Complexity of Characterization of Networks Supporting Shortest-Path Interval Routing
Tamar Eilam, Shlomo Moran, Shmuel Zaks
SIROCCO1
1995 Greedy Hot-Potato Routing on the Two-Dimensional Mesh
Ishai Ben-Aroya, Tamar Eilam, Assaf Schuster
Distributed Comput.2