Ali Jahanshahi

dblp:216/9168 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-4301-7588ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Coordinating GPU Data Centers and Power Grid Regulation Service for Exogenous Carbon Benefits
abstract
The rapid growth of AI/ML data centers has led to higher energy consumption and carbon emissions. The shift to renewable energy and growing data center energy demands can destabilize the power grid. Power grids rely on frequency regulation reserves, typically fossil-fueled power plants, to stabilize and balance the supply and demand of electricity. This paper sheds light on the hidden carbon emissions of frequency regulation service. Our work explores how modern GPU data centers can coordinate with power grids to reduce the need for fossil-fueled frequency regulation reserves. We first introduce a novel metric, Exogenous Carbon, to quantify grid-side carbon emission reductions resulting from data center participation in regulation service. We additionally introduce EcoCenter, a framework to maximize the amount of frequency regulation provision that GPU data centers can provide, and thus, reduce the amount of frequency regulation reserves necessary. We demonstrate that data center participation in frequency regulation can result in Exogenous carbon savings that can outweigh operational carbon emissions.
Ali Jahanshahi, Sara Rashidi Golrouye, Osten Anderson, Nanpeng Yu, Daniel Wong 0001
ICS1
2025 ECLIP: Energy-efficient and Practical Co-Location of ML Inference on Spatially Partitioned GPUs
abstract
As AI inference becomes mainstream, research has begun to focus on improving the energy consumption of inference servers. Inference kernels commonly underutilize a GPU’s compute resources and waste power from idling components. To improve utilization and energy efficiency, multiple models can co-locate and share the GPU. However, typical GPU spatial partitioning techniques often experience significant overheads when reconfiguring spatial partitions, which can waste additional energy through repartitioning overheads or non-optimal partition configurations. In this paper, we present ECLIP, a framework to enable low-overhead energy-efficient kernel-wise resource partitioning between co-located inference kernels. ECLIP minimizes repartitioning overheads by pre-allocating pools of CU masked streams and assigns optimal CU assignments to groups of kernels through our resource allocation optimizer. Overall, ECLIP achieves an average of 13% improvement to throughput and 25% improvement to energy efficiency.
Ryan Quach, Yidi Wang 0001, Ali Jahanshahi, Daniel Wong 0001, Hyoseung Kim 0001
ISLPED3
2024 Characterizing In-Kernel Observability of Latency-Sensitive Request-Level Metrics with eBPF
abstract
This paper explores a novel server observability approach using eBPF (extended Berkeley Packet Filter) for detailed request-level performance metrics of data center latency-sensitive applications. Utilizing eBPF system call tracing, we evaluate if syscall activity can reconstruct high-level application behaviors and bypass the need for direct userspace reporting of performance metrics. Through careful selection of eBPF events, we demonstrate that certain syscall statistics can provide robust insight into request-level metrics. In addition, we demonstrate that these metrics can also be robust to networking effects, such as packet loss. By demonstrating the ability for eBPF to provide request-level observability, we can potentially enable many non-intrusive, low-overhead use cases for feedback in system management runtime frameworks, such as resource allocation, scheduling, and power management.
Mohammadreza Rezvani, Ali Jahanshahi, Daniel Wong 0001
ISPASS2
2023 KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference Servers
abstract
Machine learning (ML) inference workloads present significantly different challenges than ML training workloads. Typically, inference workloads are shorter running and under-utilize GPU resources. To overcome this, co-locating multiple instances of a model has been proposed to improve the utilization of GPUs. Co-located models share the GPU through GPU spatial partitioning facilities, such as Nvidia’s MPS, MIG, or AMD’s CU Masking API. Existing spatially partitioned inference servers create model-wise partitions by "right-sizing" based on a model’s latency tolerance to restricting resources. We show that model-wise right-sizing is under-utilized due to varying resource restriction tolerance of individual kernels within an inference pass.We propose Kernel-wise Right-sizing for Spatial Partitioned GPU Inference Servers (KRISP) to enable kernel-wise right-sizing of spatial partitions at the granularity of individual kernels. We demonstrate that KRISP can support a greater level of concurrently running inference models compared to existing spatially partitioned inference servers. KRISP improves overall throughput by 2x when compared with an isolated inference (1.22x vs prior works) and reduce energy per inference by 33%.
Marcus Chow, Ali Jahanshahi, Daniel Wong 0001
HPCA2
2022 GPUCalorie: Floorplan Estimation for GPU Thermal Evaluation
abstract
GPUs are massively parallel architecture that consume significant power, which lead to high thermal output. Thermal constraints of GPUs are one of the major limitations in high performance, mobile and embedded applications. However, accurate thermal modeling tools for GPUs are lacking for researchers. We identify that limiting factors to further research are the absence of GPU floorplans necessary for thermal modeling, validated thermal trends, and outdated component-level power models. To this end, we present GPUCalorie, a thermal modeling methodology using specialized infrared thermography setup for measuring and validating thermal behaviors of real GPUs. We validate a floorplan of Nvidia’s GTX1050 identified through our infrared thermography setup. We validate the GPUCalorie identified floorplan against a real GTX1050 GPU, showing 10% error for the thermal map.
Marcus Chow, Ali Jahanshahi, Ana Cardenas Beltran, Sheldon X.-D. Tan, Daniel Wong 0001
ISPASS2
2022 PowerMorph: QoS-Aware Server Power Reshaping for Data Center Regulation Service
abstract
Adoption of renewable energy in power grids introduces stability challenges in regulating the operation frequency of the electricity grid. Thus, electrical grid operators call for provisioning of frequency regulation services from end-user customers, such as data centers, to help balance the power grid’s stability by dynamically adjusting their energy consumption based on the power grid’s need. As renewable energy adoption grows, the average reward price of frequency regulation services has become much higher than that of the electricity cost. Therefore, there is a great cost incentive for data centers to provide frequency regulation service. Many existing techniques modulating data center power result in significant performance slowdown or provide a low amount of frequency regulation provision. We present PowerMorph , a tight QoS-aware data center power-reshaping framework, which enables commodity servers to provide practical frequency regulation service. The key behind PowerMorph is using “complementary workload” as an additional knob to modulate server power, which provides high provision capacity while satisfying tight QoS constraints of latency-critical workloads. We achieve up to 58% improvement to TCO under common conditions, and in certain cases can even completely eliminate the data center electricity bill and provide a net profit.
Ali Jahanshahi, Nanpeng Yu, Daniel Wong 0001
ACM Trans. Archit. Code Optim.1
2021 BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems
abstract
As modern GPU workloads grow in size and complexity, there is an ever-increasing demand for GPU computational power. Emerging workloads contain hundreds or thousands of GPU kernel launches, which incur high overheads, and exhibit data-dependent behavior between kernels, which requires synchronization, leading to GPU under-utilization. Task-based execution models have been proposed to solve these issues, but they require significant programmer effort to port applications to proprietary task-based programming models in order to specify tasks and task dependencies. To address this need, we propose BlockMaestro, a software-hardware solution that combines command queue reordering, kernel-launch-time static analysis, and runtime hardware support to dynamically identify and resolve thread-block level data dependencies between kernels. Through static analysis of memory access patterns at kernel-launch-time, BlockMaestro can extract inter-kernel thread block-level data dependencies. BlockMaestro also introduces kernel pre-launching to reduce the kernel launch overheads experienced by multiple dependent kernels. Correctness is enforced by dynamically resolving thread block-level data dependency at runtime through hardware support. BlockMaestro achieves an average speedup of 51.76% (up to 2.92x) on data-dependent benchmarks, and requires minimal hardware overhead.
AmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh, Nael B. Abu-Ghazaleh, Daniel Wong 0001
ISCA3
2021 Deflection-Aware Routing Algorithm in Network on Chip against Soft Errors and Crosstalk Faults
abstract
Marching into nano-scale technology, probability of soft errors and crosstalk faults has increased by about 6-7 times. Since buffers occupy about 40-90% of the switch area, the probability of soft errors in switches is significant. We propose a deflection-aware routing algorithm (DAR) combined with an information redundancy technique to cover the soft errors and crosstalk faults in the header flow control units (FLIT). We also introduce an interleaving method along with a simple hamming code to tolerate the errors in data and tail FLITs. The proposed methods have been evaluated in both circuit and simulation level through a simulator written in C++, Booksim 2, and Synopsys Design Compiler. The evaluation results show that we can cover the soft errors and crosstalk faults with reasonable power and performance overhead of 3% and 6.5% respectively.
Hadi Zamani 0001, Zahra Shirmohammadi, Ali Jahanshahi
NAS3
2021 ICAP: Designing Inrush Current Aware Power Gating Switch for GPGPU
abstract
The leakage energy of GPGPU can be reduced by power gating the idle logic or undervolting the storage structures; however, the performance and reliability of the system degrades due to large wake up time and inrush current at time of activation. In this paper, we thoroughly analyze the realistic Break-Even Time (BET) and inrush current for various components in GPGPU architecture considering the recent design of multi-modal Power Gating Switch (PGS). Then, we introduce a new PGS which covers the current PGS drawbacks. Our redesigned PGS is carefully tailored to minimize the inrush current and BET. GPGPU-Sim simulation results for various applications, show that, with incorporating the proposed PGS into GPGPU-Sim, we can save leakage energy up to 82%, 38%, and 60% for register files, integer units, and floating units respectively.
Hadi Zamani 0001, Devashree Tripathy, Ali Jahanshahi, Daniel Wong 0001
NAS3
2019 Border Gateway Protocol Anomaly Detection Using Neural Network
abstract
Having reliable and stable connectivity to the Internet dramatically depends on how Border Gateway Protocol (BGP) can avoid bad-behaviour events by detecting them on time. Despite a lot of efforts have gone into detecting BGP anomalies during the last decade, it is still a challenging issue due to emerging new abnormal behaviours both from the attackers and network misconfigurations. In this work, we propose a Neural Network classifier to detect the abnormal BGP events caused by worm attacks in the network. The results show that our method outperforms the previous work in both generality and accuracy.
Ali Jahanshahi, Abbas Mazloumi, Hadi Zamani 0001
IEEE BigData2