Hossein SeyyedAghaei

dblp:342/4454 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 56% GPUs and heterogeneous computing · 44%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU architecture
0.812024
GPU Scale-Model Simulation · HPCA 2024
Performance modeling and evaluation › performance prediction
GPU performance prediction
0.812024
GPU Scale-Model Simulation · HPCA 2024
GPUs and heterogeneous computing › GPU architecture
multi-chip module GPU
0.812024
GPU Scale-Model Simulation · HPCA 2024
Performance modeling and evaluation
simulation
0.812024
GPU Scale-Model Simulation · HPCA 2024
Performance modeling and evaluation › parallel system performance
strong and weak scaling
0.212024
GPU Scale-Model Simulation · HPCA 2024
Performance modeling and evaluation
workload characterization
0.212024
GPU Scale-Model Simulation · HPCA 2024

Methods — techniques the papers use, named apart from their topics

regression · 0.8miss rate curve · 0.8
YearPublicationVenuePosition
2026 DyPNet-MSC: Dynamic Bandwidth Allocation in Photonic Network-on-Wafer GPU Architectures
abstract
Wafer-scale multi-GPU systems have been proposed as powerful accelerators, with Photonic Network-on-Wafer (PNoW) architectures offering clear advantages over electrical interconnects in area, bandwidth, and energy efficiency. Designing a circuit-switched photonic network that can dynamically adapt to inter-GPU traffic patterns that vary over time and in space however is an open challenge. In this paper, we propose DyPNet-MSC, a Dynamic Photonic Network-on-Wafer with Minimal Static Connectivity. DyPNetMSC provides minimal static all-to-all connectivity through dedicated waveguides while dynamically re-allocating the remaining PNoW bandwidth to cater to inter-GPU traffic demand changes during run time. For the latter, a key design choice emerges between fine-grained but costly wavelength-selective (WS) switches, versus low-overhead but coarse-grained wavelength-non-selective (WNS) switches. Our evaluation shows that DyPNet-MSC achieves the best of both worlds: DyPNet-MSC with WNS switches offers performance comparable to WS switches (within approximately 0.5%) while incurring significantly lower overhead. Furthermore, with 2 TB/s bandwidth per GPU, DyPNet-MSC improves average performance by 27% over a uniform-static $2 \mathrm{~TB} / \mathrm{s}$ photonic network, reaching a performance level surpassing a uniformstatic $4 \mathrm{~TB} / \mathrm{s}$ network by 5 percentage points. For workloads with imbalanced spatial traffic demands, DyPNet-MSC delivers up to 90% higher performance than a uniform-static $2 \mathrm{~TB} / \mathrm{s}$ network while outperforming a $4 \mathrm{~TB} / \mathrm{s}$ uniform-static network by $\mathbf{5 5}$ percentage points. Overall, this work provides new insight into designing photonic interconnection networks for next-generation wafer-scale multi-GPU systems.
Hossein SeyyedAghaei, Benyamin Eslami, Xin Wang 0130, Gunther Roelkens, Didier Colle, Mario Pickavet, Lieven Eeckhout
ISPASS2
2024 GPU Scale-Model Simulation
abstract
The continuously increasing GPU system scale and compute capabilities, i.e., increasing number of streaming multiprocessors (SMs), caches, on-chip and off-chip memory bandwidth, pose a major challenge for performance evaluation methodologies. Architectural simulation is time-consuming and resource-intensive, and because of simulator and/or simulation host infrastructure limitations, it might not even be possible to simulate large-scale systems. Scale-model simulation is a recently proposed performance prediction methodology to predict large-scale system performance based on (much smaller) scale models. Prior work in scale-model simulation for general-purpose multicore CPUs and specialized graph analytics accelerators, unfortunately, cannot be readily applied to GPUs because different GPU applications exhibit vastly different scaling behavior with system size, thereby breaking the one-size-fits-all regression models deployed in prior work. This paper proposes a GPU scale-model simulation methodology that leverages performance measurements of two scale models alongside a miss rate curve to predict GPU target system performance. A key asset of GPU scale-model simulation is that it does not require access to a simulation model of the target system, unlike prior work in simulation acceleration. Our experimental evaluation demonstrates the accuracy of GPU scale-model simulation for both strong-scaling and weak-scaling workload scenarios. Under strong scaling, the performance of a 128-SM target system is predicted within 4% error on average, and at most 17%, using 8-SM and 16-SM scale models. Under weak scaling, the performance of a 128-SM target system is estimated with an average error of 1.7%, and at most 4.5%, while yielding a 9.3× simulation time speedup. We furthermore demonstrate how scale-model simulation predicts multi-chiplet GPU performance with an average error of 2.5% (and at most 4.3%). Alternate solutions are substantially less accurate.
Hossein SeyyedAghaei, Mahmood Naderan-Tahan, Lieven Eeckhout
HPCA1
2023 Sieve: Stratified GPU-Compute Workload Sampling
abstract
To exploit the ever increasing compute capabilities offered by GPU hardware, GPU-compute workloads have evolved from simple computational kernels to large-scale programs with complex software stacks and numerous kernels. Driving architecture exploration using real workloads hence becomes increasingly challenging, up to the point of becoming intractable because of extremely long simulation times using existing architecture simulators. Sampling is a widely used technique to speed up simulation, however, the state-of-the-art sampling method for GPU-compute workloads, Principal Kernel Selection (PKS), falls short for challenging GPU-compute workloads with a large number of kernels and kernel invocations. This paper presents Sieve, an accurate and low-overhead stratified sampling methodology for GPU-compute workloads that groups kernel invocations based on their instruction count, with the goal of minimizing the execution time variability within strata. For the challenging Cactus and MLPerf workloads, we report that Sieve achieves an average prediction error of 1.2% (and at most 3.2%) versus 16.5% (and up to 60.4%) for PKS on real hardware (Nvidia Ampere GPU), while maintaining a similar simulation speedup of three orders of magnitude. We further demonstrate that Sieve reduces profiling time by a factor of 8× (and up to 98×) compared to PKS.
Mahmood Naderan-Tahan, Hossein SeyyedAghaei, Lieven Eeckhout
ISPASS2