EDBT 2026 Demo / reviewers in the wild / expert
Arya Mazaheri
dblp:173/0115
· DBLP profile ↗
13ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-5671-4710ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SuperSFL: Resource-Heterogeneous Federated Split Learning with Weight-Sharing Supernet
Abdullah Al Asif, Sixing Yu, Juan Pablo Muñoz, Arya Mazaheri, Ali Jannesari |
Euro-Par (2) | 4 |
| 2025 | Weight-Sharing NAS with Architecture-Agnostic Intermediate RepresentationabstractWeight-sharing supernet has been widely adopted in Neural Architecture Search (NAS) as a promising strategy to obtain smaller and more efficient high-performance models. However, constructing supernets requires domain expertise to design architecture-specific rules (e.g., rules for CNNs and Transformers) for generating subnets, and training a supernet demands joint optimization over a vast sample space of subnets, which is computationally expensive. This paper presents OSF (Optimized Supernet Formation), an automated and architecture-agnostic approach that transforms predefined/pretrained models into weight-sharing supernets. Specifically, we propose representing neural architectures using a high-level computational graph intermediate representation (IR) that enables both the conversion of different types of models into supernets and the extraction of executable subnets via graph traversal. To improve supernet training efficiency, we introduce a sampling strategy that prioritizes the most promising subnet candidates during training, and propose a fork-join parallel training approach with gradient accumulation that resolves write-after-write dependencies, enabling concurrent training of multiple subnet architectures with shared weights. Our empirical evaluations demonstrate that OSF successfully builds supernets from various architectures (CNNs, Transformers, SSMs, and MLPs) while achieving superior performance across language and vision benchmarks. Notably, for Vision Transformers (ViT), OSF reduces FLOPs by 49% while maintaining the accuracy, resulting in a 155% increase in throughput and 35% latency reduction. Code Open-sourced at: https://github.com/yusx-swapp/OSF Sixing Yu, Arya Mazaheri, Ali Jannesari |
HPDC | 2 |
| 2024 | Dissecting Convolutional Neural Networks for Runtime and Scalability PredictionabstractGiven the computational complexity of deep neural networks (DNN), accurate prediction of their training and inference time using performance modeling is crucial for efficient infrastructure planning and DNN development. However, existing methods often predict only the inference time and rely on exhaustive benchmarking and fine tuning, making them time consuming and restricted in scope. As a remedy, we propose ConvMeter, a novel yet simple performance model that considers the inherent characteristics of DNNs, such as architecture, dataset, and target hardware, which strongly affect their runtime and scalability. Our performance model, which has been thoroughly tested on convolutional neural networks (ConvNets), a class of DNNs widely used for image analysis, offers the prediction of inference and training time, the latter on one or more compute nodes. Experiments with various ConvNets demonstrate that our runtime predictions of inference and training phases achieved an average error rate of less than 20% and 18%, respectively, making the assessment of ConvNets regarding efficiency and scalability straightforward. Tim Beringer, Jakob Stock, Arya Mazaheri, Felix Wolf 0001 |
ICPP | 3 |
| 2024 | PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined SpeculationabstractInference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques reduce bottlenecks associated with memory bandwidth, but also increase end-to-end latency per inference run, requiring high speculation acceptance rates to improve performance. Combined with a variable rate of acceptance across tasks, speculative inference techniques can result in reduced performance. Additionally, pipeline-parallel designs require many user requests to maintain maximum utilization. As a remedy, we propose PipeInfer, a pipelined speculative acceleration technique to reduce inter-token latency and improve system utilization for single-request scenarios while also improving tolerance to low speculation acceptance rates and low-bandwidth interconnects. PipeInfer exhibits up to a $2.15 \times$ improvement in generation speed over standard speculative inference. PipeInfer achieves its improvement through Continuous Asynchronous Speculation and Early Inference Cancellation, the former improving latency and generation speed by running single-token inference simultaneously with several speculative runs, while the latter improves speed and latency by skipping the computation of invalidated runs, even in the middle of inference. Branden Butler, Sixing Yu, Arya Mazaheri, Ali Jannesari |
SC | 3 |
| 2022 | Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement LearningabstractModel compression is an essential technique for deploying deep neural networks (DNNs) on power and memory-constrained resources. However, existing model-compression methods often rely on human expertise and focus on parameters’ local importance, ignoring the rich topology information within DNNs. In this paper, we propose a novel multi-stage graph embedding technique based on graph neural networks (GNNs) to identify DNN topologies and use reinforcement learning (RL) to find a suitable compression policy. We performed resource-constrained (i.e., FLOPs) channel pruning and compared our approach with state-of-the-art model compression methods. We evaluated our method on various models from typical to mobile-friendly networks, such as ResNet family, VGG-16, MobileNet-v1/v2, and ShuffleNet. Results show that our method can achieve higher compression ratios with a minimal fine-tuning cost yet yields outstanding and competitive performance. Sixing Yu, Arya Mazaheri, Ali Jannesari |
ICML | 2 |
| 2022 | ElastiSim: A Batch-System Simulator for Malleable WorkloadsabstractAs high-performance computing infrastructures move towards exascale, the role of resource and job management systems is more critical now than ever. Simulating batch systems to improve scheduling algorithms and resource management efficiency is an indispensable option, as running large-scale experiments is expensive and time-consuming. Batch-system simulators are responsible for simulating the computing infrastructure and the types of jobs that constitute the workload. In contrast to rigid jobs, malleable jobs can dynamically reconfigure their resources during runtime. Although studies indicate that malleability can improve system performance, no simulator exists to investigate malleable scheduling policies. In this work, we present ElastiSim, a batch-system simulator supporting the combined scheduling of rigid and malleable jobs. To facilitate the simulation, we propose a malleable workload model and introduce a scheduling protocol that enables the evaluation of topology-, I/O-, and progress-aware scheduling algorithms. We validate the scaling behavior of our workload model by comparing training runtimes of various deep-learning models against the results achieved by ElastiSim. We use real-world cluster trace files to generate workloads and simulate various scheduling algorithms (FCFS, SJF, DRF, SRTF) to analyze their implications on the simulated platform. The results demonstrate that real-world executions show the same scaling behavior as our proposed workload model. We further show that ElastiSim can capture the complex interplay between emerging workloads and modern platforms to support algorithm designers by providing consistently meaningful results. ElastiSim is publicly available as an open-source project on https://github.com/elastisim. Taylan Özden, Tim Beringer, Arya Mazaheri, Hamid Mohammadi Fard, Felix Wolf 0001 |
ICPP | 3 |
| 2021 | Auto Graph Encoder-Decoder for Neural Network Pruning
Sixing Yu, Arya Mazaheri, Ali Jannesari |
ICCV | 2 |
| 2020 | Efficient Ephemeris Models for Spacecraft Trajectory Simulations on GPUs
Fabian Schrammel, Florian Renk, Arya Mazaheri, Felix Wolf 0001 |
Euro-Par | 3 |
| 2020 | Accelerating winograd convolutions using symbolic computation and meta-programmingabstractConvolution operations are essential constituents of convolutional neural networks. Their efficient and performance-portable implementation demands tremendous programming effort and fine-tuning. Winograd's minimal filtering algorithm is a well-known method to reduce the computational complexity of convolution operations. Unfortunately, existing implementations of this algorithm are either vendor-specific or hard-coded to support a small subset of convolutions, thus limiting their versatility and performance portability. In this paper, we propose a novel method to optimize Winograd convolutions based on symbolic computation. Taking advantage meta-programming and auto-tuning, we further introduce a system to automate the generation of efficient and portable Winograd convolution code for various GPUs. We show that our optimization technique can effectively exploit repetitive patterns, enabling us to reduce the number of arithmetic operations by up to 62% without compromising numerical stability. Moreover, we demonstrate in experiments that we can generate efficient kernels with runtimes close to deep-learning libraries, requiring only a minimum of programming effort, which confirms the performance portability of our approach. Arya Mazaheri, Tim Beringer, Matthew W. Moskewicz, Felix Wolf 0001, Ali Jannesari |
EuroSys | 1 |
| 2020 | Safer Parallelization
Reiner Hähnle, Asmae Heydari Tabar, Arya Mazaheri, Mohammad Norouzi 0003, Dominic Steinhöfel, Felix Wolf 0001 |
ISoLA (2) | 3 |
| 2019 | Enhancing the Programmability and Performance Portability of GPU Tensor Operations
Arya Mazaheri, Johannes Schulte, Matthew W. Moskewicz, Felix Wolf 0001, Ali Jannesari |
Euro-Par | 1 |
| 2018 | Unveiling Thread Communication Bottlenecks Using Hardware-Independent MetricsabstractA critical factor for developing robust shared-memory applications is the efficient use of the cache and the communication between threads. Inappropriate data structures, algorithm design, and inefficient thread affinity may result in superfluous communication between threads/cores and severe performance problems. For this reason, state-of-the-art profiling tools focus on thread communication and behavior to present different metrics that enable programmers to write cache-friendly programs. The data shared between a pair of threads should be reused with a reasonable distance to preserve data locality. However, existing tools do not take into account the locality of communication events and mainly focus on analyzing the amount of communication instead. In this paper, we introduce a new method to analyze performance and communication bottlenecks that arise from data-access patterns and thread interactions of each code region. We propose new hardware-independent metrics to characterize thread communication and provide suggestions for applying appropriate optimizations on a specific code region. We evaluated our approach on the SPLASH and Rodinia benchmark suites. Experimental results validate the effectiveness of our approach by finding communication locality issues due to inefficient data structures and/or poor algorithm implementations. By applying the suggested optimizations, we improved the performance in Rodinia benchmarks by up to 56%. Furthermore, by varying the input size we demonstrated the ability of our method to assess the cache usage and scalability of a given application in terms of its inherent communication. Arya Mazaheri, Felix Wolf 0001, Ali Jannesari |
ICPP | 1 |
| 2015 | Characterizing Loop-Level Communication Patterns in Shared MemoryabstractCommunication patterns extracted from parallel programs can provide a valuable source of information for parallel pattern detection, application auto-tuning, and runtime workload scheduling on heterogeneous systems. Once identified, such patterns can help find the most promising optimizations. Communication patterns can be detected using different methods, including sandbox simulation, memory profiling, and hardware counter analysis. However, these analyses usually suffer from high runtime and memory overhead, necessitating a trade off between accuracy and resource consumption. More importantly, none of the existing methods exploit fine-grained communication patterns on the level of individual code regions. In this paper, we present an efficient tool based on Disco PoP profiler that characterizes the communication pattern of every hotspot in a shared-memory application. With the aid of static and dynamic code analysis, it produces a nested structure of communication patterns based on program's loops. By employing asymmetric signature memory, the runtime overhead is around 225× while the required amount of memory remains fixed. In comparison with other profilers, the proposed method is efficient enough to be used with real world applications. Arya Mazaheri, Ali Jannesari, Abdolreza Mirzaei, Felix Wolf 0001 |
ICPP | 1 |