Amir Hossein Nodehi Sabet

dblp:217/1290 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0001-4344-3819ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 GraCFL: A Holistically Designed Vertex-Centric Graph System for CFL Reachability
abstract
Many program analyses can be formulated as context-free language (CFL) reachability problems on an edge-labeled graph.While graph systems have been proposed recently for large-scale CFL reachability analysis of system software, the design space has not yet been systematically explored, leading to sub-optimal performance.This work presents GraCFL 1 , a holistically designed graph system for CFL reachability.Inspired by the vertex-centric processing paradigm, we formalize CFL reachability using a multi-directional vertex-centric model.We then analyze this model in terms of computation redundancy, strategies for deriving new reachability, data locality, and parallelism.The analysis reveals a set of insights that guide the design of new techniques and optimizations to improve system performance.As a result of the systematic design, GraCFL demonstrates superior performance compared to state-ofthe-art graph systems, with an average 14.14× speedup over Graspan and 8.33× speedup over POCR.Its source code is available at https://github.com/AutomataLab/GraCFL.
Sakib Fuad, Amir Hossein Nodehi Sabet, Umar Farooq 0002, Zhijia Zhao 0001
ICS2
2022 Breaking the computation and communication abstraction barrier in distributed machine learning workloads
abstract
Recent trends towards large machine learning models require both training and inference tasks to be distributed. Considering the huge cost of training these models, it is imperative to unlock optimizations in computation and communication to obtain best performance. However, the current logical separation between computation and communication kernels in machine learning frameworks misses optimization opportunities across this barrier. Breaking this abstraction can provide many optimizations to improve the performance of distributed workloads. However, manually applying these optimizations requires modifying the underlying computation and communication libraries for each scenario, which is both time consuming and error-prone.
Abhinav Jangda, Amir Hossein Nodehi Sabet, Saeed Maleki, Youshan Miao, Madan Musuvathi, Todd Mytkowicz, Olli Saarikivi
ASPLOS4
2021 Scalable FSM parallelization via path fusion and higher-order speculation
abstract
Finite-state machine (FSM) is a fundamental computation model used by many applications. However, FSM execution is known to be “embarrassingly sequential” due to the state dependences among transitions. Existing solutions leverage enumerative or speculative parallelization to break the dependences. However, the efficiency of both parallelization schemes highly depends on the properties of the FSM and its inputs. For those exhibiting unfavorable properties, the former suffers from the overhead of maintaining multiple execution paths, while the latter is bottlenecked by the serial reprocessing among the misspeculation cases. Either way, the FSM parallelization scalability is seriously compromised.
Junqiao Qiu, Xiaofan Sun, Amir Hossein Nodehi Sabet, Zhijia Zhao 0001
ASPLOS3
2020 Subway: minimizing data transfer during out-of-GPU-memory graph processing
abstract
In many graph-based applications, the graphs tend to grow, imposing a great challenge for GPU-based graph processing. When the graph size exceeds the device memory capacity (i.e., GPU memory oversubscription), the performance of graph processing often degrades dramatically, due to the sheer amount of data transfer between CPU and GPU.
Amir Hossein Nodehi Sabet, Zhijia Zhao 0001, Rajiv Gupta 0001
EuroSys1
2020 Reliability Analysis for Unreliable FSM Computations
abstract
Finite State Machines (FSMs) are fundamental in both hardware design and software development. However, the reliability of FSM computations remains poorly understood. Existing reliability analyses are mainly designed for generic computations and are unaware of the special error tolerance characteristics in FSM computations. This work introduces RelyFSM -- a state-level reliability analysis framework for FSM computations. By modeling the behaviors of unreliable FSM executions and qualitatively reasoning about the transition structures, RelyFSM can precisely capture the inherent error tolerance in FSM computations. Our evaluation with real-world FSM benchmarks confirms both the accuracy and efficiency of RelyFSM.
Amir Hossein Nodehi Sabet, Junqiao Qiu, Zhijia Zhao 0001, Sriram Krishnamoorthy
ACM Trans. Archit. Code Optim.1
2018 Tigr: Transforming Irregular Graphs for GPU-Friendly Graph Processing
abstract
Graph analytics delivers deep knowledge by processing large volumes of highly connected data. In real-world graphs, the degree distribution tends to follow the power law -- a small portion of nodes own a large number of neighbors. The high irregularity of degree distribution acts as a major barrier to their efficient processing on GPU architectures, which are primarily designed for accelerating computations on regular data with SIMD executions. Existing solutions to the inefficiency of GPU-based graph analytics either modify the graph programming abstraction or rely on changes to the low-level thread execution models. The former requires more programming efforts for designing and maintaining graph analytics; while the latter couples with the underlying architectures, making it difficult to adapt as architectures quickly evolve. Unlike prior efforts, this work proposes to address the above fundamental problem at its origin -- the irregular graph data itself. It raises a critical question in irregular graph processing: Is it possible to transform irregular graphs into more regular ones such that the graphs can be processed more efficiently on GPU-like architectures, yet still producing the same results? Inspired by the question, this work introduces Tigr -- a graph transformation framework that can effectively reduce the irregularity of real-world graphs with correctness guarantees for a wide range of graph analytics. To make the transformations practical, Tigr features a lightweight virtual transformation scheme, which can substantially reduce the costs of graph transformations, while preserving the benefits of reduced irregularity. Evaluation on Tigr-based GPU graph processing shows significant and consistent speedup over the state-of-the-art GPU graph processing frameworks for a spectrum of irregular graphs.
Amir Hossein Nodehi Sabet, Junqiao Qiu, Zhijia Zhao 0001
ASPLOS1