EDBT 2026 Demo / reviewers in the wild / expert
Soheil Nazar Shahsavani
dblp:168/6301
· DBLP profile ↗
9ranked-venue papers
6as first author
3since 2021 · last 2023
0000-0001-5159-9103ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Efficient Compilation and Mapping of Fixed Function Combinational Logic onto Digital Signal Processors Targeting Neural Network Inference and Utilizing High-level SynthesisabstractRecent efforts for improving the performance of neural network (NN) accelerators that meet today’s application requirements have given rise to a new trend of logic-based NN inference relying on fixed function combinational logic. Mapping such large Boolean functions with many input variables and product terms to digital signal processors (DSPs) on Field-programmable gate arrays (FPGAs) needs a novel framework considering the structure and reconfigurability of DSP blocks during this process. The proposed methodology in this article maps the fixed function combinational logic blocks to a set of Boolean functions where Boolean operations corresponding to each function are mapped to DSP devices rather than look-up tables on the FPGAs to take advantage of the high performance, low latency, and parallelism of DSP blocks. This article also presents an innovative design and optimization methodology for compilation and mapping of NNs, utilizing fixed function combinational logic to DSPs on FPGAs employing high-level synthesis flow. Our experimental evaluations across several datasets and selected NNs demonstrate the comparable performance of our framework in terms of the inference latency and output accuracy compared to prior art FPGA-based NN accelerators employing DSPs. Soheil Nazar Shahsavani, Arash Fayyazi, Mahdi Nazemi, Massoud Pedram |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | NullaNet Tiny: Ultra-low-latency DNN Inference Through Fixed-function Combinational LogicabstractWhile there is a large body of research on efficient processing of deep neural networks (DNNs) [1]-[31], ultra-low-latency realization of these models for applications with stringent, sub-microsecond latency requirements continues to be an unresolved, challenging problem. Field-programmable gate array (FPGA)-based DNN accelerators are gaining traction as a serious contender to replace graphics processing unit/central processing unit-based platforms considering their performance, flexibility, and energy efficiency. NullaNet (2018) [32], LUTNet (2019) [33], and LogicNets (2020) [34] are among accelerators specifically designed to benefit from FPGAs' capabilities. Mahdi Nazemi, Arash Fayyazi, Amirhossein Esmaili, Atharva Khare, Soheil Nazar Shahsavani, Massoud Pedram |
FCCM | 5 |
| 2021 | A Variation-aware Hold Time Fixing Methodology for Single Flux Quantum Logic CircuitsabstractSingle flux quantum (SFQ) logic is a promising technology to replace complementary metal-oxide-semiconductor logic for future exa-scale supercomputing but requires the development of reliable EDA tools that are tailored to the unique characteristics of SFQ circuits, including the need for active splitters to support fanout and clocked logic gates. This article is the first work to present a physical design methodology for inserting hold buffers in SFQ circuits. Our approach is variation-aware, uses common path pessimism removal and incremental placement to minimize the overhead of timing fixes, and can trade off layout area and timing yield. Compared to a previously proposed approach using fixed hold time margins, Monte Carlo simulations show that, averaging across 10 ISCAS’85 benchmark circuits, our proposed method can reduce the number of inserted hold buffers by 8.4% with a 6.2% improvement in timing yield and by 21.9% with a 1.7% improvement in timing yield. Soheil Nazar Shahsavani, Massoud Pedram, Peter A. Beerel |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2020 | TDP-ADMM: A Timing Driven Placement Approach for Superconductive Electronic Circuits Using Alternating Direction Method of MultipliersabstractThis paper presents a novel timing driven global placement approach utilizing the alternating direction method of multipliers (ADMM) targeting superconductive electronic circuits. The proposed algorithm models the placement problem as an optimization problem with constraints on the maximum wirelength delay of timing-critical paths and employs the ADMM algorithm to decompose the problem into two sub-problems, one minimizing the total wirelength of the circuit and the other minimizing the delay of timing-critical paths of the circuit. Through an iterative process, a placement solution is generated that simultaneously minimizes the total wirelength and satisfies the setup time constraints. Compared to an state-of-the-art academic global placement tool, the proposed method (called TDP-ADMM) improves the worst and total negative slack for seven single flux quantum benchmark circuits by an average of 26% and 44%, respectively, with an average overhead of 1.98% in terms of total wirelength. Soheil Nazar Shahsavani, Massoud Pedram |
DAC | 1 |
| 2020 | A Timing Uncertainty-Aware Clock Tree Topology Generation Algorithm for Single Flux Quantum CircuitsabstractThis paper presents a low-cost, timing uncertainty-aware synchronous clock tree topology generation algorithm for single flux quantum (SFQ) logic circuits. The proposed method considers the criticality of the data paths in terms of timing slacks as well as the total wirelength of the clock tree and generates a (height-) balanced binary clock tree using a bottom-up approach and an integer linear programming (ILP) formulation. The statistical timing analysis results for ten benchmark circuits show that the proposed method improves the total wirelength and the total negative hold slack by 4.2% and 64.6%, respectively, on average, compared with a wirelength-driven state-of-the-art balanced topology generation approach. Soheil Nazar Shahsavani, Bo Zhang 0098, Massoud Pedram |
DATE | 1 |
| 2018 | A placement algorithm for superconducting logic circuits based on cell grouping and super-cell placementabstractThis paper presents a novel clustering based placement algorithm for single flux quantum (SFQ) family of super-conductive electronic circuits. In these circuits nearly all cells receive a clock signal and a placement algorithm that ignores the clock routing cost will not produce high quality solutions. To address this issue, proposed approach simultaneously minimizes the total wirelength of the signal nets and area overhead of the clock routing. Furthermore, construction of a perfect H-tree in SFQ logic circuits is not viable solution due to the resulting very high routing overhead and the in-feasibility of building exact zero-skew clock routing trees. Instead a hybrid clock tree must be used whereby higher levels of the clock tree (i.e., those closer to the clock source) are based on H-tree construction whereas lower levels of the clock tree follow a linear (i.e., chain-like) structure. The proposed approach is able to reduce the overall half-perimeter wirelength by 15% and area by 8% compared with state-of-the-art techniques. Soheil Nazar Shahsavani, Alireza Shafaei, Massoud Pedram |
DATE | 1 |
| 2018 | Accurate margin calculation for single flux quantum logic cellsabstractThis paper presents a novel method for accurate margin calculation of single flux quantum (SFQ) logic cells in a superconducting electronic circuit. The proposed method can be utilized as a figure of merit to estimate the robustness of a logic cell without the need for expensive Monte-Carlo simulations. This is achieved through efficient state-space exploration of all parameters in the cell structure. Using the proposed approach, distinct parameter dispersion (DPD) based yield of SFQ cells increases by 55% on average, compared with state-of-the-art techniques. Soheil Nazar Shahsavani, Bo Zhang 0098, Massoud Pedram |
DATE | 1 |
| 2017 | A thermally-aware energy minimization methodology for global interconnectsabstractAs a result of the Temperature Effect Inversion (TEI) in FinFET-based designs, gate delays decrease with the increase of temperature. In contrast, the resistive characteristic and hence delay of global interconnects increase with the temperature. However, as shown in this paper, if buffers are judiciously inserted in global interconnects, the buffer delay decrease is more pronounced than the interconnect delay increase, resulting in an overall performance improvement at higher temperatures. More specifically, this work models the delay of buffer-inserted global interconnects vs. temperature in order to derive the optimal number and size of buffers for a given interconnect length and temperature. Furthermore, the paper addresses the problem of minimizing the buffered interconnect energy consumption by changing the supply voltage level or FinFET threshold voltage, and also presents a temperature-aware optimization policy for solving this problem. Simulation results show average interconnect energy savings of 16% with no performance penalty for five different benchmarks implemented on a 14nm FinFET technology. Soheil Nazar Shahsavani, Alireza Shafaei, Shahin Nazarian, Massoud Pedram |
DATE | 1 |
| 2015 | A heuristic machine learning-based algorithm for power and thermal management of heterogeneous MPSoCsabstractIn this work, we propose a power and thermal management algorithm based on machine learning to control the thermal stresses and power consumption of the heterogeneous MPSoCs. The objectives of the proposed algorithm are increasing the performance and decreasing the spatial and temporal temperature gradients along with the thermal cycling under the power and temperature constraints. Our proposed power and thermal management method is based on a heuristic approach to speed up the convergence of the machine learning algorithm which makes it applicable for general purpose processors. Adopting Q-Learning as the machine learning algorithm, the heuristic approach aids to limit the learning space by suggesting the most appropriate actions to the agent in each decision epoch. The heuristic algorithm employs the current and previous states of the machine learning, as well as the amount of the temperature stress and power consumption of each core to determine the appropriate action for each core, independently. The proposed algorithm is evaluated on 4-core, 8-core and 16-core homogeneous and heterogeneous MPSoCs for some benchmarks in the Splash2 benchmark package. The results reveal a faster convergence of machine learning and more thermal stresses reduction. Arman Iranfar, Soheil Nazar Shahsavani, Mehdi Kamal, Ali Afzali-Kusha |
ISLPED | 2 |