Abdolah Amirany

dblp:286/2998 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-0298-6945ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
abstract
The deployment of mixture-of-experts (MoE) large language models (LLMs) presents significant challenges due to their high memory demands. These challenges become even more pronounced in multi-tenant environments, where shared resources must accommodate multiple models, limiting the effectiveness of conventional virtualization techniques. This paper addresses the problem of efficiently serving multiple fine-tuned MoE-LLMs on a single GPU. We propose a serving system that employs similarity-based expert consolidation to reduce the overall memory footprint by sharing similar experts across models. To ensure output quality, we introduce runtime partial reconfiguration, dynamically replacing non-expert layers when processing requests from different models. As a result, our approach achieves competitive output quality while maintaining throughput comparable to serving a single model, and incurs only a negligible increase in time-to-first-token (TTFT). Experiments on a server with a single NVIDIA A100 GPU (80GB) using Mixtral-8x7B models demonstrate an 85% average reduction in turnaround time compared to NVIDIA’s multi-instance GPU (MIG). Furthermore, experiments on Google’s Switch Transformer Base-8 model with up to four variants demonstrate the scalability and resilience of our approach in maintaining output quality compared to other model merging baselines, highlighting its effectiveness.
Hamid Reza Imani, Peiman Mohseni, Abdolah Amirany, Tarek A. El-Ghazawi
ICML4
2025 Synergizing spintronics and quaternary logic: a hardware accelerator for neural networks with optimized quantization algorithm
Motahareh BahmanAbadi, Abdolah Amirany, Mohammad Hossein Moaiyeri, Kian Jafari
J. Supercomput.2
2025 Balancing precision and efficiency: an approximate multiplier with built-in error compensation for error-resilient applications
Ladan Sayadi, Abdolah Amirany, Mohammad Hossein Moaiyeri, Somayeh Timarchi
J. Supercomput.2
2024 Protecting the Intellectual Property of Binary Deep Neural Networks With Efficient Spintronic-Based Hardware Obfuscation
abstract
Well-trained deep neural network (DNN) models are considered valuable assets because they require large amounts of data, expertise, and resources to achieve desired performance. Hence, protecting the intellectual property of such hard-to-develop models against unauthorized usage or model leaking is a significant concern. This paper proposes a novel key-based obfuscation method that locks the model with a significant accuracy drop when the incorrect key is applied. Due to the importance and developments of binary neural networks (BNNs) in hardware implementation of state-of-the-art DNN models, we study our method on BNNs. The proposed model protection solution leads to a higher accuracy drop with even a lower perturbation rate across different binary neural network architectures and benchmark datasets than its state-of-the-art counterpart. Furthermore, we present an efficient spintronic-based in-memory computing structure for the hardware implementation of the proposed method. We validate the proposed design using post-layout simulations based on the TSMC 40nm technology. With the same approach for hardware implementation, our proposed design provides, on average, 18%, 41%, and 40% improvements regarding the area, average power consumption, and weight modification energy per filter in the neural network structure, respectively.
Alireza Mohseni, Mohammad Hossein Moaiyeri, Abdolah Amirany, Mohammad Hadi Rezayati
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 High-Performance Spintronic Nonvolatile Ternary Flip-Flop and Universal Shift Register
abstract
Multiple-valued logic (MVL) shows considerable advantages over binary logic in certain applications because of the increased informational content of its signals, and hence reduction in interconnects. Flip-flops (FFs) are the basic elements of many systems and are widely used in microprocessors due to their high performance. This article presents two spintronic ternary retention FFs, and a nonvolatile universal ternary shift register (NUTSR) based on gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) and nonvolatile magnetic tunnel junction (MTJ). In the proposed input-aware ternary retention FF circuit, power consumption is significantly reduced by adding a magnitude comparator (MC) circuit and preventing duplicate data transfer to MTJs. Simulation results indicate that our design offers at least 22%, 40%, and 15% reductions in power consumption, backup time, and restore energy, respectively. Moreover, it eliminates the risk of data loss in the event of a sudden power outage.
Abdolah Amirany, Kian Jafari, Mohammad Hossein Moaiyeri
IEEE Trans. Very Large Scale Integr. Syst.1