Mohammad Sonji

dblp:414/4071 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0001-1253-307XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Hibiscus: End-to-end Architectural Simulation Framework for Hybrid SFQ/CMOS-Memory Compute Systems
abstract
As conventional CMOS technology approaches power and performance limits, superconducting single flux quantum (SFQ) logic offers a path to high-speed, energy-efficient computing. However, SFQ circuits require cryogenic temperatures, introducing complex challenges in memory integration and data movement between thermal zones. This paper presents an end-to-end simulation framework for hybrid SFQ/CMOS-memory systems that accurately models processor, memory, and interconnect behavior across cryogenic $(4 \mathrm{~K}, 77 \mathrm{~K})$ and room temperatures $(300 \mathrm{~K})$. The framework integrates gate-level pipelined Rapid SFQ (RSFQ) RISC-V processors, temperature-aware CryoMEM memory models, and physically grounded interconnect latency models. The simulator facilitates cross-layer design space exploration across diverse parameters such as cache placement, interconnect stack selection, and granularity. These features allow the community to identify technological gaps and re-evaluate the bottlenecks in memory-compute throughput. Our evaluations highlight the critical interplay between processor frequency and memory bandwidth, demonstrate the speedup potential of 4K SFQ caches, and quantify the impact of cryostat cabling choices on system performance.
Ryan Marsala, Yerzhan Mustafa, Prabhath Tangella, Mohammad Sonji, George Michelogiannakis, Selçuk Köse, Adwait Jog, Mehmet Esat Belviranli
ISPASS4
2026 Are We There Yet? Predicting if Executing Applications are Near Completion
Mohammad Sonji, Mohammed Baydoun, Safaa Diab, Amir Nassereldine, Pedro Bruel, Aditya Dhakal, Rolando P. Hong Enriquez, Gourav Rattihalli, Diman Zad Tootaghaj, Gallig Renaud, Barbara M. Chapman, Fatima K. Abu Salem, Eitan Frachtenberg, Dejan S. Milojicic, Izzat El Hajj
ICPE1
2025 Dissecting Performance Overheads of Confidential Computing on GPU-based Systems
abstract
Confidential computing (CC) is a critical technology for protecting data in use. By leveraging encryption and virtual machine (VM) level isolation, CC allows existing code to run without modification while offering confidentiality and integrity guarantees. However, the performance impact of CC in GPU-based systems can be significant. In this work, we present a comprehensive performance evaluation of CC guided by a simple performance model. Specifically, we start by evaluating CUDA applications with a focus on data transfer, memory management, encryption, kernel launch, and kernel execution. We also present a detailed event-level analysis of these applications, revealing that the execution times of kernels that do not use unified virtual memory (UVM) are mostly unaffected, while associated kernel launch overhead and queuing time increase significantly. On the other hand, the execution time of kernels using UVM increases drastically under CC, in addition to other launch and queuing overheads. We also study CNN training and LLM inference to see how CC overhead would affect them. Finally, we consider several optimization techniques, including kernel fusion, overlapping, and quantization, towards addressing the overheads of CC.
Mohammad Sonji, Adwait Jog
ISPASS2