Sina Darabi

dblp:298/8687 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-7082-123XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CacheCatalyst: Enhancing Web Caching for the Latency-Constrained Internet
Mohammad Hosseini 0001, Sina Darabi, Hannaneh Barahouei Pasandi, Patrick Eugster, Mahmood Choopani
NSDI2
2025 Nano-consensus: Ultra-fast, Quorum-less Coordination on the Wire
abstract
Consensus, widely regarded as the most fundamental primitive in distributed systems, lies at the core of countless services that require coordination among remote processes. Datacenter services typically achieve consensus through long-established, quorum-based algorithms such as Paxos and Raft, including recent re-adaptations for kernel bypass datapaths (e.g. smartNIC/RDMA-based consensus). While these optimizations can reduce latency to the μs-scale, they remain constrained by inherent message complexity, namely the need for acknowledgments from majority quorums to tolerate faults and arbitrary message delays. Our approach takes a step further from bare acceleration of classical primitives, focusing instead on leveraging FPGA-smartNIC and priority-queue reservation to achieve synchronous remote interactions in practice. We use synchrony to devise a novel, efficient quorum-less consensus protocol which we use to build Nano-consensus: a novel hardware consensus engine. Nano-consensus operates at network line rate and can reach consensus in 1.03μs for single-packet instances, delivering 3.82× latency and 4.8× improvements over the state of the art. We demonstrate how Nano-consensus can be integrated into distributed applications to boost both performance and consistency.
Davide Rovelli, Christian Färber, Graham McKenzie, Ali Pahlevan, Sina Darabi, Patrick Jahnke, Patrick Eugster
SoCC5
2025 A Low-latency On-chip Cache Hierarchy for Load-to-use Stall Reduction in GPUs
abstract
Memory hierarchy in Graphics Processing Units (GPUs) is conventionally designed to provide high bandwidth rather than low latency. In particular, because of the high tolerance to load-to-use latency (i.e., the time that warps wait for data fetched by memory loads), GPU L1D caches are optimized for density, capacity, and low power with latencies that are often orders of magnitude longer than conventional CPU caches. However, there are many important classes of data-parallel applications (e.g., graph, tree, priority queue processing, and sparse deep learning applications) that benefit from lower load-to-use latency than that offered by modern GPUs due to their inherent divergence and low effective Thread-Level Parallelism (TLP). This article introduces an innovative on-chip cache hierarchy that incorporates a decoupled L1D cache with reduced latency (LoTUS) and its management scheme. LoTUS is a minimally sized fully associative cache placed in each GPU subcore that captures the primary working set of data-parallel applications. It exploits conventional high-performance low-density SRAM cells and dramatically reduces load-to-use latency. We also propose an intelligent extension of LoTUS, called LoTUSage, which employs a lightweight learning-based model to predict the utility of caching requests in LoTUS. Evaluation results show that LoTUS and LoTUSage improve the average performance by 23.9% and 35.4% and reduce the average energy consumption by 27.8% and 38.5%, respectively, for the applications suffering from high load-to-use stalls with negligible area and power overheads.
Negin Mahani, Hajar Falahati, Sina Darabi, Ahmad Javadi Nezhad, Yunho Oh, Mohammad Sadrosadati, Hamid Sarbazi-Azad, Babak Falsafi
ACM Trans. Archit. Code Optim.3
2024 Uncovering Secrets of Microbursts in Datacenter Network Traffic
abstract
Designing efficient methods and policies for mitigating microbursts requires a thorough understanding of microburst characteristics and behaviors. However, the lack of detailed studies on microburst characteristics and comprehensive tools for measuring and analyzing them has been a significant challenge for researchers in this field. We introduce BurstVision, a tool that extracts various characteristics of microbursts from traffic traces. Using BurstVision, we analyze several traffic traces from various cloud datacenter applications and report on the diverse characteristics of microbursts observed. Our analysis reveals that microburst characteristics significantly vary across applications. Moreover, we discuss how these varying characteristics can influence the effectiveness of different microburst mitigation solutions. Our findings highlight the importance of considering the specific type and characteristics of microbursts in traffic when adopting a microburst mitigation solution.
Mohammad Hosseini 0001, Sina Darabi, Mohammad Nakhjiri, Patrick Eugster
CNSM2
2024 Rethinking Web Caching: An Optimization for the Latency-Constrained Internet
abstract
Caching is a fundamental web technique for reducing Page Load Time (PLT) by reusing previously fetched resources. We highlight the drawbacks of the current caching approach, especially in the context of high-speed networks where latency, rather than bandwidth, is the primary bottleneck for web performance. We discuss how the current design of web caching suffers from inefficiencies, particularly due to the latency involved in re-validation requests, which diminishes the potential benefits of caching. To address this inefficiency, we present a novel solution in which web servers proactively provide clients with the latest validation tokens for resources during the initial step of page loading, allowing browsers to use unchanged cached content without unnecessary round trips. This method significantly reduces PLT, with preliminary evaluations showing a 30% improvement.
Mohammad Hosseini 0001, Sina Darabi, Patrick Eugster, Mahmood Choopani, Amir Hossein Jahangir
HotNets2
2024 Yuz: Improving Performance of Cluster-Based Services by Near-L4 Session-Persistent Load Balancing
abstract
Large-scale services are deployed in data centers using clusters of servers, and load balancers (LB) are responsible for distributing requests for a service among its servers. Layer-4 (L4) LBs process requests faster than Layer-7 (L7) ones, but they cannot provide session-persistent load balancing, and therefore, they direct connections of an application-level session to different servers. On the other side, L7 LBs can direct all requests of an application-level session to the same server, but they have a very limited capacity because they act as a reverse-proxy and process requests at the application layer. We present “Yuz”, a stateless session-persistent load balancer that does not act as a reverse-proxy. Yuz works near layer 4, and it makes use of TLS session data instead of processing incoming requests at the application level. Our evaluations show that the request rate that can be handled by cluster-based services equipped with Yuz is twice as high as when the clusters use the best existing load balancers. Yuz also significantly reduces the average and tail of the clusters’ response time. Moreover, while each of the existing session-persistent LBs works only for a specific application, Yuz provides an application-independent session-persistent load balancing.
Mohammad Hosseini 0001, Sina Darabi, Amir Hossein Jahangir, Ali Movaghar-Rahimabadi
IEEE Trans. Netw. Serv. Manag.2
2022 Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core Resources
abstract
Graphics Processing Units (GPUs) are widely-used accelerators for data-parallel applications. In many GPU applications, GPU memory bandwidth bottlenecks performance, causing underutilization of GPU cores. Hence, disabling many cores does not affect the performance of memory-bound workloads. While simply power-gating unused GPU cores would save energy, prior works attempt to better utilize GPU cores for other applications (ideally compute-bound), which increases the GPU’s total throughput. In this paper, we introduce Morpheus, a new hardware/software co-designed technique to boost the performance of memory-bound applications. The key idea of Morpheus is to exploit unused core resources to extend the GPU last level cache (LLC) capacity. In Morpheus, each GPU core has two execution modes: compute mode and cache mode. Cores in compute mode operate conventionally and run application threads. However, for the cores in cache mode, Morpheus invokes a software helper kernel that uses the cores’ on-chip memories (i.e., register file, shared memory, and L1) in a way that extends the LLC capacity for a running memory-bound workload. Morpheus adds a controller to the GPU hardware to forward LLC requests to either the conventional LLC (managed by hardware) or the extended LLC (managed by the helper kernel). Our experimental results show that Morpheus improves the performance and energy efficiency of a baseline GPU architecture by an average of 39% and 58%, respectively, across several memory-bound workloads. Morpheus’ performance is within 3% of a GPU design that has a quadruple-sized conventional LLC. Morpheus can thus contribute to reducing the hardware dedicated to a conventional LLC by exploiting idle cores’ on-chip memory resources as additional cache capacity.
Sina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger, Mohammad Hosseini 0001, Jisung Park 0001, Juan Gómez-Luna, Onur Mutlu, Hamid Sarbazi-Azad
MICRO1
2022 OSM: Off-Chip Shared Memory for GPUs
abstract
Graphics Processing Units (GPUs) employ a shared memory, a software-managed cache for programmers, in each streaming multiprocessor to accelerate data sharing among the threads in a thread block. Although 60% of the shared memory space is underutilized, on average, there are some workloads that demand higher shared memory capacities. Therefore, improving shared memory utilization while satisfying the needs of shared memory intensive workloads is challenging. We make a key observation that the lifetime of each shared memory address is significantly shorter than the execution time of a thread block. In this paper, we first propose Off-Chip Shared Memory (OSM) that allocates shared memory space in the off-chip memory, and accelerates accesses to it via a small on-chip cache. Using an 8 KB cache for shared memory addresses, OSM provides almost the same performance as the baseline GPU that uses 96 KB on-chip shared memory. OSM improves GPU performance in two ways. First, it allocates higher shared memory capacities in the off-chip memory, and improves thread-level parallelism (TLP). Second, it designs a unified cache for shared memory and global address spaces, providing more caching space for global memory address space even for the workloads with high shared memory utilization. Our experimental results show an average 21% and 18% IPC improvement compared to the baseline and the state-of-the-art architectures.
Sina Darabi, Ehsan Yousefzadeh-Asl-Miandoab, Negar Akbarzadeh, Hajar Falahati, Pejman Lotfi-Kamran, Mohammad Sadrosadati, Hamid Sarbazi-Azad
IEEE Trans. Parallel Distributed Syst.1
2021 PF-DRAM: A Precharge-Free DRAM Structure
abstract
Although DRAM capacity and bandwidth have increased sharply by the advances in technology and standards, its latency and energy per access have remained almost constant in recent generations. The main portion of DRAM power/energy is dissipated by Read, Write, and Refresh operations, all initiated by a Precharge phase. Precharge phase not only imposes a large amount of energy consumption, but also increases the delay of closing a row in a memory block to open another one. By reduction of row-hit rate in recent workloads, especially in multi-core systems, precharge rate increases which exacerbates DRAM power dissipation and access latency. This work proposes a novel DRAM structure, called Precharge-Free DRAM (PF-DRAM), that eliminates the Precharge phase of DRAM. PF-DRAM uses the charge on bitlines from the previous Activation phase, as the starting point for the next Activation. The difference between PF-DRAM and conventional DRAM structure is limited to precharge and equalizer circuitry and simple modifications in sense amplifier, which are all limited to subarray level. PF-DRAM is compatible with the mainstream JEDEC memory standards like DDRx and HBM, with minimum modifications in memory controller. Furthermore, almost all of the previously proposed power/energy reduction techniques in DRAM are still applicable to PF-DRAM for further improvement. Our experimental results on a 8GB memory system running SPEC CPU2017 and PARSEC2.1 workloads show an average of 35.3% memory power consumption reduction (up to 54.2%) achieved by the system using PF-DRAM with respect to the system using conventional DRAM. Moreover, the overall performance is improved by 8.6%, in average (up to 24.3%). According to our analysis, all such improvements are achieved for less than 9% area overhead.
Nezam Rohbani, Sina Darabi, Hamid Sarbazi-Azad
ISCA2