Yasas Seneviratne

dblp:298/8796 · also Yasas Senevirathne · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0000-9052-4374ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HARMONI: Hierarchical ARchitecture MOdeling for LLMs with Near/In Memory Computing
abstract
LLM inference has emerged as a strongly memory-bound workload suitable for Processing-in-Memory (PIM) and Processing-near-Memory (PNM) architectures. Yet, most existing PIM-LLM evaluation frameworks are in-house, closed-source, or difficult to extend, and often fail to model end-to-end inference behavior or tensor-allocation and communication effects that critically shape performance. We introduce HARMONI, a fast, modular, memory-centric performance modeling framework designed specifically for hierarchical PIM/PNM architectures running LLM workloads. HARMONI captures the complete inference execution through a task-graph representation and supports user-defined tensor allocation with address-interleaving schemes. It models computation, communication, and queuing delays across logic nodes, enabling realistic mapping of inference kernels onto heterogeneous logic nodes. HARMONI outputs detailed, kernel-wise and resource-wise time and energy breakdowns, providing actionable visibility into architectural bottlenecks. Together, these capabilities make HARMONI a practical tool for rapidly evaluating PIM/PNM designs and DRAM standards (e.g., DDR4, DDR5, GDDR6), enabling systematic co-evaluation of inference workloads and memory-centric architecture designs.
Khyati Kiyawat, Yasas Seneviratne, Zhenxing Fan, Morteza Baradaran, Kevin Skadron
ISPASS2
2023 NearPM: A Near-Data Processing System for Storage-Class Applications
abstract
Persistent Memory (PM) technologies enable both fast memory access and recovery in case of a failure. To ensure crash-consistent behavior, programs need to enforce persist ordering and employ mechanisms that introduce additional data movements such as logging, checkpointing, and shadow-paging. The emerging near-data processing (NDP) architectures can effectively reduce this overhead. In this work, we propose NearPM, a near-data processor that accelerates common, primitive operations that are crucial to crash consistency. Using these primitives, NearPM accelerates commonly-used crash-consistency mechanisms. NearPM further reduces the synchronization overheads between the NDP and the CPU by handling ordering near memory. We propose Partitioned Persist Ordering (PPO) that ensures a correct persist ordering between CPU and NDP devices, as well as among multiple NDP devices. We prototype NearPM on an FPGA platform. NearPM executes the data-intensive operations of crash-consistency mechanisms with correct ordering guarantees, while the rest of the program runs on the CPU. We evaluate nine PM workloads, each implemented in three crash consistency mechanisms: logging, checkpointing, and shadow paging. Overall, NearPM achieves 4.3 -- 9.8× speedup in the NDP-offloaded operations and 1.22 -- 1.35× speedup in the whole applications.
Yasas Seneviratne, Korakit Seemakhupt, Sihang Liu 0001, Samira Manabi Khan
EuroSys1
2021 PMNet: In-Network Data Persistence
abstract
To guarantee data persistence, storage workloads (such as key-value stores and databases) typically use a synchronous protocol that places the network and server stack latency on the critical path of request processing. The use of the fast and byte-addressable persistent memory (PM) has helped mitigate the storage overhead of the server stack; yet, networking is still a dominant factor in the end-to-end latency of request processing. Emerging programmable network devices can reduce network latency by moving parts of the applications’ compute into the network (e.g., caching results for read requests); however, for update requests, the client still has to stall on the server to commit the updates, persistently.In this work, we introduce in-network data persistence that extends the data-persistence domain from servers to the network, and present PMNet, a programmable data plane (e.g., switch or NIC) with PM for persisting data in the network. PMNet logs incoming update requests and acknowledges clients directly without having them wait on the server to commit the request. In case of a failure, the logged requests act as redo logs for the server to recover. We implement PMNet on an FPGA and evaluate its performance using common PM workloads, including key-value stores and PM-backed applications. Our evaluation shows that PMNet can improve the throughput of update requests by 4.31× on average, and the 99th-percentile tail latency by 3.23×.
Korakit Seemakhupt, Sihang Liu 0001, Yasas Seneviratne, Muhammad Shahbaz 0001, Samira Manabi Khan
ISCA3