EDBT 2026 Demo / reviewers in the wild / expert
Kamalakkannan Kamalavasan
dblp:234/1340 · also Kamalavasan Kamalakkannan
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-7282-5165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DX100: Programmable Data Access Accelerator for IndirectionabstractIndirect memory accesses frequently appear in applications where memory bandwidth is a critical bottleneck.Prior indirect memory access proposals, such as indirect prefetchers, runahead execution, fetchers, and decoupled access/execute architectures, primarily focus on improving memory access latency by loading data ahead of computation but still rely on the DRAM controllers to reorder memory requests and enhance memory bandwidth utilization.DRAM controllers have limited visibility to future memory accesses due to the small capacity of request buffers and the restricted memorylevel parallelism of conventional core and memory systems.We introduce DX100, a programmable data access accelerator for indirect memory accesses.DX100 is shared across cores to offload bulk indirect memory accesses and associated address calculation operations.DX100 reorders, interleaves, and coalesces memory requests to improve DRAM row-buffer hit rate and memory bandwidth utilization.DX100 provides a general-purpose ISA to support diverse access types, loop patterns, conditional accesses Alireza Khadem, Kamalakkannan Kamalavasan, Zhenyan Zhu, Akash Poptani, Yufeng Gu, Jered Dominguez-Trujillo, Nishil Talati, Daichi Fujiki, Scott A. Mahlke, Galen M. Shipman, Reetuparna Das |
ISCA | 2 |
| 2022 | High throughput multidimensional tridiagonal system solvers on FPGAsabstractWe present a high performance tridiagonal solver library for Xilinx FPGAs optimized for multiple multi-dimensional systems common in real-world applications. An analytical performance model is developed and used to explore the design space and obtain rapid performance estimates that are over 85% accurate. This library achieves an order of magnitude better performance when solving large batches of systems than previous FPGA work. A detailed comparison with a current state-of-the-art GPU library for multi-dimensional tridiagonal systems on an Nvidia V100 GPU shows the FPGA achieving competitive or better runtime and significant energy savings of over 30%. Through this design, we learn lessons about the types of applications where FPGAs can challenge the current dominance of GPUs. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
ICS | 1 |
| 2021 | High-Level FPGA Accelerator Design for Structured-Mesh-Based Explicit Numerical SolversabstractThis paper presents a workflow for synthesizing near-optimal FPGA implementations of structured-mesh based stencil applications for explicit solvers. It leverages key characteristics of the application class and its computation-communication pattern and the architectural capabilities of the FPGA to accelerate solvers for high-performance computing applications. Key new features of the workflow are (1) the unification of standard state-of-the-art techniques with a number of high-gain optimizations such as batching and spatial blocking/tiling, motivated by increasing throughput for real-world workloads and (2) the development and use of a predictive analytical model to explore the design space, and obtain resource and performance estimates. Three representative applications are implemented using the design workflow on a Xilinx Alveo U280 FPGA, demonstrating near-optimal performance and over 85% predictive model accuracy. These are compared with equivalent highly-optimized implementations of the same applications on modern HPC-grade GPUs (Nvidia V100), analyzing time to solution, bandwidth, and energy consumption. Performance results indicate comparable runtimes with the V100 GPU, with over 2× energy savings for the largest non-trivial application on the FPGA. Our investigation shows the challenges of achieving high performance on current generation FPGAs compared to traditional architectures. We discuss determinants for a given stencil code to be amenable to FPGA implementation, providing insights into the feasibility and profitability of a design and its resulting performance. Kamalakkannan Kamalavasan, Gihan R. Mudalige, István Z. Reguly, Suhaib A. Fahmy |
IPDPS | 1 |