Vivek Govindasamy

dblp:348/9729 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0005-4745-3669ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Mixed-Level Modeling and Evaluation of a Cache-less Grid of Processing Cells
abstract
Modern processors experience memory contention when the speed of their computational units exceeds the rate at which new data is available to be processed. This phenomenon is well known as the memory wall and is a great challenge in computer engineering. The reason for this phenomenon is the unequal growth rate in memory access speeds compared to processor clock rates. To mitigate the memory bottleneck in classic computer architectures, a scalable parallel computing platform—the Grid of Processing Cells (GPC)—has been proposed. To evaluate its effectiveness, the GPC is modeled at the instruction level and functional level using SystemC TLM-2.0, with a focus on memory contention. Individual GPC cells can be switched between the two abstraction levels. Our mixed-level system model enables fast and accurate simulations. We test multiple streaming applications on the GPC, and analyze software-based optimization methods and their effects on the GPC, at both abstraction levels. The performance is then compared against the traditional shared memory processor architecture. Experimental results show improved execution times on the GPC primarily due to a large decrease in main memory contention.
Vivek Govindasamy, Rainer Dömer
ACM Trans. Embed. Comput. Syst.1
2025 A Quantitative Guide to Navigate Speed/Accuracy Tradeoffs in System Level Design of RISC-V Processor Grids
abstract
Even when following a well-structured top-down design methodology, system designers regularly face obstacles and pitfalls posed by speed/accuracy tradeoffs in modeling and simulation of complex hardware and software systems. Across the abstraction levels, modeling details grow exponentially, while simulation speed decreases by multiple orders of magnitude. To quantify these effects, we systematically generate, simulate, and evaluate grid-based systems-on-chip in a top-down open-source based tool flow. We map two parallel software applications onto a scalable grid of RISC-V processors and successively refine and validate the models at lower abstraction levels, namely TLM, ISS, RTL, and FPGA. Our comprehensive experimental evaluation over five abstraction levels quantifies the speed-accuracy tradeoffs in simulator build and run times. In addition to its educational value, our work can guide the system designer on an efficient path to a cycle-accurate software simulation on fully constructed hardware.
Lars Luchterhandt, Vivek Govindasamy, Christoph Scheytt, Wolfgang Müller 0003, Rainer Dömer
FDL2
2024 BusyMap, an Efficient Data Structure to Observe Interconnect Contention in SystemC TLM-2.0
abstract
For designing embedded computer architectures that meet desired performance constraints at low cost, fast and accurate simulation models are needed early in the design flow. To identify and avoid bottlenecks early at the system level, observing the contention of shared resources is critical. In this paper, we propose and evaluate a novel data structure called BusyMap that accurately reflects contention at system busses or similar inter-connect components. BusyMap is an efficient data structure that allows the system designer to accurately model and easily observe contention in IEEE loosely-timed TLM-2.0. In contrast to prior state-of-the-art, our model fully supports temporal decoupling and multiple levels of interconnect. Our experiments demonstrate the effectiveness of BusyMap with results showing high accuracy at high -speed System C simulation.
Emad Malekzadeh Arasteh, Vivek Govindasamy, Rainer Dömer
DATE2
2023 Instruction-Level Modeling and Evaluation of a Cache-Less Grid of Processing Cells
abstract
While processor speeds continue to show performance increases, memory access speeds remain significantly slower. One solution is to develop novel computer architectures, specifically designed to address the memory wall such as the cache-less Grid of Processing Cells (GPC). In this work, we model the GPC architecture using SystemC TLM-2.0 at the instruction-level with high accuracy, while retaining the characteristic high simulation speeds of transaction-level modeling. To compare the performance of our modeled GPC with existing architectures, we also model a single core and an 8-core SMP with optimized caches and run a bare-metal Canny application on the three architectures. We show promising performance evaluation results in architectures without caches.
Vivek Govindasamy, Rainer Dömer
FDL1