Santhosh Kumar Rethinagiri

dblp:39/10504 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 7 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 54% Electronic design automation · 18% Energy-efficient computing · 18%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory access patterns
irregular memory access
0.212014
APMC: advanced pattern based memory controller (abstract only) · FPGA 2014
Memory systems
memory access patterns
0.212014
APMC: advanced pattern based memory controller (abstract only) · FPGA 2014
Memory systems
memory controller
0.212014
APMC: advanced pattern based memory controller (abstract only) · FPGA 2014
Energy-efficient computing
power and energy modeling
0.212014
Power estimation tool for system on programmable chip based platforms (abstract only) · FPGA 2014
Electronic design automation
power estimation
0.212014
Power estimation tool for system on programmable chip based platforms (abstract only) · FPGA 2014
Reconfigurable computing and FPGAs
FPGA-based memory controller
0.112014
APMC: advanced pattern based memory controller (abstract only) · FPGA 2014
Embedded and real-time systems › embedded hardware platform
multicore embedded systems
0.112014
Power estimation tool for system on programmable chip based platforms (abstract only) · FPGA 2014

Methods — techniques the papers use, named apart from their topics

virtual platform simulation · 0.2functional power modeling · 0.2descriptor-based prefetching · 0.2
YearPublicationVenuePosition
2019 Visual Inertial Odometry At the Edge: A Hardware-Software Co-design Approach for Ultra-low Latency and Power
abstract
Visual Inertial Odometry (VIO) is used for estimating pose and trajectory of a system and is a foundational requirement in many emerging applications like AR/VR, autonomous navigation in cars, drones and robots. In this paper, we analyze key compute bottlenecks in VIO and present a highly optimized VIO accelerator based on a hardware-software codesign approach. We detail a set of novel micro-architectural techniques that optimize compute, data movement, bandwidth and dynamic power to make it possible to deliver high quality of VIO at ultra-low latency and power required for budget constrained edge devices. By offloading the computation of the critical linear algebra algorithms from the CPU, the accelerator enables high sample rate IMU usage in VIO processing while acceleration of image processing pipe increases precision, robustness and reduces IMU induced drift in final pose estimate. The proposed accelerator requires a small silicon footprint (1.3 mm2in a 28nm process at 600 MHz), utilizes a modest on-chip shared SRAM (560KB) and achieves 10x speedup over a software-only implementation in terms of image sample-based pose update latency while consuming just 2.2 mW power. In a FPGA implementation, using the EuRoC VIO dataset (VGA 30fps images and 100Hz IMU) the accelerator design achieves pose estimation accuracy (loop closure error) comparable to a software based VIO implementation.
Dipan Mandal, Srivatsava Jandhyala, Om Ji Omer, Gurpreet S. Kalsi, Biji George, Gopi Neela, Santhosh Kumar Rethinagiri, Sreenivas Subramoney, Lance Hacking, Jim Radford, Eagle Jones, Belliappa Kuttanna, Hong Wang 0003
DATE7
2016 Energy minimization at all layers of the data center: The ParaDIME project
Oscar Palomar, Santhosh Kumar Rethinagiri, Gulay Yalcin, J. Rubén Titos Gil, Pablo Prieto, Emma Torrella, Osman S. Unsal, Adrián Cristal, Pascal Felber, Anita Sobe, Yaroslav Hayduk, Mascha Kurpicz, Christof Fetzer, Thomas Knauth, Malte Schneegaß, Jens Struckmeier, Dragomir Milojevic
DATE2
2016 Exploring Energy Reduction in Future Technology Nodes via Voltage Scaling with Application to 10nm
abstract
DI-fusion, le Dépôt institutionnel numérique de l'ULB, est l'outil de référencementde la production scientifique de l'ULB.L'interface de recherche DI-fusion permet de consulter les publications des chercheurs de l'ULB et les thèses qui y ont été défendues.
Gulay Yalcin, Santhosh Kumar Rethinagiri, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Dragomir Milojevic
PDP2
2015 Heterogeneous Platform to Accelerate Compute Intensive Applications
abstract
Nowadays image processing applications are widely used in various industries such as traffic, safety, medical engineering, etc. In this paper, we propose a power and energy efficient heterogeneous platform to accelerate image processing applications. To achieve this efficiency, we propose a novel hybrid platform which consists of a Xilinx Zynq (ARM+FPGA) and an NVidias Jetson TK1 (ARM+GPU) coupled with PCIe card. In applications such face recognition, we optimized major tasks in detection and recognition in order to achieve a speedup of 69× when compared to sequential execution on the ARM core, 4.8× against Zynq platform (ARM+FPGA), 3.2× against NVidia platform (ARM+GPU) and 40% more energy efficient against sequential execution.
Santhosh Kumar Rethinagiri, Oscar Palomar, Javier Arias Moreno, Osman S. Unsal, Adrián Cristal
FCCM1
2015 Trigeneous Platforms for Energy Efficient Computing of HPC Applications
abstract
In this paper, we present two novel real-time heterogeneous platforms with three kinds of devices (CPU, GPU, FPGA), i.e. trigeneous platforms, for efficiently accelerating computation intensive applications in both the high-performance computing and the embedded system domains. In the high-performance computing domain, the entire platform is implemented on a workstation which consists of an Intel Xeon E5 processor, a Nvidia Tesla GPU and a Xilinx Virtex 7 FPGA. The second platform is built for achieving high-performance in the real-time embedded system domain. For this platform, we use a Xilinx Zynq and Nvidia Jetson TK1 board. In these platforms, the communication is performed using PCIe Gen3 and PCIe Gen2 cards respectively. We conducted experiments using 5 real-time and high throughput computation-data-intensive applications, namely cone beam computed tomography, face recognition, HEVC UHD decoding, number plate recognition and motion tracking. All the applications are mapped to the devices of the proposed trigeneous platforms, based on the energy efficiency of the different tasks on each device but also minimizing data transfers and maximizing parallelism. With this trigeneous platform, we are able to achieve an average speed-up of 21x compared to a CPU-GPU platform, 24x compared to a CPU-FPGA platform and 70x compared to Quad-core CPU alone execution. The proposed trigeneous platforms save 43%, 56% and 64% of the energy when compared to CPU-GPU, CPU-FPGA and quad-core CPU platforms respectively. Furthermore, we also implemented these applications by using a single programming language (OpenCL) on the trigeneous platforms and achieved 6x of speed-up on average over the quad-core setup.
Santhosh Kumar Rethinagiri, Oscar Palomar, Javier Arias Moreno, Osman S. Unsal, Adrián Cristal
HiPC1
2015 VPM: Virtual power meter tool for low-power many-core/heterogeneous data center prototypes
abstract
Power and energy consumption of data centers are steadily increasing and the work performed by the data centers is not proportional to the power dissipated, where every μA is a revenue for the entity. On the one hand, the hardware community is proposing various methodologies to address this issue such as low-power processors, heterogeneity, etc. to reduce the power of the servers. On the other hand, the software community proposes mechanisms such as virtual machines (VMs), work-load scheduling, etc. to increase the utilization of the processor. In order to properly evaluate the impact of these mechanisms, we need an accurate power monitoring and estimation tool at the hardware host level, the VM level and the system-level. This paper proposes a novel power monitoring middleware on a low-power platform at the node level (ARM Big.LITTLE) and an estimation methodology by using a simulator for future data center prototypes at any given level of virtualization. First, we built an instrumentation framework to measure the power based on hardware counter activities and with respect to current fluctuation. This allows us to build power models for the corresponding platforms, which are fed into the middleware to estimate power on the fly. Furthermore, we used the same framework for future low-power processors such as ARM Cortex-A57 and -A53 based platforms, which are integrated into the architectural simulator by providing an API to estimate power with the power model. Second, we present a machine learning-based energy efficient scheduling of the VMs that leverages VPM. The results obtained with the power monitoring middleware differ less than 2% from real board measurements and 5% when using the simulation environment regardless of the number of virtual machines used. Furthermore, we reduced 40% of energy consumption on average when compared to default scheduling of the KVM hypervisor.
Santhosh Kumar Rethinagiri, Oscar Palomar, Javier Arias Moreno, Osman S. Unsal, Adrián Cristal
ICCD1
2014 ParaDIME: Parallel Distributed Infrastructure for Minimization of Energy
abstract
Dramatic environmental and economic impact of the ever increasing power and energy consumption of modern computing devices in data centers is now a critical challenge. On one hand, designers use technology scaling as one of the methods to face the phenomenon called dark silicon (only segments of a chip function concurrently due to power restrictions). On the other hand, designers use extreme-scale systems such as teradevices to meet the performance needs of their applications which in turn increases the power consumption of the platform. In order to overcome these challenges, we need novel computing paradigms that address energy efficiency. One of the promising solutions is to incorporate parallel distributed methodologies at different abstraction levels. The FP7 project ParaDIME focuses on this objective to provide different distributed methodologies (software-hardware techniques) at different abstraction levels to attack the power-wall problem. In particular, the ParaDIME framework will utilize: circuit and architecture operation below safe voltage limits for drastic energy savings, specialized energy-aware computing accelerators, heterogeneous computing, energy-aware runtime, approximate computing and power-aware message passing. The major outcome of the project will be a processor architecture for a heterogeneous distributed system that utilizes future device characteristics for drastic energy savings. Wherever possible, ParaDIME will adopt multidisciplinary techniques, such as hardware support for message passing, runtime energy optimization utilizing new hardware energy performance counters, use of accelerators for error recovery from sub-safe voltage operation, and approximate computing through annotated code. Furthermore, we will establish and investigate the theoretical limits of energy savings at the device, circuit, architecture, runtime and programming model levels of the computing stack, as well as quantify the actual energy savings achieved by the ParaDIME approach for the complete computing stack with the real environment.
Santhosh Kumar Rethinagiri, Oscar Palomar, Anita Sobe, Thomas Knauth, Wojciech M. Barczynski, Gulay Yalcin, Yaroslav Hayduk, Adrián Cristal, Osman S. Unsal, Pascal Felber, Christof Fetzer, Julien Ryckaert, Gina Alioto
DSD1
2014 APMC: advanced pattern based memory controller (abstract only)
abstract
In this paper, we present APMC, the Advanced Pattern based Memory Controller, that uses descriptors to support both regular and irregular memory access patterns without using a master core. It keeps pattern descriptors in memory and prefetches the complex 1D/2D/3D data structure into its special scratchpad memory. Support for irregular Memory accesses are arranged in the pattern descriptors at program-time and APMC manages multiple patterns at run-time to reduce access latency. The proposed APMC system reduces the limitations faced by processors/accelerators due to irregular memory access patterns and low memory bandwidth. It gathers multiple memory read/write requests and maximizes the reuse of opened SDRAM banks to decrease the overhead of opening and closing rows. APMC manages data movement between main memory and the specialized scratchpad memory; data present in the specialized scratchpad is reused and/or updated when accessed by several patterns. The system is implemented and tested on a Xilinx ML505 FPGA board. The performance of the system is compared with a processor with a high performance memory controller. The results show that the APMC system transfers regular and irregular datasets up to 20.4x and 3.4x faster respectively than the baseline system. When compared to the baseline system, APMC consumes 17% less hardware resources, 32% less on-chip power and achieves between 3.5x to 52x and 1.4x to 2.9x of speedup for regular and irregular applications respectively. The APMC core consumes 50% less hardware resources than the baseline system's memory controller. In this paper, we present APMC, the Advanced Pattern based Memory Controller, an intelligent memory controller that uses descriptors to supports both regular and irregular memory access patterns. support of the master core. It keeps pattern descriptors in memory and prefetches the complex data structure into its special scratchpad memory. Memory accesses are arranged in the pattern descriptors at program-time and APMC manages multiple patterns at run-time to reduce access latency. The proposed APMC system reduces the limitations faced by processors/accelerators due to irregular memory access patterns and low memory bandwidth. The system is implemented and tested on a Xilinx ML505 FPGA board. The performance of the system is compared with a processor with a high performance memory controller. The results show that the APMC system transfers regular and irregular datasets up to 20.4x and 3.4x faster respectively than the baseline system. When compared to the baseline system, APMC consumes 17% less hardware resources, 32% less on-chip power and achieves between 3.5x to 52x and 1.4x to 2.9x of speedup for regular and irregular applications respectively. The APMC core consumes 50% less hardware resources than the baseline system's memory controller.memory accesses.
Tassadaq Hussain, Oscar Palomar, Osman S. Unsal, Adrián Cristal, Eduard Ayguadé, Mateo Valero, Santhosh Kumar Rethinagiri
FPGA7
2014 Power estimation tool for system on programmable chip based platforms (abstract only)
abstract
The ever increasing complexity of the applications result in the development of power hungry processors. There is a scarcity of standalone tools that have a good trade off between estimation speed and accuracy to estimate power/energy at an earlier phase of design flow. There are very few tools that addresses the design space exploration issue based on power and energy. In this paper, we propose a virtual platform based standalone power and energy estimation tool for System-on-Programmable Chip (SoPC) embedded platforms, which is independent of in-house tools. There are two steps involved in this tool development. The first step is power model generation. For the power model development, we used functional parameters to set up generic power models for the different parts of the system. This is a onetime activity. In the second step, a simulation based virtual platform framework is developed to evaluate accurately the activities used in the related power models developed in the first step. The combination of the two steps lead to a hybrid power estimation, which gives a better trade-off between accuracy and speed. The proposed tool has several benefits: it considers the power consumption of the embedded system in its entirety and leads to accurate estimates without a costly and complex material. The proposed tool is also scalable for exploring complex embedded multi-core architectures.
Santhosh Kumar Rethinagiri, Oscar Palomar, Adrián Cristal, Osman S. Unsal
FPGA1
2012 An efficient power estimation methodology for complex RISC processor-based platforms
abstract
In this contribution, we propose an efficient power estimation methodology for complex RISC processor-based platforms. In this methodology, the Functional Level Power Analysis (FLPA) is used to set up generic power models for the different parts of the system. Then, a simulation framework based on virtual platform is developed to evaluate accurately the activities used in the related power models. The combination of the two parts above leads to a heterogeneous power estimation that gives a better trade-off between accuracy and speed. The usefulness and effectiveness of our proposed methodology is validated through ARM9 and ARM CortexA8 processor designed respectively around the OMAP5912 and OMAP3530 boards. This efficiency and the accuracy of our proposed methodology is evaluated by using a variety of basic programs to complete media benchmarks. Estimated power values are compared to real board measurements for the both ARM940T and ARM CortexA8 architectures. Our obtained power estimation results provide less than 3% of error for ARM940T processor, 3.5% for ARM CortexA8 processor-based system and 1x faster compared to the state-of-the-art power estimation tools.
Santhosh Kumar Rethinagiri, Rabie Ben Atitallah, Jean-Luc Dekeyser, Eric Senn, Smaïl Niar
ACM Great Lakes Symposium on VLSI1
2011 Hybrid system level power consumption estimation for FPGA-based MPSoC
abstract
This paper proposes an efficient Hybrid System Level (HSL) power estimation methodology for FPGA-based MPSoC. Within this methodology, the Functional Level Power Analysis (FLPA) is extended to set up generic power models for the different parts of the system. Then, a simulation framework is developed at the transactional level to evaluate accurately the activities used in the related power models. The combination of the above two parts lead to a hybrid power estimation that gives a better trade-off between accuracy and speed. The proposed methodology has several benefits: it considers the power consumption of the embedded system in its entirety and leads to accurate estimates without a costly and complex material. The proposed methodology is also scalable for exploring complex embedded architectures. The usefulness and effectiveness of our HSL methodology is validated through a typical mono-processor and multiprocessor embedded system designed around the Xilinx Virtex II Pro FPGA board. Our experiments performed on an explicit embedded platform show that the obtained power estimation results are less than 1.2% of error when compared to the real board measurements and faster compared to other power estimation tools.
Santhosh Kumar Rethinagiri, Rabie Ben Atitallah, Smaïl Niar, Eric Senn, Jean-Luc Dekeyser
ICCD1