EDBT 2026 Demo / reviewers in the wild / expert
Gopalakrishnan Lakshminarayanan
dblp:97/1546 · also G. Lakshminarayanan
· DBLP profile ↗
9ranked-venue papers
1as first author
5since 2021 · last 2023
0000-0003-2753-9455ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SaHNoC: an optimal energy efficient hybrid networks-on-chip architecture
Aravindhan Alagarsamy, Sundarakannan Mahilmaran, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko |
J. Supercomput. | 3 |
| 2022 | FRDS: An efficient unique on-Chip interconnection network architecture
Aravindhan Alagarsamy, Sundarakannan Mahilmaran, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko |
Integr. | 3 |
| 2021 | Kinematic adaptive frequency sampling combined spatio temporal features for snow monitoring in aerospace applications
Parameshwaran Ramalingam, Gopalakrishnan Lakshminarayanan, Manikandan Ramachandran, Rizwan Patan |
Expert Syst. Appl. | 2 |
| 2021 | A low latency modular-level deeply integrated MFCC feature extraction architecture for speech recognition
Bibin Sam Paul S, Antony Xavier Glittas, Gopalakrishnan Lakshminarayanan |
Integr. | 3 |
| 2021 | A Real-Time Architecture for Pruning the Effectual Computations in Deep Neural NetworksabstractIntegrating Deep Neural Networks (DNNs) into the Internet of Thing (IoT) devices could result in the emergence of complex sensing and recognition tasks that support a new era of human interactions with surrounding environments. However, DNNs are power-hungry, performing billions of computations in terms of one inference. Spatial DNN accelerators in principle can support computation-pruning techniques compared to other common architectures such as systolic arrays. Energy-efficient DNN accelerators skip bit-wise or word-wise sparsity in the input feature maps (ifmaps) and filter weights which means ineffectual computations are skipped. However, there is still room for pruning the effectual computations without reducing the accuracy of DNNs. In this paper, we propose a novel real-time architecture and dataflow by decomposing multiplications down to the bit level and pruning identical computations in spatial designs while running benchmark networks. The proposed architecture prunes identical computations by identifying identical bit values available in both ifmaps and filter weights without changing the accuracy of benchmark networks. When compared to the reference design, our proposed design achieves an average per layer speedup of$\times 1.4$and an energy efficiency of$\times 1.21$per inference while maintaining the accuracy of benchmark networks. Mohammadreza Asadikouhanjani, Hao Zhang 0041, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2019 | Low-complex processing element architecture for successive cancellation decoder
Geethu Sathees Babu, Lakshmi Renuka Madala, Gopalakrishnan Lakshminarayanan, Mathini Sellathurai |
Integr. | 3 |
| 2016 | A Normal I/O Order Radix-2 FFT Architecture to Process Twin Data Streams for MIMOabstractNowadays, many applications require simultaneous computation of multiple independent fast Fourier transform (FFT) operations with their outputs in natural order. Therefore, this brief presents a novel pipelined FFT processor for the FFT computation of two independent data streams. The proposed architecture is based on the multipath delay commutator FFT architecture. It has an N/2-point decimation in time FFT and an N/2-point decimation in frequency FFT to process the odd and even samples of two data streams separately. The main feature of the architecture is that the bit reversal operation is performed by the architecture itself, so the outputs are generated in normal order without any dedicated bit reversal circuit. The bit reversal operation is performed by the shift registers in the FFT architecture by interleaving the data. Therefore, the proposed architecture requires a lower number of registers and has high throughput. Antony Xavier Glittas, Mathini Sellathurai, Gopalakrishnan Lakshminarayanan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Dynamic partial reconfigurable FFT/IFFT pruning for OFDM based Cognitive radioabstractCognitive Radio is an application in which Spectrum utilization can be improved by allowing secondary users to use the spectrum when it is not used by licensed primary users. An adaptive OFDM system for Cognitive radio has the ability to nullify unnecessary individual carriers and avoid interference to licensed primary users. A Fast Fourier Transform (FFT) block forms the core of OFDM design. But, the zero valued inputs outnumber the non-zero valued inputs in the FFT block making the standard FFT algorithms computationally inefficient due to wasted operation on zero values. To overcome this problem, several pruning algorithms have been developed. But many of them are architecturally inefficient for FPGA implementation due to complexity of the overhead operations. Moreover, these algorithms are not suitable for applications like Cognitive radio which has zero inputs in arbitrary distributions making hardware implementation to be complex. This paper presents a novel and efficient dynamically partial reconfigurable (DPR) Transform Decomposition (TD) FFT and Radix 2 based IFFT pruning for OFDM based Cognitive Radio on FPGA. Tested FPGA results on XC2VP30 for the DPR method show the configuration time improvement, good area and power efficiency. C. Vennila, Kumar Palaniappan CT, Kodati Vamsi Krishna, Gopalakrishnan Lakshminarayanan, Seok-Bum Ko |
ISCAS | 4 |
| 2005 | Optimization Techniques for FPGA-Based Wave-Pipelined DSP BlocksabstractIn this paper, techniques for efficient implementation of field-programmable gate-array (FPGA)-based wave-pipelined (WP) multipliers, accumulators, and filters are presented. A comparison of the performance of WP and pipelined systems has been made. Major contributions of this paper are development of an on-chip clock generation scheme which permits finer tuning of the frequency, a synthesis technique that reduces the area and latency by 25%, a placement utility that results in 10%-40% increase in speed and proposal of an interleaving scheme for filters that reduces the number of multipliers required by 50%. WP multipliers of size 2 /spl times/ 6 and the filters using them are found to be 11% faster and require lower power than those using pipelined multipliers. Filters with higher order WP multipliers also operate with lower power at the cost of speed. The delay-register products of such filters are found to be about 60% lower than those using the pipelined multipliers. The paper also outlines applications of these techniques for the Spartan II FPGAs and a self-tuning scheme for optimizing the speed. Gopalakrishnan Lakshminarayanan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |