EDBT 2026 Demo / reviewers in the wild / expert
Subrahmanyam Mula
dblp:173/5456
· DBLP profile ↗
14ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-5092-0524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient VLSI Architecture for Finding Maximum Value and Its Corresponding Index
Jayarani Medakal Anirudhan, C. V. Tirumala Rao, Subrahmanyam Mula |
ISCAS | 3 |
| 2026 | An Efficient VLSI Architecture for Hammerstein-Type Spline Adaptive FiltersabstractThe Hammerstein spline adaptive filter (HSAF) is a class of nonlinear adaptive filters (NAFs), known for its flexible nonlinear modeling and low complexity in applications, such as self-interference cancellation in wireless communications. This brief proposes a delayed dual-weight update reformulation for the HSAF and its efficient high-throughput and low-power architecture. We also propose hardware-efficient techniques for mapping spline interpolation and updating the spline control points in HSAF. The proposed delayed HSAF (DHSAF) architecture is synthesized using Cadence Genus in 45-nm CMOS technology. Synthesis results show that the proposed DHSAF achieves significantly higher throughput compared to the basic HSAF, with only minimal area and power overhead. Furthermore, the proposed DHSAF outperforms the state-of-the-art RFF-KLMS architecture in terms of both area and power efficiency. Pavan Kumar Ganjimala, Subrahmanyam Mula |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2026 | A Hardware-Efficient QR Algorithm and Its VLSI Architecture for Eigenvalue Decomposition of Symmetric MatricesabstractThis article proposes a novel hardware-efficient QR algorithm and its VLSI architecture to compute the eigenvalue decomposition (EVD) of symmetric matrices. By replacing the computationally expensive matrix multiplication with matrix transposition, the proposed architecture achieves a significant reduction in latency as well as area compared to the conventional QR architecture. We prove the convergence of the proposed algorithm through mathematical analysis and study its convergence rate through numerical simulations. We implemented both conventional and proposed QR architectures on both field-programmable gate array (FPGA) and application specific integrated circuit (ASIC) platforms. Implementation results on the ZCU104 Ultrascale+ FPGA demonstrate a 51% reduction in computation time and a 50% decrease in DSP usage with the proposed algorithm. Furthermore, ASIC synthesis results also demonstrate a significant reduction in computation time compared to the conventional algorithm. Vishnu P. S, Jobin Francis 0001, Subrahmanyam Mula |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | An Area-Efficient VLSI Architecture for Parallel Jacobi-based Eigenvalue Decomposition with Inherent Eigenvalue SortingabstractThe Jacobi algorithm is widely used for eigen-value decomposition (EVD), and its parallel implementation is preferred in real-time applications which demand low latency. However, the parallel implementations are computationally expensive in terms of area overhead. In this paper, we propose an area-efficient implementation of the parallel Jacobi algorithm without compromising the latency and accuracy. Additionally, the proposed architecture is inherently self-sorting, thereby avoiding the need for a dedicated hardware to sort the eigenvalues. ASIC synthesis results of the designed architecture show that the proposed method achieves a 23.4% reduction in area over the state-of-the-art implementation, with no significant impact on latency. FPGA implementation results also demonstrate optimized hardware utilization. Sandra Accamma George, Liz Maria George, Vishnu P. S, Subrahmanyam Mula |
ISCAS | 4 |
| 2024 | A proportionate type block-oriented functional link adaptive filter for sparse nonlinear systemsabstractNonlinear adaptive filters (NAFs) are used in applications like self-interference cancellation to identify nonlinear systems but at the expense of high computational complexity. When the system response is sparse, NAFs are further combined with proportionate-type algorithms to accelerate system identification which increases the computational complexity further. To alleviate the NAF complexity problem, a block-oriented functional link adaptive filter (BO-FLAF) algorithm, which has lower complexity and filter order than the FLAF was recently proposed. Since proportionate adaptation complexity is directly proportional to the filter order, a proportionate BO-FLAF is more attractive from a computation standpoint than the existing proportionate FLAF (PFLAF). Therefore, we develop a novel proportionate Hammerstein BO-FLAF (PHBO-FLAF) in this paper. Through theoretical complexity analysis we prove that the proposed PHBO trigonometric FLAF (PHBO-TFLAF) has 53.2% lesser multiplications and 7 times less logarithmic operations than the PTFLAF. We also show that the PHBO-TFLAF has faster convergence than the HBO-TFLAF and PTFLAF and models memoryless systems much better than the PTFLAF. Pavan Kumar Ganjimala, Subrahmanyam Mula |
ISCAS | 2 |
| 2024 | Performance Analysis of Hammerstein Block-Oriented Functional Link Adaptive FiltersabstractNonlinear adaptive filters (NAFs) exhibit superior modeling capabilities compared to conventional linear adaptive filters, especially in practical applications involving nonlinear input-output relationships. The functional link adaptive filter (FLAF) is an NAF that uses nonlinear functional expansions to achieve nonlinear modelling, however, at the expense of high computational complexity. In response, a low-complexity Hammerstein-type block-oriented functional link adaptive filter (HBO-FLAF) was recently developed, which requires less computation than that of the traditional FLAF. To shed more light on its behaviour and design, we provide a steady-state theoretical analysis of the HBO-FLAF in this paper. We derive the conditions for steady-state mean and mean square convergence of the weight update equations, specifically, an upper bound on the step-size parameter, an expression for the steady-state excess mean square error (EMSE) and a lower bound on the steady-state EMSE of the HBO-FLAF. Numerical simulation results show a close relation with the derived results, thus validating the theoretical analysis. Pavan Kumar Ganjimala, Vinay Chakravarthi Gogineni, Subrahmanyam Mula |
IEEE Signal Process. Lett. | 3 |
| 2023 | Algorithm and Architecture Design of Random Fourier Features-Based Kernel Adaptive FiltersabstractNumerous real-life systems exhibit complex nonlinear input-output relationships. Kernel adaptive filters, a popular class of nonlinear adaptive filters, can efficiently model these nonlinear input-output relationships. Their growing network structure, however, poses considerable challenges in terms of their hardware implementation, making them inefficient for real-time applications. Random Fourier features (RFF) facilitate the development of kernel adaptive filters with a fixed network structure. For the first time, this paper attempts to implement the RFF-based kernel least mean square (RFF-KLMS) algorithm on hardware. To this end, we propose several reformulations of the feature functions (FFs) that are computationally expensive in their native form so that they can be implemented in real-time VLSI. Specifically, we reformulate inner product evaluation, cosine, and exponential functions that appear in the implementation of FFs. With these reformulations, the proposed delayed RFF-KLMS (DRFF-KLMS) is then synthesized using 45-nm CMOS technology with 16-bit fixed-point representations. According to the synthesis results, pipelined DRFF-KLMS architectures require minimal hardware increase over the state-of-the-art conventional delayed LMS architecture while significantly improving estimation performance for the nonlinear model. Our results suggest that the cosine feature function-based DRFF-KLMS is appropriate for applications requiring high accuracy, whereas the exponential function-based DRFF-KLMS may be well suited for resource-constrained applications. Vinay Chakravarthi Gogineni, Ramesh Sambangi, Daney Alex, Subrahmanyam Mula, Stefan Werner 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | High performance VLSI architecture for the modified SORT-N algorithmabstractFast running sorting of streaming input data samples is very important in many applications such as order statistics, nonlinear filtering, MMax selective-tap adaptive filtering etc. This paper proposes a high performance VLSI architecture for the modified SORT-N algorithm for fast running sorting. Through analysis and also through synthesis results, we show that the critical path of the proposed architecture is almost independent of the sorting order N. ASIC synthesis results of the designed architecture shows that the proposed architecture has double the performance for N=1024 along with a reduction in area and power metrics compared to state-of-the-art architecture reported in literature and thus, it is potentially useful in real-time applications which have stringent throughput requirements. Pavan Kumar Ganjimala, Subrahmanyam Mula |
ISCAS | 2 |
| 2022 | Novel VLSI Architecture for Fractional-Order Correntropy Adaptive Filtering AlgorithmabstractConventional adaptive filters, which assume Gaussian distribution for signal and noise, exhibit significant performance degradation when operating in non-Gaussian environments. Recently proposed fractional-order adaptive filters (FoAFs) address this concern by assuming that the signal and noise are symmetric$\alpha $-stable random processes. However, the literature does not include any VLSI architectures for these algorithms. Toward that end, this article develops hardware-efficient architecture for fractional-order correntropy adaptive filter (FoCAF). We first reformulate the FoCAF for its efficient real-time VLSI implementation and then demonstrate that these reformulations cause negligible performance degradation under the 16-bit fixed-point implementation. Using this reformulated algorithm, we design an FoCAF architecture. Furthermore, we analyze the critical path of the design to select the appropriate level of pipelining based on the sampling rate of the application. According to the critical-path analysis, the FoCAF design is pipelined using retiming techniques to obtain delayed FoCAF (DFoCAF), which is then synthesized using$\mathbf {45}$-nm CMOS technology. Synthesis results reveal that DFoCAF architecture requires a minimal increase in hardware over the prominent least mean square (LMS) filter architecture and achieves a significant increase in the performance in symmetric$\alpha $-stable environments where LMS fails to converge. Daney Alex, Vinay Chakravarthi Gogineni, Subrahmanyam Mula, Stefan Werner 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Robust Proportionate Adaptive Filter Architectures Under Impulsive NoiseabstractThis brief proposes robust adaptive filtering algorithms and their VLSI architectures for sparse system identification under impulsive noise. Several robust algorithms are derived by combining error nonlinear adaptive filtering algorithms with proportionate adaptation. We make a comparative study of the derived algorithms and their VLSI architectures in terms of convergence rate and hardware complexity to show that the hardware overhead is negligible for the achieved improvement in robustness. Subrahmanyam Mula, Vinay Chakravarthi Gogineni, Anindya Sundar Dhar |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Algorithm and VLSI Architecture Design of Proportionate-Type LMS Adaptive Filters for Sparse System Identification
Subrahmanyam Mula, Vinay Chakravarthi Gogineni, Anindya Sundar Dhar |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | A novel framework for compressed sensing based scalable video coding
Kota Naga Srinivasarao Batta, Vinay Chakravarthi Gogineni, Subrahmanyam Mula, Indrajit Chakrabarti |
Signal Process. Image Commun. | 3 |
| 2017 | Linear Detection of a Weak Signal in Additive Cauchy NoiseabstractThe detection of a weak signal in additive Cauchy noise is of great importance in many applications. A locally optimum detector (LOD) exists for such a scenario; however, it is non-linear in nature. In general, implementation of non-linear detectors is difficult in practice, and linear detectors with good properties, such as high asymptotic relative efficiency (ARE) with respect to the LOD, are often desirable. In this paper, we propose a linear detector for a weak signal in additive Cauchy noise. The proposed test statistic is a linear combination of order statistics. For the special case of a constant signal in additive Cauchy noise, we prove the asymptotic normality of the trimmed linear detector, and show that the ARE of the trimmed linear detector with respect to the LOD is unity. Extensive simulation results are provided to demonstrate that the loss in the performance of the linear detector is very small compared with the non-linear LOD. We also discuss the hardware complexities of the LOD and the linear detector, and demonstrate the advantages of the linear detector over the LOD, in terms of hardware implementation. Siva Ram Krishna Vadali, Priyadip Ray, Subrahmanyam Mula, Pramod K. Varshney |
IEEE Trans. Commun. | 3 |
| 2017 | Algorithm and Architecture Design of Adaptive Filters With Error NonlinearitiesabstractThis paper presents a framework based on the logarithmic number system to implement adaptive filters with error nonlinearities in hardware. The framework is demonstrated through pipelined implementations of two recently proposed adaptive filtering algorithms based on logarithmic cost, namely, least mean logarithmic square (LMLS) and least logarithmic absolute difference (LLAD). To the best of our knowledge, the proposed architectures are the first attempts to implement both LMLS and LLAD algorithms in hardware. We derive error computing algorithms to realize the nonlinear error functions for LMLS and LLAD and map them onto hardware. We also propose a novel variable-α scheme to enhance the original LMLS algorithm and prove its robustness and suitability for VLSI implementations in practical applications. Detailed bit width and error analysis are carried out for the proposed VLSI fixed point implementations. Postlayout implementation results show that with an additional multiplier over conventional least mean square (LMS), 7-dB improvement in steady-state mean square deviation performance can be achieved and with the proposed variable-α scheme, 12-dB improvement can be achieved without compromising the convergence. We will show that LMLS can potentially replace LMS in practical applications, by demonstrating a proof-of-concept by extending the framework to transform domain adaptive filters. Subrahmanyam Mula, Vinay Chakravarthi Gogineni, Anindya Sundar Dhar |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |