EDBT 2026 Demo / reviewers in the wild / expert
Pramod Kumar Meher
dblp:97/6785
· DBLP profile ↗
88ranked-venue papers
27as first author
7since 2021 · last 2026
0000-0003-0992-1159ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 61 · 23 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-authorArtificial intelligence and machine learning · 7Human-computer interaction and ubiquitous computing · 6Applied, interdisciplinary, general and emerging computing · 6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Precision-Specific Efficient Designs and FPGA Implementation of Sigmoid for Machine LearningabstractThis paper presents a low-complexity design for generating the sigmoid function based on a novel piecewise linear approximation. We have proposed an iterative algorithm to break the whole argument space into the minimum number of intervals, for any given desired accuracy. The LUT address generator of the proposed design does not require any comparators since the breakpoints generated by the proposed algorithm contain only a few non-zero bits. Similarly, the slopes of the line segments are represented by a few non-zero bits, such that the multiplication with slope is realized using hardwired shift operations or a few shift-add operations without using any LUT to store the slope values and without using any conventional multiplier. For the computation of the sigmoid function with fractional accuracy of 7 bits, the proposed design involves only 6 line segments, which requires 2 adders and an LUT to store 6 intercept words. When implemented on an Xilinx Kintex-7 FPGA device, it is found to be 8 times faster than the fastest of the existing designs for the same accuracy and consumes less than half the resources used by the latter. Moreover, the proposed design for 9-bit fractional accuracy is approximately twice as fast and more resource-efficient than a recently reported design offering the same level of accuracy. Pramod Kumar Meher, Supriya Aggarwal |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Efficient Design and Implementation of Scale-Free CORDIC With Mutually Exclusive Micro-RotationsabstractIn this paper, a new approach to the design of a micro-rotation set for scale-free CORDIC is proposed. The sine and cosine functions of all the micro-rotation angles are realized by a simple shift or a shift-add operations which significantly reduces the hardware complexity. Besides, the micro-rotation set (except the first one) is designed to form mutually exclusive pairs. As a result of mutually exclusive micro-rotations, it is possible to reduce the required number of iterations to almost half for a given precision. Apart from that, the latency, as well as, the hardware complexity are also significantly reduced. A 9-bit fractional accuracy is obtained with just 5 iterations as against 13 iterations required by the conventional CORDIC. Suitable threshold angles are proposed to decide, using low-complexity comparators, whether a micro-rotation should be executed in a given iteration or can be skipped. The proposed circuits to determine the rotation conditions for different iterations involve either 2-bit or 3-bit comparators. The CORDIC circuit based on the proposed set of micro-rotations is shown to converge for any given angle of rotation. Furthermore, the proposed design involves significantly less logic, computation time, and latency than the best of the scale-free CORDIC circuits. When implemented on Xilinx FPGA (Field Programmable Gate Arrays), it requires 20% less area, offers higher operating frequency, and saves close to 19% power and 20% energy per computation over the latter. Pramod Kumar Meher, Supriya Aggarwal |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | A Flexible DA-Based Architecture for Computation of Inner Product of Variable VectorsabstractThe computation of inner products of any given pair of vectors is an indispensable requirement in several applications including artificial intelligence (AI), machine learning (ML), signal processing, image processing, communication, and many others. The throughput requirement of inner product computation varies widely for different applications. Moreover, the throughput of computation must match the requirements of the applications. It is therefore important to design flexible hardware for inner product computation that produces the desired throughput. Distributed arithmetic (DA) is a well-known approach for efficient inner product computation. This article presents an efficient DA-based architecture for computing the inner product of variable vectors, which could be tailored according to the throughput requirement of any given application and reused for different inner product lengths. The proposed designs could also be deployed to achieve a trade-off between throughput and area/energy consumption. In this article, we have used modified Booth encoding (MBE) to reduce the number of partial products and proposed a novel carry-save accumulator (CSA) for shortening the critical path delay. The proposed designs are synthesized by Cadence Genus using GPDK 90-nm technology library and place-and-route using Cadence Innovus for different inner product lengths and word lengths. As found from the postlayout synthesis results, the proposed designs offer savings of nearly 30% and 29% EPC and ADP over the bit-serial DA-based design on average for word lengths 8 and 16 and inner product lengths 8, 16, and 32, respectively. Anil Kali, Samrat L. Sabat, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | A Novel DA-Based Parallel Architecture for Inner-Product of Variable VectorsabstractComputation of the inner products is frequently used in machine learning (ML) algorithms apart from signal processing and communication applications. Distributed arithmetic (DA) has been frequently employed for area-time efficient inner-product implementations. In conventional DA-based architectures, one of the vectors is constant and known a priori. Hence, the traditional DA architectures are not suitable when both vectors are variable. However, computing the inner product of a pair of variable vectors is frequently used for matrix multiplication of various forms and convolutional neural networks. In this paper, we present a novel DA-based architecture for computing the inner product of variable vectors. To derive the proposed architecture, the inner product of any given length is decomposed into a set of short-length inner products, such that the inner product could be computed by successive accumulation of the results of short-length inner products. We have designed a DA-based architecture for the computation of the short-length inner-product of variable vectors and used that in successive clock cycles to compute the whole inner-product by successive accumulation. The post-layout synthesis results using Cadence Innovus with a GPDK 90nm technology library show that the proposed DA-based parallel architecture offers significant advantages in area-delay product and energy consumption over the bit-serial DA architecture. Anil Kali, Samrat L. Sabat, Pramod Kumar Meher |
ISCAS | 3 |
| 2024 | Low Complexity Design of Logistic Distance Metric Adaptive Filter for Impulsive Noise EnvironmentsabstractIn many practical scenarios, non-Gaussian noise contaminates the desired signal and introduces outliers. The recently proposed logistic distance metric adaptive filter (LDMAF) outperforms the existing algorithms and provides better performance in the presence of such outliers. There is a need for efficient hardware architecture for the implementation of LDMAF. This article proposes an efficient VLSI architecture of LDMAF. The implementation of error-gradient function of LDMAF puts significant implementation problem in terms of delay and cost. We introduce here an efficient tangent-based piecewise linear (TPL) approximation algorithm for implementing the corresponding architecture. The proposed approach improves the power, performance, and area (PPA) metrics over state-of-the-art implementations of other robust algorithms while meeting system performance within an acceptable deviation. Shouharda Ghosh, Pramod Kumar Meher, Dwaipayan Ray, Nithin V. George |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Low-Complexity Distributed Arithmetic-Based Architecture for Inner-Product of Variable VectorsabstractDistributed arithmetic (DA) is generally used for area-time efficient implementation of inner products, where one of the vectors is fixed and known a priori. Therefore, the conventional DA architectures cannot be used when both vectors are variable. This article proposes a novel architecture for computing inner products of variable vectors, where one of the vectors is encoded using the radix-4 modified Booth technique to reduce the logic complexity. The proposed structure for inner-product computation consists of two sections. The first Section of the architecture performs a carry-save reduction of the partial-inner-products of the same weight to two words. During every successive clock cycle, it reduces such partial-inner-products of different weights in the order of the lowest to the highest weight. In the second Section of the architecture, the pair of reduced words produced by the first Section are shift accumulated. The area, delay, and power saving are achieved by reducing the overall critical path of the structure as well as the logic complexity in both sections. The proposed architecture is synthesized by Cadence Genus using TSMC 90-nm technology library and place-and-route using Cadence Innovus for different inner-product lengths and word lengths. The postlayout synthesis results show that the proposed DA-based architecture offers significant advantages in area-delay product (ADP) and energy per computation (EPC) over the radix-4 Booth multiplier-accumulator-based architectures. Anil Kali, Samrat L. Sabat, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | An Efficient Scaling-Free Folded Hyperbolic CORDIC Design Using a Novel Low-Complexity Power-of-2 Taylor Series ApproximationabstractHyperbolic trigonometric functions are widely used in several engineering and scientific applications, including digital signal processing (DSP), communication systems, and many others. In this article, we propose a scaling-free hyperbolic coordinate rotation digital computer (CORDIC) algorithm and its architecture based on a novel power-of-2 coefficient low-complexity Taylor series approximation to implement sinh and cosh functions. CORDIC architectures are generally slow due to their high latency of computation. The proposed architecture reduces the latency and achieves the desired precision with only four iterations where an optimized angle set comprised of six CORDIC microrotations are mapped into a four-stage folded-pipeline structure leveraging mutually exclusive behavior of two pairs of microrotations. The proposed design is implemented on field-programmable gate arrays (FPGAs) Xilinx Zedboard using 65.38% less registers with ~63.63% less latency and 48.97% less power consumption compared with the best of the existing designs. The proposed design is synthesized by Synopsys Design Compiler and place and route (PnR) tool using Taiwan Semiconductor Manufacturing Company (TSMC) 65-nm CMOS process. It consumes ~76.31% less area, 68.75% less computational delay, and 68.92% less power consumption compared with the best of the existing designs. Moreover, the proposed architecture involves 46.89% less energy per output (EPO) than the best of the existing designs. The error–energy performance (EEP) and the error–area performance (EAP) of the proposed design are, respectively, ~1.25 times and ~2.8 times better than that of the best of the existing designs. Besides, the proposed architecture is also implemented and verified on a silicon chip in the TSMC 180-nm CMOS process for the validation of the algorithm and architecture. Anu Verma, Khyati Kiyawat, Bishnu Prasad Das, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | An Efficient Parallel DA-Based Fixed-Width Design for Approximate Inner-Product ComputationabstractParallel distributed arithmetic (PDA)-based structures are widely used for high-speed computation of inner product in digital signal processing (DSP) applications. In this article, we have proposed novel PDA-based structures based on an efficient truncation model. To achieve higher bit saving with relatively less truncation error, we present here a novel approach using approximate look-up tables (LUTs), adder trees (ATs), and Wallace-like shift-AT (SAT) with truncated operands to obtain hardware-efficient fixed-width PDA-based inner-product structures. We have three variants of proposed structures based on the proposed truncation approach. We find that the proposed inner-product structure-1 using approximate LUT (ALUT) and approximate AT offers nearly 20% higher bit saving, 20% saving in area-delay product (ADP) and offers relatively less truncation error than the existing structures. The proposed structure-2 using ALUT, ATs, and proposed SAT offers nearly 50% higher bit-saving, 61% ADP saving and offers nearly the same accuracy compared to the existing approximate DA-based structures. Proposed structure-3 offers nearly 60% higher bit saving and calculates outputs with almost the same or marginally less accuracy than the existing structures for higher coefficient word lengths. Basant K. Mohanty, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Analysis and Design of Unified Architectures for Zero-Attraction-Based Sparse Adaptive FiltersabstractZero-attraction-based adaptive filters are widely used for sparse system identification, where a suitable penalty function is integrated with the least mean square (LMS) framework to improve the convergence behavior of the identification process. In this brief, we have made an attempt to implement some of the most popular zero-attracting algorithms in hardware. The complexity of realization associated with these algorithms is investigated in detail. Following the above analysis, several architectural simplifications are proposed for the reduced-complexity implementation of their penalty functions. We then use these realizations to develop a set of novel design strategies for the efficient implementation of these algorithms. Simulation results show that the performance loss for the proposed algorithms is minimal compared to their standard versions. A detailed synthesis study is also carried out to validate the proposed structures, which demonstrates that the hardware overhead in the proposed designs is marginal compared to the existing delayed LMS architecture. Dwaipayan Ray, Nithin V. George, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Analysis and Design of Approximate Inner-Product Architectures Based on Distributed ArithmeticabstractDistributed arithmetic (DA) based architectures are popularly used for inner-product computation in various applications. Existing literature shows that the use of approximate DA-architectures in error resilient applications provides a significant improvement in the overall efficiency of the system. Based on precise error analysis, we find that the existing methods introduce large truncation error in the computation of the final inner-product. Therefore, to have a suitable trade-off between the overall hardware complexity and truncation error, a weight-dependent truncation approach is proposed in this paper. The overall efficiency of the structure is further enhanced by incorporating an input truncation strategy in the proposed method. It is observed that the area, time and energy efficiency of the proposed designs are superior to the existing designs with significantly lower truncation error. Evaluation in the case of noisy image smoothing application is also shown in this paper. Dwaipayan Ray, Nithin V. George, Pramod Kumar Meher |
ISCAS | 3 |
| 2019 | Low-Complexity Systolic Multiplier for GF(2m) using Toeplitz Matrix-Vector Product MethodabstractLow-complexity systolic multipliers for GF(2m) are required in several high-performance cryptographic systems. In this paper, we propose a novel design strategy to derive efficient systolic multiplier for GF(2m) based on Toeplitz Matrix-Vector Product (TMVP) approach. The proposed work is carried out through two coherent interdependent stages. (i) A novel multiplication algorithm based on TMVP method to obtain subquadratic space complexity is proposed first. (ii) The proposed algorithm is then mapped unto to a novel and efficient architecture which is optimized further to derive a low-complexity systolic structure. The complexity analysis and comparison show that the proposed design outperforms the existing work. The proposed design can thus be used in many practical cryptosystems. Jiafeng Xie, Chiou-Yng Lee, Pramod Kumar Meher |
ISCAS | 3 |
| 2019 | Novel Bit-Parallel and Digit-Serial Systolic Finite Field Multipliers Over $GF(2^m)$ Based on Reordered Normal BasisabstractEfficient implementation of finite field multipliers based on a reordered normal basis (RNB) is highly desirable in the current/emerging cryptosystems since it offers almost free realization of squaring operation. Therefore, in this paper, we propose novel bit-parallel and digit-serial finite field multipliers over GF(2m) based on RNB. By efficient transformation of the core multiplication algorithm using a unique circular shifting feature, we have derived an efficient algorithm for low-complexity systolic mapping. Both bit-parallel and digit-serial structures of the multipliers are then obtained and optimized to enhance the area-time efficiency. We have also utilized the unique feature of the proposed multiplication algorithm to obtain the systolic multipliers by Karatsubalike decomposition. Detailed analysis and comparison show the superior performance of the proposed implementation. For example, the proposed regular and Karatsuba-based bit-parallel designs involve at least 48.4% less area-delay product (ADP) and 42.2% less power-delay product (PDP) than the best existing ones (37.7% and 55.3% less ADP and PDP on field-programmable gate array platform), respectively. The proposed multipliers, because of their lower area-time complexities, can be used for efficient realization of cryptographic applications. Jiafeng Xie, Chiou-Yng Lee, Pramod Kumar Meher, Zhi-Hong Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Threshold-Guided Design and Optimization for Harris Corner Detector ArchitectureabstractHigh-speed corner detection is an essential step in many real-time computer vision applications, e.g., object recognition, motion analysis, and stereo matching. Hardware implementation of corner detection algorithms, such as the Harris corner detector (HCD) has become a viable solution for meeting real-time requirements of the applications. A major challenge lies in the design of power, energy and area efficient architectures that can be deployed in tightly constrained embedded systems while still meeting real-time requirements. In this paper, we proposed a bit-width optimization strategy for designing hardware-efficient HCD that exploits the thresholding step in the algorithm to determine interest points from the corner responses. The proposed strategy relies on the threshold as a guide to truncate the bit-widths of the operators at various stages of the HCD pipeline with only marginal loss of accuracy. Synthesis results based on 65-nm CMOS technology show that the proposed strategy leads to power-delay reduction of 35.2%, and area reduction of 35.4% over the baseline implementation. In addition, through careful retiming, the proposed implementation achieves over 2.2 times increase in maximum frequency while achieving an area reduction of 35.1% and power-delay reduction of 35.7% over the baseline implementation. Finally, we performed repeatability tests to show that the optimized HCD architecture achieves comparable accuracy with the baseline implementation (average decrease of repeatability is less than 0.6%). Bhavan A. Jasani, Siew-Kei Lam, Pramod Kumar Meher, Meiqing Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Lower Bound Analysis and Perturbation of Critical Path for Area-Time Efficient Multiple Constant MultiplicationsabstractIn this paper, a precise systematic delay model is proposed for the analysis and estimation of critical path delay of multiple constant multiplication (MCM) blocks. For the first time in literature, the mathematical derivation of lower bound of critical path delay of MCM blocks is presented and necessary conditions for achieving the lower bound of critical path delay are discussed. It is shown that the lower bound of critical path delay of MCMs is significantly smaller than that achieved by existing MCM algorithms. An improved genetic algorithm-based approach, with a heuristic algorithm to generate the initial population, is proposed to search for low complexity MCM solutions with the lower bound of critical path delay. This is the first time that design algorithms with gate-level delay control is proposed. Moreover, it is shown that using the information of lower bound of critical path delay, perturbation of timing can be applied to tradeoff the lower bound critical path delay against hardware complexity. It is shown that area-time efficient design of MCM blocks can be obtained by using the proposed techniques. Xin Lou 0001, Ya Jun Yu, Pramod Kumar Meher |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Scalable Approximate DCT Architectures for Efficient HEVC-Compliant Video CodingabstractAn approximate kernel for the discrete cosine transform (DCT) of length 4 is derived from the 4-point DCT defined by the High Efficiency Video Coding (HEVC) standard and used for the computation of DCT and inverse DCT (IDCT) of power-of-two lengths. There are two reasons for considering the DCT of length 4 as the basic module. First, it allows computation of DCTs of lengths 4, 8, 16, and 32 prescribed by the HEVC. Second, the DCTs generated by the 4-point DCT not only involve lower complexity, but also offer better compression performance. Fully parallel and area-constrained architectures for the proposed approximate DCT are proposed to have flexible tradeoff between the area and time complexities. In addition, a reconfigurable architecture is proposed where an 8-point DCT can be used in place of a pair of 4-point DCTs. Using the same reconfiguration scheme, a 32-point DCT could be configured for parallel computation of two 16-point DCTs or four 8-point DCTs or eight 4-point DCTs. The proposed reconfigurable design can support real-time coding for high-definition video sequences in the 8k ultrahigh-definition television format (7680 × 4320 at 30 frames/s). A unified forward and inverse transform architecture is also proposed where the hardware complexity is reduced by sharing hardware between the DCT and IDCT computations. The proposed approximation has nearly the same arithmetic complexity and hardware requirement as those of recently proposed related methods, but involves significantly less error energy and offers better peak signal-to-noise ratio than the others when DCTs of length more than 8 are used. A detailed comparison of the complexity, energy efficiency, and compression performance of different DCT approximation schemes for video coding is also presented. It is shown that the proposed approximation provides a better compressed-image quality than other approximate DCTs. The proposed method can perform HEVC-compliant video coding with marginal degradation of video quality and a slight increase the in bit rate, with a fraction of computational complexity of the latter. Maher Jridi, Pramod Kumar Meher |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Low-Complexity Digit-Serial Multiplier Over $GF(2^{m})$ Based on Efficient Toeplitz Block Toeplitz Matrix-Vector Product DecompositionabstractIn this paper, we have shown that a regular Toeplitz matrix-vector product (TMVP) can be transformed into a Toeplitz block TMVP (TBTMVP) using a suitable permutation matrix. Based on the TBTMVP representation, we have proposed a new (a,b)-way TBTMVP decomposition algorithm for implementing a digit-serial multiplication. Moreover, it is shown that, based on iterative block recombination, we can improve the space complexity of the proposed TBTMVP decomposition. From the synthesis results, we have shown that the proposed TBTMVP-based multiplier involves less area, less area-delay product, and higher throughput compared with the existing digit-serial multipliers. Chiou-Yng Lee, Pramod Kumar Meher, Chia-Chen Fan, Shyan-Ming Yuan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A high-performance VLSI architecture for reconfigurable FIR using distributed arithmetic
Basant K. Mohanty, Pramod Kumar Meher, Subodh Kumar Singhal, M. N. S. Swamy 0001 |
Integr. | 2 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 16 |
| 2016 | Concept, Design, and Implementation of Reconfigurable CORDICabstractThis brief presents the key concept, design strategy, and implementation of reconfigurable coordinate rotation digital computer (CORDIC) architectures that can be configured to operate either for circular or for hyperbolic trajectories in rotation as well as vectoring-modes. It can, therefore, be used to perform all the functions of both circular and hyperbolic CORDIC. We propose three reconfigurable CORDIC designs: 1) a reconfigurable rotation-mode CORDIC that operates either for circular or for hyperbolic trajectory; 2) a reconfigurable vectoring-mode CORDIC for circular and hyperbolic trajectories; and 3) a generalized reconfigurable CORDIC that can operate in any of the modes for both circular and hyperbolic trajectories. The reconfigurable CORDIC can perform the computation of various trigonometric and exponential functions, logarithms, square-root, and so on of circular and hyperbolic CORDIC using either rotation-mode or vectoring-mode CORDIC in one single circuit. It can be used in digital synchronizers, graphics processors, scientific calculators, and so on. It offers substantial saving of area complexity over the conventional design for reconfigurable applications. Supriya Aggarwal, Pramod Kumar Meher, Kavita Khare |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Area-Delay Efficient Digit-Serial Multiplier Based on k-Partitioning Scheme Combined With TMVP Block Recombination ApproachabstractShifted polynomial basis (SPB) and generalized polynomial basis (GPB) are two efficient bases of representation in binary extension fields, and are widely studied. In this paper, we use the GPB formulation to derive a new modified SPB (MSPB) representation for arbitrary irreducible trinomials and pentanomials. It is shown that the basis conversion from the MSPB to the SPB for trinomials is free of hardware cost. We have shown that multiplication based on SPB and MSPB representations can make use of Toeplitz matrix-vector product (TMVP) formulation. The existing TMVP block recombination (TMVPBR) approach is used here to derive an efficient k-partitioning TMVPBR decomposition for digit-serial double basis multiplication that can achieve subquadratic space complexity. From synthesis results, we have shown that the proposed multiplier has less area and less area-delay product compared with the existing digit-serial multipliers. We also show that the proposed multiplier using k-partitioning TMVPBR decomposition can provide a better tradeoff between time and space complexities. Chiou-Yng Lee, Pramod Kumar Meher, Chung-Hsin Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | On Efficient Retiming of Fixed-Point CircuitsabstractRetiming of digital circuits is conventionally based on the estimates of propagation delays across different paths in the data-flow graphs (DFGs) obtained by discrete component timing model, which implicitly assumes that operation of a node can begin only after the completion of the operation(s) of its preceding node(s) to obey the data dependence requirement. Such a discrete component timing model very often gives much higher estimates of the propagation delays than the actuals particularly when the computations in the DFG nodes correspond to fixed-point arithmetic operations like additions and multiplications. On the other hand, very often it is imperative to deal with the DFGs of such higher granularity at the architecture-level abstraction of digital system design for mapping an algorithm to the desired architecture, where the overestimation of propagation delay leads to unwanted pipelining and undesirable increase in pipeline overheads. In this paper, we propose the connected component timing model to obtain adequately precise estimates of propagation delays across different combinational paths in a DFG easily, for efficient cutset-retiming in order to reduce the critical path substantially without significant increase in register-complexity and latency. Apart from that, we propose novel node-splitting and node-merging techniques that can be used in combination with the existing retiming methods to achieve reduction of critical path to a fraction that of the original DFG with a small increase in overall register complexity. Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | A High-Performance FIR Filter Architecture for Fixed and Reconfigurable ApplicationsabstractTranspose form finite-impulse response (FIR) filters are inherently pipelined and support multiple constant multiplications (MCM) technique that results in significant saving of computation. However, transpose form configuration does not directly support the block processing unlike direct-form configuration. In this paper, we explore the possibility of realization of block FIR filter in transpose form configuration for area-delay efficient realization of large order FIR filters for both fixed and reconfigurable applications. Based on a detailed computational analysis of transpose form configuration of FIR filter, we have derived a flow graph for transpose form block FIR filter with optimized register complexity. A generalized block formulation is presented for transpose form FIR filter. We have derived a general multiplier-based architecture for the proposed transpose form block filter for reconfigurable applications. A low-complexity design using the MCM scheme is also presented for the block implementation of fixed FIR filters. The proposed structure involves significantly less area-delay product (ADP) and less energy per sample (EPS) than the existing block implementation of direct-form structure for medium or large filter lengths, while for the short-length filters, the block implementation of direct-form FIR structure has less ADP and less EPS than the proposed structure. Application-specific integrated circuit synthesis result shows that the proposed structure for block size 4 and filter length 64 involves 42% less ADP and 40% less EPS than the best available FIR filter structure proposed for reconfigurable applications. For the same filter length and the same block size, the proposed structure involves 13% less ADP and 12.8% less EPS than that of the existing direct-form block FIR structure. Basant K. Mohanty, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | LUT Optimization for Distributed Arithmetic-Based Block Least Mean Square Adaptive FilterabstractIn this paper, we analyze the contents of lookup tables (LUTs) of distributed arithmetic (DA)-based block least mean square (BLMS) adaptive filter (ADF) and based on that we propose intra-iteration LUT sharing to reduce its hardware resources, energy consumption, and iteration period. The proposed LUT optimization scheme offers a saving of 60% LUT content for block size 8 and still higher saving for larger block sizes over the conventional design approach. We also present here the design of a register-based LUT matrix for maximal sharing of LUT contents and full-parallel LUT-update operation. Based on the proposed design approach, we have derived a DA-based architecture for the BLMS ADF, which is scalable for larger block sizes as well as higher filter lengths. We find that the hardware complexity of the proposed structure increases less than proportionately with input block size and filter length. It offers a saving of 60% LUT-update per output and 59% LUT access per output over the recently proposed DA-based BLMS ADF structure for block size 8 and filter length 64. Besides, the proposed structure involves nearly 30% saving in the iteration period over the other for 16-bit coefficient word length. Application specific integrated circuit (ASIC) synthesis result shows that the proposed structure for block size 8 offers a saving of 48% area-delay product (ADP) and 53% energy per sample (EPS) over the existing DA-based BLMS ADF structure on average for different filter lengths, and offers 30% higher sampling rate due to its shorter iteration period. Compared with the existing DA-based LMS ADF structure, the proposed structure involves 68% less ADP and $1.6 \times $ less EPS. Basant K. Mohanty, Pramod Kumar Meher, Sujit Kumar Patel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Fine-grained pipelining for multiple constant multiplicationsabstractMultiple constant multiplication (MCM) is widely used in several digital signal processing applications. Recently, pipelining techniques have been applied to accelerate the computation of the MCM blocks. The existing pipelining techniques consider the adder stage pipelining, i.e., inserting registers between two adjacent adder stages, to reduce the adder depth. However, the critical path may be still long, even though the adder depth is minimized. In this paper, the adder stage pipelining method is analyzed at bit-level and a novel pipelining method is proposed for pipelining the adders in the MCM block. Experimental results show that the proposed pipelining method provides nearly 32% reduction of critical path over the traditional adder stage pipelining in average for several benchmark MCM blocks, while the area and power consumption are maintained. Xin Lou 0001, Pramod Kumar Meher, Ya Jun Yu |
ISCAS | 2 |
| 2015 | Critical-path optimization for efficient hardware realization of lifting and flipping DWTsabstractThe pair of scaling constants of lifting scheme for DWT computation are perfect inverse of each other while those of flipping scheme are not perfect inverse of each other. However, flipping-scheme is preferred over the lifting-scheme for area-delay efficient hardware implementation of 2-D DWT due to its smaller critical-path. In this paper we present critical-path analysis of lifting and flipping schemes and propose a low-complexity data-path for lifting based DWT. We have shown that the critical-path delay (CPD) of lifting DWT is higher by only 2 full-adder (FA) delay than that of flipping DWT. Moreover, due to the saving of multipliers, lifting scheme offers a better area delay efficient structure than the flipping scheme for parallel realization of 2-D DWT. In this paper, we propose efficient realization of both lifting and flipping DWT. Compared with the best of the available flipping-based 2-D DWT structure, the proposed lifting-based and flipping-based structures for block-size 16 involve 74% less and 73% less area-delay-product (ADP), and offer nearly 6.38 and 6.72 times higher throughput, respectively. Compared with similar lifting based existing design, the proposed lifting-based structure involves 15.38% less ADP for the same block-size. Basant K. Mohanty, Pramod Kumar Meher, Thambipillai Srikanthan |
ISCAS | 2 |
| 2015 | Efficient subquadratic parallel multiplier based on modified SPB of GF(2m)abstractToeplitz matrix-vector product (TMVP) approach is a special case of Karatsuba algorithm to design subquadratic multiplier in GF(2m). In binary extension fields, shifted polynomial basis (SPB) is a variable basis representation, and is widely studied. SPB multiplication using coordinate transformation technique can transform TMVP formulas, however, this approach is only applied for the field constructed by all trinomials or special class of pentanomials. For this reason, we present a new modified SPB multiplication for an arbitrary irreducible pentanomial, and the proposed multiplication scheme has formed a TMVP formula. Jeng-Shyang Pan 0001, Pramod Kumar Meher, Chiou-Yng Lee, Hong-Hai Bai |
ISCAS | 2 |
| 2015 | FPGA Implementation of Orthogonal Matching Pursuit for Compressive Sensing ReconstructionabstractIn this paper, we present a novel architecture based on field-programmable gate arrays (FPGAs) for the reconstruction of compressively sensed signal using the orthogonal matching pursuit (OMP) algorithm. We have analyzed the computational complexities and data dependence between different stages of OMP algorithm to design its architecture that provides higher throughput with less area consumption. Since the solution of least square problem involves a large part of the overall computation time, we have suggested a parallel low-complexity architecture for the solution of the linear system. We have further modeled the proposed design using Simulink and carried out the implementation on FPGA using Xilinx system generator tool. We have presented here a methodology to optimize both area and execution time in Simulink environment. The execution time of the proposed design is reduced by maximizing parallelism by appropriate level of unfolding, while the FPGA resources are reduced by sharing the hardware for matrix-vector multiplication across the data-dependent sections of the algorithm. The hardware implementation on the Virtex6 FPGA provides significantly superior performance in terms of resource utilization measured in the number of occupied slices, and maximum usable frequency compared with the existing implementations. Compared with the existing similar design, the proposed structure involves 328 more DSP48s, but it involves 25802 less slices and 1.85 times less computation time for signal reconstruction with N = 1024, K = 256, and m = 36, where N is the number of samples, K is the size of the measurement vector, and m is the sparsity. It also provides a higher peak signal-to-noise ratio value of 38.9 dB with a reconstruction time of 0.34 μs, which is twice faster than the existing design. In addition, we have presented a performance metric to implement the OMP algorithm in resource constrained FPGA for the better quality of signal reconstruction. Hassan Rabah, Abbes Amira, Basant K. Mohanty, Somaya Al-Máadeed, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2014 | Reconfigurable CORDIC architectures for multi-mode and multi-trajectory operationsabstractThis paper presents reconfigurable CORDIC (Coordinate Rotation Digital Computer) architectures which can be configured to operate either for circular or hyperbolic trajectories in rotation as well as vectoring-modes. We propose three reconfigurable CORDIC designs: a reconfigurable rotation-mode CORDIC that operates either for circular or hyperbolic trajectory, a reconfigurable vectoring-mode CORDIC for circular and hyperbolic trajectories, and a generalized reconfigurable CORDIC that can operate in any of the modes for both circular as well as hyperbolic trajectories. The reconfigurable CORDIC can perform the computation of various trigonometric and exponential functions, logarithms, square-root, etc. of circular and hyperbolic CORDICs using either rotation-mode or vectoring-mode of operation in one single circuit. It can be used in digital synchronizers, graphics processors, scientific calculators and many other applications, with significant area saving over that of using two CORDICs for different trajectories. Supriya Aggarwal, Pramod Kumar Meher |
ISCAS | 2 |
| 2014 | High-speed multiplier block design based on bit-level critical path optimizationabstractMultiple constant multiplications (MCM) is a popular technique to implement multiplier blocks with low hardware cost and power consumption. Research works on MCM have been on going for more than two decades. Most algorithms so far have focused on reducing the number of adders and/or adder depth to have low power and/or high speed circuit. However, low adder depth does not guarantee the low critical path examined in bit-level. In this work, we propose an algorithm to optimize the critical path of multiplier blocks in bit-level. Simulation results show that the critical path delay can be reduced by using the proposed algorithm. Xin Lou 0001, Ya Jun Yu, Pramod Kumar Meher |
ISCAS | 3 |
| 2014 | Area-delay efficient architecture for MP algorithm using reconfigurable inner-product circuitsabstractMatching pursuit (MP) algorithm is popularly used as a low-cost alternative to the orthogonal matching pursuit (OMP) algorithm for the reconstruction of signal from compressively sensed samples. In this paper, we have proposed an efficient scheduling of computation along with a novel reconfigurable inner-product (IP) unit and buffer units to provide regular inflow of input, and storage of intermediate results for efficient implementation of MP algorithm. The proposed reconfigurable IP unit can compute inner-products of different lengths by reusing the arithmetic components with very low reconfiguration overhead. We have customized on-chip buffers to exploit desired level of parallelism with lower latency and higher hardware utilization efficiency. The proposed structure of MP algorithm for the reconstruction of compressively sensed data is found to involve nearly 15% less critical path delay, 7% less area, 10% less reconstruction time, and 17% less area-delay-product (ADP) than the existing structure. Pramod Kumar Meher, Basant K. Mohanty, Thambipillai Srikanthan |
ISCAS | 1 |
| 2014 | A novel DA-based architecture for efficient computation of inner-product of variable vectorsabstractDistributed arithmetic (DA) has been widely used for area-time efficient implementation of inner-products, where one of the vectors is fixed and known a priori. Computation of inner-product of a pair of variable vectors, however, is required very often for matrix-multiplication of different forms, and implementation of digital filters of unknown coefficients and variable lengths. But the possibility of using DA for the computation of inner-product of variable vectors is yet to be explored. In this paper, we analyse the design issues relating to DA-based implementation of inner-product of variable vectors, and derive a novel area-time efficient flexible solution for the bit-parallel DA-based implementation of inner-product of variable vectors and variable inner-product lengths. It is found that the proposed structures are nearly 34% faster than the conventional multiplier-based implementation in average for different inner-product lengths (N = 8, 16, 32 and 64) and for input word-lengths, L = 8 and L = 16. Moreover, proposed designs offer saving of nearly 22% and 36% area-delay product (ADP) and saving of nearly 16% and 24% power delay product (PDP) over the multiplier-based designs for L = 8 and L = 16, respectively, in average, for various inner-product lengths. Pramod Kumar Meher |
ISCAS | 1 |
| 2014 | Area-delay-power-efficient architecture for folded two-dimensional discrete wavelet transform by multiple lifting computationabstractMultiple lifting computation could be performed for block processing of two‐dimensional (2D) discrete wavelet transform (DWT) by combined‐lifting (CLF) or separated‐lifting (SLF) approaches. CLF and SLF have the same computational complexities but they differ by their register requirements. In this study, the authors have chosen CLF for row processing and SLF for column processing, and suggested an efficient scheduling scheme for the computation of block‐based lifting 2D DWT. Based on this approach, the authors have derived a parallel‐pipeline structure for high‐throughput implementation of one‐level lifting 2D DWT. The authors have partitioned the multilevel 2D DWT computation appropriately and mapped that to a folded structure where the frame‐buffer size is independent of input block size. The proposed structure requires 3 N on‐chip memory words, which is the lowest among all the existing similar structures. Compared with the best of the existing block‐based structures for the one‐level DWT, the proposed structure involves less on‐chip memory words, requires the same number of multipliers and adders and offers the same throughput rate. The application specific integrated circuit (ASIC) synthesis result shows that the proposed structure involves significantly less area‐delay‐product and less energy per image than those of the best of the available designs. Basant K. Mohanty, Pramod Kumar Meher |
IET Image Process. | 2 |
| 2014 | Efficient Integer DCT Architectures for HEVCabstractIn this paper, we present area- and power-efficient architectures for the implementation of integer discrete cosine transform (DCT) of different lengths to be used in High Efficiency Video Coding (HEVC). We show that an efficient constant matrix-multiplication scheme can be used to derive parallel architectures for 1-D integer DCT of different lengths. We also show that the proposed structure could be reusable for DCT of lengths 4, 8, 16, and 32 with a throughput of 32 DCT coefficients per cycle irrespective of the transform size. Moreover, the proposed architecture could be pruned to reduce the complexity of implementation substantially with only a marginal affect on the coding performance. We propose power-efficient structures for folded and full-parallel implementations of 2-D DCT. From the synthesis result, it is found that the proposed architecture involves nearly 14% less area-delay product (ADP) and 19% less energy per sample (EPS) compared to the direct implementation of the reference algorithm, on average, for integer DCT of lengths 4, 8, 16, and 32. Also, an additional 19% saving in ADP and 20% saving in EPS can be achieved by the proposed pruning algorithm with nearly the same throughput rate. The proposed architecture is found to support ultrahigh definition 7680 × 4320 at 60 frames/s video, which is one of the applications of HEVC. Pramod Kumar Meher, Basant K. Mohanty, Khoon Seong Lim, Chuohao Yeo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Area-Delay-Power Efficient Fixed-Point LMS Adaptive Filter With Low Adaptation-DelayabstractIn this paper, we present an efficient architecture for the implementation of a delayed least mean square adaptive filter. For achieving lower adaptation-delay and area-delay-power efficient implementation, we use a novel partial product generator and propose a strategy for optimized balanced pipelining across the time-consuming combinational blocks of the structure. From synthesis results, we find that the proposed design offers nearly 17% less area-delay product (ADP) and nearly 14% less energy-delay product (EDP) than the best of the existing systolic structures, on average, for filter lengths N=8, 16, and 32. We propose an efficient fixed-point implementation scheme of the proposed architecture, and derive the expression for steady-state error. We show that the steady-state mean squared error obtained from the analytical result matches with the simulation result. Moreover, we have proposed a bit-level pruning of the proposed architecture, which provides nearly 20% saving in ADP and 9% saving in EDP over the proposed structure before pruning without noticeable degradation of steady-state-error performance. Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Flexible integer DCT architectures for HEVCabstractIn this paper, we present high throughput and power-efficient architectures for the implementation of integer DCT of different lengths to be used in upcoming High Efficiency Video Coding (HEVC). We have shown that efficient matrix-multiplication schemes could be used to derive parallel architectures for 1-D integer DCT of different lengths. Apart from that we have proposed three different flexible architectures which could be used for implementing the DCT of any of the prescribed lengths such as 4, 8, 16 and 32, each having particular advantage in terms of area, delay, or power. The proposed architectures can provide higher throughput at a lower operating frequency than the existing architectures for HEVC. Furthermore, it can support Ultra-High-Definition (UHD) 7680×4320 @30fps video which is one of the applications of HEVC. Pramod Kumar Meher |
ISCAS | 2 |
| 2013 | Zero-quantised discrete cosine transform coefficients prediction technique for intra-frame video encodingabstractOne promising solution to reduce the computational complexity of discrete cosine transform (DCT) is to identify the redundant computations and to get rid of them. In this study, the authors present a new method to predict zero‐quantised DCT coefficients for efficient implementation of intra‐frame video encoding by identifying such redundant computations. Traditional methods use the Gaussian statistical model of residual pixels to predict all‐zero or partial‐zero blocks. The proposed method is based on two key ideas. At first, the bounds of DCT coefficients are derived from the intermediate signals of the Loeffler DCT algorithm instead of calculating the sum of absolute difference (SAD) of residual pixels. The sufficiency conditions are then suitably chosen to predict the zero‐quantised coefficients to reduce the arithmetic complexity without degrading the video quality. Simulation results are found to validate the analytical model and show that the proposed prediction eliminates more redundant computations than the existing methods. Moreover, the authors have derived a pipelined VLSI architecture of the proposed prediction scheme which offers a saving of more than 63 and 91% of multiplications of the second stage of one‐dimensional DCT for high and low bit‐rate intra‐video encoding, respectively. Maher Jridi, Pramod Kumar Meher, Ayman Alfalou |
IET Image Process. | 2 |
| 2013 | Hardware-Efficient Realization of Prime-Length DCT Based on Distributed ArithmeticabstractThis paper presents an efficient decomposition scheme for hardware-efficient realization of discrete cosine transform (DCT) based on distributed arithmetic. We have proposed an efficient design for the implementation of cyclic convolution based on a group distributed arithmetic (GDA) technique where the read-only memory size could be reduced over the existing GDA-based design. The proposed structure for DCT implementation, based on the new decomposition scheme and proposed design of GDA-based cyclic convolution, involves significantly less area complexity than the existing one. For example, to implement the DCT of transform length N = 17, the proposed design needs a lookup table of 128 words, while the existing design for N = 16 requires a lookup table of 256 words. From the synthesis results, it is found that proposed design involves significantly less area, gives higher throughput, and consumes less power compared to the existing designs of nearly the same or lower lengths. Jiafeng Xie, Pramod Kumar Meher |
IEEE Trans. Computers | 2 |
| 2013 | Memory-Efficient High-Speed Convolution-Based Generic Structure for Multilevel 2-D DWTabstractIn this paper, we have proposed a design strategy for the derivation of memory-efficient architecture for multilevel 2-D DWT. Using the proposed design scheme, we have derived a convolution-based generic architecture for the computation of three-level 2-D DWT based on Daubechies (Daub) as well as biorthogonal filters. The proposed structure does not involve frame-buffer. It involves line-buffers of size 3(K-2)M/4 which is independent of throughput-rate, whereKis the order of Daubechies/biorthogonal wavelet filter andMis the image height. This is a major advantage when the structure is implemented for higher throughput. The structure has regular data-flow, small cycle periodTMand 100% hardware utilization efficiency. As per theoretical estimate, for image size 512 × 512, the proposed structure for Daub-4 filter requires 152 more multipliers and 114 more adders, but involves 82 412 less memory words and takes 10.5 times less time to compute three-level 2-D DWT than the best of the existing convolution-based folded structures. Similarly, compared with the best of the existing lifting-based folded structures, proposed structure for 9/7-filter involves 93 more multipliers and 166 more adders, but uses 85 317 less memory words and requires 2.625 times less computation time for the same image size. It involves 90 (nearly 47.6%) more multipliers and 118 (nearly 40.1%) more adders, but requires 2723 less memory words than the recently proposed parallel structure and performs the computation in nearly half the time of the other. Inspite of having more arithmetic components than the lifting-based structures, the proposed structure offers significant saving of area and power over the other due to substantial reduction in memory size and smaller clock-period. ASIC synthesis result shows that, the proposed structure for Daub-4 involves 1.7 times less area-delay-product (ADP) and consumes 1.21 times less energy per image (EPI) than the corresponding best available convolution-based structure. It involves 2.6 times less ADP and consumes 1.48 times less EPI than the parallel lifting-based structure. Basant K. Mohanty, Pramod Kumar Meher |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | CORDIC Designs for Fixed Angle of RotationabstractRotation of vectors through fixed and known angles has wide applications in robotics, digital signal processing, graphics, games, and animation. But, we do not find any optimized coordinate rotation digital computer (CORDIC) design for vector-rotation through specific angles. Therefore, in this paper, we present optimization schemes and CORDIC circuits for fixed and known rotations with different levels of accuracy. For reducing the area- and time-complexities, we have proposed a hardwired pre-shifting scheme in barrel-shifters of the proposed circuits. Two dedicated CORDIC cells are proposed for the fixed-angle rotations. In one of those cells, micro-rotations and scaling are interleaved, and in the other they are implemented in two separate stages. Pipelined schemes are suggested further for cascading dedicated single-rotation units and bi-rotation CORDIC units for high-throughput and reduced latency implementations. We have obtained the optimized set of micro-rotations for fixed and known angles. The optimized scale-factors are also derived and dedicated shift-add circuits are designed to implement the scaling. The fixed-point mean-squared-error of the proposed CORDIC circuit is analyzed statistically, and strategies for reducing the error are given. We have synthesized the proposed CORDIC cells by Synopsys Design Compiler using TSMC 90-nm library, and shown that the proposed designs offer higher throughput, less latency and less area-delay product than the reference CORDIC design for fixed and known angles of rotation. We find similar results of synthesis for different Xilinx field-programmable gate-array platforms. Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Low Latency Systolic Montgomery Multiplier for Finite Field $GF(2^{m})$ Based on PentanomialsabstractIn this paper, we present a low latency systolic Montgomery multiplier over$GF(2^{m})$based on irreducible pentanomials. An efficient algorithm is presented to decompose the multiplication into a number of independent units to facilitate parallel processing. Besides, a novel so-called “pre-computed addition” technique is introduced to further reduce the latency. The proposed design involves significantly less area-delay and power-delay complexities compared with the best of the existing designs. It has the same or shorter critical-path and involves nearly one-fourth of the latency of the other in case of the National Institute of Standards and Technology recommended irreducible pentanomials. Jiafeng Xie, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Low-Complexity Multiplier for GF(2m) Based on All-One PolynomialsabstractThis paper presents an area-time-efficient systolic structure for multiplication overGF(2m) based on irreducible all-one polynomial (AOP). We have used a novel cut-set retiming to reduce the duration of the critical-path to one XOR gate delay. It is further shown that the systolic structure can be decomposed into two or more parallel systolic branches, where the pair of parallel systolic branches has the same input operand, and they can share the same input operand registers. From the application-specific integrated circuit and field-programmable gate array synthesis results we find that the proposed design provides significantly less area-delay and power-delay complexities over the best of the existing designs. Jiafeng Xie, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Efficient architectures for VLSI implementation of 2-D discrete Hadamard transformabstractIn this paper, we present three different structures, namely the transposition-free structure, the folded structure and the pipeline structure for 2-D discrete Hadamard transform (DHT). The transposition-free structure and pipeline structure produce one column of output during each clock cycle, while the folded structure requires two clock cycles for that. The folded structure uses one 1-D DHT module for both row and column processing, while the pipeline structure processes rows and columns concurrently using two separate 1-D DHT modules. Interestingly, the transposition-unit of the pipeline structure involves nearly the same number of registers as the folded structure, and offers twice the throughput of the other. The transposition-free structure is less area-time efficient than pipeline structure due to its relatively less efficient serial-output processors. ASIC synthesis result shows that the pipeline structure involves 47.4% less area-delay product (ADP) and 53.74% less energy per sample (EPS) than the folded structure, and involves slightly less ADP and consumes 31.67% less EPS than the transposition-free design. Basant K. Mohanty, Pramod Kumar Meher, Subodh Kumar Singhal |
ISCAS | 2 |
| 2012 | Low-latency area-delay-efficient systolic multiplier over GF(2m) for a wider class of trinomials using parallel register sharingabstractSystolic structures for finite field multiplication involve large number of registers for parallel implementation, while bit-serial implementations require a large computation time, which increases along with the order of the field. In this paper, we present a novel scheme for the decomposition of the multiplication over GF(2m) based on irreducible trinomials into several independent units that facilitates maximal resister sharing and low-latency parallel implementation. It is shown that the proposed design involves significantly less area-delay complexity compared with the best of the corresponding existing systolic designs, and could be used for a wider class of trinomials. Jiafeng Xie, Pramod Kumar Meher |
ISCAS | 2 |
| 2012 | Area-Time Efficient Scaling-Free CORDIC Using Generalized Micro-Rotation SelectionabstractThis paper presents an area-time efficient CORDIC algorithm that completely eliminates the scale-factor. By suitable selection of the order of approximation of Taylor series the proposed CORDIC circuit meets the accuracy requirement, and attains the desired range of convergence. Besides we have proposed an algorithm to redefine the elementary angles for reducing the number of CORDIC iterations. A generalized micro-rotation selection technique based on high speed most-significant-1-detection obviates the complex search algorithms for identifying the micro-rotations. The proposed CORDIC processor provides the flexibility to manipulate the number of iterations depending on the accuracy, area and latency requirements. Compared to the existing recursive architectures the proposed one has 17% lower slice-delay product on Xilinx Spartan XC2S200E device. Supriya Aggarwal, Pramod Kumar Meher, Kavita Khare |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | High-Throughput Interpolator Architecture for Low-Complexity Chase Decoding of RS CodesabstractIn this paper, a high-throughput interpolator architecture for soft-decision decoding of Reed-Solomon (RS) codes based on low-complexity chase (LCC) decoding is presented. We have formulated a modified form of the Nielson's interpolation algorithm, using some typical features of LCC decoding. The proposed algorithm works with a different scheduling, takes care of the limited growth of the polynomials, and shares the common interpolation points, for reducing the latency of interpolation. Based on the proposed modified Nielson's algorithm we have derived a low-latency architecture to reduce the overall latency of the whole LCC decoder. An efficiency of at least 39%, in terms of area-delay product, has been achieved by an LCC decoder, by using the proposed interpolator architecture, over the best of the previously reported architectures for an RS(255,239) code with eight test vectors. We have implemented the proposed interpolator in a Virtex-II FPGA device, which provides 914 Mb/s of throughput using 806 slices. Francisco Garcia-Herrero, Ma José Canet, Javier Valls-Coquillat, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | A high-speed FIR adaptive filter architecture using a modified delayed LMS algorithmabstractIn this paper, we present a modified delayed least means square (DLMS) adaptive algorithm to achieve lower adaptation-delay. Besides, we have proposed an efficient pipelined architecture for the implementation of this adaptive filter. We have shown that the proposed DLMS adaptive filter can be implemented by a pipelined inner-product computation unit for calculation of feedback error, and a pipelined weight-update unit consisting of N parallel multiply accumulators, for filter order N. From the synthesis results we find that the existing direct-form structure of [8] involves nearly 50% more area-delay product (ADP) and nearly 74% more energy per sample (EPS) than the proposed one, in average, for filter orders N = 8,16 and 32. The best of the existing systolic structures [7], similarly, involves nearly 43% more ADP and nearly 35% higher EPS than the proposed one for the same filter orders. Pramod Kumar Meher, Megha Maheshwari |
ISCAS | 1 |
| 2011 | Efficient coefficient partitioning for decomposed DA-based inner-product computationabstractThe Look-Up-Table (LUT) size grows exponentially with increasing number of coefficients in a straight forward implementation of memory based Distributed Arithmetic (DA) computation. To avoid exponential blow-up, the common practice breaks the entire set of coefficients into multiple groups, each represented by a much smaller LUT, and forms the result by summing up the outputs from these LUTs. A detailed inspection shows that the overall word-size of the LUTs could be minimized by proper reordering and grouping of the coefficients, thus reduces the total memory usage of the LUTs. A fast heuristic partitioning algorithm for this purpose is devised and analyzed in this paper, showing up to 16% resource reduction compared to in-order grouping. The proposed technique can be applied orthogonally to offset binary coded (OBC) LUTs for resource efficient DA implementation as well. Pramod Kumar Meher |
ISCAS | 2 |
| 2011 | MCM-based implementation of block FIR filters for high-speed and low-power applicationsabstractBlock finite impulse response (FIR) digital filters have potential for high-speed and low-power realization through parallel processing. In this paper, we suggest an efficient implementation of block FIR filters using multiple constant multiplication (MCM) technique. Constant multiplication methods are widely used for reducing computational complexity of implementation of FIR filters. Sub-expression sharing for single constant multiplications can be performed for direct-form as well as transposed direct-form structures of FIR filters, while MCM techniques are not applicable to the direct-form FIR structure. On the other hand block FIR filters cannot be implemented in transposed direct-form. In this paper, we have shown that MCM can be used for direct-form implementation block FIR filters. Experimentation on block filters for filter orders 8 and 16 of different block lengths indicates that, compared to sample-by-sample MCM based transposed direct-form filters, the maximum sampling rate and energy delay product may be improved by up to 3.5 and 7.6 times respectively due to aggressive parallelization of block processing. It is also found that by using MCMs, up to 20% total area may be reduced compared to straightforward block filter implementation. Pramod Kumar Meher |
VLSI-SoC | 1 |
| 2011 | High-throughput pipelined realization of adaptive FIR filter based on distributed arithmeticabstractIn this paper, we propose an efficient pipelined architecture for high-speed adaptive filter based on distributed arithmetic (DA). We have shown that the sampling period could be substantially reduced by using carry-save accumulation instead of shift-accumulation for DA-based inner-product implementation for the computation of filter output. Unlike the existing design, the proposed design does not involve any lookup table (LUT). It involves half the number of registers compared to the existing DA-based design to store the sum of different combinations of input samples. The proposed design involves nearly 17% more hardware but offers nearly 7 times throughput and nearly 14 times less energy per sample, in average for filter orders N = 8, 16 and 32 over the existing DA-based design for adaptive filter. Pramod Kumar Meher |
VLSI-SoC | 1 |
| 2011 | A Self-Configurable Systolic Architecture for Face Recognition System Based on Principal Component Neural NetworkabstractAn efficient self-configurable systolic architecture is proposed in this paper for very large scale integration implementation of a face recognition system. The proposed system applies principal component neural network (PCNN) with generalized Hebbian learning for extracting eigenfaces from the face database. It demonstrates a recognition performance of more than 85% when evaluated on the benchmark Yale and FRGC databases containing images with varying illumination and expression. Unlike the existing face recognition systems, the proposed approach not only recognizes the faces using computed eigenfaces, but also updates eigenfaces automatically whenever the face database changes. The challenge, however, lies in hardware realization of the PCNN-based face recognition system. In the presence of computation-intensive steps of varying nature, it is not straightforward to map the overall computation to a single systolic architecture. A primary contribution of this paper from the architecture point of view is an optimized mapping of fine-grained systolized signal flow graphs (SFGs) for each individual step of the algorithm on to a single self-configurable linear systolic array by appropriate merging of the computations pertaining to different nodes of different SFGs. The architecture has the flexibility of processing face images and databases of any size and it is easily scalable with the number of eigenfaces to be computed. The proposed PCNN-based systolic face recognition system has been implemented and evaluated on a Xilinx ML403 evaluation platform with Virtex-4 XC4VFX12 FPGA. The FPGA-based design for a reasonably large-sized face database can process more than 400 faces in a video image frame which is fast enough for video surveillance in busy public places and sensitive locations. N. Sudha, A. R. Mohan, Pramod Kumar Meher |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | An improved SOM-based visualization technique for DNA microarray data analysisabstractEffective and meaningful visualization techniques are quite important for multidimensional DNA microarray gene expression data analysis. Elucidating the cluster properties of these multidimensional data are often complex. Patterns, hypotheses on the relationships, and ultimately of the function of the gene can be analyzed and visualized by non-linear reduction of the multidimensional data to a lower dimension. In this paper, an improved SOM visualization technique named Improved Side Intensity Modulated (ISIM) Self-Organizing Map (SOM) has been proposed and compared with other SOM based visualization techniques. On different datasets, ISIM-SOM is found to offer better cluster boundary, simplicity and clarity. Jagdish C. Patra, Jacob A. Abraham, Pramod Kumar Meher, Goutam Chakraborty |
IJCNN | 3 |
| 2010 | DNA microarray analysis using Equalized Orthogonal MappingabstractGene expression data obtained from DNA microarray experiments consists of expression levels of thousands of genes of only a few samples. Thus, accurate analysis of these datasets for classification of cancer is a big challenge. In this paper, we apply a novel Equalized Orthogonal Map (EOM) as a dimension reduction technique to produce topologically correct map of DNA microarray data for visualization and classification of cancer types. Effectiveness of the EOM has been investigated for visualization and self organization using a benchmark microarray dataset and its performance was compared with the Kohonen's Self Organizing Map (SOM). EOM was able to produce 100% accuracy for both FL and CLL samples of the lymphoma dataset. Furthermore, EOM is computationally more efficient, e.g., it took only 3 seconds for training of FL-CLL samples of Alizadeh et al. dataset, whereas SOM took more than 8 seconds. With extensive simulation results we have illustrated superiority of the EOM over SOM in terms of quality of maps, classification accuracy and computational complexity. Jagdish C. Patra, Nyttle V. George, Pramod Kumar Meher |
IJCNN | 3 |
| 2010 | An optimized lookup-table for the evaluation of sigmoid function for artificial neural networksabstractIn this paper, we present an efficient design of lookup-table (LUT) for the evaluation of hyperbolic tangent sigmoid to be used for the hardware implementation of artificial neural networks. Besides, we have suggested an LUT optimization scheme which maximizes the number of argument values in a sub-domain corresponding to each LUT word for a specified limit of accuracy. We have shown that the hardware-complexity of the proposed LUT implementation could be significantly reduced by using simplified combinational circuits for selective sign-conversion and efficient design of range decoder by logic subexpression sharing. From the synthesis results, we find that the proposed design involves comparable delay, but requires less than one-fourth of the area and area-delay complexity compared with the existing LUT-based implementations. Pramod Kumar Meher |
VLSI-SoC | 1 |
| 2010 | Novel input coding technique for high-precision LUT-based multiplication for DSP applicationsabstractIn this paper, we present a novel input-coding scheme for high-precision lookup-table (LUT)-based implementation of constant multiplications by input operand decomposition. Besides, we have described an efficient LUT design for the multiplication of input sub-words where the input coding technique is combined with the odd-multiple-storage technique to achieve the reduction of LUT size by a factor of ~ 4 over the conventional technique. Compared with the antisymmetric product coding (APC) scheme, the input coding scheme involves significantly less area and less time overheads. The proposed LUT-multiplier and the existing one are coded in VHDL and synthesized by Synopsys Design Compiler using 90 nanometer CMOS library. The proposed one is found to offer more than 28% saving of area-delay product over the existing LUT multiplier, in average, for word-sizes 8, 16 and 32. Pramod Kumar Meher |
VLSI-SoC | 1 |
| 2010 | An improved common subexpression elimination method for reducing logic operators in FIR filter implementations without increasing logic depth
A. Prasad Vinod 0001, Edmund M.-K. Lai, Douglas L. Maskell, Pramod Kumar Meher |
Integr. | 4 |
| 2010 | Parallel and Pipeline Architectures for High-Throughput Computation of Multilevel 3-D DWTabstractIn this paper, we present a throughput-scalable parallel and pipeline architecture for high-throughput computation of multilevel 3-D discrete wavelet transform (3-D DWT). The computation of 3-D DWT for each level of decomposition is split into three distinct stages, and all the three stages are implemented in parallel by a processing unit consisting of an array of processing modules. The processing unit for the first level decomposition of a video stream of frame-size (M × N) consists ofQ/2 processing modules, whereQis the number of input samples available to the structure in each clock cycle. The processing unit for a higher level of decomposition requires 1/8 times the number of processing modules required by the processing unit for its preceding level. ForJlevel 3-D DWT of a video stream, each of the proposed structures involvesJprocessing units in a cascaded pipeline. The proposed structures have a small output latency, and can perform multilevel 3-D DWT computation with 100% hardware utilization efficiency. The throughput rate of proposed structures areQ/7 time higher than the best of the corresponding existing structures. Interestingly, the proposed structures involve a frame-buffer ofO(MN) while the frame-buffer size of the existing structures isO(MNR) . Besides, the on-chip storage and the frame-buffer size of the proposed structure is independent of the input-block size and this favors to derive highly concurrent parallel architecture for high-throughput implementation. The overall area-delay products of proposed structure are significantly lower than the existing structures, although they involve slightly more multiplier-delay product and more adder-delay product, since it involves significantly less frame-buffer and storage-word-delay product. The throughput rate of the proposed structure can easily be scaled without increasing the on-chip storage and frame-memory by using more number of processing modules, and it provides greater advantage over the existing designs for higher frame-rates and higher input block-size. The full-parallel implementation of proposed scalable structure provides the best of its performance. When very high throughput generated by such parallel structure is not required, the structure could be operated by a slower clock, where speed could be traded for power by scaling down the operating voltage and/or the processing modules could be implemented by slower but hardware-efficient arithmetic circuits. Basant K. Mohanty, Pramod Kumar Meher |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Concurrent Error Detection in Bit-Serial Normal Basis Multiplication Over GF(2m) Using Multiple Parity Prediction SchemesabstractNew bit-serial architectures with concurrent error detection capability are presented to detect erroneous outputs in bit-serial normal basis multipliers over GF(2m) using single and multiple-parity prediction schemes. It is shown that different types of normal basis multipliers could be realized by similar architectures. The proposed architectures can detect errors with nearly 100% probability. Chiou-Yng Lee, Pramod Kumar Meher, Jagdish C. Patra |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | An optimized design for serial-parallel finite field multiplication over GF(2m) based on all-one polynomialsabstractIn this paper, we derive a recursive algorithm for finite field multiplication over GF(2m) based on irreducible all-one-polynomials (AOP), where the modular reduction of degree is achieved by cyclic-left-shift without any logic operations. A regular and localized bit-level dependence graph (DG) is derived from the proposed algorithm and mapped into an array architecture, where the modular reduction is achieved by a serial-in parallel-out shift-register. The multiplier is optimized further to perform the accumulation of partial products by the T flip flops of the output register without XOR gates. It is interesting to note that the optimized structure consists of an array of (m+1) AND gates between an array of (m+1) D flip flops and an array of (m+1) T flip flops. The proposed structure therefore involves significantly less area and less computation time compared with the corresponding existing structures. Pramod Kumar Meher, Yajun Ha, Chiou-Yng Lee |
ASP-DAC | 1 |
| 2009 | Computationally efficient FLANN-based intelligent stock price prediction systemabstractWe propose a computationally efficient and effective novel neural network for predicting the next-day's closing price of US stocks in different sectors: technology, energy and finance. In this paper we used a computationally efficient functional link artificial neural network (FLANN) in making stock price prediction. We modeled the trend in stock price movement as a dynamic system and apply FLANN to predict the stock price behavior. In addition to historical pricing data, we considered other financial indicators such as the industrial indices and technical indicators, for better accuracy. We showed its superior performance by comparing with a multilayer perceptron (MLP)-based model through several experiments based on different performance metrics, namely, computational complexity, root mean square error, average percentage error and hit rate. Jagdish C. Patra, Nguyen C. Thanh, Pramod Kumar Meher |
IJCNN | 3 |
| 2009 | New Approach to LUT Implementation and Accumulation for Memory-based MultiplicationabstractA new approach to look-up-table (LUT) implementation for memory-based multiplication is presented, where the memory-size is reduced to half at the cost of some increase in combinational circuit complexity. The proposed design offers a saving of nearly 42% area and 38% area-delay product (ADP) at the cost of 6% increase in computational delay for memory-based multiplication of 8-bit inputs with 16-bit coefficient. For high-precision multiplication, a shift-save-accumulation scheme is proposed to accumulate the LUT outputs corresponding to the segments of input-operand, which requires nearly 1.5 times more area, but offers more than twice the throughput and nearly two-third the ADP of direct shift-accumulation approach. Pramod Kumar Meher |
ISCAS | 1 |
| 2009 | Scalable Serial-parallel Multiplier over GF(2m) by Hierarchical Pre-reduction and Input DecompositionabstractThis paper presents a novel serial-parallel architecture for finite field multiplications over GF(2m) defined by irreducible trinomials as field polynomials. By recursive decomposition of one of the operands, and hierarchical pre-reduction of the other, it is possible to feed multiple bits in parallel to the serial-parallel structure. The level of parallelism could be doubled after each level of decomposition of the input operand, when high throughput rate is required. One of the key features of the proposed design is that its clock-period remains invariant with the digit-size. The area-complexity of the proposed design increases linearly with the digit-size, which is unlike some of the existing architectures, where area-complexity increases quadratically with the digit-size. Although the proposed structure involves more area compared with some of the existing architectures, since the clock-period of the proposed design is small, it involves significantly less area-delay complexity than the others. Pramod Kumar Meher, Chiou-Yng Lee |
ISCAS | 1 |
| 2009 | Nonlinear channel equalization for wireless communication systems using Legendre neural networks
Jagdish C. Patra, Pramod Kumar Meher, Goutam Chakraborty |
Signal Process. | 2 |
| 2009 | Extended Sequential Logic for Synchronous Circuit Optimization and Its ApplicationsabstractIn this paper, we present a new approach for the extension of sequential logic functionality ofDflip-flop in order to perform an additional Boolean function simultaneously along with its usual bit-storage function. We show that a combinational function of the form (amiddotb), (a+b) , (a+[(b)]), or ([(a)] middotb) which occurs frequently in a feedforward path with aDflip-flop could be implemented efficiently by aDflip-flop with RESET or SET provision. Similarly, (aoplusb) or ((amiddotb) oplusc) in the feedback loop with aDflip-flop could be implemented by aTflip-flop by suitable modification of the clock. The use of such extended sequential logic is found to result in a significant reduction in critical path and saving in area complexity over the direct implementation. Moreover, we present a simple approach for the construction of CMOSTflip-flop by modification of clock signal ofDflip-flop, which is found to be more efficient than theTflip-flop derived fromJKflip-flop. The extended sequential logic is used for the implementation of finite-field multiplication overGF(2m) and carry-save addition of real numbers. In both these cases, the use of extended logic is found to offer a substantial saving in area and time complexity over the conventional implementations. Pramod Kumar Meher |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | On Efficient Implementation of Accumulation in Finite Field Over GF(2m) and its ApplicationsabstractFinite field accumulation is the simplest of all the finite field operations, but at the same time, it is one of the most frequently encountered operations in finite field arithmetic. In this paper, we present a simple but highly useful modification of the conventional hardware implementation of accumulation in finite field overGF(2m) . The critical path, as well as, the hardware-complexity are reduced in the proposed design by performing the accumulation operation usingmnumber ofTflip-flops instead of using a combination ofmnumber of XOR gates with equal number ofDflip-flops in dependent loop structures. The conventional design is found to involve nearly 39% more area, 53% more delay, and 40% more maximum ac power consumption compared with the proposed accumulator. The proposed finite field accumulator is used further for the implementation of serial/parallel polynomial-basis finite field multiplication and bit-serial inter-conversion between polynomial basis representation and normal basis representation overGF(2m). The area-time complexity of the proposed bit-serial/parallel multiplier is less than half of the best of the corresponding existing structures. The structure proposed for digit-serial/parallel multiplication for trinomials is found to involve nearly 56% less area-time complexity compared with the best of the corresponding existing multipliers; and the existing design of bit-serial basis conversion is found to involve nearly twice area-time complexity compared with the proposed design using the proposed finite field accumulator. Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Systolic and Non-Systolic Scalable Modular Designs of Finite Field Multipliers for Reed-Solomon CodecabstractIn this paper, we present efficient algorithms for modular reduction to derive novel systolic and non-systolic architectures for polynomial basis finite field multipliers overGF(2m) to be used in Reed-Solomon (RS) codec. Using the proposed algorithm for unit degree reduction and optimization of implementation of the logic functions in the processing elements (PEs), we have derived an efficient bit-parallel systolic design for finite field multiplier which involves nearly two-thirds of the area-complexity of the existing design having the same time-complexity. The proposed modular reduction algorithms are also used to derive efficient non-systolic serial/parallel designs of field multipliers overGF(28) with different digit-sizes, where the critical path and the hardware-complexity are further reduced by optimizing the implementation of modular reduction operations and finite field accumulations. The proposed bit-serial design involves nearly 55% of the minimum of area, and half the minimum of area-time complexity of the existing bit-serial designs. Similarly, the proposed digit-serial/parallel designs involve significantly less area, and less area-time complexities compared with the existing designs of the same digit-size. By parallel modular reduction through multiple degrees followed by appropriate logic-level sub-expression sharing; a hardware-efficient regular and modular form of a balanced-tree bit-parallel non-systolic multiplier is also derived. The proposed bit-parallel non-systolic pipelined design involves less than 65% of the area and nearly two-thirds of the area-time complexity of the existing bit-parallel design for a RS codec, while the non-pipelined form offers nearly 25% saving of area with less time-complexity. Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Efficient Bit-Parallel Multipliers in Composite FieldsabstractHardware implementation of multiplication in finite field GF(2m) based on sparse polynomials is found to be advantageous in terms of space-complexity as well as the time-complexity. In order to design multipliers for the composite fields, we have found another permutation polynomial to convert irreducible polynomials into like-trinomials of the forms (x2+ x + 1)m+ (x2+ x + 1)n+ 1, (x2+ x)m+ (x2+ x)n+ 1 and (x4+ x + 1)m+ (x4+ x + 1)n+ 1. The proposed bit-parallel multiplier over GF(24m) is found to offer a saving of about 33% multiplications and 42.8% additions over the corresponding existing architectures. Chiou-Yng Lee, Pramod Kumar Meher |
APSCC | 2 |
| 2008 | Efficient systolization of cyclic convolution for systolic implementation of sinusoidal transformsabstractThis paper presents an algorithm to convert composite-length cyclic convolution into a block cyclic convolution sum of small matrix-vector products, even if the co-factors of convolution-length are not mutually prime. It is shown that by using optimal short-length convolution algorithms, the block-convolution could be computed from a few short-length cyclic and cyclic-like convolutions, when one of the co-factors belongs to {2, 3, 4, 6, 8}. A generalized systolic array is derived for cyclic-like convolution, and used that for the computation of long-length convolutions. The proposed structure for convolution-length N= 2L involves nearly the same hardware and half the time-complexity as the direct implementation; and the structure for N= 4L involves sime12.5% more hardware and one-fourth the time-complexity of the latter. The structures for N=2L and N=4L, respectively, have the same and sime12.5% less area-time complexity as the corresponding existing prime-factor systolic structures, but unlike the latter type, do not involve complex input/output mapping; and could be used even if the co-factors of convolution-length are not relatively prime. Pramod Kumar Meher |
ASAP | 1 |
| 2008 | Fully-pipelined efficient architectures for FPGA realization of discrete Hadamard transformabstractFully-pipelined simple modular structures are presented in this paper for efficient hardware realization of discrete Hadamard transform (HT). From the kernel matrix of HT, we have derived four different pipelined modular designs for transform length N = 4. It is shown further that the HT of transform-length N = 8 can be obtained from two 4-point HT modules, and similarly, the HT of transform-length N=16 can be obtained from four 4-point HT modules. Long-length transforms may, however, be computed from these short-length modules as N-point transforms can be computed from 2M number of M point HT-modules, where M = N1/2. The proposed architectures are coded in VHDL, simulated by Xilinx ISE tool for validation and testing; and synthesized thereafter to be implemented in different FPGA devices, e.g., Virtex-E, Virtex-II Pro and Virtex-4. From the synthesis result, it is found that the proposed designs involve considerably less number of slices and provide significantly higher best-achievable-frequency compared with the existing architectures for FPGA implementation of HT. Pramod Kumar Meher, Jagdish C. Patra |
ASAP | 1 |
| 2008 | Concurrent systolic architecture for high-throughput implementation of 3-dimensional discrete wavelet transformabstractIn this paper, we present a novel systolic architecture for high-throughput computation of 3-dimensional (3-D) discrete wavelet transform (DWT). The entire 3-D DWT computation is decomposed into three distinct stages and implemented concurrently in a linear array of fully pipelined processing elements (PE). The proposed structure for 3-D DWT provides higher throughput than the existing architecture; and involves nearly half or less the number of multipliers and adders; and less on-chip memory (when normalized for unit throughput rate) than the other. Most importantly, the proposed one does not require any frame buffer unlike the other to perform inter-frame DWT computation. The proposed structure has a small latency and can perform 3-D DWT computation with 100% hardware unitization efficiency. Basant K. Mohanty, Pramod Kumar Meher |
ASAP | 2 |
| 2008 | Throughput-scalable hybrid-pipeline architecture for multilevel lifting 2-D DWT of JPEG 2000 coderabstractIn this paper, we propose a pipelined-architecture for high-throughput computation of multilevel lifting 2D discrete wavelet transform (DWT). The multilevel DWT computation is shared by the proposed devices based on pyramid algorithm (PA) and recursive pyramid algorithm (RPA), where the PA-based devices compute the lower order subands and the higher order subbands are computed by an RPA-based device. The hardware- and time-complexities of the proposed structure are compared with those of the existing recursive architectures for performance evaluation. Compared with the best of the existing recursive architectures, the proposed one has nearly 16 times less average computation time (ACT) for the 2D DWT of input size 512 x 512 for S=32, where S is half of the input rate of the structure. Moreover, it involves less number of multipliers and adders than the others when normalized for unit throughput rate. The proposed design offers nearly 100% utilization efficiency for S=32, and 94% efficiency for S=8. The latency of the structure is very small (which is of the order of a few cycles), and involves a small on-chip storage and less number of data/pipeline registers. Basant K. Mohanty, Pramod Kumar Meher |
ASAP | 2 |
| 2008 | Discrete tchebichef transform-A fast 4x4 algorithm and its application in image/video compressionabstractDiscrete Tchebichef transform (DTT), derived from a discrete class of the popular Chebyshev polynomials, is a novel orthogonal transform that has high energy compaction and de-correlation properties. Therefore, in this paper, DTT is examined and treated for transform coding applications. A framework is laid to derive an approximation- free integer representation of DTT to meet the current application requirements. A fast algorithm is further proposed for multiplier-free computation of DTT. The image compression performance of the 4- point DTT is found to be superior to that of the 4-point discrete cosine transform (DCT) and integer cosine transform (ICT), the integer approximation of DCT. It is shown that the fast DTT is easily derived, has low complexity, does not involve approximations and can be carried out within the same dynamic range. Hence, DTT can be used for image and data compression applications. Since the image compression performance and computational simplicity of DTT are found to be significantly better than that of ICT, the use of DTT in place of ICT for transform coding in the H.264/AVC looks promising. Sujata Ishwar, Pramod Kumar Meher, M. N. S. Swamy 0001 |
ISCAS | 2 |
| 2008 | Determination of QSAR of aldose reductase inhibitors using an RBF networkabstractIncreasingly drug discovery is turning to in-silico methods to find potential enzyme or protein-protein inhibitors. The quantitative structure-activity relationship (QSAR) is commonly used to find potential inhibitors in the search for new drugs. In this paper we propose to use a radial basis function (RBF) network to determine the QSAR of aldose reductase inhibitors (ARIs). We find that the RBF network shows promising results of predicting the bioactivities of the ARIs. Jagdish C. Patra, Rowena Wai Sim Cheong, Pramod Kumar Meher, Goutam Chakraborty |
SMC | 3 |
| 2008 | Legendre-FLANN-based nonlinear channel equalization in wireless communication systemabstractIn this paper, we present the result of our study on the application of artificial neural networks (ANNs) for adaptive channel equalization in a digital communication system using 4-quadrature amplitude modulation (QAM) signal constellation. We propose a novel single-layer Legendre functional-link ANN (L-FLANN) by using Legendre polynomials to expand the input space into a higher dimension. A performance comparison was carried out with extensive computer simulations between different ANN-based equalizers, such as, radial basis function (RBF), Chebyshev neural network (ChNN) and the proposed L-FLANN along with a linear least mean square (LMS) finite impulse response (FIR) adaptive filter-based equalizer. The performance indicators include the mean square error (MSE), bit error rate (BER), and computational complexities of the different architectures as well as the eye patterns of the various equalizers. It is shown that the L-FLANN exhibited excellent results in terms of the MSE, BER and the computational complexity of the networks. Jagdish C. Patra, Wei Chiat Chin, Pramod Kumar Meher, Goutam Chakraborty |
SMC | 3 |
| 2008 | Robust CRT-based watermarking technique for authentication of image and documentabstractThe advent of the Internet and the wide availability of computers, scanners, and printers make digital data acquisition, exchange, and transmission as simple tasks. However, making digital data accessible to others through networks also creates opportunities for malicious parties to make salable copies of copyrighted content without permission of the content owner. Digital watermarking techniques have been proposed as a solution to the problem of copyright protection of multimedia data in network environments. In this paper, we propose a novel CRT-based technique for digital watermarking and also look into increasing the capacity of watermark embedding in the host images. The proposed CRT-based technique, besides being computationally efficient, is also more resistant to different types of attacks together with a significant reduction in processing time and minimal distortion of the original image. Experimental results have shown that the performance of the new scheme is superior in both the quality of the extracted watermark and processing times in comparison to the two existing SVD-based watermarking techniques. Jagdish C. Patra, Alagappan Karthik, Pramod Kumar Meher, Cédric Bornand |
SMC | 3 |
| 2008 | Support vector machine application in drug discovery of aldose reductase inhibitorsabstractUsing support vector machine (SVM) function approximation, in this paper, we present the quantitative structure-activity relationship (QSAR) among the known aldose reductase inhibitors (ARIs). The two physical descriptors of a molecule, namely the electronegativity and the molar volume are evaluated by SVM. SVM is found to work better than multi-layer perceptron (MLP). Jagdish C. Patra, Pramod Kumar Meher |
SMC | 3 |
| 2008 | Development of intelligent sensors using Legendre functional-link artificial neural networksabstractDifferent types of sensors are used to control and monitor complex systems in many applications, where the environmental parameters, e.g., temperature, humidity, etc., undergo large variations. In such conditions, the sensor's output may be erroneous and the system being controlled may malfunction. The need of intelligent sensors arise in such situations. These sensors should be capable of compensating for the adverse effects of the environmental conditions on the sensor output and linearization of sensor response, in order to provide correct readout. In this paper, we propose a novel computationally efficient Legendre functional-link artificial neural network (L-FLANN) to develop a smart sensor that can compensate for the adverse effects of the environmental conditions. By taking two types of environmental models and a pressure sensor, we have shown with extensive computer simulations that the proposed smart sensor is computationally efficient with respect to a multi-layer perceptron (MLP)-based sensor model and capable of satisfactory linearization of sensor output. This smart sensor produces only plusmn0.5% full-scale error between the actual and estimated output for the two selected environmental models under a temperature variation of -50 to 200degC. Jagdish C. Patra, Pramod Kumar Meher, Goutam Chakraborty |
SMC | 2 |
| 2008 | Content-based image retrieval using orthogonal moments heuristicallyabstractA number of approaches have been used recently for image retrieval using color features. Use of orthogonal moments instead of normal moments and histograms has been found to be more effective. We propose a scheme to further improve the efficiency by using the most differentiating moments heuristically. The performance for moments calculated using four different orthogonal polynomials is compared and a study is conducted to see if the four types of moments can be used together. Jagdish C. Patra, Sonaabh Sood, Pramod Kumar Meher, Cédric Bornand |
SMC | 3 |
| 2008 | Parallel and Pipelined Architectures for Cyclic Convolution by Block Circulant Formulation Using Low-Complexity Short-Length AlgorithmsabstractFully pipelined parallel architectures are derived for high-throughput and reduced-hardware realization of prime-factor cyclic convolution using hardware-efficient modules for short-length rectangular transform (RT). Moreover, a new approach is proposed for the computation of block pseudocyclic convolution using a block cyclic convolution of equal length along with some correction terms, so that the block pseudocyclic representation of cyclic convolution for non-prime-factor-length (N=rP, whenrandPare not mutually prime) could be computed efficiently using the algorithms and architectures of short-length cyclic convolutions. Low-complexity algorithms are derived for efficient computation of those error terms, and overall complexities of the proposed technique are estimated forr=2, 3, 4, 6, 8 and 9. The proposed algorithms are used further to design high-throughput and reduced-hardware structures for cyclic convolution where the cofactors are not relatively prime. The proposed structures for high-throughput implementation are found to offer a reduction of nearly 50%-75% of area-delay product over the existing structures for several convolution-lengths. Low-complexity structures for input/output addition units of short length convolutions are derived and used them along with high-throughput modules for hardware-efficient realization of multifactor convolution, which offers nearly 25%-75% reduction of area-delay complexity over the existing structures for various non-prime-factor length convolutions. Pramod Kumar Meher |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Systolic Formulation for Low-Complexity Serial-Parallel Implementation of Unified Finite Field Multiplication over GF(2m)abstractIt presents a high-throughput hardware-efficient semi-systolic linear array for a serial-parallel implementation of finite field multiplier over GF(2m) using bidirectional modulo reduction technique. Necessary recurrence relations are formulated and a pair of dependence graphs (DG) are designed for least significant bit (LSB) and most significant bit (MSB) elimination algorithms for modular reduction. Both the DGs are merged together and mapped into a fully-pipelined linear array architecture consisting of to number of processing elements (PEs), which performs one field multiplication in every (m/2) cycles. The structure of each PE is optimized further to be implemented by a pair of AND gates, three XOR gates and a pair of latches. The duration of a cycle amounts to T = TA+ Tx3+ TL, where TA,Tx3and TL, are respectively the delays of a 2-input AND gate, a three-input XOR gate and a latch. The proposed design is found to have significantly low area-time complexity compared with the existing serial-parallel structures for finite filed multiplications. It is shown that the proposed multiplier can also be used for the Montgomery multiplication in binary field. Pramod Kumar Meher |
ASAP | 1 |
| 2007 | DNA Microarray Data Analysis: Effective Feature Selection for Accurate Cancer ClassificationabstractAccurate classification of DNA microarray data is vital for cancer diagnosis and treatment. For greater accuracy, a preferable strategy is to make a decision based on the result of a single classifier that is trained with various aspects of data space. It is a difficult task to create an optimal classifier for DNA analysis that deals with only a few samples with large number of features. Usually, different feature sets are provided for classifiers to learn. If the feature sets provide similar information, the classifiers trained from them cannot improve the performance because they will make the same error and there is no possibility of compensation. In this paper, we adopt correlation analysis of feature selection methods as a guideline for selection of features for classifiers to learn. We use a negative correlation method for generation of feature sets those are mutually exclusive. Each classifier is learned from different features sets based on correlation analysis to classify cancer precisely. In this way, we evaluated the performance with two benchmark datasets. Experimental results show that classifiers, which have learned from different feature sets that are negatively correlated with each other, produce the best recognition rates on the two benchmark datasets. Jagdish C. Patra, Goh P. Lim, Pramod Kumar Meher, Ee Luang Ang |
IJCNN | 3 |
| 2006 | Field Programmable Gate Array Implementation of a Neural Network-based Intelligent Sensor SystemabstractA multi-layer perceptron neural network with floatingpoint number system is implemented on a field programmable gate array (FPGA). IEEE-754 32-bit single precision floatingpoint number is used to represent values in the neural network accurately. The neural network forms the core of an intelligent sensor system which has the ability to mitigate the nonlinear influence on the sensor output by external disturbances. Training is performed on the neural network to approximate the response characteristics of a sensor for different level of disturbances so as to compensate for the nonlinearity. The intelligent sensor system is implemented on Celoxica RC203E development board which contains a Xilinx Virtex-II FPGA chip. A custom-built intelligent light intensity sensor is used for experimentation and the neural network is able to achieve a maximum full-scale (FS) error of plusmn1.5% under the nonlinear influence caused by the varying distance between the sensor and the light source. In terms of root mean squared error (RMSE), it is able to achieve a RMSE of 0.0052 Jagdish C. Patra, Han Yang Lee, Pramod Kumar Meher, Ee Luang Ang |
ICARCV | 3 |
| 2006 | A New SOM-based Visualization Technique for DNA Microarray DataabstractWe present a new side-intensity modulated self-organizing map (SIM-SOM) for improved visualization of multidimensional data. We have utilized DNA microarray dataset [15] for this purpose. A Gene Signature which contains a set of most informative genes was extracted using the discrimination factor method [16]. Next, the reduced dimensional microarray data are visualized using the proposed SIM-SOM. In addition to providing clear and unambiguous cluster boundary, the SIM-SOM can help in discovering new subtypes of cancer. Jagdish C. Patra, Ee Luang Ang, Pramod Kumar Meher, Qin Zhen |
IJCNN | 3 |
| 2006 | Financial Prediction of Major Indices using Computational Efficient Artificial Neural NetworksabstractTwo computational efficient artificial neural networks (ANNs) for the prediction of major financial indices are proposed. First, we propose a single layer functional link artificial neural network (FLANN) for this purpose. FLANN has a simple structure in which the nonlinearity is introduced by the functional expansion of the input pattern using trigonometric polynomials. The second ANN proposed is a Chebyshev neural network (chNN) in which the functional expansion is carried out using Chebyshev polynomials. Performance comparison of the two ANNs with regards to a multilayer perceptron (MLP) were carried out through extensive computer simulations. It is shown that the proposed ANNs outperform the MLP for the prediction of the three financial indices. Jagdish C. Patra, Weineng Lim, Pramod Kumar Meher, Ee Luang Ang |
IJCNN | 3 |
| 2006 | A new approach to secure distributed storage, sharing and dissemination of digital imageabstractIn this paper, we present a novel technique for secure distributed storage and dissemination of digital images using Chinese remainder theorem (CRT). The proposed technique not only involves significantly low computational complexity but also imposes various stages of difficulty to an intruder who intends to extract the original image from the encrypted sub-blocks. The proposed CRT-based technique is, therefore, useful for secure archival of large volume of image data and sharing of digital images by the users from different locations Pramod Kumar Meher, Jagdish C. Patra |
ISCAS | 1 |
| 2006 | Low-complexity technique for secure storage and sharing of biomedical imagesabstractIn this paper, we present an apparently trivial but powerful technique for secure storage and sharing of medical images using simple bit-level scrambling. Unlike the conventional cryptographic techniques, it involves very low computational complexity. Unless some one knows the keys for deciphering, cannot extract the original image from the encrypted blocks of image by any kind of attack. The proposed technique is useful for secure storage and sharing of large volume of medical images for real-time healthcare applications Pramod Kumar Meher, Jagdish C. Patra, M. R. Meher |
ISCAS | 1 |
| 2006 | A novel neural network-based linearization and auto-compensation technique for sensorsabstractAn artificial neural network (NN)-based technique for smart sensors operating in harsh environments is proposed that automatically calibrates, linearizes and compensates for the adverse effects due to nonlinear response characteristics and complex nonlinear dependency of the sensor characteristics on the environmental parameters. To show the potential of the proposed NN-based technique, we provide simulation results of a smart capacitive pressure sensor (CPS) operating under a wide temperature range. Jagdish C. Patra, Ee Luang Ang, Pramod Kumar Meher |
ISCAS | 3 |
| 2006 | Highly concurrent reduced-complexity 2-D systolic array for discrete Fourier transformabstractA simple two-dimensional (2-D) architecture is derived for highly concurrent systolization of the discrete Fourier transform (DFT). The concurrency of computation has been enhanced, and complexity is minimized by the proposed algorithm where an N-point DFT is computed via four inner-products of real-valued data of length ap(N/2). The proposed structure offers significantly lower latency, higher throughput, and involves nearly half the minimum of the area-time complexity of the existing multiplier-based DFT structures. It is found that the 2-D DFT using proposed one-dimensional (1-D) structure has nearly half the area-complexity and less than one-fourth of area-time complexity of the existing structures. Besides, it is also found to have nearly half the area-time complexity of the existing multiplierless DFT structure. Unlike some of the existing structures, the proposed one can be used for the DFT of any transform-length and does not involve tag-bit control Pramod Kumar Meher |
IEEE Signal Process. Lett. | 1 |
| 2006 | Systolic Designs for DCT Using a Low-Complexity Concurrent Convolutional FormulationabstractA reduced-complexity convolutional formulation is presented for systolic implementation of the discrete cosine transform, where N-point transform can be computed by four numbers of nearly (N/4)-point circular-convolution-like operations. The proposed algorithm not only provides a reduction of computational complexity by four times over the conventional formulation, where N-point transform is computed via (N-1)-point cyclic convolution, but also leads to concurrent pipelined execution in linear systolic arrays. It is shown that the multiplications in the processing elements can be implemented by lookup-tables using dual-port ROM. Two variants of systolic structures using ROM-based multipliers are presented for efficient implementation of the proposed algorithm. The proposed structures are found to offer significant saving of hardware, require less latency, and yield more throughput over the existing structures. Apart from simplicity and regularity, the proposed structures would also have flexibility of implementation by CORDIC circuits and canonical-signed-digit-based multipliers as well Pramod Kumar Meher |
IEEE Trans. Circuits Syst. Video Technol. | 1 |