EDBT 2026 Demo / reviewers in the wild / expert
M. B. Srinivas
dblp:38/5326 · also Mandalika B. Srinivas, Srinivas B. Mandalika, Srinivas Bala Mandalika
· DBLP profile ↗
30ranked-venue papers
0as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5Computer networks · 3Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Architecture slack exploitation for phase classification and performance estimation in server-class processors
Diyanesh Chinnakkonda, Karthick Rajamani, M. B. Srinivas |
J. Parallel Distributed Comput. | 3 |
| 2022 | Techniques to Improve Write and Retention Reliability of STT-MRAM Memory SubsystemabstractSpin transfer torque magneto-resistive random-access memory (STT-MRAM) has many advantages, such as scalability, persistence, practically infinite endurance, and fast access speed, that make it a promising and emerging technology for memory. However, this technology has multiple reliability issues, such as read and write reliability, higher write power, and long write latency, etc. At elevated temperatures, these issues exacerbate further. As the temperature increases massively in the latest compute nodes, we need to study and understand the effect of temperature on STT-MRAM memory writes and reliability. In this article, we propose the temperature-aware memory controller (MC) and device architecture techniques specific to STT-MRAM technology, which can improve write reliability, retention reliability, and memory power without sacrificing the performance. Our simulation results show that the proposed techniques cumulatively improve the write bit error rate (BER) on an average by$603\times $, increase retention reliability by 65%, along with 27% power reduction and 5.8% improved system performance over the baseline STT-MRAM-based memory subsystem. Saravanan Sethuraman, T. Venkata Kalyan, M. B. Srinivas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | A 12-bit, 1.1-GS/s, Low-Power Flash ADCabstractIn this article, an efficient architecture for a low-power, high-resolution flash analog-to-digital converter (flash ADC) is presented. It operates at 12-bit resolution with a sampling frequency of 1.1 GS/s. The architecture is a segmented one consisting of three subflash ADCs that we call SADC1, SADC2, and SADC3. During the operation, SADC1 detects the MSB bit, and SADC2 detects the intermediate bits, while SADC3 detects the finer bits of the flash ADC. Furthermore, SADC1 has a 1-bit resolution, SADC2 has a 6-bit resolution, and SADC3 has a 5-bit resolution. The binary operation related to 12-bit resolution is achieved by combining the subflash ADCs, thermometer digital outputs, and an encoder. The ADC has been fabricated in Global Foundry (GF) 65-nm standard CMOS process. The measured performance parameters show a differential nonlinearity (DNL) of ±0.24 LSB, an integral nonlinearity (INL) of ±0.45 LSB, a signal-to-noise and distortion ratio (SNDR) of 64.55 dB, a spurious-free dynamic range (SFDR) of 73.9 dB, and an effective number of bits (ENOB) of 10.43 bits. Also, power consumption is found to be 15.10 mW at 1.1-GS/s sampling frequency and 1.2-V supply voltage. The ADC achieves an FOM of 9.95 fJ/conversion-step (c-s) at the Nyquist input frequency and occupies a core area of 0.084 mm2. Mahesh Kumar Adimulam, M. B. Srinivas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Temperature Aware Adaptations for Improved Read Reliability in STT-MRAM Memory SubsystemabstractSpin-transfer torque magneto-resistive random-access memory (STT-MRAM) is an exciting new emerging technology, being considered as a strong candidate to fill the gaps in the existing memory hierarchy between DRAM and the secondary memory. STT-MRAM has adequate endurance. However, unresolved write switching and read reliability issues still exist at the functional operating temperature corners. One biggest challenge is that the read bit error rate (RBER) is not at an acceptable level for system reliability across the wide operating temperature range. We present an STT-MRAM memory subsystem that is fully compatible with existing DDR-based DIMM designs and evaluate read disturb and read sense bit-error rate (BER) under various operating temperature conditions. We propose temperature aware adaptive techniques for reliable reads at the rank level. The proposed temperature adaptation technique improves overall reliability of the DDR4 STT-MRAM-based memory subsystem with an optimal read current considering an acceptable 64-byte cacheline BER. Our full system simulations show 1000× order of improvements toward a cell raw read disturb BER along with 5% reduction in memory power and less than 1% impact on overall system performance. Saravanan Sethuraman, T. Venkata Kalyan, Karthick Rajamani, Chitra K. Subramanian, Kyu-Hyoun Kim, Hillery C. Hunter, M. B. Srinivas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2020 | A 7-Cell, Stackable, Li-Ion Monitoring and Active/Passive Balancing IC With In-Built Cell Balancing Switches for Electric and Hybrid VehiclesabstractElectric vehicles (EVs) and hybrid EVs need stacked lithium-ion (Li-ion) cells to achieve the required high voltage (HV). Cell monitoring and balancing in a stackable battery system is necessary to compensate for accumulative discharge mechanisms and keep the individual cells in the same state of charge. In this article, an integrated circuit consisting of a 12-bit successive approximation register (SAR) analog-to-digital converter is designed to measure, monitor, and balance Li-ion cell voltages with a measured accuracy of ±7 mV. These stacked cells are compared simultaneously with a reference voltage to balance the cells. Balancing switches for charging/discharging of the cells are integrated within the circuit to support a balancing current of up to 100 mA, which reduces the number of external components for balancing significantly. The circuit supports both active and passive balancing. A synchronous voltage mode level shifting circuit is implemented for communication between the stacked integrated circuits (ICs) to eliminate external components. Internal linear drop outs (LDOs) (3 and 5 V) power most of the blocks in the IC. The design is fabricated in HV 0.35 μm complementary metal oxide semiconductor (CMOS) technology and found to consume a quiescent current of 17 μA. Veeresh Babu Vulligaddala, Sandeep Vernekar, Sudhakar Singamala, Ravikumar Adusumalli, Vijay Ele, Manfred Brandl, M. B. Srinivas |
IEEE Trans. Ind. Informatics | 7 |
| 2018 | Improved designs of digit-by-digit decimal multiplier
Syed Ershad Ahmed, Ch. Santosh Varma, M. B. Srinivas |
Integr. | 3 |
| 2017 | Optimizing the Reversible Circuits Using Complementary Control Line Transformation
Sai Phaneendra P., Chetan Kumar Vudadha, M. B. Srinivas |
RC | 3 |
| 2017 | An ESOP Based Cube Decomposition Technique for Reversible Circuits
Sai Phaneendra P., Chetan Kumar Vudadha, M. B. Srinivas |
RC | 3 |
| 2017 | A low power, programmable bias inverter quantizer (BIQ) flash ADCabstractIn this paper, a low power, programmable bias inverter quantizer (BIQ) flash ADC for communication and biopotential signal processing applications is presented. The comparator of the proposed BIQ flash ADC is designed by using the digital inverter with cascode PMOS and NMOS bias transistors in the top and bottom. The voltage range of the bias transistors will control different switching threshold values of the inverter. The BIQ flash ADC increases the operating frequency due to reduced input gate load capacitance, while the area and power consumption of BIQ ADC are reduced due to smaller device dimensions compared to conventional comparator based flash ADC and inverter based ADCs. The BIQ ADC is programmable for 5-bit, 6-bit, 7-bit and 8-bit resolutions by using control bits which scales the power consumption based upon ADC operation. The BIQ flash ADC results are compared with conventional comparator based flash ADC and inverter based ADC for different resolutions. The proposed ADC is designed in 90nm standard CMOS process which occupies a core area of 0.0676 mm2. The performance parameters of the BIQ flash ADC design are found to be, differential non-linearity (DNL) of ±0.30 LSB, integral non-linearity of ±0.51 LSB, signal-to-noise-and-distortion ratio (SNDR) of 46.18 dB, effective number of bits (ENOB) of 7.38 at 1.0 V supply voltage. The power consumption of this ADC at DC/lower input frequencies is 220 nW and at higher input frequencies upto 2 GHz it is 18.5 mW. Mahesh Kumar Adimulam, Krishna Kumar Movva, Amit Kapoor, M. B. Srinivas |
VLSI-SoC | 4 |
| 2016 | An Iterative Logarithmic Multiplier with Improved PrecisionabstractRecent studies have demonstrated the potential for achieving higher area and power saving with approximate computation in error tolerant applications involving signal and image processing. Multiplication is a major mathematical operation in these applications which when performed in logarithmic number system results in faster and energy efficient design. In this paper, the authors present a method which combines the Mitchell's approximation and hardware truncation scheme in a novel way resulting in an iterative multiplier with improved precision and area. Further, proposed truncation approach and fractional predictor significantly reduce the overall hardware requirement of the multiplier. Experimental results prove the superiority of the proposed multiplier over previous designs. Syed Ershad Ahmed, Sanket Kadam, M. B. Srinivas |
ARITH | 3 |
| 2011 | A Unified Architecture for BCD and Binary Adder/SubtractorabstractThe need to have hardware support for decimal arithmetic is increasing in recent years because of the growth in the decimal data processing in commercial, financial and internet based applications. In this paper a new architecture for efficient Binary coded decimal (BCD) addition/subtraction is presented that can be reconfigured to perform binary addition/subtraction. The architecture is mainly designed, keeping in mind the signed magnitude format. The proposed architecture avoids the usage of additional 2's complement and 10's complement circuitry, for correcting the results to sign magnitude format. The architecture is run-time reconfigurable to facilitate both BCD and Binary operations. Simulation results show that the proposed architecture is 13.6% better in terms of delay than the existing design. Chetan Kumar V., Sai Phaneendra P., Syed Ershad Ahmed, Sreehari Veeramachaneni, N. Moorthy Muthukrishnan, M. B. Srinivas |
DSD | 6 |
| 2010 | A low power, variable resolution two-step flash ADCabstractIn this paper, a new low power and configurable resolution two step flash ADC is proposed. Comparators of conventional flash ADC are replaced with CMOS inverters whose threshold can be varied dynamically. A novel peak-detector circuit is employed to achieve variable resolution as well as to switch the unused parallel inverters to standby mode. Linear reduction in resolution leads to exponential reduction in power. The ADC is capable of operating at 8-bit, 10-bit, and 12-bit precision and at a supply voltage of 2.5V; it consumes 16mW at 12-bit, 12mW at 10-bit and 8mW at 8-bit resolution. The sampling frequency ranges from 0.5 to 1.0 GSPS, and the ADC has a DNL Mahesh Kumar Adimulam, Krishna Kumar Movva, Sreehari Veeramachaneni, N. Moorthy Muthukrishnan, M. B. Srinivas |
ACM Great Lakes Symposium on VLSI | 5 |
| 2010 | Effects of channel SNR in mobile cognitive radios and coexisting deployment of cognitive wireless sensor networksabstractIn this paper, we describe the cognitive radios sharing the spectrum with licensed users and its effects on operational coexistence with unlicensed users. Due to the unlicensed spectrum band growing needs and usage by many IEEE 802.11 protocols, normal wireless radio operation sees high interference leading to high error rates on operational environments. We study the licensed bands and the characteristics of the unlicensed bands in general and more specific to radio characterization of individual radios and cognitive deployment of sensor networks and its effect on lifetime. The cognitive radio signals detection algorithm for this probabilistic model for the unlicensed users, uses a mobility model which takes into account the threshold variable ratio Eb/Noand also calculates the lower-bound of the combined value of secondary user interference for overlapping frequencies with the primary user. By using simulation, we detect the primary user when the radio frequencies are known a priori and compare it when the frequencies are unknown. In our analysis we exploit the similarity measure seen at each sub-channel frequency, which are due to multiple paths of the same reflected signal by maximizing the correlated information of the correlation matrix. For the general case the co-variance matrix for blind source separation, we use ICA de-correlation methods and show that cognitive radio can efficiently identify users in complex situations. The effects of large deployment and cognitive sensor network are studied for a family of 802.15.4 radios adapting to power-aware algorithms. Vasanth Iyer, S. Sitharama Iyengar, Garimella Rama Murthy, Nandan Parameswaran, Dhananjay Singh 0001, M. B. Srinivas |
IPCCC | 6 |
| 2009 | Multi-hop scheduling and local data link aggregation dependant Qos in modeling and simulation of power-aware wireless sensor networksabstractIn this study of wireless sensor networks (WSN) protocols, the application Qos, system, and protocol performance metrics are measured for a large scalable wireless deployment using a typical wireless radio and an energy model. As there are many different types of WSN algorithms, we have categorized it into pro-active, re-active, and query driven information processing. A typical Qos is based on the useful lifetime of sensor nodes, after which reliability of the sensor data cannot be guaranteed and typically, a threshold such as a percentage of the sensor drains out of energy or a minimum through-put of real-time data from the sensor network is expected, which is used to compare the Qos of the routing algorithm. The results from lifetime based Qos, measured in simulation seconds, for the implemented protocols show that with varying sampled data sources for a BE Qos multi-hop deployment and varying percentage of cluster heads in a time- synchronized deployment, the lifetime is based on network size and protocol invariant. However, low sensing ranges result in dense networks, and therefore, it becomes necessary to achieve an efficient medium-access protocol subjected to power constraints. Scalability of sensor network applications are based on energy energy-harvesting techniques in which the various layers of the network inter-operate and extend the system network lifetime, the battery residual power per node, and the application reliability in terms of cross-layer energy savings. In this study, we have extended the lifetime metrics from a constant metrics into a break down of how much percentage of time is spent for Tx, Rx, and Idle tasks, respectively. This helps one to highlight the cross-layer energy dissipation per node and how the performance of an algorithm differs in terms of duty-cycling. Furthermore, we have shown that the energy savings due owing to the distributed algorithms in a large sensor network will not be practical without a complimentary lower-layer MAC. We show have demonstrated that the Qos is very much related to the ambient conditions, namely, are the Rx and Idle modes. From these preliminary results, we have added a new category of WSN protocols which are based ion the renewable energy resources, namely, the Fusion Ambient Renewable Measuring Sensors (FARMS). The study of sensor FARMS -harvesting applications allows one to measure the impact on Idle, Sleep, and renewable energy cycles as well as their unique deployment (density) needs, as all the sensor are not active(Rx) at all times. We have also shown that the efficiency of cross-layer Qos performance of routing algorithms with MAC losses has a long tail which is similarly observed in Power Law. In this sensor network model we like to show the complexity of clustering, messaging and data rate in terms of O(√(N) log N), O(N) and O(log2N) where N is the number of nodes. Vasanth Iyer, S. Sitharama Iyengar, Garimella Rama Murthy, Bertrand Hochet, Vir V. Phoha, M. B. Srinivas |
IWCMC | 6 |
| 2008 | Generic sub-space algorithm for generating reduced order models of linear time varying vlsi circuitsabstractWe present an algorithm for reducing large VLSI circuits to much smaller ones with similar input-output behavior. A key feature of our method, called generic subspace, is that it is capable of reducing linear time-varying systems. This enables it to capture frequency-translation and sampling behavior, important in communication subsystems such as mixers, RF components and switched-capacitor filters. Reduction is obtained by projecting the original system described by linear differential equations into subspace of a lower dimension. Experiments have been carried out using Cadence Design Simulator which indicates that the proposed sub-space technique achieves more % reduction with less CPU time than the other model order reduction techniques existing in literature. We also present applications to RF circuit subsystems, obtaining size reductions and evaluation speedups of orders of magnitude with insignificant loss of accuracy. J. V. R. Ravindra, M. B. Srinivas |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | A Statistical Model for Estimating the Effect of Process Variations on Delay and Slew Metrics for VLSI InterconnectsabstractFor optimizations like placement, interconnect synthesis and static timing analysis, efficient interconnect delay computation is critical for RC networks. Because of its simple closed form and fast evaluation, the Elmore delay model has been widely used. The other delay metrics PRIMO and H-gamma match the first three circuit moments to the probability density function (PDF) of a Gamma statistical distribution. Although these methods demonstrate impressive accuracy compared to other delay metrics, their implementations tend to be challenging. In this paper simple and efficient two-parameter analytic expressions for both delay and slew, based on Erlang distribution (ERD) function, are presented under process variations. The effectiveness of the proposed metrics for RC trees is proved through experimental results. J. V. R. Ravindra, M. B. Srinivas |
DSD | 2 |
| 2007 | Area Efficient High Speed Architecture of Bruun's FFT for Software Defined RadioabstractFast Fourier Transform (FFT) is one of the most basic and essential operation performed in software defined radio (SDR). Therefore designing a universal, reconfigurable FFT computation block with low area, delay and power requirement is very important. Recently it is shown that Bruun's FFT is ideally suited for SDR even when operating with higher bit precision to maintain same NSR. In this paper, authors have proposed a new architecture for Bruun's FFT using a distributed approach for incrementing the number of bits (precision) with successive stages of FFT. It is also shown that proposed architecture further reduces the hardware requirement of Bruun's FFT with negligible changes in it's NSR. The proposed design makes Bruun's FFT, a better option for most practical cases in SDR. A detailed comparison of Bruun's traditional and proposed hardware architectures for same NSR is carried out and results of FPGA and ASIC implementations are provided and discussed. Shashank Mittal, Mohammed Zafar Ali Khan, M. B. Srinivas |
GLOBECOM | 3 |
| 2007 | Bus-encoding technique to reduce delay, power and simultaneous switching noise (SSN) in RLC interconnectsabstractInductance effects cannot be neglected in global interconnect lines as well as in circuits operating at higher frequencies. This paper presents a new spatio-temporal bus-encoding technique to minimize simultaneous switching noise as well as reduce delay and power dissipation in on-chip buses where inductance effects are dominating. Simulation experiments are carried out to find out the delay and SSN reduction for interconnect lines of different lengths (2mm, 5mm and 10mm) at various technology nodes (180nm, 130nm, 90nm and 65nm). Results obtained show that that the proposed bus-encoding scheme provides a delay reduction of about 54% to 73% with respect to the worst case delay. In addition, encoding is combined with wire shaping and its impact on further delay reduction is observed to be 4% to 26%. Further, when encoding was combined with wire shaping and repeater insertion, an additional delay reduction of 9% to 33% is observed. Concerning SSN, the encoding scheme is tested with various SPEC'95 benchmarks and it is found that SSN is reduced by about 33% on an average compared with the un-encoded data. Finally, energy minimization of about 13% on an average is achieved by the application of new spatio-temporal encoding scheme as reflected by the SPEC'95 bench mark tests. Chittarsu Raghunandan, K. S. Sainarayanan, M. B. Srinivas |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Novel architectures for efficient (m, n) parallel countersabstractParallel counters are key elements in many arithmetic circuits, especially fast multipliers. In this paper, novel architectures and designs for high speed, low power (3,2), (7,3), (15,4) and (31,5) counters capable of operating at ultra-low voltages are presented. Based on these counters, a generalized architecture is derived for large (m, n) parallel counters. The proposed architecture lays emphasis on the use of multiplexers and a combination of CMOS and transmission gate logic in arithmetic circuits that result in high speed and efficient design. The proposed counter designs have been compared with existing designs and are shown to achieve an improvement of about 45% in delay and a reduction of about 25% in power consumption. Sreehari Veeramachaneni, Lingamneni Avinash, Kirthi M. Krishna, M. B. Srinivas |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Area Efficient Bus Encoding Technique for Minimizing Simultaneous Switching Noise (SSN)abstractInductance effects cannot be neglected in circuits with higher operating frequencies. With shrinking technologies, as spacing between the interconnect lines decreases, simultaneous switching noise (SSN) or M*di/dt noise increases due to inductive coupling. The encoding techniques for minimizing crosstalk considering only RC effects are not suitable for RLC circuits. In this paper we propose a new bus encoding technique to minimize the simultaneous switching noise and delay in RLC interconnect lines. Hardware implementation details are given for encoder and decoder. Spice simulations are carried out for delay analysis on 2 mm and 5 mm interconnect lines at various technology nodes (130 nm, 90 nm and 65 nm). Proposed encoding scheme is tested with various SPEC'95 benchmarks and it is found that SSN is reduced by about 31% on an average compared to the un-encoded data Chittarsu Raghunandan, K. S. Sainarayanan, M. B. Srinivas |
ISCAS | 3 |
| 2007 | Novel High-Speed Redundant Binary to Binary converter using Prefix NetworksabstractFast addition and multiplication are of paramount importance in many arithmetic circuits and processors. The use of redundant number system for efficient implementation of these operations has been widely discussed in literature. A redundant binary to binary converter lies directly within the critical path of any operations in this number system, thereby dictating the performance of the overall circuit. In this paper, a new redundant binary to binary converter is proposed using the logic of prefix adders. Though carry propagation is still present in the proposed implementation, the latency has been reduced to O (log n) by the use of sparse-tree networks. The architecture of the proposed converter has been compared (both qualitatively as well as quantitatively) with the existing designs and is shown to achieve an efficiency of 52% in the overall delay and reduction of 36% in power-delay product. Sreehari Veeramachaneni, Kirthi M. Krishna, Lingamneni Avinash, Reddy Puppala Sreekanth, M. B. Srinivas |
ISCAS | 5 |
| 2007 | A bit-sliced, scalable and unified montgomery multiplier architecture for RSA and ECCabstractThis paper presents a reconfigurable, bit-sliced, scalable Montgomery multiplier architecture which can operate in both prime and binary fields, that is, GF(p) and GF(2n). It can be configured for any bit length thus making it applicable for emerging elliptic curve cryptography (ECC) as well as widely used RSA cryptosystems. Existing word-based, scalable multiplier architectures perform well for key sizes in RSA (but not ECC) as they result in higher computational time. Limited utility of word-based architectures for ECC precisions, which are in general not equal to an integer multiple of word-size, is discussed and a new bit-sliced architecture to improve the performance in terms of delay is proposed. The new bit-sliced, scalable architecture computes the Montgomery multiplication with fewer clock cycles compared to existing architectures by configuring them at bit-level rather than at word-level, without compromising on the performance. Synthesis results (Mentor Graphic’s Leonardo Spectrum) are compared with that of other scalable architectures and discussed. M. Sudhakar, Ramachandruni Venkata Kamala, M. B. Srinivas |
VLSI-SoC | 3 |
| 2007 | New and Improved Architectures for Montgomery Modular Multiplication
M. Sudhakar, Ramachandruni Venkata Kamala, M. B. Srinivas |
Mob. Networks Appl. | 3 |
| 2006 | Low Power Hierarchical Multiplier and Carry Look-Ahead ArchitectureabstractThis paper proposes a novel 8x8 multiplier architecture based on Wallace Tree, efficient in terms of power and regularity without significant increase in delay and area. The idea involves the generation of partial products in parallel using AND gates. The addition of these partial products is done using Wallace Tree which is hierarchically divided into levels. There will be a significant reduction in the power consumption, since power is provided only to the level that is involved in computation and thereby rendering the remaining two levels switched off (by employing a control circuitry). Furthermore, to improve the speed of addition at the 3rd level of computation, a novel carry look-ahead adder (CLA) is also proposed which is better than the recently proposed CLA architecture when compared its efficiency in terms of area/speed. The efficiency of the proposed multiplier is also tested by embedding it in higher width partition multipliers. Himanshu Thapliyal, Gopi Neela, K. K. Pavan Kumar, M. B. Srinivas |
AICCSA | 4 |
| 2006 | Novel Reversible Multiplier Architecture Using Reversible TSG GateabstractIn the recent years, reversible logic has emerged as a promising technology having its applications in low power CMOS, quantum computing, nanotechnology, and optical computing. The classical set of gates such as AND, OR, and EXOR are not reversible. Recently a 4 * 4 reversible gate called “TSG” is proposed. The most significant aspect of the proposed gate is that it can work singly as a reversible full adder, that is reversible full adder can now be implemented with a single gate only. This paper proposes a NXN reversible multiplier using TSG gate. It is based on two concepts. The partial products can be generated in parallel with a delay of d using Fredkin gates and thereafter the addition can be reduced to log2N steps by using reversible parallel adder designed from TSG gates. A 4x4 architecture of the proposed reversible multiplier is also designed. It is demonstrated that the proposed multiplier architecture using the TSG gate is much better and optimized, compared to its existing counterparts in literature; in terms of number of reversible gates and garbage outputs. Thus, this paper provides the initial threshold to building of more complex system which can execute more complicated operations using reversible logic. Himanshu Thapliyal, M. B. Srinivas |
AICCSA | 2 |
| 2006 | An Efficient Reconfigurable Montgomery Multiplier Architecture for GF(n)abstractIn this paper, the authors propose an efficient reconfigurable Montgomery multiplier for Galois prime field GF(n) that employs carry-save addition. The multiplier can operate for any operand length 'k' where 1 Ramachandruni Venkata Kamala, M. Sudhakar, M. B. Srinivas |
DSD | 3 |
| 2006 | A novel, coupling driven, low power bus coding technique for minimizing capacitive crosstalk in VLSI interconnectsabstractIn current VLSI technology, interconnects have become the predominant source of power dissipation. Particularly in DSM technology, the spacing between interconnects is very less leading to the dominance of coupling capacitance over self capacitance. In 0.13 /spl mu/m technology, it has been found that 75% of the power consumption is due to the coupling capacitance whereas only 25% is due to self capacitance. Thus, earlier schemes which concentrated on minimizing the substrate capacitances are not valid in these buses. Considering these aspects, this paper proposes a novel encoding technique that reduces the switching activity due to capacitive coupling resulting in lower power overhead. The simulation results show that the power dissipation in a bus is reduced by about 23% with this encoding technique. The encoding and decoding circuits have been designed and the power consumed has been compared with those of other coding schemes. K. S. Sainarayanan, J. V. R. Ravindra, M. B. Srinivas |
ISCAS | 3 |
| 2006 | High-Throughput Montgomery Modular MultiplicationabstractThe efficiency of public key encryption schemes like RSA and elliptic curve cryptography can be improved using fast modular multiplication schemes. In this paper, the authors propose an efficient Montgomery modular multiplication technique that employs multi-bit shifting and carry-save addition to perform long-integer arithmetic and hence conventional lengthy additions required at each stage are avoided. The corresponding hardware realization is optimal in terms of delay and offers high data throughput compared to the recently proposed designs while it occupies slightly more area. The optimization is technology independent and thus should suit well for not only FPGA implementation but also ASIC. The design has been evaluated on Virtex2 series FPGA for practical bit lengths of 512,1024 and 2048 bit Ramachandruni Venkata Kamala, M. B. Srinivas |
VLSI-SoC | 2 |
| 2003 | Speech encoding and encryption in VLSIabstractAbstract-In this work, an attempt has been made to design and synthesize speech encoding and encryption as a system-on-chip. The novelty of this design is that it uses wavelet decomposition for data compression and perpetual audio masking to keep quantization noise level to a minimum. The encryption is done by implementing RSA algorithm in hardware. 1. K. Kalyan Chakravarthy, M. B. Srinivas |
ASP-DAC | 2 |
| 2003 | Design of a digital CDMA receiverabstractIn this work, an attempt has been made to design a Digital Signal Code Division Multiple Accesses Receiver using VHDL [1] [2] and synthesize using Mentor Graphics tools. The receiver is designed for a maximum of three users as per the design specification given for the contest. The novelty of this design is that it uses four independent processes to achieve the concurrent operation for the entire functionality. The simulation and synthesis are carried out using Mentor Graphics' FPGA tools and an optimized circuit in terms of area and timing is brought out. I. Vijay Kumar, M. B. Srinivas |
ASP-DAC | 2 |