EDBT 2026 Demo / reviewers in the wild / expert
Poras T. Balsara
dblp:93/947
· DBLP profile ↗
40ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0003-1263-787XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adaptive PLL-based Sensorless Control for CSI-Fed PMSM Drives Used in Submersible PumpsabstractThis study examines the critical role of precise rotor position estimation in field-oriented control (FOC) of current source inverter (CSI)-fed Permanent Magnet Synchronous Motor (PMSM) drives for submersible pump applications. Despite advancements, challenges such as inverter nonlinearity and parameter variations introduce significant errors in rotor position estimation under dynamic and steady state conditions. This adversely affects the reliability and efficiency of the drive system. The proposed adaptive PLL aims to optimize PMSM acceleration by dynamically adjusting PLL bandwidth, effectively mitigating estimation errors caused by various non-ideal factors in steady state condition. Simulation and experimental results validate the effectiveness of the proposed approach. Milad Bahrami-Fard, Majid Ghasemi Korrani, Mohammad Rastegar, Poras T. Balsara, Babak Fahimi |
IECON | 4 |
| 2024 | Bifurcation Analysis under Various Modulation Techniques in PWM-Controlled DC-DC ConvertersabstractThe type of carrier signal waveform plays a pivotal role in the dynamic response of the PWM-controlled DC-DC converters. To fully understand the nonlinear behavior of DC-DC converters, this paper aims to dissect the bifurcation process in a PWM-controlled buck converter under various modulations. To this end, various carrier signal waveforms, including single-edge modulation, symmetrical and asymmetrical dual-edge modulations, and various phase-shifted PWMs are considered. Next, using discrete-time modeling, the dynamic analysis of the system at various phase shifts is performed. Inspection of the bifurcation diagram and eigenvalue trajectories, obtained by shifting the gate signal within the switching period, reveals that the converter experiences period-1, period-2, and period-4 orbits within one switching period. The results of this analysis are validated through simulation studies and experiments. Mohammad Hassan Ghaderi, Milad Bahrami-Fard, Nasim Rashidirad, Babak Fahimi, Poras T. Balsara |
IECON | 5 |
| 2024 | Dynamic Modeling of a Double-Stator Switched Reluctance Motor (DSSRM) Using a Lumped Parameter Circuit ModelabstractThis paper introduces a dynamic lumped circuit model for a Double-Stator Switched Reluctance Motor (DSSRM). The inductance dependency on rotor position is modeled using a truncated Fourier series, with the coefficients derived from measurements taken at aligned, unaligned, and midway rotor positions. This approach allows for precise modeling of the inductance variation with rotor position, enhancing the accuracy and performance of the DSSRM simulation, and the key advantage of the proposed model is its ability to predict the entire dynamic performance of the drive system while requiring minimal measurements. The model’s accuracy is validated by comparing its results with those obtained from Finite Element Analysis, demonstrating the reliability and effectiveness of the proposed modeling technique. Behnam Mosammam, Milad Bahrami-Fard, Babak Fahimi, Poras T. Balsara |
IECON | 4 |
| 2023 | Fully Distributed Control of Microgrids Using Multi-Agent ApproachabstractIn an islanded microgrid, the main objective of a distribution operator is to manage the generation load balance based on the available resources. With the rapid integration of renewable energy and low inertia-based resources in today's power grid, it is becoming challenging to achieve optimal operation by just using conventional control techniques. This work aims to create a fully distributed control system that can efficiently manage the functioning of a microgrid (MG). The MG system should be able to maintain nominal frequency and voltage as well as synchronize between the inverters and power grid autonomously. To achieve this, we implement a hybrid control architecture employing distributed implementation with a hierarchy via a Multi-Agent System (MAS). Additionally, we propose a secondary control for frequency restoration and a control algorithm for autonomous load connection/disconnection. To develop the microgrid model under study MATLAB Simulink was used, and the control architecture was validated using several test cases through controller-in-the-loop hardware testing (C-HIL) using Raspberry Pi as an external controller. Vaibhav Uttam Pawaskar, Poras T. Balsara, Babak Fahimi, Ghanshyamsinh Gohil |
IECON | 2 |
| 2022 | A new submodule structure with parallel capacitor connection in modular multilevel convertersabstractThere have been multiple attempts to improve the submodule (SM) structure, the fundamental block in Modular Multilevel Converter (MMC). The focus has been mainly on submodule voltage ripple reduction, system efficiency improvement, and short circuit blocking capabilities. This paper is one such attempt and proposes a new submodule structure to achieve some of the above-mentioned features. The proposed submodule structure is a three-level topology with two capacitors and five switches. There is a provision in the proposed submodule to connect these two capacitors in parallel during the intermediate voltage level, which can be advantageous in reducing the submodule voltage ripple. The proposed submodule can be seen as a replacement for two Half-Bridge submodules. The lesser number of devices and the parallel paths during an intermediate voltage level reduce the total device losses. The simulation validation of the proposed submodule structure is presented along with a detailed comparison with the existing submodule topologies. G. Veera Bharath, Ghanshyamsinh Gohil, Poras T. Balsara |
ISCAS | 3 |
| 2017 | Portable impedance measurement device for sweat based glucose detectionabstractThe future of disease diagnostics and health care wearables lies in the development of low-cost sensors that can detect minute traces of pathogens or antigens from body fluids. Developments in nanotechnology and biomedical research have already shown us that a nanosensor can be specifically tailored to detect a specific biomolecule. These sensors would allow patients to run point of care diagnostic tests, thereby saving time and cost of running clinical tests and can give early stage disease diagnosis and help physicians to provide personalized treatment. This work involves the development of a configurable electronic sensor platform that will interface with these sensors. The device is tested by quantification of glucose from sweat using a nanosensor developed in the Biomedical Microdevices and Nanotechnology Lab in the University of Texas at Dallas. The platform can be easily configured to run Electrochemical Impedance Spectroscopy based detection test for other biomolecules by using sensor tailored for it. Athul Asokan Thulasi, Dinesh Bhatia, Poras T. Balsara, Shalini Prasad |
BSN | 3 |
| 2016 | Effect of sampling time and sampling instant on the frequency response of a boost converterabstractDC-DC converters like boost, buck-boost and flyback converters have unstable internal dynamics and are classified as non-minimum phase systems. The linear approximation of these DC-DC converters around the equilibrium point shows a zero that maps to the open right-half plane (RHP) of the complex s-plane. This paper models the boost converter using State Space Averaging (SSA) and Enhanced State Space Averaging (ESSA) to show changes in bode response as function of sampling time. Further, we use Discrete Modeling Analysis (DMA) to model the changes in transfer function with change in sampling instant when sampled during off-time of the switch. The result shows that significant bandwidth improvement can be achieved when sampling instant is rightly chosen during the off-time. Such a choice is easily achieved with a digital implementation for a DC-DC converter. Sameer Arora, Poras T. Balsara, Dinesh K. Bhatia |
IECON | 2 |
| 2016 | A Wideband Digital-to-Frequency Converter with Built-In Mechanism for Self-Interference Mitigation
Imran Bashir, Robert Bogdan Staszewski, Oren E. Eliezer, Poras T. Balsara |
J. Electron. Test. | 4 |
| 2012 | Alien crosstalk mitigation in vectored DSL systems for backhaul applicationsabstractThe performance of digital subscriber line (DSL) systems, such as ADSL and VDSL is limited by crosstalk. Suppression of in-domain far-end self crosstalk using vectoring technology enables very high bidirectional data rates over twisted-pairs of copper wires. However, the performance of vectored DSL systems is severely degraded in the presence of alien or out-of-domain crosstalk that arises from sources that lie outside the vectored DSL system and share the same cable binder. In this paper, we propose a practical, non-iterative, high performance algorithm for alien crosstalk mitigation. Simulation results and complexity analysis corresponding to a vectored VDSL2 system in presence of alien crosstalk are presented to illustrate the significant performance gains of the proposed algorithm and its low implementation complexity. Aditya Awasthi, Naofal Al-Dhahir, Oren E. Eliezer, Poras T. Balsara |
ICC | 4 |
| 2011 | Multi-clock domain analysis and modeling of all-digital frequency synthesizersabstractAll-digital phase-locked loops (ADPLLs) are inherently multi-rate systems with time-varying behavior. In support to this statement accurate multi-clock domain models of ADPLL frequency synthesizers are presented. Their analytically derived phase transfer characteristics accurately predict second order effects (such as spectral aliasing) that are not captured using conventional modeling approaches. The results are validated through simulations using well accepted time-domain modeling techniques of an RF ADPLL. Ioannis L. Syllaios, Poras T. Balsara |
ISCAS | 2 |
| 2010 | A Generic Scalable Architecture for Min-Sum/Offset-Min-Sum Unit for Irregular/Regular LDPC DecoderabstractThe most common algorithm used in iterative decoding of low-density parity check (LDPC) codes is based on a generic class of the sum-product algorithm, which has a nonlinear dependence on the log(tanh()) function. The implementation based on fixed precision has substantial loss of accuracy and is computationally expensive with full precision. A suboptimal version of belief propagation called theoffset-min-sumalgorithm is generally used in hardware implementation. This paper proposes a generic scalable architecture for minimum search during check-node operation in theoffset-min-sumalgorithm applicable to regular as well as irregular LDPC codes with check node of any degreed. For an LDPC code with maximum check node degreed, the proposed architecture consists of 2(d-2) 2 × 1 multiplexers and 3(d-2) two-input compare-and-select units (CSUs). This has latency of [2⌈log2(d)⌉-2]tdcwhen ⌈log2(d)⌉-log2(d)2(4/3) else [2⌈log2(d)⌉-3]tdc, withtdcrepresenting the delay of a two-input CSU. The proposed architecture has been implemented ford= 20 using a TSMC 0.18-μm CMOS process. Venkata K. Kidambi Srinivasan, Chitranjan K. Singh, Poras T. Balsara |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | A Low Power and Low Quantization Noise Digital Sigma-Delta Modulator for Wireless TransmittersabstractDigital sigma-delta (ΣΔ) modulators are used extensively in CMOS wireless SoC designs to achieve high-resolution data conversion while controlling the quantization noise spectrum. This paper presents an implementation of a 90nm CMOS digital low-pass ΣΔ modulator, which has lower quantization noise and lower power consumption compared to other recent structures. A conventional digital ΣΔ structure uses a 1-bit quantizer and generates very high quantization noise at higher frequencies. In this work, we present a low-pass digital ΣΔ architecture with a multi-bit quantizer, which achieves very low in-band as well as out-of-band quantization noise levels. It is shown that the structure can be run at half the frequency while meeting the required noise performance and essentially delivering a better power-performance trade-off. The proposed architecture, along with the original 1-bit quantizer structure, has been synthesized in a 90nm CMOS process. Area and power consumption results are presented and a comparison between a commonly used structure and the proposed one is provided. Viral K. Parikh, Poras T. Balsara, Oren E. Eliezer, Jaimin Mehta |
ISCAS | 2 |
| 2007 | A Low Area and Low Power Digital Band-Pass Sigma-Delta Modulator for Wireless TransmittersabstractThe digital sigma-delta (ΣΔ) modulators are used extensively in CMOS wireless SoC designs to achieve high-resolution data conversion while controlling the quantization noise spectrum. This paper presents an implementation of a 90nm CMOS digital band-pass ΣΔ modulator, running at 900 MHz. The conventional band-pass ΣΔ structures required to achieve such noise shaping are hardware intensive and do not meet the timing requirements when synthesized in 90nm technology using a static CMOS implementation. The loop-unrolling concept was recently presented as a solution for getting required rate of operation. However, this approach requires larger area and consumes higher power. In this work, we present a new architecture to achieve the necessary noise shaping at required rate of operation. The proposed structure has the different Noise Transfer Function (NTF) compared to conventional band-pass structures and hence it gives more control over the location of zeros. It also meets timing requirements of 900 MHz across all PVT corners and shows significant saving in area and power. Viral K. Parikh, Poras T. Balsara, Oren E. Eliezer, Jaimin Mehta |
ISCAS | 2 |
| 2007 | Effect of Word-length Precision on the Performance of MIMO SystemsabstractIn this paper, we investigate the effect of quantization noise and roundoff errors involved in finite-precision (FP) signal processing on the performance of multiple-input multiple-output (MIMO) receivers under flat-fading channel conditions. We simulated the performance degradation with FP for six transmission schemes, namely the Alamouti transmit diversity scheme, spatial multiplexing (SM) under maximum-likelihood (ML) detection, SM under ordered successive interference cancellation (OSIC) for (2 times 2) MIMO, orthogonal space-time block code (STBC), quasi-orthogonal STBC for (4 times 2) MIMO and interference cancellation with two (2 times 2) Alamouti scheme. We also provide a procedure to analyze roundoff errors in FP MIMO receivers and quantify the effective decision-point SNR. The quantified variance of accumulated quantization noise and decision-point SNR for FP Alamouti scheme corroborates the simulation result. Finally, we provide the minimum word-length precision requirements for these schemes. Chitranjan K. Singh, Naofal Al-Dhahir, Poras T. Balsara |
ISCAS | 3 |
| 2007 | On the Reconfigurability of All-Digital Phase-Locked Loops for Software Defined RadiosabstractA new all-digital phase-locked loop (ADPLL) for wireless applications has recently been proposed and commercially demonstrated. It replaces conventional phase/frequency detector and charge pump with a time-to-digital converter (TDC). Analog frequency tuning of a VCO is replaced with an all- digital tuning of a digitally-controlled oscillator (DCO). Due to its digital intensive structure, the ADPLL is well suited for single-chip radio solutions fabricated using state-of-the- art low cost and power nanometer-scale CMOS processes. Being integrated with a digital signal processor (DSP), the ADPLL parameters can be properly controlled and seamlessly reconfigured using the available on-chip DSP unit making the ADPLL a software defined radio (SDR) platform. In this paper, we present a DSP based technique for the fully dynamic control of the ADPLL settling performance that allows the loop band width to be seamlessly widened or narrowed allowing for fast frequency acquisition or tracking with excellent phase noise and spurious performance, respectively. The arbitrary and dynamic control of the frequency synthesizer loop bandwidth will address the dynamically varying nature of a multi-radio multi- standard SDR environment. Ioannis L. Syllaios, Poras T. Balsara, Robert Bogdan Staszewski |
PIMRC | 2 |
| 2007 | Iterative (TURBO) IQ Imbalance Estimation and Correction in BICM-ID for Flat Fading ChannelsabstractTURBO principle has been exploited gainfully to implement many receiver functions. RF front-end impairments are a serious issue in high spectral efficient applications. IQ imbalance is one of these impairments and in this work, we study the issue of IQ imbalance correction using baseband signal processing techniques. In particular, we propose an estimation technique based on EM algorithm. Such a technique is developed rather intuitively for the case of a Bit Interleaved Coded modulation - Iterative Detection (BICM-ID) receiver for burst mode communications. The resulting TURBO IQ Decorrelator is embedded in the BICM-ID loop, is blind in the sense that it does not require any training symbols or tones. Performance is simulated for 64QAM under flat fading channel conditions. Raghunath Cherukuri, Poras T. Balsara |
VTC Fall | 2 |
| 2006 | Reconfigurable CAM Architecture for Network Search EnginesabstractA novel reconfigurable content addressable memory, called RCAM, is proposed that supports on-the-fly reconfiguration between CAM and TCAM. The area overhead of the proposed RCAM cell is only 5.6% when compared to conventional TCAM. This overhead is compensated by area saving due to removal of the priority encoder. Other features of our architecture include reconfigurability, and better overall performance and power. To achieve these we incorporated two novel techniques: (i) a hybrid CAM/TCAM architecture that allows user to pre-define CAM/TCAM cell behavior in each bit or word position and ultimately curtail the overall power consumptions of memory unit: and (ii) a wired-AND technique by which we can completely eliminate the sorting requirement and thus significantly reduce the update time. A 4 Kb RCAM architecture was implemented using 0.18 mum CMOS technology. The simulations indicate a search time of 6.15 ns, i.e. capability of handling about five OC-192 at wire speed. Mehrdad Nourani, Deepak S. Vijayasarathi, Poras T. Balsara |
ICCD | 3 |
| 2006 | A generalized signal reconstruction method for designing interpolation filtersabstractA generalized signal reconstruction method is presented for the development of digital FIR interpolation filters, applicable to an analog reconstruction filter impulse response with finite duration that extends to an integer number of the input sampling periods. The technique produces efficient interpolation filter structures for both polynomial and non-polynomial based impulse responses, and could minimize implementation requirements (area, power) for filter designs involving sampling rate conversion. Ioannis L. Syllaios, Poras T. Balsara, Oren E. Eliezer |
ISCAS | 2 |
| 2005 | Exploiting temporal idleness to reduce leakage power in programmable architecturesabstractOne of the biggest challenges that programmable devices like FPGAs are facing in ultra deep sub-micron regime is the exponential rise in leakage power consumption. As technology shrinks below 90nm, a new design paradigm has to evolve to tackle the issue of leakage power consumption. In this work we focus on a new design methodology for reducing leakage power by exploiting temporal locality in designs and accordingly group them into. clusters that can be switched on and off. We propose a Power State Controller based method, which controls the switching of the clusters from one state to another. We show our technique using Data Flow Graphs where temporal locality can be effectively explored. Our results show that substantial leakage savings can be achieved if temporal idleness of designs can be exploited effectively. Rajarshee P. Bharadwaj, Rajan Konar, Poras T. Balsara, Dinesh Bhatia |
ASP-DAC | 3 |
| 2005 | FPGA Architecture for Standby Power Management
Rajarshee P. Bharadwaj, Rajan Konar, Dinesh Bhatia, Poras T. Balsara |
FPT | 4 |
| 2005 | Ripple-Precharge TCAM A Low-Power Solution for Network Search EnginesabstractA novel low power ripple-precharge ternary CAM (RP-TCAM) architecture is proposed for applications in longest prefix matching tasks. The main motivation behind this research is to reduce the dynamic power consumption in TCAM due to frequent charging and discharging of the highly capacitive match line. This issue is addressed by exploiting the fact that when we compare only the first four bits of incoming packet's destination address we can identify up to 80% mismatches in the forwarding table. A selective precharge scheme was devised exploiting the above fact wherein the match line is charged only when there is an exact match in the first four bits of TCAM word, thereby significantly reducing the number of transitions in the match line. The parasitics for simulation were extracted from the layout implemented for a 64 /spl times/ 32 RP-TCAM architecture using 0.18/spl mu/m technology. This structure has 1.71% less area and 80% less power when compared to the conventional TCAM of equal storage size and functionality. Our RP-TCAM architecture has a search time of 1.86ns. Deepak S. Vijayasarathi, Mehrdad Nourani, Mohammad J. Akhbarizadeh, Poras T. Balsara |
ICCD | 4 |
| 2005 | SoC with an integrated DSP and a 2.4-GHz RF transmitterabstractWe present a system-on-chip (SoC) that integrates a TMS320C54x digital signal processor (DSP), which is commonly used in cellular phones, with a multigigahertz digital RF transmitter that meets the Bluetooth specifications. The RF transmitter is tightly coupled with the DSP and is directly mapped to its address space. The transmitter architecture is based on an all-digital phase-locked loop (ADPLL), which is built from the ground up using digital techniques and digital creation flow that exploit high speed and high density of a deep-submicrometer CMOS process while avoiding its weaker handling of voltage. The frequency synthesizer features a wideband frequency modulation capability. As part of the digital flow, the digitally controlled oscillator (DCO) and a class-E power-amplifier are created as ASIC cells with digital I/Os. All digital blocks, including the 2.4-GHz logic, are synthesized from VHDL and auto routed. The use of VHDL allows for a tight and seamless integration of RF with the DSP. To take advantage of the direct DSP-RF coupling and to demonstrate a software-defined radio (SDR) capability, a DSP program is written to perform modulation of the GSM standard. The chip is fabricated in a baseline 130-nm CMOS process with no analog extensions and features high logic gate density of 150 kgates per mm/sup 2/. The RF transmitter area occupies only 0.54 mm/sup 2/, and the current consumption (including the companion DSP) is 49 mA at 1.5-V supply and 4 mW of RF output. This proves attractiveness and competitiveness of the "digital RF" approach, whose goal is to replace RF functions with high-speed digital logic gates. Robert Bogdan Staszewski, Roman Staszewski, John L. Wallberg, Tom Jung, Chih-Ming Hung, Jinseok Koh, Dirk Leipold, Kenneth Maggio, Poras T. Balsara |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2004 | PCAM: A Ternary CAM Optimized for Longest Prefix Matching TasksabstractAn optimized ternary CAM concept is introduced for application in the longest prefix matching tasks of the Internet search engines. It employs w+1 RAM bits for a word of size w. A conventional TCAM needs 2w RAM bits for the same word size. Based on this concept an 8 bit prefix-CAM cluster is designed out of 9 SRAM bits, four of which merge to store a 32-bit IPv4 prefix. A complete prefix-CAM module employs 22% less transistors than a conventional TCAM, for equal storage size and equal functionality. We confirm the 22% area saving by implementing the layouts for prefix-CAM and TCAM words. Our design also reduces interconnect area by reducing address decode lines. Mohammad J. Akhbarizadeh, Mehrdad Nourani, Deepak S. Vijayasarathi, Poras T. Balsara |
ICCD | 4 |
| 2001 | Challenges in integrated CMOS transceivers for short distance wirelessabstractThis paper presents recent trends in the area of integrated CMOS transceiver design for short distance wireless applications. This application is characterized by very low-cost and low-power so-lutions. Current challenges and recent trends are described and digital-oriented design opportunities for increasing integration out-lined. Signal processing approaches applied to the front-end elec-tronics find an increasing emphasis and are extremely viable. 1. Khurram Muhammad, Robert Bogdan Staszewski, Poras T. Balsara |
ACM Great Lakes Symposium on VLSI | 3 |
| 2001 | Speed, power, area, and latency tradeoffs in adaptive FIR filtering for PRML read channelsabstractIn this paper, we describe area and power reduction techniques for a low-latency adaptive finite-impulse response filter for magnetic recording read channel applications. Various techniques are used to reduce area and power dissipation while speed and latency remain as the main performance criteria for the target application. The proposed parallel transposed direct form architecture operates on real-time input data samples and employs a fast, low-area multiplier based on selection of radix-8 premultiplied coefficients in conjunction with one-hot encoded bus leading to a very compact layout and reduced power dissipation. Area, speed, and power comparisons with other low-power implementation options are also shown. The proposed filter has been fabricated using a 0.18-/spl mu/m L-effective CMOS technology and operates at 550 MSamples/s. Trading off filter latency to improve speed is also discussed. Khurram Muhammad, Robert Bogdan Staszewski, Poras T. Balsara |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | Reconfigurable Array Media Processor (RAMP)abstractThis paper presents the architecture of a Reconfigurable Array Media Processor (RAMP). RAMP features a 2D array of coarse-grained configurable logic blocks (CLBs) connected together by local and global inter-connects. The CLBs on RAMP provide a 4-bit ALU, 2/spl times/2 bit parallel multiply function, 4-bit barrel shifter, two 4-bit registers and a local programmable control unit. RAMP is capable of partial run-time reconfiguration and supports block-mode reconfiguration. The novel features of this device include two programmable high-speed clocks available to each CLB, scalable parallel multiplier, on-chip memory/registers. RAMP can be used to implement high-performance computational kernals of video, audio and signal processing functions. Matrix multiplication, FIR filters and Inverse DCT functions are used as examples to demonstrate the capabilities of the RAMP architecture. Kamlesh Rath, Sirisha Tangirala, Patrick Friel, Poras T. Balsara, Jose Flores, John P. Wadley |
FCCM | 4 |
| 2000 | Low power techniques and design tradeoffs in adaptive FIR filtering for PRML read channelsabstractIn this paper, we describe area and power reduction techniques for a low-latency adaptive finite-impulse response filter for magnetic recording read channel applications. Various techniques are used to reduce area and power dissipation while speed remains as the main performance criterion for the target application. A parallel transposed direct form architecture operates on real-time input data samples and employs a fast, low-area multiplier based on selection of radix-8 pre-multiplied coefficients in conjunction with one-hot encoded bus leading to a very compact layout and reduced power dissipation. Area, speed and power comparisons with other low-power implementation options are also shown. The proposed filter has been fabricated using a 0.18 µm L-effective CMOS technology and operates at 550 MSamples/s. Khurram Muhammad, Robert Bogdan Staszewski, Poras T. Balsara |
ISLPED | 3 |
| 2000 | High-performance energy-efficient D-flip-flop circuitsabstractThis paper investigates performance, power, and energy efficiency of several CMOS master-slave D-flip-flops (DFF's). To improve performance and energy efficiency, a push-pull DFF and a push-pull isolation DFF are proposed. Among the five DFF's compared, the proposed push-pull isolation circuit is found to be the fastest with the best energy efficiency. Effects of using a double-pass-transistor logic (DPL) circuit and tri-state push-pull driver are also studied. Last, metastability characteristics of the five DFP's are also analyzed. Uming Ko, Poras T. Balsara |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | High performance low power array multiplier using temporal tilingabstractDigital multipliers are a major source power dissipation in digital signal processors. Array architecture is a popular technique to implement these multipliers due to its regular compact structure. High power dissipation in these structures is mainly due to the switching of a large number of gates during multiplication. In addition, much power is also dissipated due to a large number of spurious transitions on internal nodes. Timing analysis of a full adder, which is a basic building block in array multipliers, has resulted in a different array connection pattern that reduces power dissipation due to the spurious transition activity. Furthermore, this connection pattern also improves the multiplier throughput. This array pattern is based on creating a compact tiled structure, wherein the shape of a tile represents the delay through that tile. That is, a compact structure created using these tiles is nothing but a structure with high throughput. Such a temporal tiling technique can also be applied to other digital circuits. Based on our simulation studies, a temporally tiled array multiplier achieves 50% and 35% improvements in delay and power dissipation compared to a conventional array multiplier. Improvement in delay can be traded for power using voltage reduction techniques. Shivaling S. Mahant-Shetti, Poras T. Balsara, Carl Lemonds |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | LAPLUS: An Efficient, Effective and Stable Switch Algorithm for Flow Control of the Available Bit Rate ATM ServiceabstractLAPLUS is a novel switch algorithm for flow control of the available bit rate (ABR) asynchronous transfer mode (ATM) service. It ensures a steady-state rate allocation satisfying the MCR-plus-equal-share criterion. It only requires constant-time processing and two tags to be stored per flow. It is naturally able to take peak cell rates of flows into account. LAPLUS solves in a novel way the problem of selecting a measurement interval. The solution allows it to contain queue growth and keep utilization high on one hand and control low speed flows and operate with stability on the other. We describe results of simulation study of LAPLUS. The results show it to be fair, responsive and stable. Sharat Prasad, Kamran Kiasaleh, Poras T. Balsara |
INFOCOM | 3 |
| 1998 | Energy optimization of multilevel cache architectures for RISC and CISC processorsabstractIn this paper, we present the characterization and design of energy-efficient, on chip cache memories. The characterization of power dissipation in on-chip cache memories reveals that the memory peripheral interface circuits and bit array dissipate comparable power. To optimize performance and power in a processor's cache, a multidivided module (MDM) cache architecture is proposed to conserve energy in the bit array as well as the memory peripheral circuits. Compared to a conventional, nondivided, 16-kB cache, the latency and power of the MDM cache are reduced by a factor of 1.9 and 4.6, respectively. Based on the MDM cache architecture, the energy efficiency of the complete memory hierarchy is analyzed with respect to cache parameters in a multilevel processor cache design. This analysis was conducted by executing the SPECint92 benchmark programs with the miss ratios for reduced instruction set computer (RISC) and complex instruction set computer (CISC) machines. Uming Ko, Poras T. Balsara, Ashwini K. Nanda |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1996 | Design techniques for high performance, energy efficient control logicabstractThis paper investigates delay, power and area of critical components in designing energy-efficient control logic. To improve performance and energy efficiency, a split-slave dual-path (SSDP) register is proposed which improves the energy efficiency of the prior art by 30%. For multiplexers (MUX) three MUXes are proposed and compared to existing solutions. The proposed MUXes improve performance by 50% or power by 22%. The impact of scaling supply voltage alone and scaling threshold voltage with supply voltage on delay and power is also examined. Uming Ko, Anthony M. Hill, Poras T. Balsara |
ISLPED | 3 |
| 1996 | Leap frog multiplierabstractArray multipliers are popular due to their regular compact structure. Timing analysis of a full adder has resulted in a different array connection pattern that provides improved throughput for the multiplier while reducing its power dissipation from spurious transitions. The paper details the new array design and some results obtained. Shivaling S. Mahant-Shetti, Carl Lemonds, Poras T. Balsara |
ISLPED | 3 |
| 1995 | Hierarchy embedded differential image for progressive transmission using lossless compressionabstractAlgorithms for constructing differential images with hierarchical data structure are presented. The data structures are simple, efficient, and ideal for viewing images in progressive transmission using lossless compression. Unlike conventional pyramidal structures, the total number of nodes required to build the structure is the same as the number of pixels in an image at the same time its hierarchy is preserved. These structures are constructed using subsampling or mean-sampling methods for predictors with block sizes of 2/spl times/2 or 3/spl times/3. Experiments were conducted to compare these structures in terms of their first order entropy and RMS errors in the reconstruction process. Results indicate that the mean-sampling with circular-difference method yields the lowest entropy, comparable to that with 1-D lossless DPCM predictive coding. Lastly, hardware for the efficient construction and access of the hierarchical structures is discussed and evaluated.> Whoi-Yul Kim, Poras T. Balsara, David T. Harper III, Jon Wong Park |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | An architecture for a DSP field-programmable gate arrayabstractThis paper describes an application specific architecture for field-programmable gate arrays (FPGAs). Emphasis is placed on the logic module architecture and channel segmentation for the FPGAs targeted for application areas related to digital signal processing (DSP). The proposed logic module architecture is well-suited for efficient implementation of frequently used logic functions in the DSP application area. This is mainly because it is possible to implement most of these functions using one logic module, which results in a reduction in both the net lengths and the number of antifuses used. The performance improvements are achieved by customizing the logic module architecture and the programmable interconnect to suit the requirements of DSP applications.> M. Agarwala, Poras T. Balsara |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | Short-circuit power driven gate sizing technique for reducing power dissipationabstractOne major challenge in low-power technology is how to reduce overall power dissipation of a given subsystem without impacting its performance. In this paper we present a technique that can be applied to the nonspeed-critical nets in a circuit in order to reduce overall power dissipation. This technique involves a study of short-circuit power dissipation as a function of input signal slews and output load conditions, to aid in making a judicious choice of drive strengths for various gates in a circuit. The resulting low-power solution does not degrade the original performance and yields a circuit which occupies less silicon area. The technique described here can be incorporated into any power optimization or synthesis tool. Lastly, we present the savings in power and area for a 32-b carry lookahead adder which was designed using the technique described here.> Uming Ko, Poras T. Balsara |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1992 | Intermediate-level vision tasks on a memory array architecture
Poras T. Balsara, Mary Jane Irwin |
Mach. Vis. Appl. | 1 |
| 1991 | Digit Serial Multipliers
Poras T. Balsara, Robert Michael Owens, Mary Jane Irwin |
J. Parallel Distributed Comput. | 1 |
| 1987 | Systolic & semi-systolic digit serial multipliersabstractDigit serial data transmission can be used to an advantage in the design of special purpose processors where communication issues dominate and where digit pipelining can be used to maintain high data rates. VLSI signal processing is one such problem domain. We propose designs of systolic and semi-systolic digit serial multipliers. These multipliers are programmable i.e. one operand is pre-stored in the multiplier and the other operand is fed in a digit serial fashion. The VLSI implementation of the systolic multiplier is also given. This systolic multiplier is used in our VLSI signal processing system. Poras T. Balsara, Robert Michael Owens |
IEEE Symposium on Computer Arithmetic | 1 |
| 1986 | Design and implementation of real time video processorabstractThis paper describes the design and implementation of an arithmetic unit for a video filter. The central unit of the video filter consists of six identical chips called Common Arithmetic Unit (CAU's), each of which contains three Common Arithmetic Cells (CAC's). These 64-pin CAU's are assembled on a board in a pipelined architecture to realize real time performance. The throughput rate for the chip is 11.3 Mhz. A constant time pipelined adder design has been proposed and implemented. The absolute delay is stillO(\logn). The areaO(n\logn)and absolute delayO(\logn)for our adder are within a constant factor of the optimal bounds. Shishpal Rawat, Poras T. Balsara, Mary Jane Irwin, Tom Mackowiak |
ICASSP | 2 |