EDBT 2026 Demo / reviewers in the wild / expert
Sudeb Dasgupta
dblp:94/537
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-4044-1594ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 252.8-GOPS and 10.1-TOPS/W Split DP-8T SRAM-Based Analog CIM Macro for 9-bit Signed MAC Operations With High Signal MarginabstractWe present a reconfigurable and scalable current-based analog compute-in-memory (CIM) macro based on a split dual-port 8T (DP-8T) SRAM bitcell. This bitcell can perform MAC operations between signed 9-bit inputs and signed 9-bit weights and logical operations. Throughput is enhanced by sensing MAC and logical operations from both true and complementary bitcell sides via dual bitline sensing, without requiring external circuitry to decode the MAC output from the complementary side. To improve the signal margin and reduce the number of ADCs, we split the 8-bit magnitudes of inputs and weights into 2-bit groups and perform efficient weight encoding with a 2C–1C network combined with analog shift-and-add using 4C and 1C at each bitline. The proposed$64\times 128$macro performs 4096 signed MAC operations in 4 cycles with a latency of 16.2 ns. The design is implemented in TSMC 65-nm technology at 1.2 V. The architecture performs linear MAC operations across different inputs, weights, and process corners. It delivers a throughput of 252.83 GOPS, an MAC energy efficiency of 10.11 TOPS/W, and a signal margin of 26 mV for signed MAC operations. In the logical compute mode, the architecture performsnor,and,nand, andorBoolean operations in a single cycle, as well asexnorandexorusing additional logical gates, achieving a throughput of 3276.8 GOPS and a latency of 1.25 ns at 0.8 V. The work achieves inference accuracies of 99.1%, 91.65%, and 71.8% on MNIST, CIFAR-10, and CIFAR-100, respectively. Abhishek Goel, Cheena Singhal, Sparsh Mittal, Sudeb Dasgupta |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A 1.91 POPS/W Energy-Efficient SRAM Based Signed Multi-Bit Time Domain CIM ArchitectureabstractIn this paper, we present an area and energy efficient multi-bit signed SRAM based time-domain compute-in-memory (TDCIM) architecture. We have proposed an area efficient time-domain multiplication bit-cell (TDMC) for signed/unsigned operation. The proposed TDMC is 1.5× times area efficient than state-of-the-art. The work uses a time domain computing approach that uses combination of delay and pulse width to represent the equivalent multiplication and accumulation (MAC) operation performed. The proposed TDCIM focuses on improving the energy efficiency of the MAC operation. A time to digital converter (TDC) is used instead of analog to digital converter (ADC) to save power and area consumption. It achieves 1.3× higher energy efficiency compared to state of the art. The designed architecture has been implemented using a 28 nm FDSOI STM technology, resulting in the development of an 8 Kb TDCIM macro. This work achieves a normalized energy efficiency of 1910.45 TOPS/W at VDDof 0.8 V. The architecture achieves an inference accuracy of 91.20% on the CIFAR-10 dataset with 8-bit precision in inputs and weights. Subhradip Chakraborty, Dinesh Kushwaha, Himanshu Ranjan, Sudeb Dasgupta |
ISCAS | 4 |
| 2025 | A Methodology for Datapath Energy Prediction and Optimization in Near Threshold Voltage RegimeabstractIn this article, we propose a method for sizing an arbitrary combinational datapath to minimize its energy consumption. Our method involves deriving expressions for the components of energy consumption at both the stage and path levels. In this work, we identify overshoot energy ($E_{\text {OS}}$) consumption as a previously unreported component contributing to energy consumption, particularly significant in the near/sub-threshold voltage regime. We determine that this$E_{\text {OS}}$consumption is proportional to the input and output transition times and size of a logic gate at a particular stage of a datapath. We also observe that, for a given number of stages (N) and path effort, the total energy consumption is optimized when the stage effort (f) in a datapath is kept constant. Based on our observations and derivations of all the energy components and the requirement for a constant “f” in the datapath, we develop a method to minimize the energies of a logic circuit while maintaining the timing closure requirement. We determine that the non-critical paths (NCPs) must be sized to a minimum “f” while maintaining the timing requirements. We verified our models on several ISCAS and EPFL benchmark circuits with an average reduction of 28.1% (41.2%) and 19.2% (28.4%) in energy consumption [figure of merit (FoM)], respectively. The proposed methodology predicts the total energy consumption at a stage and path level of N-stage logic, with only one-time SPICE simulation on a single stage, with a maximum error of 1.3% and 1.62%, respectively, against SPICE simulations. The simulations are performed in Synopsys HSPICE environment with ST Microelectronics 65 nm CMOS and 28 nm FDSOI technology nodes, resulting in a very good agreement with the developed methodology. Mahipal Dargupally, Lomash Chandra Acharya, Arvind K. Sharma, Sudeb Dasgupta, Bulusu Anand |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | SRAM-Based Hybrid Analog Compute-In-memory Architecture to Enhance the Signal MarginabstractThis manuscript proposes an SRAM-based hybrid analog compute-in-memory (CIM) architecture to enhance the signal margin. This hybrid architecture presents fully differential current-based and C-2C charge-sharing-based multiplication and accumulation (MAC) CIM schemes for 4-bit MAC operation. The MAC operation of the filter weight's least significant bits (w0and w1) is implemented in the current-based CIM. However, the MAC operation of the most significant bits (w2and w3) is implemented in the charge-based CIM. The proposed architecture achieves a 4.37× enhancement in signal margin compared to the state-of-the-art. The energy efficiency and throughput of the proposed architecture are 1551.5 TOPS/W and 512 GOPS, respectively, at 0.9 V supply voltage and 250 MHz frequency. A convolutional neural network (CNN) is implemented on the proposed architecture, and the inference accuracy for the MNIST and CIFAR-10 data sets is 98.6 % and 86 %, respectively. The proposed architecture is scalable for multi-bit MAC operation and implemented in 28 nm CMOS technology. Dinesh Kushwaha, Rajiv V. Joshi, Sudeb Dasgupta, Bulusu Anand |
ISCAS | 3 |
| 2024 | Interface Trap Analysis in Multi-Fin FinFET Technology: a Crucial Reliability Issue in Digital ApplicationabstractThis paper focuses on investigating the impact of NBTI degradation on p-FinFETs, a major reliability concern. The impact of donor type interface traps on single and multi-fin FinFET structures are considered for both device and circuit aspects. The degradation of device performance occurs due change in threshold voltage (Vth) and drain current caused by traps. This study extends to analyze the effect of traps on the Voltage Transfer Characteristic (VTC) and transient behavior in an inverter, followed by a study of the Ring Oscillator (RO). We precisely evaluate parameter changes and differentiate the impact among 1-Fin, 2-Fin, and 3-Fin structures. Under NBTI, device reliability concern, the End of Lifetime (EOL) is achieved at a threshold voltage shift of 50 mV, occurring at a trap concentration of 1.25×1012cm-2. While the change in Vth is consistent across single and multi-fin FinFETs parameters however Subthreshold Slope (SS), transconductance (gm) and DIBL (Drain Induced Barrier Lowering) exhibit more significant variations and degradation, particularly in 2-Fin and 3-Fin structures. In terms of Ring Oscillator (RO) performance, the impact is less pronounced in 1-Fin structures, resulting in finer performance compared to 2-Fin and 3-Fin structures. Jyoti Patel, Sankalp Rai, Sudeb Dasgupta |
ISCAS | 4 |
| 2024 | Switching Activity Factor-Based ECSM Characterization (SAFE): A Novel Technique for Aging-Aware Static Timing AnalysisabstractWe propose switching activity factor-based effective current source model (SAFE) for aging-aware static timing analysis (STA), a new technique for estimating the timing performance of digital circuits. SAFE is based on the development of device-level variation-aware analytical timing models of stacked and multistage logic cells (commonly employed transistor topologies in a synthesized netlist of a random logic path), which drastically reduces the recharacterization efforts of the standard cells. The models developed are derived as a function of input transition time$(T_{R})$and load capacitance$(C_{L})$. The timing performance of a standard cell degrades with threshold voltage$(V_{\mathrm {th}})$degradation in a MOS device due to various aging mechanisms. SAFE, makes the entire STA process aging aware by updating its model coefficients with$V_{\mathrm {th}}$degradation caused by aging. It is achieved by proposing a method for estimating$V_{\mathrm {th}}$degradation under various stress conditions, including static, dynamic, and asymmetric, that applies to any process design kit (PDK). To consider asymmetric aging, we have developed a method to find effective switching activity factor$(\alpha _{\mathrm {eff}})$for N-stage stacked and N-stage parallel logic which is used to find the value of switching activity factor$(\alpha)$at intermediate nodes in pipelined logic circuits. Our simulations are performed in Mentor Graphics Eldo SPICE environment using STMicroelectronics 28 and 65-nm CMOS process. The proposed technique provides a high-simulation accuracy (2.5% average error) when compared with SPICE simulations. Finally, we achieved a ~98.14% reduction in the required number of simulations using SAFE when compared with a completely SPICE/Aging simulation-based approach. Lomash Chandra Acharya, Arvind K. Sharma, Neeraj Mishra, Khoirom Johnson Singh, Mahipal Dargupally, Nayakanti Sai Shabarish, Ajoy Mandal, Ramakrishnan Venkatraman, Sudeb Dasgupta, Bulusu Anand |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2023 | Aging-Aware Timing Model of CMOS Inverter: Path Level Timing Performance and Its Impact on the Logical EffortabstractA static timing analysis (STA) methodology based on an effective current source model (ECSM) is proposed for the first time for estimating the aging-aware path-level timing performance and its impact on the logical effort of a CMOS inverter for digital timing closure in pre-stress and post-stress conditions. Degradation in the threshold voltage$(V_{\mathrm{ th}})$of PMOS occurs due to temporal variability mechanisms (aging), such as negative bias temperature instability, resulting in delay degradation of a standard cell. Therefore, we proposed a technique to make the STA process aware of this degradation by developing device-level variation aware (with aging) timing models of CMOS inverters to represent threshold-crossing points (TCPs) in an ECSM.libs file as a function of stress time ($t$). A device-level approach for$V_{\mathrm{ th}}$degradation into different aging conditions, such as static and dynamic, is developed for a given process design kit to update TCPs in a (.libs) file as a function of$t$. A python-based tool is being developed to estimate the path-level timing performance of digital circuits in pre- and post-stress conditions. Again, we developed a technique for relating the inverter’s logical effort with$t$to resize a near-critical path in pre-stress conditions for achieving digital timing closure in pre- and post-stress conditions. The verification and validation of the proposed model with different benchmark circuits are performed using a parasitic extracted netlist in the Eldo SPICE environment with the 65-nm CMOS process technology. Finally, our model reduces the number of SPICE/Stress simulations by 98.13% compared to the previously reported only simulation-based techniques. Lomash Chandra Acharya, Arvind K. Sharma, Neeraj Mishra, Khoirom Johnson Singh, Mahipal Dargupally, Nayakanti Sai Shabarish, Ajoy Mandal, Ramakrishnan Venkatraman, Sudeb Dasgupta, Bulusu Anand |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2022 | A 65nm Compute-In-Memory 7T SRAM Macro Supporting 4-bit Multiply and Accumulate Operation by Employing Charge SharingabstractIn this work, we propose an energy-efficient 64$\times $ 64 compute-in-memory (CIM) SRAM macro using a 7T bit-cell in 65nm CMOS UMC PDK. It supports 4-bit inputs, 4-bit weights & 4-bit outputs and performs 4-bit MAC operations. It also supports multiple row activations performing 1024 4b$\times $4b multiply and accumulate (MAC) operations in one clock cycle. Inputs are realized by the number of pulses on the read wordline (RWL), which discharges read bitline (RBL) according to bitwise multiplication of weights & inputs. Outputs of 4 columns storing 4-bit weights are then combined via charge sharing to perform a binary-weighted average representing MAC operation, further quantized by a flash analog to digital converter (ADC) giving 4-bit output. The proposed CIM macro achieves an energy efficiency of 28.9 TOPS/W and throughput of 212.9 GOPS operating at supply voltage 1 V with a 2 GHz clock frequency. Dinesh Kushwaha, Ritik Raj, Ashish Joshi, Jwalant Mishra, Rajat Kohli, Sandeep Miryala, Rajiv V. Joshi, Sudeb Dasgupta, Bulusu Anand |
ISCAS | 10 |
| 2022 | Significance of Organic Ferroelectric in Harnessing Transient Negative Capacitance Effect at Low Voltage Over Oxide FerroelectricabstractThe concept of leveraging the transient negative capacitance (TNC) effect in a ferroelectric (FE) is a relatively new addition to the field of nanoelectronics. Until now, there has been no comparison of organic and oxide FE-based metal-FE-metal (MFM) devices in harnessing the TNC effect. As a result, we introduce an external resistor-MFM(R-MFM) series circuit to investigate the role of organic and oxide FEs in harnessing the TNC effect at low supply voltages. The multidomain Ginzburg-Landau-Khalatnikov theory is used to model the FE materials in a technology computer-aided design environment. We show that: (i) organic FE-based R-MFM series circuit can harness the TNC effect at just 1 V whereas an oxide FE-based R-MFM series circuit cannot; (ii) the coercivity of an organic FE is 77.39% lower than its counterpart, oxide FE; (iii) the remanent polarization of an organic MFM (1.2$\mu$C/cm2) is very close to the channel charge density of a CMOS transistor (1.6$\mu$C/cm2) making it helpful in addressing capacitance matching issues in NC transistor; (iv) an organic FE-based R-MFM series circuit dissipates 78.89% less energy than an oxide FE-based R-MFM series circuit; (v) the TNC effect and time are justified by its dependence on R. Finally, this article suggests that an organic FE-based MFM could be used as a gate stack of any transistor to achieve sub-60 mV/decade switching energy, making it ideal for ultra-low voltage NC transistors. Khoirom Johnson Singh, Lomash Chandra Acharya, Bulusu Anand, Sudeb Dasgupta |
ISCAS | 4 |
| 2022 | Phase Noise Analysis of Separately Driven Ring OscillatorsabstractIn this paper, for the first time, the phase noise analysis of a Multi-loop Skew based Single Ended Oscillator (MSSROs) is derived and validated. Compared to the three stages of conventional ring oscillators (CROs), SDROs provide an equivalent oscillation frequency with improved phase noise with increasing stages. The primary distinction between these two designs (SDRO and three-stage CROs) is the inherent skew offset between the PMOS/NMOS gates caused by the unique connection. This skew offset is the fundamental cause of delay cell noise suppression; the SDROs have loosely coupled oscillators that run concurrently, forming multiple 3-stages of separately driven Ring Oscillators. As a result, a shaping function is derived in terms of skew offset, and simulating these with varying skew offset results in suppressing behavior. Additionally, we derived phase noise for a skew-based design and validated it in PDKs of 180nm and 65 nm. We plotted the thermal (flicker) noise contribution and found that increasing the number of stages leads to an approximately 1-2 dB reduction in phase noise while maintaining the same NMOS/PMOS size ratio. Finally, a 2-3 dB reduction in phase noise is achieved in MSSROs by incorporating the shaping function into phase noise equations. Neeraj Mishra, Anchit Proch, Lomash Chandra Acharya, Jeffrey Prinzie, Sudipto Chakraborty, Rajiv V. Joshi, Sudeb Dasgupta, Bulusu Anand |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2021 | Harnessing Maximum Negative Capacitance Signature Voltage Window in P(VDF-TrFE) Gate StackabstractIn this paper, the observation of transient negative capacitance signature (NCS) in an organic ferroelectric gate stack (OFEGS) at minimum supply voltage (Vs) of ±0.5 V is investigated employing a well-calibrated Ginzburg-Landau-Khalatnikov (GLK) model in the environment of Sentaurus technology computer-aided design (STCAD). We observe an 88.62 to 94.76 % reduction in the average coercive voltage (Vc) of the proposed OFEGS, which is still a significant challenge for the conventional ferroelectric (FE) lead zirconate titanate. We study the resistor-OFEGS (RCofe) series network behaviors in response to a bipolar and unipolar triangular signal. Our findings prove that the presence of NCS is directly correlated with the FE polarization (FEP) switching and not because of any extrinsic defects in the system. The various impacts of Vs, GLK parameters, R, dipole switching resistivity (Rofe) variations on the NCS response are investigated. The proposed OFEGS can harness the NCS effect at ±0.5 V with minimum energy dissipation of 4.81 × 10-16J, a challenge for the oxide FE-based gate stacks. Calibrating R to the maximum limit, we can capture the S-shaped ideal Landau path where the NCS is maximum with a small deviation of about ±0.006 V at zero FEP. Finally, an OFEGS based Landau transistor is implemented, providing a minimum subthreshold swing (SSmin) of 38.21 mV/decade, which is 36.32 % lesser than the fundamental SSminlimitation of 60 mV/decade. Therefore, the proposed OFEGS with a minimum Vs and remanent polarization (Pr=3D 1.244 μC/cm2) could be used as a gate stack for designing sub-60 mV/decade transistor technology. Khoirom Johnson Singh, Bulusu Anand, Sudeb Dasgupta |
ISCAS | 3 |