Shengqi Yu

dblp:268/1673 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-8059-5793ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HotReRAM: A Performance-Power-Thermal Simulation Framework for ReRAM-Based Caches
abstract
This article proposes a comprehensive thermal modeling and simulation framework, HotReRAM, for resistive RAM (ReRAM)-based caches that is verified against a memristor circuit-level model. The simulation is driven by power traces based on cache accesses for detailed temperature modeling over time. HotReRAM models power at a fine-grain level and generates temperature traces for different cache regions together with detailed analyses of thermal stability, retention time and write latency. Combining HotReRAM with gem5, a full-system simulator, and NVSim, a power simulator, for ReRAM enables temporal and spatial modeling of crucial ReRAM characteristics. This integration allows designers and architects to analyze various cache characteristics within a single cache bank and address thermal-induced issues when designing ReRAM caches. Our simulation results for an 8-MiB ReRAM cache show that the spatial thermal variance can be as high as 7 K for a single cache bank, whereas the temporal thermal variance is more than 40 K. Such temperature variances impact retention time with a standard deviation of 3.9–10.2 for a set of benchmark applications, where the write latency can increase by up to 14.5%.
Shounak Chakraborty 0001, Thanasin Bunnam, Jedsada Arunruerk, Sukarn Agarwal, Shengqi Yu, Rishad A. Shafik, Magnus Själander
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin Machines
abstract
In-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively.
Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik
ISLPED4
2023 Approximate digital-in analog-out multiplier with asymmetric nonvolatility and low energy consumption
abstract
Many modern compute-intensive applications require arithmetic results (usually multiplication) to be represented as analog signals. Using digital multipliers followed by digital-to-analog conversion (DAC) results in high energy and performance costs. This is because digital multipliers have costly carry propagation, and DAC circuits add associated conversion costs. Another concern, especially for arithmetic on the edge, is the need for nonvolatile operands in the face of power uncertainty. To deal with this, nonvolatile memory technologies have been combined with in-memory computing. This paper proposes a mixed-signal multiplier which directly generates an analog product based on two digital input operands. Fundamental to the design are transistor-memristor cells, organized in a crossbar structure. Using analog resistive partial product accumulation in the crossbar, the approximate multiplier eliminates the need for carry propagation and an explicit DAC. It also provides asymmetric nonvolatility making memristor writing a rare event, extending the application significance of the method. The design is shown to be functionally correct up to 4-bit, and achieves 8× to over 300× speedup, competitive peak-power and orders of magnitude energy reduction, compared with existing full-digital memristor-based multipliers and low-power multiplication DAC solutions.
Shengqi Yu, Fei Xia 0001, Rishad A. Shafik, Domenico Balsamo, Alexandre Yakovlev
Integr.1
2022 Editable asynchronous control logic for SAR ADCs
abstract
This paper presents a novel design method for asynchronous control logic targeting successive approximation register (SAR) analog-to-digital converters (ADCs). This work is based on modeling the control logic for SAR ADCs using signal transition graphs (STGs). Different from conventional synchronous controllers, the proposed method results in asynchronous controllers driven by the causality of signals rather than relying on clocks to control the conversion process. Moreover, the proposed asynchronous control logic can be modularized through the handshake protocol, making it possible to build ADCs of arbitrary precision based on single-bit control units. This work results in a formal, model-based asynchronous design flow for SAR ADC control, which is shown to produce resulting circuits of similar speeds but great power efficiency improvements.
Fei Xia 0001, Gang Mao, Shengqi Yu, Rishad A. Shafik, Alexandre Yakovlev
ISCAS4
2021 Optimized Multi-Memristor Model based Low Energy and Resilient Current-Mode Multiplier Design
abstract
Multipliers are central to modern compute-intensive applications, such as signal processing and artificial intelligence (AI).However, the complex logic chain in conventional multipliers, particularly due to cascaded carry propagation circuits, contributes to high energy and performance costs.This paper proposes a novel current-mode multiplier design that reduces the carry propagation chain and improves the current amplification.Fundamental to this design is a one transistor multi-memristor (1TxM) cell architecture.In each cell, transistor can be switched ON/OFF to determine the cell selection, while the high/low resistive states of memristors determine the corresponding cell output current when selected.The memristor states as well as biasing configurations in each memristor are suitably optimized through a new memristor model.The number of memristors implementing this model in each cell is suitably determined depending on the cell significance to achieve the required amplification.Consequently, the design reduces the need to have current mirror circuits in each current path, while also ensuring high resilience in transitional bias voltages.Parallel cell currents are then directed to a common current accumulation path to generate the multiplier output without requiring any carry propagation chain.We carried out a wide range of experiments to extensively validate our multiplier design in Cadence Virtuoso analogue design environment for functional and parametric properties.The results show that the proposed multiplier reduces up to 85% latency and 99% energy cost when compared with the recently proposed approaches.
Shengqi Yu, Rishad A. Shafik, Thanasin Bunnam, Kaiyun Chen, Alexandre Yakovlev
DATE1
2020 Current-Mode Carry-Free Multiplier Design using a Memristor-Transistor Crossbar Architecture
abstract
Multipliers are a major energy and delay contributor in modern compute-intensive applications due to their complex logic architecture. As such, designing multipliers with reduced energy and faster speed has remained a thoroughgoing challenge. This paper presents a novel, carry-free multiplier, which is suitable for a new-generation of energy-constrained applications. The multiplier circuit consists of an array of memristor-transistor cells that can be selected (i.e., turned ON or OFF) using a combination of DC bias voltages based on the operand values. When a cell is selected it contributes to current in the array path, which is then amplified by current mirrors with variable transistor gate sizes. The different current paths are connected to a node for analogously accumulating the currents to produce the multiplier output directly. This removes the need for latency-sensitive carry propagation stages, typically seen in traditional multipliers. We conduct a number of experiments to validate the functional and parametric properties. Our experiments showed that proposed multiplier achieves 51.44% savings in energy at a similar accuracy when compared with recently proposed approaches.
Shengqi Yu, Ahmed Soltan, Rishad A. Shafik, Thanasin Bunnam, Fei Xia 0001, Domenico Balsamo, Alexandre Yakovlev
DATE1