EDBT 2026 Demo / reviewers in the wild / expert
Yuqi Wang 0004
dblp:20/1168-4
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0002-4850-5126ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Reliable and High-Speed 6T Compute-SRAM Design With Dual-Split-VDD Assist and Bitline Leakage CompensationabstractCompute SRAM (CSRAM) can be configured to execute efficient in-memory logic computation and search operations, which is made possible by the multirow activation scheme. Nevertheless, simultaneous multirow activation introduces notorious compute access disturbance in 6T-based CSRAM designs. Existing solutions aim to mitigate the disturbance issue via weakening the access transistor strength, clamping the sensed bitline voltage swing, or substituting with read-decoupled SRAM bitcell designs. However, these techniques all come with significant design overheads of performance degradation in computational read and write accesses and increased layout area. This article presents multiple circuit-level techniques for the design and optimization of a reliable and high-speed 6T-based CSRAM. First, we propose a novel dual-split-$V_{\text {DD}}$(DSV)-assisted scheme for mitigating the compute access disturbance in 6T SRAM and simultaneously improving the computational read access performance. Second, we propose a leakage-compensated asymmetrical differential sense amplifier (LCAD-SA) to further improve the compute access performance. Third, we propose a DSV-assisted columnwise write scheme for accelerating the write performance. The proposed 6T CSRAM was implemented in the 28-nm CMOS, achieving a 1.18-GHz peak operating frequency, which is a$2.36\times $throughput improvement compared with that of the state-of-the-art CSRAM designs. Yuqi Wang 0004, Yajun Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | A Reliable 8T SRAM for High-Speed Searching and Logic-in-Memory OperationsabstractTo efficiently implement searching and logic functions with the SRAM-based in-memory computing (IMC), we need to perform computations on bitlines (BLs) (called compute access) via multiple wordline (WL) activations. However, this may cause prominent read disturbance when the IMC is implemented with the standard 6 T SRAM. To address this reliability issue, existing solutions adopt either auxiliary assistance circuits or alternative bitcell topologies, but they lead to substantial overheads of the access speed or array density. In this article, we propose a novel 8T compute SRAM (CSRAM) for reliable and high-speed in-memory searching and compound logic-in-memory computations. Our 8T CSRAM features a pair of pMOS access transistors and split-WLs dedicated to the compute access. A thorough circuit-level analysis reveals that the pMOS-based compute access port is essential for significantly mitigating the read disturbance. Moreover, we propose an elevated precharge voltage scheme and a low-skewed inverter-based sensing amplifier to improve the sensing speed. We have validated the proposed 8T CSRAM design in a 16 Kb array with a 28-nm CMOS technology. Compared to the state-of-the-art 8 T CSRAM, results show that our design is not only reliable but also 3.1 times faster, with a maximum operating frequency upping to 2.44 GHz. Yuqi Wang 0004, Yuhao Shu, Weixiong Jiang, Yajun Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | Analysis and Optimization Strategies Toward Reliable and High-Speed 6T Compute SRAMabstractIn-SRAM Computation improves the throughput and energy-efficiency of data-intensive applications by utilizing parallelism and reducing the data transfers. However, when multiple wordlines are accessed simultaneously, a short-circuit path will likely incur dynamic read disturbance and generate extra direct current in 6T Compute SRAM (CSRAM). In order to mitigate this issue, existing works either degrade the access speed, use area-hungry bitcells, or incur architecture-level overheads. In this paper, we first perform a comprehensive circuit-level analysis of the dynamic read disturbance issues of 6T SRAM for the first time and find that such disturbance can be efficiently avoided by maintaining the bitline voltage at a high level. Second, we propose a novel energy-efficient, reconfigurable sense amplifier design that is able to achieve fast and reliable sensing when the bitline voltage level is high for the compute access. Third, we propose an adaptive wordline control scheme that keeps the bitline voltage at a high level to eliminate the dynamic read disturbance and the sneaky direct current pathway. Both the new sense amplifier and adaptive wordline control are also optimized to support the normal read access efficiently. We have validated our design in a 55nm CMOS technology. Experimental results show that our design not only reliably addresses the read disturbance and the extra direct current, but also operates 19% faster than the state-of-the-art design using an advanced 28nm FDSOI technology. Yuqi Wang 0004, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Energy-Efficient Arbitrary Precision Multi-Bit Multiplication with Bi-Serial In/Near Memory ComputingabstractRecent works show that multi-bit multiplications can be achieved in multi-cycles via serial in/near memory computing, aiming at reducing data transfers hence improving energy efficiency. However, the pure serial approach suffers from a long latency and leaves additional room to optimize energy efficiency. We propose three new techniques to develop a novel SRAM structure to realize multi-bit multiplication with bi-serial in/near memory computing. Firstly, we use 2-bit binary numbers (bi-serial) as the smallest unit of operation working in the mixed-signal mode to achieve arbitrary precision. This significantly reduces latency compared to the pure serial approaches. Secondly, we take the unique advantages of bi-serial (2-bit by 2-bit) multiplication and reduce the number of voltage bands that we need to differentiate from seven to four. This reduces the voltage swing on analog bit-lines (ABLs), in this way, reducing the dynamic power and improving accuracy. Thirdly, we optimize a low-cost voltage comparator based on the inverter chain to reduce static power further. When normalized to 8-bit by 8-bit multiply operation and implemented a 16KB SRAM array in SMIC 55-nm CMOS technology, the energy efficiency of our design is 1.47 TOPS/W, which is 2.7 times better than the state-of-the-art. Yuqi Wang 0004, Yu Pu, Yajun Ha |
ISCAS | 1 |