EDBT 2026 Demo / reviewers in the wild / expert
Yajuan He
dblp:34/1869
· DBLP profile ↗
14ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-6081-313XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Stable Approximate Pruning Framework with Heuristic Optimization
Lulin Cai, Dongxiao Yu, Pufan Luo, Benhao Pan, Yajuan He |
ISCAS | 5 |
| 2026 | Static Segment-Based Approximate Booth Multiplier with Probabilistic Error Compensation
Dongxiao Yu, Yuhang Ren, Chen Qin, Lulin Cai, Yajuan He |
ISCAS | 6 |
| 2023 | An 8T SRAM Based Digital Compute-In-Memory Macro For Multiply-And-Accumulate AcceleratingabstractCompute-in-memory (CIM) has been a promising technology to reduce the data movement energy and latency, which is the bottleneck of Von Neumann architecture. Digital approaches in CIM macro have many advantages compared with analog counterparts, such as programmability and inference precision. However, previous works with digital approaches generally employ complex SRAM bit-cells and computational components, which cause a large area overhead. In this paper, we propose a new 8T SRAM bit-cell to reduce the overall area of the SRAM array, which is able to implement the 1-bit multiplication without read-disturb issue. Additionally, the interleaving adder tree and dual supply voltage strategy are employed for further reduction on area and power consumption of computational circuits. Besides, a result combination circuit is designed to increase the bit-precision flexibility. A 16Kb SRAM CIM macro with proposed techniques is designed in a 40-nm CMOS technology. The simulation results show that our work achieves 820 GOPS throughput and 94 TOPS/W energy efficiency with 4-b of both input and weight. It achieves 1.3× higher energy efficiency and 70% area reduction when compared to the recent state-of-the-art works. Hongyang Luo, ZeYang Peng, Xingchen Chao, Yajuan He |
ISCAS | 5 |
| 2021 | Design of Approximate Multiplierless DCT with CSD Encoding for Image ProcessingabstractThis paper presents a multiplierless discrete cosine transforms (DCT) design with approximate canonical signed digit (CSD) encoding for image processing. Two approximation strategies on CSD encoding are proposed in constant multiplication. Based on these two coding approaches, an approximate DCT architecture is presented by taking advantage of the correlation between adjacent pixels of image data. Higher frequency coefficients are gradually ignored from the calculation due to the energy compaction property. Four approximate DCT architectures are thus proposed representing different accuracy levels. The proposed DCTs are implemented using 0.18μm standard CMOS process. The simulation results indicate that the proposed ADCT reduces 51.5% power and 30.0% area with a PSNR penalty of 1.6dB when compared with the traditional design. For lossy applications which allow lower computational accuracy, ADCT-III achieves 70.0% and 43.9% reduction on the power consumption and area, respectively, at a cost of 11.7 dB accuracy loss. Lulin Cai, Yiduan Qian, Yajuan He, Wen Feng |
ISCAS | 3 |
| 2021 | A 40nm 1Mb 35.6 TOPS/W MLC NOR-Flash Based Computation-in-Memory Structure for Machine LearningabstractComputation-in-memory (CIM) is a feasible method to overcome "Von-Neumann bottleneck" with high throughput and energy efficiency. In this paper, we proposed a 1Mb Multi-Level (MLC) NOR Flash based CIM (MLFlash- CIM) structure with 40nm technology node. A multi-bit readout circuit was proposed to realize adaptive quantization, which comprises a current interface circuit, a multi-level analog shift amplifier (AS-Amp) and an 8-bit SAR-ADC. When applied to a modified VGG-16 Network with 16 layers, the proposed MLFlash-CIM can achieve 92.73% inference accuracy under CIFAR-10 dataset. This CIM structure also achieved a peak throughput of 3.277 TOPS and an energy efficiency of 35.6 TOPS/W with 4-bit multiplication and accumulation (MAC) operations. Sitao Zeng, Zhiguo Zhu, Zhaolong Qin, Chen Wang 0130, Jingjing Li 0001, Sanfeng Zhang 0001, Yajuan He, Chunmeng Dou, Xin Si, Meng-Fan Chang, Qiang Li 0021 |
ISCAS | 8 |
| 2019 | Design of an Energy-Efficient Approximate Compressor for Error-Resilient MultiplicationsabstractDigital Multiplier is a fundamental component in many digital signal processing (DSP) systems, which takes up the most part of the computational resources. As many DSP applications have an inherent tolerance for inexact computations, approximate multiplication is considered as an appropriate substitution to obtain energy-performance-accuracy tradeoffs, especially in those applications that require high energy-efficiency in computing. Meanwhile, reducing the supply voltage is proved to be an efficient way to further lower the total energy consumption. In this paper, a novel approximate 4-2 compressor and its circuit implementation is proposed for error-resilient multiplication with a low supply voltage. Simulation results indicate that the approximate multiplier with our proposed approximate 4-2 compressor consumes the least energy per operation with the same computational accuracy when compared with other multipliers for the operand length of 8 bits. It achieves 26.7% reduction on energy-delay product (EDP) when compared with the exact multiplication. Xilin Yi, Haoran Pei, Ziji Zhang 0001, Yajuan He |
ISCAS | 5 |
| 2019 | A Subthreshold 10T SRAM Cell with Enhanced Read and Write OperationsabstractThis paper presents a novel 10T static random access memory (SRAM) cell for ultra-low voltage operations. A built-in read assist scheme is proposed to eliminate the read disturb issue. It significantly improves the read stability. In addition, a data-aware write assist technique is also proposed to enhance the write ability of the SRAM cell. Simulations are performed in a 40-nm standard CMOS technology at a 0.4V supply voltage. The experimental results indicate that the proposed 10T SRAM cell exhibits the best write margin and large read static noise margin. When compared with the conventional 6T SRAM cell, it achieves 2.95× larger read static noise margin and 1.93× higher write margin. Jiubai Zhang, Xiaoqing Wu, Xilin Yi, Jiaxun Lv, Yajuan He |
ISCAS | 5 |
| 2019 | A Half-Select Disturb-Free 11T SRAM Cell With Built-In Write/Read-Assist Scheme for Ultralow-Voltage OperationsabstractThis paper presents a half-select disturb-free 11T static random access memory (SRAM) cell for ultralow-voltage operations. The proposed SRAM cell is well suited for bit-interleaving architecture, which helps to improve the soft-error immunity with error correction coding. The read static noise margin (RSNM) and the write margin (WM) are significantly improved due to its built-in write/read-assist scheme. The experimental results in a 40-nm standard CMOS technology indicate that at a 0.5-V supply voltage, RSNM of the proposed SRAM cell is$19.8\times $and$0.96\times $as that of 6T and 8T SRAM cells with min-area, respectively. It achieves$11.84\times $and$9.56\times $higher WM correspondingly. As a result, a lower minimum operation voltage is obtained. In addition, its leakage power consumption is reduced by 53.3% and 44.5% when compared with 6T and 8T SRAM cell with min-area, respectively. Yajuan He, Jiubai Zhang, Xiaoqing Wu, Xin Si, Shaowei Zhen, Bo Zhang 0027 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | An Energy-Efficient Approximate DCT for Wireless Capsule Endoscopy ApplicationabstractWireless capsule endoscopy is widely used as an efficient way to receive images of gastrointestinal tract for medical diagnostics. Due to its energy compaction property for correlated image pixels, DCT) is highly preferred in image compression for this real time application. To accommodate the tight power budget and small size condition in the wireless capsule endoscopy application, an energy efficient DCT based on CMOS implementation is proposed in this paper. Three approximating methods are combined and applied in a multiplier-less DCT architecture. The simulation results indicate that the proposed DCT reduces 53.85% energy and 25.32% area with a PSNR penalty of 1.9dB when compared with the integer DCT. Compared with CORDIC-Loeffler DCT, it achieves 19.38% reduction on energy dissipation and 14.4dB more PSNR at the cost of only 1.39% area overhead. Ziji Zhang 0001, Yiduan Qian, Qiang Li 0021, Yajuan He |
ISCAS | 5 |
| 2018 | Optimal Slope Ranking: An Approximate Computing Approach for Circuit PruningabstractIn recent years, approximate computing has been widely used in low power digital circuits design. As it incurs error in computation, there are always tradeoffs between hardware performances and computational accuracy. Thus, the circuit implementation can vary from applications to applications and even in the same application, which makes difficult to evaluate the design efficiency in terms of delay, power and computational accuracy. In this paper, approximate efficiency (AE) is proposed as a new design metric dedicated in energy-efficient approximate computation. An automatic pruning approach is presented based on quickly estimated AE, which leads to an efficient approximate circuit design and a significant reduction on energy dissipation. A 32-bit adder is demonstrated using proposed method and compared with other approximation methods, including direct truncation and significance-activity pruning method. For legitimate comparisons, same computational accuracy is obtained for all imprecise adders. The simulation results show that assessed on mean absolute error, the proposed adder outperforms the others. With given error threshold, it can reach 31.15% energy-delay-product reduction compared with the most competitive contender. Ziji Zhang 0001, Yajuan He, Xilin Yi, Qiang Li 0021, Bo Zhang 0027 |
ISCAS | 2 |
| 2015 | A fast and energy efficient binary-to-pseudo CSD converterabstractThe canonical signed digit (CSD) coding is widely used in digital arithmetic operations due to its property that there is no adjacent nonzero digits in the encoded numbers. However, the benefits of the CSD coding may be faded because of the recursive conversion process from the binary representations. This paper presents a novel pseudo CSD coding method, which takes the merits of CSD, while simplifies the conventional conversion process. The simulation results indicate that the proposed converter can achieve at least 31.8% speed improvement and 42.9% energy reduction for a 16-bit binary operand at 1.2V in a 0.13-μm CMOS technology. It could run even faster than the competitors when the operand length increases. Yajuan He, Ziji Zhang 0001, Bin Ma 0010, Shaowei Zhen, Ping Luo 0005, Qiang Li 0021 |
ISCAS | 1 |
| 2013 | Blind-LMS based digital background calibration for a 14-Bit 200-MS/s pipelined ADCabstractA 14-bit and 200-MS/s SHA-less pipelined ADC is implemented by 0.13 μm CMOS process with blind least mean square (BLMS) calibration technique which corrects errors of this pipelined ADC with fast, low gain and inaccurate opamps. Using skip and fill approach, we employ an interpolation filter and a front-end DAC to make the pipelined ADC self-calibratable in the background. Incorporated a 18 stages and 1.5 bit/stage structure, the simulation shows the ADC achieves an SINAD of 86 dB, an SFDR of 107 dB with a 90.55 MHz input signal. Yajuan He, Qiang Li 0021 |
VLSI-SoC | 1 |
| 2013 | Digital Error Corrector for Phase Lead-Compensated Buck Converter in DVS ApplicationsabstractModern low-power system on a chip needs direct current converter with dynamic voltage scaling (DVS) ability for core power supply. The converter output should be accurate voltage across the full load current and voltage scaling range. An integrated buck converter for DVS application is proposed in this brief. Voltage mode phase lead compensation is implemented in the converter, with much smaller passive components than conventional type-III compensation. To improve accuracy, the output voltage error accompanied with load current and reference voltage caused by finite loop gain in analog control loop is corrected by the digital error corrector. The output voltage is compared by two comparators whose threshold voltage is about 10 mV above and below the reference voltage, respectively. The duty cycle is slightly adjusted by finite state machine according to outputs of the two comparators. Experimental results show that the converter is well regulated over an output range of 0.7-1.8 V, with step voltage of 25 mV. When load current suddenly changes between 170 and 500 mA, the overshoot and undershoot voltage are 32 and 50 mV, respectively. Load regulation is maintained about 1% throughout the full load range. The voltage error is within ±10 mV in the voltage scaling range. Shaowei Zhen, Ping Luo 0005, Yajuan He, Bo Zhang 0027 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | A low-power, high-speed RB-to-NB converter for fast redundant binary multiplierabstractIn this paper, a power-delay efficient redundant binary (RB) to 2's complement number converter for RB multiplication is presented. A new conversion algorithm is proposed to fully exploit the redundancy of RB encoding for a VLSI efficient implementation. The hierarchical linear expansion of the carry equation creates a regular multi-level parallel structure which is well suited for implementation with a logarithmic depth hybrid carry-lookahead/carry-select (CLA/ CSL) adder. A special add-one circuit is also incorporated into the CSL circuit to further reduce its logic complexity. A 64-bit reverse converter is designed using TSMC 0.18 /spl mu/m CMOS Technology. Pre-layout HSPICE simulation of the proposed design shows that it is capable of completing a 64-bit conversion in 761 ps and dissipates merely 0.34 mW at a data rate of 100MHz and a supply voltage of 1.8V. Yajuan He, Chip-Hong Chang |
ISCAS | 1 |