Kanglin Xiao

dblp:235/0143 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0002-1552-1613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 A 28nm 64Kb SRAM based Inference-Training Tri-Mode Computing-in-Memory Macro
abstract
Many computing-in-memory (CIM) macros achieve local inference with forward propagation (FP), and some CIM macros also support backward propagation (BP) computation. However, they can not calculate the weight change related to the learning rate and forward propagation input. these macros can not support backward propagation training algorithm completely. In this paper, we proposed a 28nm 64Kb SRAM based CIM macro, which supports a more complete backward propagation training algorithm. This macro supports three computing modes. A multiply unit (MU) supports FP and BP modes. A multiply circuit (MC) supports three-inputs-multiplication (TIM) mode for the weight change analog computing. MC uses the principle of charge sharing which has a high resistance to process variation and perfect linearity. In FP and BP modes, this macro achieves an energy efficiency of 42.1TOPS/W with 2-bit input, 8-bit weight and 14-bit output multiplication and accumulation operations (MAC). In TIM mode, this macro achieves an energy efficiency of 59.4 - 2222TOPS/W with multiplication of 3 inputs and 1 output.
Nanbing Pan, Xiaoxin Cui, Kanglin Xiao, Qingyu Guo, Yuan Wang 0001
ISCAS4
2022 A Computing-in-Memory SRAM Macro Based on Fully-Capacitive-Coupling With Hierarchical Capacity Attenuator for 4-b MAC Operation
abstract
In this work, we present a fully capacitive-coupling-based SRAM computing-in-memory (CIM) macro aimed at improving the energy efficiency and throughput of edge devices running multi-bit multiply-and-accumulate (MAC) operations. The proposed architecture is built around a customized 9T1C bit-cell in charge-domain computation in a 28nm technology. The proposed design supports 8192 $4{\mathrm{b}}\times 4{\mathrm{b}}$ MAC operations simultaneously. A 4-bit input is generated by DAC, while a 4-bit weight is achieved by a hierarchical capacity attenuator array without additional sharing switches, long sharing time, and complicated controlling signal. To minimize the expensive AD conversion, an input sparsity sensing scheme is proposed, allowing to skip redundant comparators. Access time is 4 ns with 0.9 V power supply at room temperature. The proposed design achieves energy efficiency of 666 TOPS/W and throughput of 4096 GOPS.
Kanglin Xiao, Xiaoxin Cui, Nanbing Pan, Xin'an Wang, Yuan Wang 0001
ISCAS1
2021 An SNN-Based and Neuromorphic-Hardware-Implementable Noise Filter with Self-adaptive Time Window for Event-Based Vision Sensor
abstract
Event-based dynamic vision sensors (DVS), inspired by biological vision systems, lead to new sensing and computing paradigms. The novel sensors output the sensed signal alone with many noise events asynchronously. Data-preprocessing for filtering these noises is significant before utilizing the data in applications such as classification, tracking and motion-data extraction. This paper describes a fully spike-based and neuromorphic-hardware-implementable neural network with a signal-oriented self-adaptive filtering time window for filtering the noise events robustly in the data captured by DVS. In particular, the simple leaky integrate-and-fire (LIF) neuron model is adopted as the basic elements of the network out of the purpose of hardware-friendly. Experiments based on both synthesized data and authentically-captured data are designed for quantitative comparison with traditional DVS noise filters to verify the outperformance of the proposed filter. The main contribution of this work is that the proposed spiking neural network (SNN) based filter achieves higher signal-noise-ratio (SNR) compared to traditional noise filters and performances more robust in the tolerance for changing signals.
Kanglin Xiao, Xiaoxin Cui, Kefei Liu 0002, Xiaole Cui, Xin'an Wang
IJCNN1
2021 Design and Implementation of a Temperature Self-Compensation Balanced Hybrid Ring Oscillator BHRO
abstract
The ring oscillator (RO) is applied to many modern circuits given its simplicity, low-area-cost, low-power. However, the temperature-drifting propagation delay of logic gates makes the RO a difficult solution for frequency reference design in deep submicron process. This work presents the design and implementation of a temperature self-compensation architecture based on a balanced hybrid ring oscillator (BHRO) for precise temperature compensation and clock-on-chip applications. The proposed BHRO is formed by both PTAT (proportional to the absolute temperature) and CTAT (complementary to absolute temperature) delay cells to implements RO-internal compensation. The features include temperature-self-compensation, process insensitive and low-area-cost. Four different test chips are fabricated in the 0.13μm CMOS process. The measurement result exhibits two performance-friendly BHRO architectures with a best temperature coefficient of 31 ppm/C over -55 to 80 , which is among the lowest to our best knowledge. The output compensated frequency is verified to be adjustable varying from 6.95 MHz to 26.5MHz, which are the highest as we have known.
Kanglin Xiao, Bo Wang 0016, Changpei Qiu, Xin'an Wang
ISCAS1