Yinyin Lin

dblp:49/479 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 2 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeepPiC: xPU-PIM Cluster Architecture with Adaptive Resource-Aware Task Orchestration for DeepSeek-Style MoE Inference
abstract
The success of DeepSeek has driven demand for deploying high-performance inference clusters. However, due to its Transformer-based autoregressive structure, DeepSeek remains severely bandwidth-bound, limiting the scalability of traditional xPU (e.g., GPU/TPU). While DRAM-based processing-inmemory (PIM) offers a promising solution to overcome memory bottlenecks, its use in inference clusters for DeepSeek remains underexplored due to three challenges: (1) non-trivial inter-device communication overhead; (2) the need for expert parallelism in the mixture-of-experts (MoE) module; and (3) lack of efficient task offloading to PIM. To this end, we propose DeepPiC, a novel xPU-PIM cluster architecture designed for DeepSeek-style models with multi-latent attention (MLA) and MoE modules. DeepPiC introduces a heterogeneous xPU+HBM-PIM device to accelerate low arithmetic intensity operations. It can seamlessly replace conventional xPU devices without any modification to clusterlevel interconnect topology. However, DeepPiC cannot fully realize its performance potential under static scheduling, which fails to adapt to shifting compute and memory demands driven by multidimensional variability (model heterogeneity, cluster-scale volatility, runtime dynamics). This induces inter-device communication overhead and intra-device underutilization. Thus, we propose Adaptive Resource-Aware Task Orchestration (ARTO), a two-phase strategy that decouples global model partitioning from local task assignment by dynamically coordinating (1) crossdevice parallelism optimization and (2) intra-device xPU/PIM mapping. Evaluated on DeepSeek V3-671B using H20-, A100-, and $\mathbf{H 2 0 0}$-Cluster ($\mathbf{H 2 0}$ serves as a compute-limited alternative to high-end GPUs), DeepPiC (H20+HBM-PIM) achieves up to $\mathbf{3} \times \mathbf{, 2} \times$ and $\mathbf{1. 3} \times$ speedup over $\mathbf{H 2 0}$-, A100-, and $\mathbf{H 2 0 0}$-Cluster at small batch sizes, while maintaining $\mathbf{7 4 \%}$ and $\mathbf{5 4 \%}$ of A100and $\mathbf{H 2 0 0}$-Cluster performance at large batch sizes. These results demonstrate that DeepPiC enables low-end xPU to approach or even exceed premium ones by fundamentally overcoming memory bottlenecks via adaptive scheduling that orchestrates PIM and xPU heterogeneous resources.
Manni Li, Zijian Huang 0017, Wending Zhao, Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong
ASP-DAC6
2026 MPiCO: Memory-Pool-Based XPU-PIM Cluster over Optical I/O with Load-Imbalance-Aware Assignment and Execution-Site-Matching Mapping Strategies for MoE Inference
abstract
We first propose MPiCO, a memory-pool-based XPU–PIM cluster over Optical I/O, together with Load-Imbalance-Aware Assignment (LIAA) and Execution-Site-Matching Mapping (ESMM) strategies. Confining processing-in-memory (PIM) to a small set of HBMs in a hybrid HBM–DDR pool, MPiCO cuts PIM cost and offsets the resulting performance loss by eliminating inter-XPU communication overhead. LIAA resolves MoE load imbalance via dynamic assignment of warm experts to XPU/PIM, and ESMM avoids PIM-induced bandwidth loss by aligning address mapping: interleaved for XPU, PIM-friendly mapping dedicated to PIM-dies. On DeepSeek-V3 671B, MPiCO with LIAA and ESMM achieves a 2.4 × speedup and 3.5 × higher energy efficiency over H20-Electric I/O (EIO) cluster, 3 × lower PIM cost than H20-EIO with local PIM, and a 1.8 × speedup over a state-of-the-art MoE platform.
Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong
ACM Great Lakes Symposium on VLSI2
2026 ATSGRU: Attention-Sparse Gated Recurrent Unit for Computationally Efficient Wideband Digital Predistortion of Quadrature Digital Power Amplifiers
abstract
Digital predistortion (DPD) is a widely used technique for enhancing signal quality in modern radio frequency (RF) power amplifiers (PAs). However, the strong performance of deep neural network (DNN)-based DPD models is often offset by their prohibitive computational complexity, which limits their practical deployment in wideband systems. This paper presents an attention-sparse gated recurrent unit (ATSGRU)—a novel neural architecture designed for computationally efficient wideband DPD in quadrature digital PAs (DPAs). The ATSGRU integrates the attention mechanism that evaluates the temporal relevance of input features and prunes redundant components, thereby simplifying the model structure and reducing computational load. The proposed method is validated on a custom 28-nm CMOS DPA chip. Experimental results demonstrate that the proposed ATSGRU achieves superior linearization with a favorable balance between accuracy and complexity compared with the state-of-the-art (SOTA) DPD model, reducing multiply-accumulate (MAC) operations by 54% while maintaining comparable performance. These results highlight its strong potential for efficient and scalable wideband DPD applications.
Wending Zhao, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
ACM Great Lakes Symposium on VLSI4
2026 ICDL: Inverse Compensation Direct Learning with Model-Accelerator Co-Optimization enabling Real-time Inference in Precision Motion Control
Manni Li, Longbin Jiang, Zijian Huang 0017, Wending Zhao, Yinyin Lin
ISCAS7
2025 Unmask Tampering: Efficient Document Tampering Localization under Recapturing Attacks with Real Distortion Knowledge
Changsheng Chen 0001, Yinyin Lin, Bin Li 0011, Jiwu Huang
CCS3
2025 Linearization of Quadrature Digital Power Amplifiers by Neural Network of ULR_LSTM: Unsupervised Learning Residual LSTM
abstract
For the first time, this paper presents an unsupervised learning residual long short-term memory (ULR_LSTM) neural network to develop a digital predistortion (DPD) method for the linearization of digital power amplifiers (DPAs). Our method eliminates the need for iterative learning control (ILC) to obtain the ideal input of the DPA required by state-of-the-arts (SOTAs), which leads to high computational complexity and extensive training time. We perform behavioral modeling of the DPA using the R_LSTM network. After determining the optimal behavioral model architecture, the corresponding DPD model is obtained through an inverse training process. A 15-bit transformer-based quadrature DPA chip incorporating Class-G and IQ-cell-sharing techniques was implemented in a 28nm CMOS process to validate our proposed method. Experimental results demonstrate outstanding linearization performance comparing to prior arts, achieving an error vector magnitude (EVM) of -40.4dB for the 802.11ax 40MHz 64QAM signal.
Luyi Guo, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
DATE8
2025 Digital Predistortion for Quadrature Digital Power Amplifiers Using Deep Neural Network of AT_LSTM: Attention LSTM
Wending Zhao, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
ACM Great Lakes Symposium on VLSI8
2025 CPSnB: Compressing and Processing Spatial Similarity near Memory Bank for DNNs
abstract
Near memory bank processing (NMBP) architecture only benefits memory-bound operations of DNNs in terms of energy consumption. Drawing on the insight that data compression can reduce the compute density of operators, transforming compute-bound operations into memory-bound operations, We propose CPSnB, a NMBP architecture combined with preserving numerical jump-spatial similarity compression (PNJ-SSC) method. CPSnB provides a tiling strategy for optimizing operators of different DNN models. Compared to the systolic host-side accelerator and existing dense and sparse NMBP, CPSnB significantly reduces energy consumption. Analysis of the experimental results indicates that a 60% compression ratio of activation can enhance the versatility of CPSnB in processing DNN operators to 22.3 times.
Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
ISCAS7
2025 APCPU: Adaptive-Pooling Compression Processing Unit for Energy-Efficient DNNs Processing
abstract
Integrating compression in the multiply-and-accumulate (MAC) path can significantly improve the energy efficiency of DNN operators. However, existing unstructured sparse compression (USSC) methods struggle to effectively compress activations with low sparsity. Computing core processing USSC face challenges such as load imbalance and complex index control circuit design. Based on insights into local spatial correlation, a block-wise adaptive-pooling compression (APC) method is proposed to achieve a high compression ratio for activations. Furthermore, this paper proposes an APCPU to integrate APC into the MAC path with minimal overhead, facilitating highly energy-efficient sparse processing of DNN operators. Leveraging a hybrid data flow design to achieve load balancing results in speedups of 1.25× to 1.33×. The experiment results show that the APCPU achieves energy savings of 1.35× and 1.27× compared to JPZ-PU, and 2.63× and 2.71× compared to CSC-PU when evaluated on AlexNet and Bert.
Wang Wang, Wending Zhao, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
ISCAS7
2025 GPOS: A General and Precise Offloading Strategy for High Generality of DNN Acceleration by OCP and NDP Co-Optimizing
abstract
The arithmetic intensity (ArI) of different DNNs can be opposite. This challenges the generality of single acceleration architectures, including both dedicated on-chip processing (OCP) and near-data processing (NDP). Neither architecture can simultaneously achieve optimal energy efficiency and performance for operators with opposite ArI. It is relatively straightforward to think of combining the respective advantages of OCP and NDP. However, few publications have addressed their real-time co-optimization, primarily due to the lack of a quantifiable offloading method. Here, we propose GPOS, a general and precise offloading strategy that supports high generality of DNN acceleration. GPOS comprehensively considers the complex interactions between OCP and NDP, including hardware configurations, dataflow (DF), DNN model, and interdie data movements (DMs). Three quantifiable indicators—ArI, execution cost (Ex-cost), and DM-cost—are employed to precisely evaluate the impacts of these interactions on energy and latency. GPOS adopts a four-step flow with progressive refinement: each of the first three steps focuses on a single indicator at the operator level, while the final step performs context-based calibration to address operator interdependencies and avoid offsetting NDP benefits. Narrowing down offloading candidates in step 1 and step 3 significantly accelerates real-time quantitative analysis. Optimized mapping techniques and NDP-input stationary DF are proposed to reduce Ex-cost and extend operator types supported by NDP. Next, for the first time, sparsity—one of the most popular methods for energy optimization that can alter data reuse or ArI—is quantitatively investigated for its impacts on offloading using GPOS. Our evaluations include representative DNNs, including GPT-2, Bert, RNN, CNN, and MLP. GPOS achieves the minimum energy and latency for each benchmark, with geometric mean speedups of 49.0% and 94.1%, and geometric mean energy savings of 45.8% and 89.2% over All-OCP and All-NDP, respectively. GPOS also reduces offloading analysis latency by a geometric mean of 92.7% compared to the evaluation that traverses each operator and its relative combinations. On average, sparsity further improves performance and energy efficiency by increasing the number of operators offloaded to NDP. However, for DNNs where all operators exhibit either very high or very low ArI, the number of offloaded operators remains unchanged, even after sparsity is applied.
Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 A 0.09-pJ/Bit Logic-Compatible Multiple-Time Programmable (MTP) Memory-Based PUF Design for IoT Applications
abstract
The Internet of Things (IoT) allows devices to interact for real-time data transfer and remote control. However, IoT hardware devices have been shown security vulnerabilities. Edge device authentications, as a crucial process for IoT systems, generate and use unique IDs for secure data transmissions. Conventional authentication techniques, computational and heavyweight, are challenging and infeasible in IoT due to limited resources in IoTs. Physical unclonable functions (PUFs), a lightweight hardware-based security primitive, were proposed for resource-constrained applications. We propose a new PUF design for resource-constrained IoT devices based on low-cost logic-compatible multiple-time programmable (MTP) memory cells. The structure includes an array of MTP differential memory cells and a PUF extraction circuit. The extraction method uses the random distribution of BL current after programming each memory cell in logic-compatible MTP memory as the entropy source of PUF. Responses are obtained by comparing the current values of two memory cells under a certain address by challenge, forming challenge–response pairs (CRPs). This scheme does not increase hardware consumption and circuit differences on edge devices and is intrinsic PUF. Finally, 200 PUF chips were fabricated by CSMC based on the 0.153-$\mu $m MCU single-gate CMOS process. The performance of the logic-compatible MTP memory cell and its PUF was evaluated. A logic-compatible MTP cell has good programming erase efficiency and good durability and retention. The uniqueness of the proposed PUF is 50.29%, the uniformity is 51.82%, and the reliability is 93.61%.
Shuming Guo, Yinyin Lin, Yao Li 0018, Chongyan Gu, Weiqiang Liu 0001, Yijun Cui
IEEE Trans. Very Large Scale Integr. Syst.2
2025 Digital Predistortion for Wide Dynamic Power Range Quadrature Switched-Capacitor Power Amplifiers Using Self-Adaptive Residual LSTM Neural Network
Luyi Guo, Yicheng Li 0002, Yinyin Lin, Yun Yin, Hongtao Xu
IEEE Trans. Very Large Scale Integr. Syst.4
2024 LauWS: Local Adaptive Unstructured Weight Sparsity of Load Balance for DNN in Near-Data Processing
abstract
Memory wall issue has become the overwhelming bottleneck of future systems due to the explosive parameter growth and low computing density large language model (LLM). Near-data processing (NDP) could alleviate data traffic and energy consumption, but the storage demand of LLM is still enormous. Weight sparsity is helpful for reducing data capacity. Unstructured sparsity sacrifices less accuracy compared to structured one, but the random non-zero values distribution in NDP leads to load imbalance among parallel processing units. Here we propose LauWS which is seamlessly combined into various prior arts of sparsity. LauWS follows the local characteristics of feature distribution in weight matrix for various models, preserving even tiny features and discarding non-feature values as far as possible region by region. That is the key for LauWS achieving a trade-off between high prune ratio (PR) and less accuracy loss (AL). Evaluations are carried out based on a GDDR6-based bank-NDP system. The typical optimization compared to the no-prune includes 38% speedup at 0.8PR with no AL for MLP, 22.7% speedup at 0.5PR with no AL for GPT-2, 23.6% speedup at 0.5PR with the lowest perplexity for OPT-125m.
Wang Wang, Manni Li, Yinyin Lin, Guhyun Kim, Yosub Song, Chengchen Wang, Xiankui Xiong
ISCAS6
2022 Statistical Observations of Three Co-Existing NBTI Behaviors in 28 nm HKMG by On-Chip Monitor With Less Recovery Impact
abstract
An on-chip digital sensor has been demonstrated in 28nm High-k Metal Gate (HKMG) for bias temperature instability (BTI) statistical characterization with the benefits: fast statistical measurement, less recovery impact (Toff-stress@around 15ns, Fast Period Sampling (FPS) @around 300ns), and high resolution (0.1mV of$\Delta $Vth). As far as we know, it is the first time to statistically observe the very early stage of trap recovery of individual device in practical scenario, e.g., static random-access memory (SRAM). We find that three Negative BTI (NBTI) recovery behaviors, 2/3/4-step with clear transition slope, co-exist in HKMG devices. Our further analysis ascribes the phenomena to co-existing of four types of defects in 28nm HKMG Devices Under Test (DUTs). Three types are recoverable and one unrecoverable. The transition slope instead of steep drop between steps is the aggregative effects of one certain type of recoverable defect contained across DUTs. More types of defects lead to more Vth shift. But the contribution percentage of unrecoverable defect remains quite close, while recoverable defects dominate the Vth degradation. Only when the Toff-stress is less than the starting point of 1st recover step (within 1$\mu \text{s}$in our case), accurate and consistent Vth degradation data can be achieved.
Yarong Fu, Wang Wang, Manni Li, Yinyin Lin
IEEE Trans. Circuits Syst. I Regul. Pap.8
2017 A small area and low power true random number generator using write speed variation of oxidebased RRAM for IoT security application
abstract
A true random number generator using write speed variation of oxide-based RRAM is proposed for the first time. The signal of this physical unclonable function (PUF) is strong with long duration to be easily and accurately captured by simple circuit, of which the advantage is attributed to the mechanism that the speed variation amplifies the fluctuation of oxygen vacancy trap and de-trap. Some function parts of normal RRAM IP can be reused as entropy source cells and implementation circuit. The variation of write end point is monitored by a self-adaptive write drive circuit to trig a counter, and then serialized into a bit stream. The test chips, which are AlOx/WOx bilayer back-end RRAM fabricated in 0.18 Um logic process, passed all NIST tests with advantages of small area, low power, and not using post-processor corrector. Enough bits can be generated within the endurance limitation to ensure usual Internet of Things (IoT) security application.
Yinyin Lin, Yarong Fu, Xiaoyong Xue, B. A. Chen
ISCAS2
2017 Retention-Aware Hybrid Main Memory (RAHMM): Big DRAM and Little SCM
abstract
Hybrid memory comprised of a big SCM and a little DRAM (BSLD) is widely studied to address the growing power consumption challenge of pure DRAM. However, the performance degradation, limited endurance and immature mass production of ultra-high-density SCM are still the painful points of BSLD. Here we propose a Retention-Aware Hybrid Main Memory (RAHMM) architecture with a big DRAM and a little SCM (BDLS) for the first time. DRAM is refreshed at a much longer interval by using SCM to store the small quantity of leaky tail bits in DRAM. A two-step search technology combined with outcome forecasting is put forward to get ultra-fast read access, as well as to diminish the power and performance overheads. A hidden buffer strategy (HBS) is proposed to optimize write performance and endurance hurt. The experimental results show 45 percent reduction of power consumption and 30 percent performance optimization, which are significantly improved compared to that of both serial and parallel BSLD with a counterpart capacity
Weiliang Jing, Yinyin Lin, Beomseop Lee, Sangkyu Yoon, Yuan Du, Bomy Chen
IEEE Trans. Computers3
2016 A compact pico-second in-situ sensor using programmable ring oscillators for advanced on chip variation characterization in 28nm HKMG
abstract
An all-digital on-chip sensor using programmable ring oscillator (PRO) for advanced on chip variation (AOCV) characterization is proposed and verified in 28nm HKMG node. Bypassing technique combined with statistical testing based on array of PROs enables resolution of 1 pico-second. Multiple cells under test (CUTs) are embedded into one single PRO through programmability to get mask area cost effectiveness. Extra variation is eliminated by symmetrical duplicate structure. Test results indicate that there is a flex point around 7-stage on the curves of delay sigma/n vs. stage number. Local variation dominates and decreases significantly with the increase of stage number before 7-stage point. For small dimension, inverter is more sensitive to on-chip-variation than NAND. But no same trend is observed for large dimension. Curves of delay average vs. stage number and delay sigma/n vs. stage number among dies based on the same type of cell indicate a good uniformity.
Yinyin Lin, Xiaoyong Xue
ISCAS1
2016 Novel 3D horizontal RRAM architecture with isolation cell structure for sneak current depression
abstract
Both 3D Vertical RRAM (VRRAM) and 3D Horizontal RRAM (HRRAM) architecture suffer from the issue of serious sneaking current, which leads to read or write disturbance and unacceptable power consumption waste, severely limiting its spatial stack-ability. In this work, the power consumption caused by sneaking current is separated out from the total set power consumption in HRRAM architecture. Then isolation cell structure is proposed to suppress the sneaking current. Simulation results show that total power consumption is reduced by about 30% with our proposed structure. Meanwhile, disturbance and read margin also show improvements.
Xiaoyong Xue, Yinyin Lin, Jaehwang Sim
ISCAS5
2016 Low-Power Variation-Tolerant Nonvolatile Lookup Table Design
abstract
Emerging nonvolatile memories (NVMs), such as MRAM, PRAM, and RRAM, have been widely investigated to replace SRAM as the configuration bits in field-programmable gate arrays (FPGAs) for high security and instant power ON. However, the variations inherent in NVMs and advanced logic process bring reliability issue to FPGAs. This brief introduces a low-power variation-tolerant nonvolatile lookup table (nvLUT) circuit to overcome the reliability issue. Because of large ROFF/RON, 1T1R RRAM cell provides sufficient sense margin as a configuration bit and a reference resistor. A single-stage sense amplifier with voltage clamp is employed to reduce the power and area without impairing the reliability. Matched reference path is proposed to reduce the parasitic RC mismatch for reliable sensing. Evaluation shows that 22% reduction in delay, 38% reduction in power, and the tolerance of variations of 2.5× typical RONor ROFFin reliability are achieved for proposed nvLUT with six inputs.
Xiaoyong Xue, Yinyin Lin, Ryan Huang, Qingtian Zou, Jingang Wu
IEEE Trans. Very Large Scale Integr. Syst.3
2015 3D vertical RRAM architecture and operation algorithms with effective IR-drop suppressing and anti-disturbance
abstract
We propose co-optimization of VRRAM cell structure and array architecture as well as IR-drop-aware read/write algorithms to overcome issues of disturbance and IR drop from long wire. A bi-directional diode (2D) access device is combined with one resistor to form 2D1R cell. A dummy reference plane is inserted into array to set up the same IR drop path of reference cell with that of selected cell. Consequently, the same IR drop effect can be cancelled during read. The model for disturbance analysis is put forward. Voltage dropped on un-selected bit lines is the key parameter to suppress set disturbance. Set disturbance is significantly suppressed even when number of RRAM layers increases to 64. Set voltage has to meet corresponding requirements in order to minimize the disturbance risk.
Yinyin Lin, Xiaoyong Xue, B. A. Chen
ISCAS1