Jing Kou

dblp:352/2109 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0007-9419-1723ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FALCON: A Fast and Low-Power Current-Mode Near-Sensor-Computing Architecture for Real-Time Edge Visual Processing
Jing Kou, Jinyao Mi, Junda Zhao, Junzhan Liu, Wang Kang 0001
DATE2
2026 InFuzz: An Efficient and Lightweight In-Memory-Computing Cryptographic Fuzzy Extractor for IoT Security
Jing Kou, Wang Kang 0001
ISCAS1
2026 FABS-CIM: Unlocking A/D Conversion Bottlenecks of Bit-Serial Computing-In-Memory with Analog Shift-and-Addition and In-Situ Batch Normalization
Junda Zhao, Jing Kou, Junzhan Liu, Wang Kang 0001
ISCAS4
2025 Accuracy Is Not Always We Need: Precision-Aware Bayesian Yield Optimization
abstract
Integrated circuit yield optimization plays a vital role in ensuring reliable semiconductor manufacturing, directly impacting both product quality and production costs. Current approaches to yield optimization face two fundamental challenges that limit their practical effectiveness. First, yield estimation requires intensive computational resources. Second, traditional black-box optimization methods inefficiently allocate these resources across design candidates. Most existing approaches compound these issues by performing detailed yield estimations uniformly across all candidates, regardless of their potential quality. To address these limitations, we introduce a novel precision-aware yield optimization framework that intelligently adapts computational resource allocation based on each design candidate’s predicted performance. Our approach moves beyond simple simulation counting by incorporating a Figure of Merit (FoM) as a continuous quality metric. By combining a Continuous AutoRegression model to characterize the relationship between true yield and precision levels with a sophisticated multi-fidelity acquisition strategy, our framework achieves optimal resource distribution. Experimental validation on four industry-standard benchmark circuits demonstrates that our method converges with fewer than 1,000 simulations, reducing simulation costs by over $10 \times$ while achieving better final designs and robustness than state-of-the-art high-fidelity approaches.
Jing Kou, Zidong Chen, Haiyan Qin, Wang Kang 0001, Wei W. Xing
DAC1
2025 Multi-Agent Yield Analysis For Circuit Design
abstract
Semiconductor yield estimation presents a critical challenge in modern manufacturing, directly impacting production costs and market competitiveness. Traditional estimation methods, particularly Monte Carlo simulation, while reliable, become computationally prohibitive for complex modern circuits. Contemporary approaches, including importance sampling and machine learning techniques, face fundamental limitations in consistency across circuit topologies and practical validation. This work introduces YieldAgent, a novel Large Language Model (LLM)-powered framework that revolutionizes yield estimation through dynamic integration of multiple analytical strategies. YieldAgent employs a three-layer agent architecture to analyze circuit characteristics and historical data, optimizing estimation methods while balancing computational efficiency and precision. The framework incorporates Retrieval-Augmented Generation for domain knowledge integration and Tree-structured Parzen Estimators for dynamic hyperparameter optimization. Experimental validation across 12nm and 40nm technology nodes demonstrates that YieldAgent reduces computational overhead by up to $2.9 \times$ while maintaining or exceeding state-of-the-art accuracy. The system’s ability to adapt across different circuit topologies and technology nodes establishes a new paradigm for scalable, intelligent yield estimation in electronic design automation.
Haiyan Qin, Jing Kou, Wang Kang 0001, Wei W. Xing
DAC2
2025 Towards Accurate Characterization of In-Memory Computing Non-Idealities: A Physics & Data Co-Driven Generative Framework
abstract
Analog in-memory computing (IMC) promises unprecedented energy efficiency for deep learning acceleration, but suffers from non-idealities that severely degrade inference accuracy in fabricated chips. Therefore, accurate modeling of these non-idealities becomes significant. In this work, we present PDGM-IMC, the first physics and data co-driven generative framework for IMC non-idealities characterizing. Unlike traditional physical models that fail to model complex non-ideality behaviors, or black-box neural networks that lack interpretability and generalization, PDGM-IMC leverages normalizing flows with custom transformations directly derived from device physics principles. This novel approach enables explicit modeling of complex probability distributions, spatial correlations, and die-to-die variations that previous methods could not capture. Validated on multiple dies of a commercial eFlash-based IMC SoC, PDGM-IMC improves modeling accuracy by 4.6× for the input circuit and IMC array and by 2.0× for the output circuit, significantly outperforming existing approaches. By extracting the statistical signature of fabricated chips, PDGM-IMC enables accurate pre-silicon prediction of post-silicon behavior, fundamentally transforming hardware-aware neural network optimization for analog accelerators. The source code and the pre-trained models are publicly available at https://github.com/BUAA-BASIC-Lab/PDGM-IMC.
Jing Kou, Guangyao Wang, Saiya Wang, Yuexi Lv, Xinghao Cui, Wei. W. Xing, Wang Kang 0001
ICCAD1
2025 A 0.88 e‾rms 8-Mpixel 3D-Stacked Low Temporal-Noise CMOS Image Sensor With Auto-Zero Single-Slope ADC, Fast Correlated Multi-Sampling, Row-Wise Noise Reduction, and Dark Current Non-Uniformity Calibration Techniques
abstract
This paper presents a low temporal noise, low-power, 8-Mpixel, rolling-shutter (RS)-type, back-illuminated CMOS image sensor (CIS) employing through silicon via (TSV) 3D-stack technology. To achieve temporal noise less than 1erms-, we explored auto-zero (AZ) column single-slope (SS) ADC and fast correlated multi-sampling (CMS) techniques. The pixel signal was sampled two times by the readout circuits using a 9-bits ADC, resulting in a 10-bits digital output. To enhance image quality in low light conditions, we adopted a parity column counter (PCC) for power supply stabilization and H-banding elimination, and employed row-wise noise reduction (RWNR) and dark-current non-uniformity calibration (DCNUC) techniques for reducing row-wise noise and improving image uniformity. Our CIS chip was fabricated using a 55nm 1P4M (pixel substrate) and a 55nm 1P5M (logic substrate) CIS 3D stacked process. The die area is ~3.99*3.45 mm2with 1.008-μm pixel pitch and the total energy consumption is 170mW under a 2.8V analog-VDD and a 1.2V digital-VDD. The chip achieves a temporal noise of only ~0.88erms-, fixed pattern noise (FPN) of ~25.08μVrms, row-wise noise of ~5.5μVrmsand an energy efficiency figure-of-merit (FoM) of ~0.6erms-*nJ/step at a frame rate of 60 frames per second (FPS).
Wang Kang 0001, Jing Kou, Liangchen Li, He Zhang 0011, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4