EDBT 2026 Demo / reviewers in the wild / expert
Yijiao Wang
dblp:289/4297
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-0159-2231ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NFGen: Normalizing Flow-Based Joint Generative Model for Variability-Aware Design Technology Co-OptimizationabstractAs transistor sizes continue shrinking, impacts of variability has become ever more paramount in circuit design and manufacturing. Their accurate representations in model cards help save design margins and provide appropriate guidelines in design technology co-optimization (DTCO). To address such a challenge, we propose a novel machine learning framework, Normalizing Flow-Based Joint Generative Model (NFGen), which generates a comprehensive model library from a limited number of model cards. Unlike traditional generative methods that focus on the marginal distribution of model card parameters, NFGen is the first model to approximate their joint distribution, which includes information on their correlation and thus enables closer representation of variability effects. In addition, we introduce two similarity metrics to rigorously evaluate the quality of generated model cards. Experimental results show that NFGen reduces overall error by 2x to 8x compared to state-of-the-art methods, validating its superiority in variability-aware DTCO. Zhenxing Dou, Yijiao Wang, Peng Wang 0022, Runsheng Wang, Weisheng Zhao 0001, A. Asenov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | An FD-SOI-Based Compact In-Pixel Computing Architecture Enabling Real-Time Feature ExtractionabstractTo empower resource-limited edge devices in artificial intelligence (AI) and Internet of Things (IoT) applications, it is essential to overcome challenges posed by restricted area resources and the high latency demands of transmitting and processing substantial sensory data. In-pixel computing addresses these challenges effectively, and the Fully Depleted Silicon-On-Insulator (FD-SOI)-based pixel, which relies on an FD-SOI transistor whose current is made photosensitive to light by applying a negative back-gate voltage, shows significant potential with its compact structure and in-situ computation capability. In this paper, for the first time, we present an FD-SOI-based chip-level architecture for in-pixel computing. Our design implements programmable, massively parallel convolution with low latency using a pulse-width modulation (PWM) input encoding scheme. Furthermore, the proposed compact 1P1T (1 Phototransistor 1 Transistor) pixel design, integrated with an improved single-slope analog-to-digital converter (SS ADC), greatly enhances area efficiency. Validated through simulation in a 22nm FD-SOI process, the design achieves 990 frames/s under typical outdoor illumination conditions, with the figure of merit (FoM) of 16.86pJ/pixel/frame. In addition, the proposed architecture has been evaluated on hand gesture recognition (5697 training and 633 validation images across six categories), achieving an accuracy of 97.48%. The results demonstrate that, compared to state-of-the-art designs, our approach achieves a$6\times $improvement in in-pixel convolution speed and a$7.3\times $reduction in area overhead. Yijiao Wang, Jiayao Wu, Zhongzhen Tong, Xinrui Duan, Yiming Shi, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Mitigating methodology of hardware non-ideal characteristics for non-volatile memory based neural networks
Lixia Han, Peng Huang 0004, Yijiao Wang, Haozhang Yang, Jinfeng Kang |
Sci. China Inf. Sci. | 3 |
| 2025 | CIMUS: 3D-Stacked Computing-in-Memory Under Image Sensor Architecture for Efficient Machine VisionabstractComputational image sensors with CNN processing capabilities are emerging to alleviate the energy-intensive and time-consuming data movement between sensors and external processors. However, deploying CNN models onto these computational image sensors faces challenges from the limited on-chip memory resources and insufficient image processing throughput. This work proposes a 3D-stacked NAND flash-based computing-in-memory under image sensor architecture (CIMUS) to facilitate the complete deployment of CNN model. To fully leverage the potential of high bandwidth from the 3D-stacked integration, we design a novel distributed CNN mapping and dataflow to process the full focal plane image in parallel, which senses and recognizes ImageNet tasks with >1000fps. To tackle the computational error of inputs “0” in 3D NAND flash-based CIM, we propose an input-independent offset compensation method, which reduces the average vector-matrix multiplication (VMM) error by 48%. Evaluation results indicate that CIMUS architecture achieves a 9.8× improvement in CNN inference speed and a 33× boost in energy efficiency compared to the state-of-the-art computational image sensor in the ImageNet recognition task. Lixia Han, Haozhang Yang, Ao Shi, Guihai Yu, Yijiao Wang, Yanzhi Wang 0001, Jinfeng Kang, Peng Huang 0004 |
IEEE Trans. Computers | 9 |
| 2025 | MS-SCIM: A Mixed-Signal Stochastic Computing-in-Memory Paradigm for Information SecurityabstractStochastic Computing (SC), an emerging paradigm with advantages in hardware cost and fault tolerance, is well-suited for applications in image processing and information security. However, the existing SC paradigms suffer from significant hardware costs due to the conversion between binary and stochastic sequences, and the long latency caused by low computational parallelism. In this work, we demonstrate a Mixed-Signal Stochastic Computing-In-Memory (MS-SCIM) paradigm utilizing spin orbit torque magnetic random access memory (SOT-MRAM) arrays, in order to realize a energy-efficient and conversion-less SC method for the first time. The main contributions include: 1) The inherent stochastic switching behaviors of spintronic devices are exploited to enable the SOT-MRAM array to serve both as a parallel true random number generator (TRNG) and a CIM cell. 2) A high-parallelism mixed signal stochastic CIM paradigm is proposed to accelerate SC-based edge detection algorithm. The whole process achieves binary outputs without the conversion circuits and the energy efficiency achieves 446 Tops/W. 3) Based on the results of MS-SCIM, a novel image steganography method using stochastic bit streams is leveraged for information security, which enables lossless embedding and extraction of secret image information in a$156\times 156$size, enhancing both capacity and undetectability. Pengxu Wang, Yijiao Wang, Jialiang Yin, Jiayao Wu, Xinrui Duan, Zhaohao Wang, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Efficient Discrete Temporal Coding Spike-Driven In-Memory Computing Macro for Deep Neural Network Based on Nonvolatile MemoryabstractNonvolatile memory (NVM) based neural network can directly perform in situ computation in memory to significantly reduce energy consumption resulting from the data movement. However, the energy consumption by the analog-to-digital converter (ADC) restricts the efficiency of the mixed-signal in-memory computing macro. The rate coding spike-driven in-memory computing macro can increase the energy efficiency via eliminating the ADC, but the improvement is limited because substantial energy is consumed for the coding of multiple spikes. In this work, we propose a discrete temporal coding spike-driven in-memory computing macro, including input coding scheme, weight mapping method, and improved leaky integrate-and-fire (LIF) neuron circuit, to perform the efficient forward inference of deep neural networks based on NVM array. We then optimize the designment of the proposed in-memory computing macro to mitigate the neural network accuracy loss due to the nonlinearity of the LIF neuron and voltage drop caused by interconnect resistance. Because the temporal coding scheme reduces spike numbers and the improved-LIF circuit simultaneously integrates two bit-lines current corresponding to positive and negative weight, the proposed macro achieves 46.63TOPS/W energy efficiency and 1.92TOPS throughput for 3bit temporal coding precision. Lixia Han, Peng Huang 0004, Yijiao Wang, Jinfeng Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |