EDBT 2026 Demo / reviewers in the wild / expert
Yakun Zhou
dblp:49/3658
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High Energy Efficiency Spatial Parallel Stochastic Computing for Precision Scalable Neural Processing UnitabstractA precision-scalable neural processing unit, considering the quantization-sensitive of each neural network layer, has large hardware redundancy in multiplication units and shift logics. In this paper, we explore a spatial parallel stochastic computing (SPSC) precision-scalable architecture to reduce hardware redundancy. The conventional SPSC multiplier has large accuracy loss, so we analyze the error component and propose an error compensation SC (ECSC) multiplier. To reduce the hardware cost, the bitbrick (BB) in fusion-based precision-scalable architecture is replaced by stochastic bitbrick (STOBB), which is composed of a fine-grained ECSC multiplier. The proposed SC architecture supports 2/4/8-bit operation. Our design is synthesized under SMIC 55nm CMOS technology. Experiments show that our design achieves a 2.048 TOPS/W energy efficiency as well as 2.21× area efficiency in comparison to the state-of-the-art designs. Yakun Zhou, Jienan Chen |
ISCAS | 1 |
| 2025 | A 4.86-pJ/b Energy-Efficient Fully Parallel Stochastic LDPC Decoder With Two-Stage Shared MemoryabstractThe complex calculations of the low-density parity-check (LDPC) decoder result in significant energy and hardware consumption. To solve the challenge, this brief describes a fully parallel stochastic LDPC decoder with a two-stage shared memory (TSM) variable node (VN). To enhance cost efficiency, our design incorporates a shared low-cost random number generator (RNG) for all 2160 channels. We introduce a TSM VN function, which demonstrates faster convergence and reduced hardware overhead in comparison with the existing methods. We have taped out the (2160, 1760) stochastic LDPC decoder in the 55-nm process. The measure results exhibit that the proposed design achieves a throughput of 57.6 Gb/s, an efficiency of 33.68 Gb/s/mm2, and a power efficiency of 4.86 pJ/bit, underlining superior performance in terms of decoding throughput, hardware efficiency, and energy conservation. Yakun Zhou, Jienan Chen, Yizhuo Zhou, Zihan Xia 0002, Chuan Zhang 0001, Runsheng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | A Hardware Efficient Matrix Multiplications Scheme with Dynamic Precisions and Dimensions for Massive MIMO SystemsabstractMatrix multiplication serves as the primary operation in massive multiple-input multiple-output (MIMO). However, with the continuous advancement of MIMO technology, the escalating computational complexity and the necessity for adaptable matrix multiplication in MIMO communication pose a formidable challenge. In this paper, we propose the Joint Serial-Parallel Dataflow (JSPD) mapping method for dynamic precision matrix multiplication, which relies on the utilization of Spatio-Temporal Transforms (STT) and Principal-Auxiliary Matrices (PAM). The data is initially processed serially and is mapped onto Processing Element (PE) arrays through STT transforms. Subsequently, both high-precision and low-precision components of the data are computed selectively in parallel, employing a combination of principal and auxiliary matrices. The matrix multiplication process is mapped onto the PE arrays, incorporating output-stationary and output-flow modes. Compared with the traditional fixed structure PE arrays, the JSPD mapping method improves the computational flexibility while increasing the PE cell work share ratio by 7.57% ∼ 21.86%. Meanwhile, JSPD reduces the hardware overhead by a maximum of 68% and improves hardware efficiency by 35.77%. Qiuyu Cheng, Yakun Zhou, Chentao Liang, Zuofeng Zhang, Jienan Chen |
ISCAS | 2 |
| 2023 | Hardware Efficient Reconfigurable Logic-in-Memory Circuit Based Neural Network ComputingabstractWith the explosive growth in processing and data storage capability, computing in memory (CIM) technology is considered a feasible method to mitigate the memory wall. Re-cently, a floating-gate field-effect transistor (FGFET) technology with single-layer MoS2 channel was proposed, which can flexibly transfer the memory and logic gate by setting the corresponding voltage. In this paper, based on the new FGFETs devices, we propose a hardware-efficient reconfigurable logic-in-memory (LIM) circuit to perform neural network (NN) computing. We first represent the basic FGFET array as a matrix form. Thereby, the adders, multipliers, registers, and PE units are designed by the matrix function of the arrays. The proposed circuits can dynamically transfer the memory to computing logic online when fewer data are required to be stored. A theoretical experiment based on the VGG16 network exhibits an efficiency increase of 85.35% in the same logic area compared with the traditional method. Tianchi Liu 0005, Yizhuo Zhou, Yakun Zhou, Jienan Chen |
ISCAS | 3 |
| 2021 | Dynamic Service Migration with Partially Observable Information in Mobile Edge ComputingabstractService migration, determining when, where and how to migrate the ongoing service, is of paramount importance in mobile edge computing (MEC) for provisioning high quality of service to mobile users. With respect to high network dynamics and stringent delay requirements, service migration is a rather challenging issue in MEC. In this paper, we formulate service migration as a partially observable Markov decision process (POMDP) based on the fact that an edge server can only obtain partial users' information, or the information of its own serving users. A learning-based intelligent service migration algorithm, named iSMA, is proposed to minimize the long-term service delay of all users. iSMA consists of two function modules, a latent space model and a cross-entropy planning algorithm, where the latent space model is used to infer the full state of the environment based on the partial information observed, and the cross-entropy planning algorithm is used to search the best service migration strategy. Numerical results show that our proposed iSMA reduces the service delay by about 58% when compared with a well-known deep learning-based solution. Yakun Zhou, Yao Sun 0002, Siyu Chen 0018, Jienan Chen, Gang Feng 0004 |
GLOBECOM | 2 |
| 2018 | Frequency Support and Stability Analysis for an Integrated Power System with Wind FarmsabstractIn order to handle the challenges associated with large capacities wind farms integrated in the power grid, different techniques such as optimization algorithms, artificial intelligence and others are employed and studied. Recently, many researches focus on the frequency regulation problem by showing the role of kinetic energy stored in the rotor side of wind turbine in supporting the grid frequency. The main challenge of the inertia emulation scheme is to release the maximum possible amount of kinetic energy without violating operating limits of the wind turbines and keeping in mind the grid integration standards for wind farms. However, no selection criteria for the control parameters has yet been given. The objective in this paper is to consider the constraints and limitations on the rotor speed and buffered power in DFIG wind turbines and also to highlight the influence of the controller parameters on the stability of the power grid and the operation of the wind farms. The stability of the power system with high penetration of wind farms is analyzed based on Kuramoto theory and a sufficient stability condition is derived and mathematically proved. Finally the stability condition is verified in a test power system with a large integration of wind power. Bashar Mousa Melhem, Yakun Zhou, Steven Liu |
IECON | 2 |