VLDB 2026 Research / reviewers in the wild / expert
Xinyu Wang 0027
dblp:68/1277-27
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-3676-2408ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perception-Core: Reconfigurable Energy-Efficient Domain-Specific Architecture with Multimodal Fusion for AIoT
Xinyu Wang 0027, Yuanhua Deng, Haixiang Ren, Hui Chen 0015, Li Li 0003 |
ISCAS | 1 |
| 2026 | Thermal-Aware 3D-IC Floorplan Based On TSV-Coordination Simulated Annealing
Xingjie Zou, Linfeng Wu, Heng Zhang 0025, Xinyu Wang 0027, Li Li 0003 |
ISCAS | 5 |
| 2026 | An Adaptive Congestion-aware Approximate Communication (ACAC) scheme and implementation for Network-on-Chip systems
Shize Zhou, Wenjie Fan 0004, Yongqi Xue, Shiping Li, Songfeng Deng, Jinlun Ji, Tong Cheng, Xinyu Wang 0027, Li Li 0003 |
Integr. | 9 |
| 2026 | Fast Modular Reduction Algorithm and Reconfigurable Domain-Specific Architecture Design Based on Generalized Mersenne PrimesabstractModular arithmetic enjoys a broad spectrum of applications. In recent years, the advancement of post-quantum cryptography (PQC) has imposed growing demands on the flexibility and scalability of domain-specific accelerators. This article presents a fast modular reduction algorithm based on generalized Mersenne (GM) primes, which employs approximate scaling and iterative compression approaches to rapidly converge the quotient value while reducing the precomputation complexity to rely solely on the modulus itself. Building upon this algorithm, we have designed a reconfigurable modular reduction array using multiplier units with smaller word length. Operating at 1GHz, the proposed array achieves an area reduction of 45.14% compared to Barrett-based structures, and 13.93% compared to Montgomery-based structures. The array has been integrated into a complete number-theoretic transform (NTT) acceleration architecture. The resulting reconfigurable GM/general modular reduction domain-specific architecture elevates the chip frequency to$0.42\sim 1$GHz under the same process technology. Under identical test conditions, it improves area efficiency by 10.6% and reduces energy consumption by 17.0% compared to the state-of-the-art ASIC design. When compared to the latest field-programmable gate array (FPGA) implementations, it achieves a reduction in area-time product (ATP) by 7.2%~18.4% for Kyber and by 13.4% for Dilithium. These results strongly demonstrate the notable advantages of the hardware-friendly GM algorithm. Xinyu Wang 0027, Guoqiang He, Congyi Sun, Kai Chen 0034, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2026 | QSNNA: An Energy-Efficient Quaternary Spiking Neural Network Accelerator for Seizure Detection
Heng Zhang 0025, Linfeng Wu, Linxiang Wang, Youbin Luo, Haochuan Pan, Xinyu Wang 0027, Guoqiang He, Qinyu Chen, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | FAS-NoC: A Real-Time Fused Approximation Scheme Coordinating Communication and Computation for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture widely used in neural network accelerators. However, data-intensive applications like machine learning place significant demands on the NoC’s communication and computation, and often have a degree of resilience to data noise, which allows to use approximation techniques to reduce execution time and energy consumption for both computation and communication, under the constraints of acceptable quality loss. Traditional approximate NoCs do not consider the data distribution characteristics of the neural networks, resulting in a lower approximate rate. Moreover, these schemes do not take the synergistic optimization of computation and communication, which limits reductions in execution time. In this paper, we propose a Fused Approximation Scheme of NoC (FAS-NoC) that incorporates the characteristics of data distribution in neural networks. FAS-NoC includes an approximate compression and recovery scheme based on data hierarchy, and uses a congestion-aware scheme to adjust the approximate rate of the node. Additionally, by leveraging the characteristics of recovered data after approximate communication, the scheme optimizes the design of computing units within the computing array. FAS-NoC collaboratively optimizes approximate communication and computation, organically integrating the two aspects. Compared with the state-of-the-art approximate framework ACDC (ACDC_ABDTR and ACDC_APPROX), the execution time of FAS-NoC is reduced by$48.94\%$and$47.72\%$, respectively. The experimental results show that the additional area overhead of the FAS-NoC only accounts for$0.81\%$of the original node, the additional power consumption overhead only accounts for$0.75\%$of the original node. Chuanzhu Liu, Wenjie Fan 0004, Heng Zhang 0025, Chenyang Dai, Congyi Sun, Xinyu Wang 0027, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | A Hierarchical Parallel Discrete Gaussian Sampler for Lattice-Based CryptographyabstractDiscrete Gaussian sampling is one of the important components in lattice-based cryptosystems which are promising candidates for post-quantum cryptographic algorithms. For sufficient security and satisfactory performance, the Knuth-Yao algorithm is an efficient way to implement discrete Gaussian samplers. Nevertheless, most polynomials in lattice-based cryptography have 256 coefficients or more, which suffers from long latency to complete the sample generation. In this paper, the first parallel discrete Gaussian sampler with hierarchical structure is proposed, while keeping statistical distance to the actual distribution. Based on the imbalanced visiting frequency of the probability matrix, a three-stage generation strategy is adopted with hierarchical bit search units (BSUs) that can greatly reduce area consumption of the repeated costly lookup tables. Besides the architecture improvement, a lowest-set-bit scanning scheme is introduced to BSUs. Moreover, the parallelism of our design provides obfuscation ability against side-channel attacks (SCAs). A practical hardware implementation of discrete Gaussian distributions with $\sigma$=3.33 on the Xilinx Virtex-5 XC5VLX30 FPGA device spends 26.12 ns on average to generate 256 samples, consuming 994 slices. Results have verified its advantages of area efficiency over the state-of-the-arts (SOAs). Sirui Shen, Wenqing Song, Xinyu Wang 0027, Xinyu Shao, Zhonghai Lu, Li Li 0003 |
ISCAS | 3 |
| 2022 | An Energy Efficient STDP-Based SNN Architecture With On-Chip LearningabstractIn this paper, we propose a spike-time based unsupervised learning method using spiking-timing dependent plasticity (STDP). A simplified linear STDP learning rule is proposed for the energy efficient weight updates. To reduce unnecessary computations for the input spike values, a stop mechanism of the forward pass is introduced in the forward pass. In addition, a hardware-friendly input quantization scheme is used to reduce the computational complexities in both the encoding phase and the forward pass. We construct a two-layer fully-connected spiking neuron network (SNN) based on the proposed method. Compared to general rate-based SNNs trained by STDP, the proposed method reduces the complexity of network architecture (an extra inhibitory layer is not needed) and the computations of synaptic weight updates. According to the fixed-point simulation with 9-bit synaptic weights, the proposed SNN with 6144 excitatory neurons achieves 96% of recognition accuracy on MNIST dataset without any supervision. An SNN processor that contains 384 excitatory neurons with on-chip learning capability is designed and implemented with 28 nm CMOS technology based on the proposed low complexity methods. The SNN processor achieves an accuracy of 93% on MNIST dataset. The implementation results show that the SNN processor achieves a throughput of 277.78k FPS with$0.50~\mu \text{J}$/inference energy consuming in inference mode, and a throughput of 211.77k FPS with$0.66~\mu \text{J}$/learning energy consuming in learning mode. Congyi Sun, Haohan Sun, Jianing Han, Xinyuan Wang 0007, Xinyu Wang 0027, Qinyu Chen, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | High-Throughput Portable True Random Number Generator Based on Jitter-Latch StructureabstractUnder the requirement of highly reliable encryption, the design of true random number generators (TRNGs) based on field-programmable gate arrays (FPGAs) is receiving increased attention. Although TRNGs based on ring oscillators (ROs) and phase-locked loops (PLLs) have the advantages of small resource overhead and high throughput, there are problems such as instability of randomness and poor portability. To improve the randomness, portability, and throughput of a random number generator, we design a TRNG whose randomness is generated by the oscillation of self-timed rings (STRs) and accurately extracted by a jitter-latch structure. The portability of the structure is verified by electronic design automation (EDA) tools. Under the condition of 0°C-80°C ambient temperature and 1.0 ± 0.1 V output voltage, the proposed structure is tested many times on Xilinx Spartan-6 and Virtex-6 FPGAs with an automatic routing mode. Theoretical analysis shows that this method can effectively improve the coverage of jitter and reduce the migration phenomenon. Experimental results show excellent performance in randomness, robustness, and portability, and the throughput reaches 100 Mbps. Xinyu Wang 0027, Huaguo Liang, Maoxiang Yi, Zhengfeng Huang, Haochen Qi, Yingchun Lu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Pure Digital Scalable Mixed Entropy Separation Structure for Physical Unclonable Function and True Random Number GeneratorabstractThis study presents a pure digital scalable mixed entropy separation structure for the physical unclonable function (PUF) and true random number generator (TRNG), which is implemented on Xilinx field-programmable gate arrays (FPGAs). The mixed entropy separation structure in this study is to solve the problems of unstable output in the existing PUF structure and the poor scalability of the TRNG and PUF design. The proposed design has the following innovations: 1) concept of sensitive entropy and the corresponding processing method are proposed for the first time; 2) design does not need to modify the internal design of the entropy source, and it is suitable for most entropy source arrays; and 3) adjustable PUF bit width and TRNG throughput enhance the scalability of the architecture under different requirements. The structure is simulated and validated on two Xilinx FPGAs and tested under nominal working conditions. The results show that the 512-bit PUF entropy source after treatment is 92.7% more stable than the 1024-bit entropy source before treatment, the random number after treatment passes all types of randomness tests, and the minimum entropy is more than 0.8. Under various conditions within the ranges of 0 °C–80 °C and 0.8–1.2 V, the TRNG output remains stable after treatment; the maximum intra-Hamming distance of the treated PUF is 5.1028%, and the average intra-Hamming distance is 2.7065%. Yingchun Lu, Xinyu Wang 0027, Maoxiang Yi, Zhengfeng Huang, Huaguo Liang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |