Jian-Yi Meng

dblp:144/6717 · also Jianyi Meng · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Democratizing and Accelerating Hardware Verification with Software-Native Optimization
Yunlong Xie, Zhicheng Yao, Fangyuan Song, Junyue Wang, Haojin Tang, Yinan Xu 0001, Ziyuan Gao, Duan Yu, Jiayi Rao, Junyu Yue, Yunqi Lu, Zechen Yang, Xu An, Qi Ge, Jiuyue Ma, Jian-Yi Meng, Kan Shi, Dan Tang 0002, Sa Wang, Yungang Bao
ISCA23
2025 ARV-Q: An Adaptive RISC-V Vector Processor for Unified Support of Post-Quantum Standards and Side-Channel Protection on the Edge
abstract
Under the threat of quantum computers, the public-key cryptosystems need to transition to the Post-Quantum Cryptography (PQC) standards. However, this migration process is hindered by the diverse mathematical structures of PQC standards as well as the side-channel attacks, especially for the resource-constrained and physically accessible edge devices. To address this issue, we present ARV-Q, an adaptive RISC-V vector processor that efficiently offers unified support of PQC standards and side-channel security enhancement. Firstly, we propose an adaptive RISC-V-based computing paradigm to adapt to PQC algorithms across diverse mathematical bases, the core of which is a crypto extension supporting all PQC standards plus schemes in the fourth round in NIST standardization process. This crypto extension can well cooperate with the RISC-V Vector Extension (RVV) and is capable of adaptive operator-type support. Second, we design a highly resource-efficient vector crypto engine featuring versatile Butterfly Units and multi-Selected-Element-Width-adaptive modular arithmetic constructs, achieving high hardware utilization with configurable parameter settings. The crypto engine is integrated into the RISC-V core via an agile extension interface capable of bridging all types of register transfers. Besides, due to the capability of cooperative hybrid vector computing of RVV and crypto extension, ARV-Q can flexibly adapt to side-channel attacks with no hardware overhead while maintaining high performance. ARV-Q is implemented in 22nm process, and post-layout simulations are conducted. Results outperform the state-of-the-art counterparts in primary PQC standards with 1.21-5.14× better latency and more than 4.41× better area efficiency.
Yifan Zhao 0007, Honglin Kuang, Xinglong Yu, Ziyi Hao, Jian-Yi Meng, Jun Han 0003
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 Enhancing RISC-V Vector Extension for Efficient Application of Post-Quantum Cryptography
abstract
We present a cryptography extension built on RISC-V Vector Extension for efficient application of lattice-based post-quantum cryptography, offering custom instructions that can perform vectorized operations on polynomials of variable length and data width. We use micro-operation architecture to simplify the execution of variable-latency vector instructions and propose fracturable modular arithmetic units to support operations on variable coefficient width. On this basis, a vector unit is designed, achieving significant speed-up compared to the state-of-the-art counterparts for number-theoretic-transform-based polynomial multiplication. This cryptography extension is further integrated into the gem5 simulator to evaluate CRYSTALS-Kyber and CRYSTALS-Dilithium; results outperform the state-of-the-art implementations with more than 2.3 × improvement in cycle count.
Yifan Zhao 0007, Honglin Kuang, Chen Chen 0058, Jian-Yi Meng, Jun Han 0003
ASAP6
2022 Predicting the Output Structure of Sparse Matrix Multiplication with Sampled Compression Ratio
abstract
Sparse general matrix multiplication (SpGEMM) is a fundamental building block in numerous scientific applications. One critical task of SpGEMM is to compute or predict the structure of the output matrix (i.e., the number of nonzero elements per output row) for efficient memory allocation and load balance, which impact the overall performance of SpGEMM. Existing work either precisely calculates the output structure or adopts upper-bound or sampling-based methods to predict the output structure. However, these methods either take much execution time or are not accurate enough. In this paper, we propose a novel sampling-based method with better accuracy and low costs compared to the existing sampling-based method. The proposed method first predicts the compression ratio of SpGEMM by leveraging the number of intermediate products (denoted as FLOP) and the number of nonzero elements (denoted as NNZ) of the same sampled result matrix. And then, the predicted output structure is obtained by dividing the FLOP per output row by the predicted compression ratio. We also propose a reference design of the existing sampling-based method with optimized computing overheads to demonstrate the better accuracy of the proposed method. We construct 623 test cases with various matrix dimensions and sparse structures to evaluate the prediction accuracy. Experimental results show that the absolute relative errors of the proposed method and the reference design are 1.30% and 7.93%, respectively, on average, and 25% and 158%, respectively, in the worst case.
Zhaoyang Du, Yijin Guan, Tianchan Guan, Dimin Niu, Nianxiong Tan, Xiaopeng Yu 0002, Hongzhong Zheng, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001
ICPADS8
2020 Xuantie-910: Innovating Cloud and Edge Computing by RISC-V
abstract
This article consists only of a collection of slides from the author's conference presentation.
Chen Chen 0058, Xiaoyan Xiang, Chang Liu 0021, Yunhai Shang, Ren Guo, Dongqi Liu 0003, Ziyi Hao, Chunqiang Li, Yu Pu, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001, Xiaoning Qi
Hot Chips Symposium13
2020 Xuantie-910: A Commercial Multi-Core 12-Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension : Industrial Product
abstract
The open source RISC-V ISA has been quickly gaining momentum. This paper presents Xuantie-910, an industry leading 64-bit high performance embedded RISC-V processor from Alibaba T-Head division. It is fully based on the RV64GCV instruction set and it features custom extensions to arithmetic operation, bit manipulation, load and store, TLB and cache operations. It also implements the 0.7.1 stable release of RISCV vector extension specification for high efficiency vector processing. Xuantie-910 supports multi-core multi-cluster SMP with cache coherence. Each cluster contains 1 to 4 core(s) capable of booting the Linux operating system. Each single core utilizes the state-of-the-art 12-stage deep pipeline, out-of-order, multi-issue superscalar architecture, achieving a maximum clock frequency of 2.5 GHz in the typical process, voltage and temperature condition in a TSMC 12nm FinFET process technology. Each single core with the vector execution unit costs an area of 0.8 mm2, (excluding the L2 cache). The toolchain is enhanced significantly to support the vector extension and custom extensions. Through hardware and toolchain co-optimization, to date Xuantie-910 delivers the highest performance (in terms of IPC, speed, and power efficiency) for a number of industrial control flow and data computing benchmarks, when compared with its predecessors in the RISC-V family. Xuantie-910 FPGA implementation has been deployed in the data centers of Alibaba Cloud, for applicationspecific acceleration (e.g., blockchain transaction). The ASIC deployment at low-cost SoC applications, such as IoT endpoints and edge computing, is planned to facilitate Alibaba’s end-to-end and cloud-to-edge computing infrastructure.
Chen Chen 0058, Xiaoyan Xiang, Chang Liu 0021, Yunhai Shang, Ren Guo, Dongqi Liu 0003, Ziyi Hao, Chunqiang Li, Yu Pu, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001, Xiaoning Qi
ISCA13
2019 A Granular Resampling Method and Adaptive Speculative Mechanism-Based Energy-Efficient Architecture for Multiclass Heartbeat Classification
abstract
This brief presents an energy-efficient design with cascaded structure aiming at multiclass heartbeat classification. support vector machine-based granular resampling method is put forward to obtain a hybrid classifier which includes a low-complexity model (LCM) to identify most easy-to-learn heartbeats and a high-accuracy classifier to discriminate the remained. The hybrid classifier combined with one-versus-all strategy is employed to achieve a multiclass classification model. An adaptive speculative mechanism based on the occurrence regularity of electrocardiogram abnormities is proposed to lower the complexity and computation burden of the multiclass classification model. The corresponding energy-efficient hardware architecture is designed and its architecture optimizations include memory segmentation to reduce energy consumption and time domain reuse to save resources. Implemented in 40-nm CMOS process, the design occupies 0.135 mm2area. It consumes 2.60-48.99 nJ/classification under 1-V voltage supply and 1 MHz operating frequency. Results show that the design provides an average prediction speedup by 60.66% and a significant energy dissipation reduction by 55.26% per beat compared with a high-accuracy model without LCMs.
Yin Xu 0003, Feiteng Li, Jian-Yi Meng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Relative ordering learning in spiking neural network for pattern recognition
Zhitao Lin, De Ma, Jian-Yi Meng, Linna Chen
Neurocomputing3
2017 Eliminating Timing Errors Through Collaborative Design to Maximize the Throughput
abstract
In advanced technology nodes, large timing margins must be added to allow for worse process, voltage, temperature, and aging variations. The error detection and correction (EDAC) technique effectively eliminates these margins by timing speculation, but the high design complexity and large hardware cost make many existing EDAC systems unsuitable for commercial processors. Based on the instruction-level locality of timing errors, a collaborative EDAC approach is proposed to address this issue. The hardware layer adopts simple and low cost EDAC circuits to ensure correct operation when timing error occurs, while a runtime software layer prevents recurring errors of the same instruction by sending timing error alarms to the hardware layer. Cooperation of both layers, accompanied with the proposed profile-guided timing error avoidance algorithm, eliminates more than 95% of errors with small runtime overhead. This significantly improves overall performance and alleviates pressure on the EDAC circuits. Experimental results based on the three-stage commercial CK802 processor in SMIC 40LL process present that the approach has improved the peak performance of the baseline EDAC system (Razor-Lite + half-frequency replay) by 8% and reduced the energy consumption by 25%, with less than 1.4% area overhead.
Zhan-Hui Li, Tao-Tao Zhu, Zhi-Jian Chen, Jian-Yi Meng, Xiaoyan Xiang, Xiaolang Yan
IEEE Trans. Very Large Scale Integr. Syst.4
2017 A Variation-Tolerant Near-Threshold Processor With Instruction-Level Error Correction
abstract
Timing error resilience is a promising alternative to eliminate margins and improve energy efficiency in subthreshold and near-threshold processors. However, the existing techniques have some limitations, such as uncontaminated architecture registers (ARs), strict timing constraints on error consolidation and propagation, and high design complexity. To address these limitations, a new timing error resilience technique based on sacrificial instruction-level registers is proposed. It dynamically captures and incrementally records the changes of ARs at each instruction boundary. Once a timing error occurs, it only needs to restore the changed ARs to a preerror state. Then, the erroneous instruction can be safely reexecuted. This technique is applicable to different processors. The 32-bit embedded processor employing the proposed technique is demonstrated in a 40-nm CMOS technology. This variation-tolerant processor operates at 27.4 MHz under 0.6 V with 8.7% total area overhead compared with the baseline without timing error resilience. At the same throughput, the proposed technique achieves 44% and 27% energy benefits compared with the baseline and the canary technique, respectively.
Sheng Wang 0005, Chen Chen 0058, Xiaoyan Xiang, Jian-Yi Meng
IEEE Trans. Very Large Scale Integr. Syst.4
2017 An Energy-Efficient and Wide-Range Voltage Level Shifter With Dual Current Mirror
abstract
This brief presents an energy-efficient level shifter (LS) to convert a subthreshold input signal to an above-threshold output signal. In order to achieve a wide range of conversion, a dual current mirror (CM) structure consisting of a virtual CM and an auxiliary CM is proposed. The circuit has been implemented and optimized in SMIC 40-nm technology. The postlayout simulation demonstrates that the new LS can achieve voltage conversion from 0.2 to 1.1 V. Moreover, at the target voltage of 0.3 V, the proposed LS exhibits an average propagation delay of 66.48 ns, a total energy per transition of 72.31 fJ, and a static power consumption of 88.4 pW, demonstrating an improvement of 6.0×, 13.1×, and 89.0×, respectively, when compared with Wilson CM-based LS.
Zhenqiang Yong, Xiaoyan Xiang, Chen Chen 0058, Jian-Yi Meng
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Error-Resilient Integrated Clock Gate for Clock-Tree Power Optimization on a Wide Voltage IOT Processor
abstract
Energy-efficiency optimization occupies an important position in the Internet of Things application. The error-resilience technique has begun to emerge and brought the performance and energy benefits as a new vision for alternative computing, because it eliminates the overconstrained margin in current processor design flow and protects the system from process, supply voltage, temperature, and aging variations through an error-resilient mechanism rather than expensive guardbands. However, as a traditional clock-tree power optimization technique, the clock gating mechanism cannot work in such a system when it faces the timing violation problem. In this paper, we propose an error-resilient integrated clock gate (ERICG) and its automatic integration methodology in error detection and correction (EDAC) system design flow. ERICG can provide the ability of in situ timing EDAC with only four additional transistors compared with a conventional integrated clock gate. The SPICE simulation shows that it is a metastable-hardened cell and can work well in the wide voltage operation (0.5~ 1.1 V) including the near-threshold region. We implement it in a commercial C-SKY CK802 processor based on an SMIC 40-nm technology. The result shows that it improves the energy efficiency by 68% compared with the non-EDAC design and lowers the total power by 28.72% over the conventional EDAC design at 0.6 V.
Tao-Tao Zhu, Jian-Yi Meng, Xiaoyan Xiang, Xiaolang Yan
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Value locality based storage compression memory architecture for ECG sensor node
Chaojun Zhao, Chen Chen 0058, Jian-Yi Meng
Sci. China Inf. Sci.4
2014 Preservation of local linearity by neighborhood subspace scaling for solving the pre-image problem
abstract
An important issue involved in kernel methods is the pre-image problem. However, it is an ill-posed problem, as the solution is usually nonexistent or not unique. In contrast to direct methods aimed at minimizing the distance in feature space, indirect methods aimed at constructing approximate equivalent models have shown outstanding performance. In this paper, an indirect method for solving the pre-image problem is proposed. In the proposed algorithm, an inverse mapping process is constructed based on a novel framework that preserves local linearity. In this framework, a local nonlinear transformation is implicitly conducted by neighborhood subspace scaling transformation to preserve the local linearity between feature space and input space. By extending the inverse mapping process to test samples, we can obtain pre-images in input space. The proposed method is non-iterative, and can be used for any kernel functions. Experimental results based on image denoising using kernel principal component analysis (PCA) show that the proposed method outperforms the state-of-the-art methods for solving the pre-image problem.
Sheng-Kai Yang, Jian-Yi Meng, Haibin Shen
J. Zhejiang Univ. Sci. C2