Jinming Lyu

dblp:364/6195 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
circuit simulation
0.812024
On Model Order Reduction and Exponential Integrator for Transient Circuit Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Electronic design automation › circuit simulation › transient analysis
exponential integrator
0.812024
On Model Order Reduction and Exponential Integrator for Transient Circuit Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Electronic design automation › circuit simulation
model order reduction
0.812024
On Model Order Reduction and Exponential Integrator for Transient Circuit Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Electronic design automation › circuit simulation
transient analysis
0.812024
On Model Order Reduction and Exponential Integrator for Transient Circuit Simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024

Methods — techniques the papers use, named apart from their topics

rational krylov subspace projection · 0.8moment matching · 0.8krylov subspace approximation · 0.8
YearPublicationVenuePosition
2025 C2OPU: Hybrid Compute-in-Memory and Coarse-Grained Reconfigurable Architecture for Overlay Processing of Transformers
abstract
Transformer-based models have shown huge success in natural language processing (NLP) with increasing model size and attention mechanism. However, this makes Von Neumann architecture based accelerators memory-bound such that the accelerators cannot leverage all the advantages of Transformer-based models. Although computing-in-memory (CIM) processors have emerged to tackle this problem through in-situ computing, the mismatch of computing patterns and low computing precision of CIM make it still challenging to accelerate Transformers. In this paper, we propose C2OPU, a hybrid dual-core processor to accelerate Transformers with hardware and software co-optimization. The dual-core architecture uses CIM arrays to accelerate weight-stationary vector-matrix multiplications, which accounts for the main computation complexity of Transformers. Meanwhile, a coarse-grained reconfigurable architecture (CGRA) is used to address the issues of computing pattern mismatch and low precision of the CIM. In addition, we propose an accuracy-bound workload allocation strategy, which considers non-ideal characteristics in analog computing, to balance throughput and accuracy. Furthermore, C2OPU provides a compiler to automatically determine optimal system configurations when Transformer model changes. Experimental results show that C2OPU achieves an average speedup of 145.41×, 4.73×, 4.70x and 3.85×, and 1.37x compared to CPU, GPU, Science23, Nature23, and VLSI24, respectively, on ten different Transformer models.
Siyuan Miao, Lingkang Zhu, Shaoqiang Lu, Jinming Lyu, Lei He 0001
FCCM5
2024 On Model Order Reduction and Exponential Integrator for Transient Circuit Simulation
abstract
Model order reduction (MOR) has long been a mainstream strategy to accelerate large scale transient circuit simulation. Exponential integrator (EI) based on Krylov subspace approximation methods, on the other hand, are more recently developed for a similar goal. This article aims to examine in-depth the underlying relationship between model order reduction (MOR) and exponential integrator (EI) that are commonly seen as two separate methods. The main finding is that EI can be viewed as a moment-matching MOR in the time-domain. Specifically, EI, under certain conditions, is equivalent to performing moment-matching MOR based on rational Krylov subspace projection at each time step with a single input vector and a selected expansion point, then advancing the reduced system one step in the time-domain. The equivalence is mathematically proved under different settings and numerically verified in the experiments. Their differences in the transient circuit analysis context are also elaborated from various perspectives. It is hoped that these new insights would benefit the future development of this classical EDA topic.
Cong Wang 0040, Dongen Yang, Jinming Lyu, Cheng Zhuo, Quan Chen 0007
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3