EDBT 2026 Demo / reviewers in the wild / expert
Jiawen Cheng
dblp:270/4186
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing General Sparse Linear-Equation Solvers via Nested-Dissection-Based Parallel Scheduling and Randomized Linear AlgebraabstractSparse linear-equation solvers are indispensable to circuit simulation. They provide the mathematical engine that faithfully forecasts the dynamic response of analog circuits. In this invited paper, we present two novel techniques to speed up these solvers. The first is a parallel LU factorization driven by a new task scheduling strategy. Based on the nested dissection approach for matrix reordering, we derive a task assignment/scheduling strategy which largely reduces synchronization and develops more parallelism. Thus, a more efficient parallel sparse LU factorization algorithm (named SubtreeLU) is obtained. It outperforms both PARDISO and CKTSO in computational speed while remaining similar robustness. The second technique is a practical randomized GMRES algorithm. By implementing the Gram-Schmidt process with an extremely efficient random sketched linear-least-squares kernel, we obtain a fast randomized Arnoldi procedure that orthogonormalizes the Krylov subspace basis. Coupled with on-the-fly residual error estimates, this yields a practical randomized GMRES that is provably stable and runs remarkably faster than the standard GMRES on a wide range of circuit and field simulation benchmarks. Wenjian Yu, Jiawen Cheng |
ASP-DAC | 2 |
| 2026 | A Parallel Mixed-Precision GMRES-IR Solver for Ill-Conditioned Equations in Device SimulationabstractEfficient and reliable device simulation remains a critical challenge for modern electronic design automation (EDA), where ill-conditioned sparse linear equation systems are often solved. Traditional linear matrix solvers struggle to balance accuracy, performance, and scalability concurrently in the presence of ill-conditioning. In this work, we propose a parallel solver framework that integrates mixed-precision iterative refinement with GMRES algorithm and novel architecture-aware optimizations on modern CPUs. Our approach leverages vectorization, parallel scheduling, and memory hierarchy optimizations to accelerate Krylov subspace methods while preserving numerical robustness. Comprehensive evaluation on matrices arising from realistic device simulation problems demonstrates that our solver achieves 5.4× speedup on average, compared to the high-precision direct solver baseline, while maintaining solution accuracy within given tolerances. Moreover, the proposed mixed-precision GMRES-IR solver attains further 3.3× parallel speedup with 8 threads, demonstrating its parallel efficiency. Jiawen Cheng, Ding Gong, Wenjian Yu |
DATE | 1 |
| 2026 | Efficient Parallel ILU Factorization and Forward/Backward Substitution with Application to Large-Scale Nonlinear Circuit SimulationabstractEfficient techniques are proposed for parallel incomplete LU (ILU) factorization and forward/backward substitution, for the sparse matrices with the same sparsity pattern. These parallel algorithms are then used as a preconditioner for the generalized minimal residual (GMRES) algorithm to obtain an ILU-GMRES solver for large-scale circuit simulation. The novelty of the parallel ILU and substitution algorithms includes a subtree-based task scheduling scheme, a nested dissection-based approach for generating task queues, and the task packing and reverse-order execution techniques for forward/backward substitution. Experiments on 43 matrices dumped from circuit simulation show that the 8-thread parallel ILU with threshold (ILUT) factorization and forward/backward substitution with the proposed techniques achieve 4.2 \(\times\) and 3.3 \(\times\) parallel speedups on average, respectively. The proposed parallel ILUT-GMRES solver runs 5.3 \(\times\) , on average, faster than PARDISO on these benchmarks. When integrated into Ngspice, it enables up to 2.3 \(\times\) and 1.7 \(\times\) faster execution of a step of Newton-Raphson iteration than the commercial parallel HSPICE, for the DC analysis and time integration stages, respectively. Jiawen Cheng, Shan Shen, Zhenya Zhou, Wenjian Yu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | A Practical Randomized GMRES Algorithm for Solving Linear Equation System in Circuit SimulationabstractEfficient solver for general linear equations is of significance for EDA problems. The generalized minimal residual (GMRES) method, which can solve general linear equations efficiently, is one of the most widely-used fundamental algorithms. Randomized Arnoldi process, which leverages sketched least-squares solver to orthogonalize Krylov subspace basis, has shown potential to promote the effectiveness of Arnoldi process, the core step in GMRES. However, how to make it more efficient, and utilize it to develop a practical GMRES solver is still an open problem. In this work, we aim at obtaining a practically-useful randomized GMRES algorithm (named PRGMRES) for solving general sparse linear equations. Firstly, an efficient estimator of residual error based on a modified randomized Gram-Schmidt process and a double-tolerance scheme are proposed to enable a practical restarted GMRES algorithm which terminates at a solution satisfying the specified accuracy tolerance. Then, a linear-time-complexity sketching algorithm based on Rademacher matrices is proposed to facilitate fast and robust random sketching. After that, incremental solution of the sketched least-squares problems, and the theoretical analysis supporting smaller sketching size are presented. Based on the above proposed techniques and theoretical results, the PRGMRES algorithm, which has stronger theoretically-supported stability and efficiency, is proposed. Numerical experiments on various circuit simulation problems validate the efficiency and effectiveness of the proposed algorithm. Jiawen Cheng, Wenjian Yu |
ASP-DAC | 2 |
| 2025 | SubtreeLU: High-Performance Parallel Sparse LU Factorization for Circuit SimulationabstractSolving sparse linear systems via LU factorization remains a critical performance bottleneck in SPICE-based circuit simulation. Existing solvers such as NICSLU and CKTSO have introduced various improvements, but they continue to face limitations such as synchronization overheads, conservative dependency estimation, and under-utilization of supernodal structures. In this work, we propose SubtreeLU, a high-performance parallel sparse LU factorization framework tailored for circuit simulation. SubtreeLU introduces a novel scheduling framework that either collapses or partitions the separator tree generated by nested dissection to organize computation into a private-pipeline structure. This enables efficient support for both pivoting and non-pivoting modes. Supernodal methods are integrated to further accelerate numerical updates. Extensive experiments on 46 circuit matrices demonstrate that SubtreeLU consistently outperforms state-of-the-art solvers in terms of accuracy, runtime, and scalability. The proposed SubtreeLU solver is available at https://numbda.cs.tsinghua.edu.cn/download.html. Jiawen Cheng, Wenjian Yu |
ICCAD | 1 |
| 2025 | PCB-AM: Enhanced Defect Detection for PCBs via Attention-Guided Modules
Shuai Wang 0083, Jiawen Cheng, Qiushuang Yu, Jianjian Chen |
ICIC (13) | 2 |
| 2024 | Nested Dissection Based Parallel Transient Power Grid Analysis on Public Cloud Virtual MachinesabstractAccurate and efficient transient analysis of power grids (PGs) poses a large challenge of computation for nowadays integrated circuit design. In this work, we propose to leverage the public cloud computing to do PG transient analysis while preserving security. A multi-level distributed parallel LU factorization and forward/backward substitution approach based on nested dissection is then proposed to guarantee accuracy and robustness. Experimental results show that the proposed algorithm can achieve an average 2.06X speedup over NICSLU and 2.85X over conventional domain decomposition method based parallel approach. And, it exhibits good scalability with up to 6.0X parallel speedup on large-scale PGs with 4 cloud computer nodes. Jiawen Cheng, Wenjian Yu |
ASPDAC | 1 |
| 2024 | Neural Canvas: Supporting Scenic Design Prototyping by Integrating 3D Sketching and Generative AIabstractWe propose Neural Canvas, a lightweight 3D platform that integrates sketching and a collection of generative AI models to facilitate scenic design prototyping. Compared with traditional 3D tools, sketching in a 3D environment helps designers quickly express spatial ideas, but it does not facilitate the rapid prototyping of scene appearance or atmosphere. Neural Canvas integrates generative AI models into a 3D sketching interface and incorporates four types of projection operations to facilitate 2D-to-3D content creation. Our user study shows that Neural Canvas is an effective creativity support tool, enabling users to rapidly explore visual ideas and iterate 3D scenic designs. It also expedites the creative process for both novices and artists who wish to leverage generative AI technology, resulting in attractive and detailed 3D designs created more efficiently than using traditional modeling tools or individual generative AI platforms. Yulin Shen 0001, Yifei Shen 0002, Jiawen Cheng, Chutian Jiang, Mingming Fan 0001, Zeyu Wang 0003 |
CHI | 3 |
| 2024 | LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresabstractStencil computations play a pivotal role in numerous scientific and industrial applications, yet their efficient execution on specialized hardware accelerators like Tensor Core Units (TCUs) remains a challenge. This paper introduces LoRAStencil1, a novel stencil computing system designed to mitigate memory access redundancies on TCUs through low-rank adaptation. We first identify a nuanced form of this redundancy, dimension residue, specific to TCUs. Then LoRAStencil leverages orchestrated mathematical transformations to decompose stencil weight matrices into smaller rank-1 matrices, facilitating efficient data gathering along residual dimensions. It comprises three key components: memory-efficient Residual Dimension Gathering to facilitate more data reuse, compute-saving Pyramidal Matrix Adaptation to exploit the inherent low-rank characteristics, and performance-boosting Butterfly Vector Swapping to circumvent all data shuffles. Comprehensive evaluations demonstrate that LoRAStencil address dimension residues effectively, which outperforms state-of-the-arts with up to a 2.16x speedup, offering promising advancements for efficient tensorized stencil computation on TCUs by Low-Rank Adaptation. Yiwei Zhang 0009, Kun Li 0016, Jiawen Cheng, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
SC | 4 |
| 2024 | Music genre classification based on res-gated CNN and attention mechanism
Changjiang Xie, Huazhu Song, Kaituo Mi, Zhouhan Li, Jiawen Cheng, Honglin Zhou, Haofeng Cai |
Multim. Tools Appl. | 7 |
| 2023 | Machine-learning-driven Architectural Selection of Adders and Multipliers in Logic SynthesisabstractDesigning high-performance adders and multiplier components for diverse specifications and constraints is of practical concern. However, selecting the best architecture for adder or multiplier, which largely affects the performance of synthesized circuits, is difficult. To tackle this difficulty, a machine-learning-driven approach is proposed for automatic architectural selection of adders and multipliers. It trains a machine learning model for classification through learning a number of existing design schemes and their performance data. Experimental results show that the proposed approach based on a multi-perception neural network achieves as high as 94% prediction accuracy with negligible inference time. On a CPU server, the proposed approach runs about 4× faster than a brute-force approach trying four candidate architectures and consumes 10%~20% less runtime than the DesignWare datapath generator for obtaining the optimal adder/multiplier circuit. The adder (multiplier) generated with the proposed approach achieves performance metrics close to the optimal and has 1.6% (5.2%) less area and 2.2% (7.1%) more worst negative slack averagely than that generated with the DesignWare datapath generator. Our experiment also shows that the proposed approach is not sensitive to the size of training subset. Jiawen Cheng, Yun Shao 0008, Guanghai Dong, Songlin Lyu, Wenjian Yu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2020 | Social Data Assisted Multi-Modal Video Analysis For Saliency DetectionabstractVideo saliency should be taken into consideration to facilitate optimization of the end-to-end video production, delivery and consumption ecosystem to improve user experience at lowered cost. Although recent studies have significantly increased the accuracy of saliency prediction, the approaches are mostly video-centric, without considering any prior "bias" that viewers may have with regard to the video contents. In this paper, we propose a novel learning-based multi-modal method for optimizing user-oriented video analysis. In particular, we generate a face-popularity mask using face recognition results and popularity information obtained from social media, and combine it with conventional content-only saliency analysis to produce multi-modal popularity-motion features. A convolutional long short-term memory (ConvL- STM) network discovers temporal correlation of human attention across frames. Experiments show that our method outperforms the state-of-the-art video saliency prediction approaches in representing human viewing preferences in real world applications, and demonstrate the necessity as well as the potential for integrating user bias information into attention detection. Jiangyue Xia, Jingqi Tian, Jiankai Xing, Jiawen Cheng, Jiangtao Wen, Zhengguang Li, Jian Lou 0003 |
ICASSP | 4 |