EDBT 2026 Demo / reviewers in the wild / expert
Jianan Mu
dblp:314/8846
· DBLP profile ↗
27ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0001-8513-0792ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 5 first-author · 24 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VecTEE: Compact TEE Metadata Caching for Efficient Secure Vector Computing
Husheng Han, Tianyun Ma, Xinyao Zheng, Jianan Mu, Zidong Du, Xing Hu 0001, Qi Guo 0001 |
APPT | 5 |
| 2026 | FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) imposes substantial memory demands, presenting significant challenges for efficient hardware acceleration. Near-Memory Processing (NMP) has emerged as a promising architectural solution to alleviate the memory bottleneck. However, the irregular memory access patterns and flexible dataflows inherent to FHE limit the effectiveness of existing NMP accelerators, which fail to fully utilize the available near-memory bandwidth. In this work, we propose FlexMem, a near-memory accelerator featuring high-parallel computational units with varying memory access strides and interconnect topologies to effectively handle irregular memory access patterns. Furthermore, we design polynomialand ciphertext-level dataflows to efficiently utilize near-memory bandwidth under varying degrees of polynomial parallelism and enhance parallel performance. Experimental results demonstrate that FlexMem achieves $1.26 \times$ performance improvement over the state-of-the-art near-memory architectures in end-to-end benchmarks, with on average 95.7% of near-memory bandwidth utilization. Shangyi Shi, Husheng Han, Jianan Mu, Xinyao Zheng, Ling Liang 0003, Zidong Du, Xiaowei Li 0001, Xing Hu 0001 |
ASP-DAC | 3 |
| 2026 | Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToTabstractLarge language models (LLMs) hold promise for automating integrated circuit (IC) engineering using register transfer level (RTL) hardware description languages (HDLs) like Verilog. However, challenges remain in ensuring the quality of Verilog generation. Complex designs often fail in a single generation due to the lack of targeted decoupling strategies, and evaluating the correctness of decoupled sub-tasks remains difficult. While the chain-of-thought (CoT) method is commonly used to improve LLM reasoning, it has been largely ineffective in automating IC design workflows, requiring manual intervention. The key issue is controlling CoT reasoning direction and step granularity, which do not align with expert RTL design knowledge. This paper introduces VeriBToT, a specialized LLM reasoning paradigm for automated Verilog generation. By integrating Top-down and design-for-verification (DFV) approaches, VeriBToT achieves self-decoupling and self-verification of intermediate steps, constructing a Backtrack Tree of Thought with formal operators. Compared to traditional CoT paradigms, our approach enhances Verilog generation while optimizing token costs through flexible modularity, hierarchy, and reusability. Zhiteng Chao, Yonghao Wang, Tenghui Hua, Husheng Han, Tianmeng Yang, Jianan Mu, Bei Yu 0001, Rui Zhang 0040, Jing Ye 0001, Huawei Li 0001 |
DATE | 8 |
| 2026 | He2: A Communication-Light Heterogeneous Architecture for Efficient Fully Homomorphic Encryption
Shangyi Shi, Husheng Han, Zhaoxuan Kan, Jianan Mu, Tenghui Hua, Xinyao Zheng, Ling Liang 0003, Zidong Du, Xing Hu 0001 |
ISCA | 5 |
| 2026 | DomSim: Hardware-Aware Hybrid Fault Simulation With Dominator Tree-Guided PartitioningabstractGate-level fault simulation is a critical step in design for test and functional safety verification of the chip design process, essential to ensuring circuit reliability. As chip complexity grows for mission-critical applications such as autonomous vehicles, medical devices, and military systems, the efficiency of fault simulation increasingly becomes a bottleneck in the chip’s time-to-market. However, existing methods often suffer from computational redundancy, inefficiencies in memory access, or failure to optimize performance for specific CPU hardware platforms. This paper proposes DomSim, a hardware-aware hybrid fault simulation method that combines compiled simulation and event-driven simulation with an optimized computation-to-memory-access ratio. By utilizing circuit information and hierarchical structure provided by dominator trees, DomSim achieves high-quality circuit partitioning, optimizing hardware resource utilization and memory access locality. Furthermore, a parameter adjustment strategy tailored to hardware capabilities and circuit characteristics enables adaptive optimization. Extensive experiments show that DomSim surpasses a commercial tool by 10.29× on average. Further experiments demonstrate that DomSim exhibits good adaptability across different hardware platforms and circuits, highlighting the superiority of our method. Hui Wang 0152, Zizhen Liu, Jianan Mu, Shengwen Liang, Zhongkai Yu, Zheng Liang 0003, Jiaping Tang, Jing Ye 0001, Xiaowei Li 0001, Bei Yu 0001, Huawei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | ETPG: Efficient Transition Fault Simulation via Dual-Strategy Pattern Parallelism and Gate RestructuringabstractWith the advancement of integrated circuit (IC) technology, the sensitivity to delay defects has significantly increased, rendering Transition Fault (TF) testing crucial for ensuring chip quality. However, as the complexity of IC designs increases, existing pattern parallelization methods are not flexible in detecting multi-cycle faults. In addition, the growing demand for simulation memory exacerbates inefficient memory access, becoming another critical bottleneck. This paper introduces ETPG (Efficient Transition fault simulation via dual-strategy Pattern parallelism and Gate restructuring), a novel TF simulation algorithm based on multi-dimensional optimization. The key innovations include an adaptive dual-strategy pattern parallel strategy that dynamically optimizes parallelization based on test pattern characteristics, enhancing efficiency and multi-cycle fault detection capability; a dual-dimension gate restructuring method that optimizes memory storage order, significantly reducing memory access time, particularly beneficial for large-scale circuits; and a collaborative mechanism between pattern processing and circuit storage optimization, achieving comprehensive performance improvements at both algorithmic and memory access levels. Experimental results demonstrate ETPG's significant performance improvements across various circuit scales, particularly for larger circuits. Compared to the synopsys commercial tool testmax (TMAX), ETPG achieves average speedups of 2.846× for circuits below 100k gates and 4.428× for circuits above 100k gates. Hui Wang 0152, Zizhen Liu, Jianan Mu, Jiaping Tang, Huawei Li 0001, Jing Ye 0001, Xiaowei Li 0001 |
ASP-DAC | 5 |
| 2025 | PastATPG: A Hybrid ATPG Framework for Better Test Compaction with Partial Assignment SATabstractIn automatic test pattern generation (ATPG), SAT-based methods are typically used to complement structural approaches, especially for addressing hard-to-detect faults. However, as the size and complexity of circuits grow, SAT-based ATPG faces challenges like pattern inflation and excessive runtime, limiting its overall performance. The key problem lies in the fact that current mainstream SAT solvers perform complete assignments for all primary inputs of the fault’s transitive fanin cone without considering the detection of other faults, making test compaction extremely difficult and time consuming. In this paper, a novel SAT solver PA-MiniSat is proposed, which is capable of generating partial assignments for solving variables and significantly reduces the number of specified bits in test cubes. As an extension of MiniSat, it employs a full-literal watching technique and a circuit-adapted heuristic branching strategy, achieving overall improved performance in ATPG. Based on PA-MiniSat, a hybrid ATPG framework PastATPG is proposed for better test compaction, which tightly integrates structural algorithms with the SAT solver into the unified test compaction flow. Experimental results demonstrate that our method outperforms other SAT solvers in pattern compaction and, in some cases, even surpasses commercial ATPG tools in terms of speed. The code is available at https://github.com/sklp-eda-lab/PastATPG. Zhiteng Chao, Xindi Zhang 0001, Jianan Mu, Zizhen Liu, Shengwen Liang, Shaowei Cai 0001, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001 |
DAC | 4 |
| 2025 | MOSS: Multi-Modal Representation Learning on Sequential CircuitsabstractDeep learning has significantly advanced Electronic Design Automation (EDA), with circuit representation learning emerging as a key area for modeling the relationship between a circuit’s structure and functionality. Existing methods primarily use either Large Language Models (LLMs) for Register Transfer Level (RTL) code analysis or Graph Neural Networks (GNNs) for netlist modeling. While LLMs excel at high-level functional understanding, they struggle with detailed netlist behavior. GNNs, however, face challenges when scaling to larger sequential circuits due to long-range information dependencies and insufficient functional supervision, leading to decreased accuracy and limited generalization. To address these challenges, we propose MOSS, a multimodal framework that integrates GNNs with LLMs for sequential circuit modeling. By enhancing D-type Flip-Flop (DFF) node features with embeddings from fine-tuned LLMs on RTL code, we focus the GNN on critical anchor points, reducing reliance on long-range dependencies. The LLM also provides global circuit embeddings, offering efficient supervision for functionality-related tasks. Additionally, MOSS introduces an adaptive aggregation method and a two-phase propagation mechanism in the GNN to better model signal propagation and sequential feedback within the circuit. Experimental results demonstrate that MOSS significantly improves the accuracy of functionality and performance predictions for sequential circuits compared to existing methods, particularly in larger circuits where previous models struggle. Specifically, MOSS achieves a $\mathbf{9 5. 2 \%}$ accuracy in arrival time prediction. Jianan Mu, Tianmeng Yang, Silin Liu, Yihan Wen, Hui Wang 0152, Zhiteng Chao, Husheng Han, Zizhen Liu, Shengwen Liang, Jing Ye 0001, Bei Yu 0001, Xiaowei Li 0001, Huawei Li 0001 |
DAC | 3 |
| 2025 | EPICS: Efficient Parallel Pattern Fault Simulation for Sequential Circuits via Strongly Connected ComponentsabstractAs functional safety of electronic chips gains importance in autonomous vehicles and aerospace, standards like ISO 26262 mandate high diagnostic coverage, requiring extensive gate-level fault simulations. However, for large-scale industrial sequential circuits, these simulations are time-consuming, creating a significant bottleneck in chip development. Prior approaches have focused on reducing computational complexity and optimizing CPU hardware usage by minimizing redundant computations during fault propagation and leveraging bit-level parallel processing capabilities. Techniques like parallel-pattern and event-driven simulations have improved performance in combinational circuits but face limitations in sequential circuits due to timing dependencies within loops. The challenge lies in parallelizing simulations across different cycles without violating these dependencies, which is exacerbated by the complex feedback structures in SCCs. In this work, we propose a novel parallel-pattern fault simulation framework that combines loop fusion with efficient event traversal to accelerate sequential circuit simulations. By compiling simple loops into larger nodes, we reduce the number of feedback events without introducing excessive redundancy. For larger SCCs, we develop specialized algorithms for selecting loop entrance nodes based on indegree analysis and implement the lazy propagation strategy for internal nodes. This approach minimizes simulation events caused by inaccurate predictions and reduces overhead associated with false event propagation. We integrate these techniques into our simulation framework, EPICS, which strategically mixes compiled and event-driven simulations to optimize performance. Experimental results demonstrate that EPICS achieves a $5.94 \times$ speedup over state-of-the-art commercial tool while maintaining the same fault coverage. Hui Wang 0152, Jianan Mu, Yihan Wen, Zizhen Liu, Shengwen Liang, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001 |
DAC | 3 |
| 2025 | ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution RedundancyabstractAs intelligent computing devices increasingly integrate into human life, ensuring the functional safety of the corresponding electronic chips becomes more critical. A key metric for functional safety is achieving a sufficient fault coverage. To meet this requirement, extensive time-consuming fault simulation of the RTL code is necessary during the chip design phase. The main overhead in RTL fault simulation comes from simulating behavioral nodes (always blocks). Due to the limited fault propagation capacity, fault simulation results often match the good simulation results for many behavioral nodes. A key strategy for accelerating RTL fault simulation is the identification and elimination of redundant simulations. Existing methods detect redundant executions by examining whether the fault inputs to each RTL node are consistent with the good inputs. However, we observe that this input comparison mechanism overlooks a significant amount of implicit redundant execution: although the fault inputs differ from the good inputs, the node's execution results remain unchanged. Our experiments reveal that this overlooked redundant execution constitutes nearly half of the total execution overhead of behavioral nodes, becoming a significant bottleneck in current RTL fault simulation. The underlying reason for this overlooked redundancy is that, in these cases, the true execution paths within the behavioral nodes are not affected by the changes in input values. In this work, we propose a behavior-level redundancy detection algorithm that focuses on the true execution paths. Building on the elimination of redundant executions, we further developed an efficient RTL fault simulation framework, Eraser. Experimental results show that compared to commercial tools, under the same fault coverage, our framework achieves a 3.9 × improvement in simulation performance on average. Jiaping Tang, Jianan Mu, Silin Liu, Zizhen Liu, Leyan Wang, Shengwen Liang, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
DATE | 2 |
| 2025 | VIRTUAL: Vector-based Dynamic Power Estimation via Decoupled Multi-Modality LearningabstractDynamic power analysis in digital integrated circuits (ICs) conventionally relies on gate-level synthesis and simulation, creating a critical bottleneck in iterative design flows. We propose VIRTUAL, a multi-modality learning framework for rapid post-synthesis dynamic power estimation directly from Register-Transfer Level (RTL) implementations and input waveform vectors, eliminating the need for gate-level synthesis and extensive simulation. By decoupling features from input port waveforms and RTL implementations, VIRTUAL employs a transformer-based encoder to extract temporal patterns from input port waveforms and a graph neural network (GNN) to capture structural and functional dependencies within RTL implementations. Through self-supervised contrastive learning across sequential and graph modalities, the framework learns robust power-relevant representations with minimal labeled data. Subsequently, VIRTUAL refines the multi-modality embeddings using a lightweight fusion module and a power prediction head, enabling dynamic power estimation for fixed clock periods within seconds or minutes. Experimental evaluations demonstrate approximately a Pearson correlation coefficient (PCC) of 0.842 and a mean absolute percentage error (MAPE) of 23.43%, while achieving 14.27×speedup compared to traditional gate-level power analysis workflows. Experimental results across diverse RTL designs and input port waveforms validate that our proposed learning-based approach maintains the accuracy while significantly reducing design iteration time, transforming hours of synthesis and simulation into minutes of direct prediction. Yuntao Lu, Yihan Wen, Jianan Mu, Huawei Li 0001, Bei Yu 0001 |
ICCAD | 5 |
| 2025 | RIROS: A Parallel RTL Fault SImulation FRamework with TwO-Dimensional Parallelism and Unified ScheduleabstractWith the rapid development of safety-critical applications such as autonomous driving and embodied intelligence, the functional safety of the corresponding electronic chips becomes more critical. Ensuring chip functional safety requires performing a large number of time-consuming RTL fault simulations during the design phase, significantly increasing the verification cycle. To meet time-to-market demands while ensuring thorough chip verification, parallel acceleration of RTL fault simulation is necessary. Due to the dynamic nature of fault propagation paths and varying fault propagation capabilities, task loads in RTL fault simulation are highly imbalanced, making traditional single-dimension parallel methods, such as structural-level parallelism, ineffective. Through an analysis of fault propagation paths and task loads, we identify two types of tasks in RTL fault simulation: tasks that are few in number but high in load, and tasks that are numerous but low in load. Based on this insight, we propose a two-dimensional parallel approach that combines structural-level and fault-level parallelism to minimize bubbles in RTL fault simulation. Structural-level parallelism combining with work-stealing mechanism is used to handle the numerous low-load tasks, while fault-level parallelism is applied to split the high-load tasks. Besides, we deviate from the traditional serial execution model of computation and global synchronization in RTL simulation by proposing a unified computation/global synchronization scheduling approach, which further eliminates bubbles. Finally, we implemented a parallel RTL fault simulation framework, RIROS. Experimental results show a performance improvement of 7.0× and 11.0× compared to the state-of-the-art RTL fault simulation and a commercial tool. Jiaping Tang, Jianan Mu, Zizhen Liu, Tenghui Hua, Silin Liu, Jing Ye 0001, Huawei Li 0001 |
ICCAD | 2 |
| 2025 | FicGCN: Unveiling the Homomorphic Encryption Efficiency from Irregular Graph Convolutional NetworksabstractGraph Convolutional Neural Networks (GCNs) have gained widespread popularity in various fields like personal healthcare and financial systems, due to their remarkable performance. Despite the growing demand for cloud-based GCN services, privacy concerns over sensitive graph data remain significant. Homomorphic Encryption (HE) facilitates Privacy-Preserving Machine Learning (PPML) by allowing computations to be performed on encrypted data. However, HE introduces substantial computational overhead, particularly for GCN operations that require rotations and multiplications in matrix products. The sparsity of GCNs offers significant performance potential, but their irregularity introduces additional operations that reduce practical gains. In this paper, we propose FicGCN, a HE-based framework specifically designed to harness the sparse characteristics of GCNs and strike a globally optimal balance between aggregation and combination operations. FicGCN employs a latency-aware packing scheme, a Sparse Intra-Ciphertext Aggregation (SpIntra-CA) method to minimize rotation overhead, and a region-based data reordering driven by local adjacency structure. We evaluated FicGCN on several popular datasets, and the results show that FicGCN achieved the best performance across all tested datasets, with up to a $4.10\times$ improvement over the latest design. Zhaoxuan Kan, Husheng Han, Shangyi Shi, Tenghui Hua, Xiaowei Li 0001, Jianan Mu, Xing Hu 0001 |
ICML | 7 |
| 2025 | Bridging Layout and RTL: Knowledge Distillation based Timing PredictionabstractAccurate and efficient timing prediction at the register-transfer level (RTL) remains a fundamental challenge in electronic design automation (EDA), particularly in striking a balance between accuracy and computational efficiency. While static timing analysis (STA) provides high-fidelity results through comprehensive physical parameters, its computational overhead makes it impractical for rapid design iterations. Conversely, existing RTL-level approaches sacrifice accuracy due to the limited physical information available. We propose RTLDistil, a novel cross-stage knowledge distillation framework that bridges this gap by transferring precise physical characteristics from a layout-aware teacher model (Teacher GNN) to an efficient RTL-level student model (Student GNN), both implemented as graph neural networks (GNNs). RTLDistil efficiently predicts key timing metrics, such as arrival time (AT), and employs a multi-granularity distillation strategy that captures timing-critical features at node, subgraph, and global levels. Experimental results demonstrate that RTLDistil achieves significant improvement in RTL-level timing prediction error reduction, compared to state-of-the-art prediction models. This framework enables accurate early-stage timing prediction, advancing EDA’s “left-shift” paradigm while maintaining computational efficiency. Our code and dataset will be publicly available at https://github.com/sklp-eda-lab/RTLDistil. Yihan Wen, Jianan Mu, Jing Ye 0001, Bei Yu 0001, Huawei Li 0001 |
ICML | 4 |
| 2025 | TESLA: Testability Enhancement for Shift-Left Automation via Multi-LLM CollaborationabstractThe "Shift-Left" Design-for-Test (DFT) paradigm has gained significant attention in recent years, enabling early-stage testability enhancement at the Register Transfer Level (RTL) to optimize Power-Performance-Area-Testability (PPAT) trade-offs and accelerate Time-to-Market (TTM). However, existing methods struggle to perform quantitative testability analysis at the RTL stage, particularly in Partial Scan Selection (PSS) and Test Point Insertion (TPI), due to the lack of structured netlist representations and cross-stage optimization. To address this challenge, we propose TESLA, a multi-LLM collaboration framework that autonomously performs PSS and TPI at the RTL stage. TESLA leverages the semantic understanding capabilities of Large Language Models (LLMs) to analyze RTL Verilog code and optimize testability without requiring synthesis. Two key data augmentation strategies are introduced for efficient Instruction Tuning: (1) back-annotating heuristic PSS results from the synthesized netlist to RTL, and (2) utilizing advanced LLMs guided by DFT knowledge to generate synthetic RTL TPI training data. Furthermore, we integrate Direct Preference Optimization (DPO) to refine LLM decision-making, incorporating real feedback from commercial EDA tools to align optimization objectives with practical testability metrics. The experimental results demonstrate that our proposed approach achieves better test coverage compared to other RTL stage PSS and TPI combination schemes on the majority of circuits in the RTLLM benchmark, while also reducing the number of patterns for a significant portion of the circuits. On the larger, hierarchical OpenCores benchmark, our approach surpasses the solution combining heuristic PSS and commercial DFT tool’s TPI, achieving improvements on the same two test metrics. Zhiteng Chao, Rengang Zhang, Hongqin Lyu, Wenxing Li, Zizhen Liu, Jianan Mu, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001 |
ITC | 8 |
| 2025 | QiMeng-CodeV-R1: Reasoning-Enhanced Verilog GenerationabstractLarge language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automatically generating hardware description languages (HDLs) like Verilog from natural-language (NL) specifications, however, poses three key challenges: the lack of automated and accurate verification environments, the scarcity of high-quality NL-code pairs, and the prohibitive computation cost of RLVR. To this end, we introduce CodeV-R1, an RLVR framework for training Verilog generation LLMs. First, we develop a rule-based testbench generator that performs robust equivalence checking against golden references. Second, we propose a round-trip data synthesis method that pairs open-source Verilog snippets with LLM-generated NL descriptions, verifies code–NL–code consistency via the generated testbench, and filters out inequivalent examples to yield a high-quality dataset. Third, we employ a two-stage "distill-then-RL" training pipeline: distillation for the cold start of reasoning abilities, followed by adaptive DAPO, our novel RLVR algorithm that can reduce training cost by adaptively adjusting sampling rate. The resulting model, CodeV-R1-7B, achieves 68.6 \% and 72.9 \% pass@1 on VerilogEval v2 and RTLLM v1.1, respectively, surpassing prior state-of-the-art by 12$\sim$20 \%, while even exceeding the performance of 671B DeepSeek-R1 on RTLLM. We have released our model, training code, and dataset to facilitate research in EDA and LLM communities. Yaoyu Zhu, Han-Qi Lyu, Chongxiao Li, Jianan Mu, Yang Zhao 0013, Pengwei Jin, Shuyao Cheng, Shengwen Liang, Xishan Zhang, Rui Zhang 0040, Zidong Du, Qi Guo 0001, Xing Hu 0001, Yunji Chen |
NeurIPS | 8 |
| 2025 | HighTPI: A Hierarchical Graph Based Intelligent Method for Test Point InsertionabstractAs integrated circuits grow in complexity, test point insertion (TPI) has become vital for enhancing testability and improving reliability in design for test (DFT). Recent studies have shown the effectiveness of deep learning-based TPI using graph neural networks (GNNs) in improving test quality. However, the high cost of collecting training data, incomplete capture of the intrinsic characteristics of circuits, and the vast search space in large circuits hinder the performance of existing intelligent approaches. This paper introduces HighTPI, a two-stage learning approach for TPI to effectively reduce the number of test patterns, which leverages hierarchical graph representation by constructing a hypergraph based on hypernodes in fanout-free regions (FFRs). HighTPI better captures multi-fanout reconvergence information while lowering the cost of obtaining ground-truth labels due to the smaller scale of the FFR-based hypergraph. Two specialized GNNs are designed in stage I to select candidate insertion points for observation and control points, respectively. This integration of expert knowledge through supervised learning helps guide the reinforcement learning process in stage II, mitigating the challenges of sparse rewards and a large decision space. The experimental results demonstrate that HighTPI outperforms other TPI methods in terms of the trade-off between pattern reduction and fault coverage enhancement. Zhiteng Chao, Hongqin Lyu, Minjun Wang, Wenxing Li, Zizhen Liu, Jianan Mu, Shengwen Liang, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001 |
VTS | 8 |
| 2024 | TensorTEE: Unifying Heterogeneous TEE Granularity for Efficient Secure Collaborative Tensor ComputingabstractHeterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is considered a promising solution because of its comparatively lower overhead. However, existing heterogeneous TEE designs are inefficient for collaborative computing due to fine and different memory granularities between CPU and NPU. 1) The cacheline granularity of CPU TEE intensifies memory pressure due to its extra memory access, and 2) the cacheline granularity MAC of NPU escalates the pressure on the limited memory storage. 3) Data transfer across heterogeneous enclaves relies on the transit of non-secure regions, resulting in cumbersome re-encryption and scheduling. Husheng Han, Xinyao Zheng, Yuanbo Wen 0001, Yifan Hao 0001, Erhu Feng, Ling Liang 0003, Jianan Mu, Xiaqing Li, Tianyun Ma, Pengwei Jin, Xinkai Song, Zidong Du, Qi Guo 0001, Xing Hu 0001 |
ASPLOS (4) | 7 |
| 2024 | Accelerating Sequential Circuit Simulation with Spatial Locality Enhancement and Redundant Event ReductionabstractFast simulation is vital for efficient digital design, especially for safety-critical applications, where functional safety verification is paramount. However, existing gate-level event-driven simulators often encounter performance challenges attributed not only to inefficient memory access, but also to redundancy events in sequential elements during event-driven algorithms. In this paper, we introduce a memory-efficient, low-redundancy event-driven simulation framework to accelerate sequential circuit simulation. Firstly, we propose an event-based memory layout approach that fully considers memory access characteristics within and between logic levels to enhance the spatial locality of simulators. Secondly, we present an event trace approach tailored for flip-flops to reduce event redundancies that hinder simulator performance. Comparative experiments demonstrate that our proposed optimization strategies deliver an average performance improvement of 1.9× for logic simulation and 1.4× for fault simulation. Jiaping Tang, Zizhen Liu, Jianan Mu, Wenxing Li, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001 |
ATS | 3 |
| 2024 | Alchemist: A Unified Accelerator Architecture for Cross-Scheme Fully Homomorphic EncryptionabstractThe use of cross-scheme fully homomorphic encryption (FHE) in privacy-preserving applications present to be a new challenge to hardware accelerator design. Existing accelerator architectures with customized polynomial-level operator abstraction fail to efficiently handle hybrid FHE schemes due to the mismatch between computational demands and available hardware resources under various parameter settings. In this work, we propose a new accelerator architecture that consists of a novel finer-grained low-level operator, i.e., Meta-OP, that not only mathematically supports a diverse range of polynomial operations, but is also hardware-friendly for accelerator design without complex topological logic. We then design a new slot-based data management scheme to efficiently handle the distinct memory access patterns over the Meta-OP. With a slot-based data management approach, Alchemist can accelerate both arithmetic and logic FHE workloads with high hardware utilization rates. In the experiment, we show that Alchemist is up to 24,829X faster than CPU. For arithmetic FHE, compared with the SOTA ASIC accelerators, Alchemist achieves a 29.4X performance per area improvement on average. For logic FHE, compared with the SOTA ASIC accelerators, Alchemist achieves a 7.0X overall speed up on average. Jianan Mu, Husheng Han, Shangyi Shi, Jing Ye 0001, Zizhen Liu, Shengwen Liang, Meng Li 0004, Mingzhe Zhang 0005, Song Bian 0001, Xing Hu 0001, Huawei Li 0001, Xiaowei Li 0001 |
DAC | 1 |
| 2024 | DDP-Fsim: Efficient and Scalable Fault Simulation for Deterministic Patterns with Two-Dimensional ParallelismabstractFault simulation is a fundamental component in the design for testability (DFT) processes, especially in automatic test pattern generation (ATPG). Various approaches have been proposed to enhance the efficiency of fault simulation on multi-core systems. However, these approaches have not taken full consideration of the intrinsic characteristics of deterministic patterns. Deterministic patterns are generated by ATPG and are predominantly employed in practical applications rather than random patterns. In this paper, we introduce DDP-Fsim, a fast and scalable fault simulator on multi-core systems. DDP-Fsim capitalizes on the distinctive nature of deterministic patterns, wherein a small subset of patterns can effectively detect the majority of faults. Initially, DDP-Fsim parallels in fault dimension by dynamically scheduling fanout-free regions (FFR) to handle easy-to-detect faults. Subsequently, it parallels in pattern dimension by dynamically scheduling patterns to address the remaining hard-to-detect faults. Experiments demonstrate that on a 24-core system, DDP-Fsim is 10× faster than the commercial tools for full-scan circuits and deterministic patterns. Additionally, DDP-Fsim with 24 cores achieves an average speed-up of 16× compared to its single-core execution, while the commercial tools with 24 cores achieves only 3×-6× speed-up than their single-core execution. This indicates the significantly superior scalability for DDP-Fsim. Jianan Mu, Zizhen Liu, Jiaping Tang, Hui Wang 0152, Yonghao Wang, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
ICCAD | 3 |
| 2024 | Efficient Functional Safety Method for Gate-Level Fine-Grained Digital Circuits with ISO-26262abstractIn applications such as automotive chips that require high service responsiveness, ensuring the functional safety of electronic systems is crucial. The prevalent method involves conducting Failure Modes, Effects, and Diagnostic Analysis (FMEDA) and fault simulation at the design verification stage to assess safety levels. However, existing approaches primarily analyze at the register transfer level (RTL), which does not reflect the actual structure of chips where faults occur at the gate level, resulting in inaccuracies. This is due to the slower analysis speed at the gate level, making it challenging to balance precision with speed, thus defaulting to RTL for simulation. To address these challenges, we propose an innovative method for functional safety analysis and verification that integrates advanced gate-level fault simulation technology with FMEDA techniques. Our approach is based on an enhanced gate-level FMEDA framework, enabling deeper and more accurate safety performance analysis. Through experimental verification, our method has proven to be over 3 times faster than commercial tools in fault simulation, significantly enhancing the reliability and speed of the functional safety process. Ultimately, our research provides rapid and precise safety analysis and verification at the gate level for high-risk applications like automotive chips, offering robust technical support and practical guidelines for advancing functional safety technology in this sector. Hui Wang 0152, Jianan Mu, Zizhen Liu, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
ITC-Asia | 3 |
| 2023 | Configurable and High-Level Pipelined Lattice-Based Post Quantum Cryptography Hardware Accelerator DesignabstractNumber Theoretic Transform (NTT) and Secure Hash Algorithm 3 (SHA3), are the two main operators in the lattice-based Post-Quantum Cryptography (PQC) algorithms. Lattice-based PQC algorithms have different parameter settings, e.g., the length and modulus of NTT polynomials and the different hash functions. Motivated by the demands for more versatile NTT and SHA3 hardware accelerators, we implement the NTT and SHA3 designs that can accommodate to different parameters at run-time. Furthermore, to reduce the running cycles of the whole NTT operation and whole SHA3 operation including data transferring and calculation, we propose a pipelined architecture to optimize the gap between data transfer and calculation process in high-level. The designed configurable accelerators can be embedded in SoC to accelerate different lattice-based PQC algorithms efficiently. The experimental results show that our high-level pipelined and configurable NTT and SHA3 designs have good area-time efficiency. In specific, for the NTT design, our architecture is 4.1 times more area-time efficient compared with the state-of-the-art. For SHA3, our architecture is 1.4 times more area-time efficient over the existing configurable SHA3 designs. Jianan Mu, Huajie Tan, Min Cai, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
ATS | 1 |
| 2023 | Energy-efficient NTT Design with One-bank SRAM and 2-D PE ArrayabstractIn Number Theoretic Transform (NTT) operation, more than half of the active energy consumption stems from memory accesses. Here, we propose a generalized design method to improve the energy efficiency of NTT operation by considering the effect of processing element (PE) geometry and memory organization on the data flow between PEs and memory. To decrease the number of data bits that are required to be accessed from the memory, a two-dimensional (2-D) PE array architecture is used. A pair of ping-pong buffers are proposed to transposed swap the coefficients to enable a single bank of memory to be used with the 2-D PE array to reduce the average memory bit access energy without compromising the throughput. Our experimental results show that this design method can produce NTT accelerators with up to 69.8% saving in average energy consumption compared with the existing designs based on multi-bank SRAM and one-bank SRAM with one-dimensional PE array with the same number of PEs and total memory size. Jianan Mu, Huajie Tan, Haotian Lu 0002, Chip-Hong Chang, Shengwen Liang, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
DATE | 1 |
| 2023 | Online Reliability Evaluation Design: Select Reliable CRPs for Arbiter PUF and Its VariantsabstractPhysical Unclonable Function (PUF) is a hardware security primitive with broad application prospects. Variants of the arbiter PUF have been proposed to resist modeling attacks. However, their low reliability issue limits their applications. To solve the low reliability issue, this paper proposes an Online Reliability Evaluation (ORE) design for the arbiter PUF and its variants. Moreover, a corresponding machine learning method to select reliable Challenge Response Pairs (CRPs) for applications is proposed. Based on the ORE design, a small number of CRPs and their reliability levels are collected during the enrollment phase. Then they are trained to build reliability models for predicting the responses and reliability levels of other challenges. Since the ORE design does not change the security structures of the arbiter PUF and its variants, the resistance to modeling attacks of PUF designs equipped with it is maintained. Compared to the previous work that tests 100,000 times per CRP, our design is time-saving in the enrollment phase since each CRP is only tested three times for training reliability models. The proposed design is implemented under the 40nm process. Experimental results on real chips show that all the CRPs selected by our reliability models are indeed reliable for applications, verifying the effectiveness of our method. Chaofang Ma, Jianan Mu, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001 |
ETS | 2 |
| 2023 | Scalable and Conflict-Free NTT Hardware Accelerator Design: Methodology, Proof, and ImplementationabstractNumber theoretic transform (NTT) is useful for the acceleration of polynomial multiplication, which is the main performance bottleneck in the next-generation cryptographic schemes. Different NTT-based cryptographic algorithms have different security settings. The diverse application scenarios introduce different cost-performance tradeoffs and hardware constraints. Motivated by the emerging demand for more versatile NTT hardware accelerators, we propose a new design methodology that can generate area-efficient and high-performance NTT accelerators for any length and modulus of NTT polynomials and single processing element (PE) or PE array with a varying number of layers. The proposed NTT accelerator architecture pivots on a conflict-free memory access pattern for adaptation to different combinations of security and PE array configuration parameters. The proposed memory access pattern is formally proved to be conflict-free for any parametric configurations. The criterion for read-after-write conflict without pipeline stall is also established. Our proposed design methodology can produce NTT accelerators with single PE or multilayer PE array for different polynomial size and modulus, with hardware area and computational efficiency comparable to accelerators customized for a fixed set of parameters. Our proposed methodology produces parameterized accelerator with higher scalability than the existing parameterized accelerator design. On average, the accelerators generated by our proposed method are 71.4% more area-time efficient. Up to 30.7% area-time reduction over the most area-time efficient state-of-the-art scalable NTT accelerator can be achieved for the same security parameters. Jianan Mu, Wen Wang 0007, Yizhong Hu, Chip-Hong Chang, Junfeng Fan, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | A Voltage Template Attack on the Modular Polynomial Subtraction in KyberabstractKyber is one of the four final Key Encapsulation Mechanism (KEM) competitors of the National Institute of Standards and Technology PostQuantum Cryptography standardization competition. This paper reveals the vulnerability of Kyber under a voltage template side channel attack: the modular polynomial subtraction operation in Kyber.CCAKEM.Dec. In this paper, by splicing data under different selected ciphertexts, a small number of traces are required to recover the secret key. Experiments show that the recovering accuracy of secret key achieves 100% when using 330 traces, and it still achieves 98% when only using 44 traces. Jianan Mu, Zongyue Wang, Jing Ye 0001, Junfeng Fan, Huawei Li 0001, Xiaowei Li 0001, Yuan Cao 0003 |
ASP-DAC | 1 |