EDBT 2026 Demo / reviewers in the wild / expert
Xing Li 0023
dblp:26/379-23
· DBLP profile ↗
24ranked-venue papers
4as first author
24since 2021 · last 2025
0000-0003-1669-5954ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MTLSO: A Multi-Task Learning Approach for Logic Synthesis OptimizationabstractElectronic Design Automation (EDA) is essential for IC design and has recently benefited from AI-based techniques to improve efficiency. Logic synthesis, a key EDA stage, transforms high-level hardware descriptions into optimized netlists. Recent research has employed machine learning to predict Quality of Results (QoR) for pairs of And-Inverter Graphs (AIGs) and synthesis recipes. However, the severe scarcity of data due to a very limited number of available AIGs results in overfitting, significantly hindering performance. Additionally, the complexity and large number of nodes in AIGs make plain GNNs less effective for learning expressive graph-level representations. To tackle these challenges, we propose MTLSO - a Multi-Task Learning approach for Logic Synthesis Optimization. On one hand, it maximizes the use of limited data by training the model across different tasks. This includes introducing an auxiliary task of binary multi-label graph classification alongside the primary regression task, allowing the model to benefit from diverse supervision sources. On the other hand, we employ a hierarchical graph representation learning strategy to improve the model's capacity for learning expressive graph-level representations of large AIGs, surpassing traditional plain GNNs. Extensive experiments across multiple datasets and against state-of-the-art baselines demonstrate the superiority of our method, achieving an average performance gain of 8.22% for delay and 5.95% for area. Faezeh Faez, Raika Karimi, Yingxue Zhang 0001, Xing Li 0023, Lei Chen 0031, Mingxuan Yuan, Mahdi Biparva |
ASP-DAC | 4 |
| 2025 | EDGE: DBMS-Empowered Boolean Decomposition for GIG SynthesisabstractBoolean decomposition is a powerful technique in logic synthesis that breaks down Boolean functions into simpler components. Decomposition-based logic synthesis yields high-quality results and is particularly effective when combined with small-window optimization methods in Gate-Inverter Graphs (GIG). However, the efficiency limitations of current methods have constrained their applicability in handling large and complex logic. To address this challenge, we propose a novel framework, called EDGE, which leverages modern database techniques to accelerate Boolean decomposition, thereby achieving improved synthesis results while maintaining high efficiency. Experimental results demonstrate a runtime speedup of up to $21 \times$ and an overall reduction in node count of at least $15 \%$ compared to state-of-the-art synthesis methods. Ruofei Tang, Xuliang Zhu, Lei Chen 0002, Xing Li 0023, Mingxuan Yuan, Jianliang Xu |
DAC | 5 |
| 2025 | ELF: Efficient Logic Synthesis by Pruning Redundancy in RefactoringabstractIn electronic design automation, logic optimization operators play a crucial role in minimizing the gate count of logic circuits. However, their computation demands are high. Operators such as refactor conventionally form iterative cuts for each node, striving for a more compact representation - a task which often fails $98 \%$ on average. Prior research has sought to mitigate computational cost through parallelization. In contrast, our approach leverages a classifier to prune unsuccessful cuts preemptively, thus eliminating unnecessary resynthesis operations. Experiments on the refactor operator using the EPFL benchmark suite and 10 large industrial designs demonstrate that this technique can speedup logic optimization by $3.9 \times$ on average compared with the state-of-the-art ABC implementation. Dimitris Tsaras, Xing Li 0023, Zhiyao Xie, Mingxuan Yuan |
DAC | 2 |
| 2025 | Maximum Fanout-Free Window Enumeration: Towards Multi-Output Sub-Structure SynthesisabstractPeephole optimization is commonly used in And-Inverter Graphs (AIGs) optimization algorithms. The efficiency of these algorithms heavily relies on the enumeration process of sub-structures. One common sub-structure is the cut, known for its efficient enumeration method and single-output characteristic. However, an increasing number of optimization algorithms now target sub-structures that incorporate multiple outputs. In this paper, we explore Maximum Fanout-Free Windows (MFFWs), a novel sub-structure with a multi-output nature, as well as its practical applications and enumeration algorithms. To accommodate various algorithm execution processes, we propose two different enumeration styles: Dynamic and Static. The Dynamic approach provides flexibility in adapting to changes in the AIG structure, whereas the Static method ensures efficiency as long as the AIG structure remains unchanged during execution. We apply these methods to rewriting and technology mapping to improve their runtime performance. Experimental results on pure enumeration and practical scenarios show the scalability and efficiency of the proposed MFFW enumeration methods. Ruofei Tang, Xuliang Zhu, Lei Chen 0002, Xing Li 0023, Xin Huang 0001, Mingxuan Yuan, Jianliang Xu |
DATE | 4 |
| 2025 | Circuit Synthesis based on Hierarchical Conditional Diffusion
Xinyi Zhou 0010, Xing Li 0023, Yingzhao Lian, Lei Chen 0031, Mingxuan Yuan, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Circuit Transformer: A Transformer That Preserves Logical EquivalenceabstractImplementing Boolean functions with circuits consisting of logic gates is fundamental in digital computer design. However, the implemented circuit must be exactly equivalent, which hinders generative neural approaches on this task due to their occasionally wrong predictions. In this study, we introduce a generative neural model, the “Circuit Transformer”, which eliminates such wrong predictions and produces logic circuits strictly equivalent to given Boolean functions. The main idea is a carefully designed decoding mechanism that builds a circuit step-by-step by generating tokens, which has beneficial “cutoff properties” that block a candidate token once it invalidate equivalence. In such a way, the proposed model works similar to typical LLMs while logical equivalence is strictly preserved. A Markov decision process formulation is also proposed for optimizing certain objectives of circuits. Experimentally, we trained an 88-million-parameter Circuit Transformer to generate equivalent yet more compact forms of input circuits, outperforming existing neural approaches on both synthetic and real world benchmarks, without any violation of equivalence constraints.
Code: https://github.com/snowkylin/circuit-transformer Xihan Li 0001, Xing Li 0023, Lei Chen 0002, Mingxuan Yuan, Jun Wang 0012 |
ICLR | 2 |
| 2025 | KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM InferenceabstractKV cache quantization can improve Large Language Models (LLMs) inference throughput and latency in long contexts and large batch-size scenarios while preserving LLMs effectiveness. However, current methods have three unsolved issues: overlooking layer-wise sensitivity to KV cache quantization, high overhead of online fine-grained decision-making, and low flexibility to different LLMs and constraints. Therefore, we theoretically analyze the inherent correlation of layer-wise transformer attention patterns to KV cache quantization errors and study why key cache is generally more important than value cache for quantization error reduction. We further propose a simple yet effective framework KVTuner to adaptively search for the optimal hardware-friendly layer-wise KV quantization precision pairs for coarse-grained KV cache with multi-objective optimization and directly utilize the offline searched configurations during online inference. To reduce the computational cost of offline calibration, we utilize the intra-layer KV precision pair pruning and inter-layer clustering to reduce the search space. Experimental results show that we can achieve nearly lossless 3.25-bit mixed precision KV cache quantization for LLMs like Llama-3.1-8B-Instruct and 4.0-bit for sensitive models like Qwen2.5-7B-Instruct on mathematical reasoning tasks. The maximum inference throughput can be improved by 21.25% compared with KIVI-KV8 quantization over various context lengths. Our code and searched configurations are available at https://github.com/cmd2001/KVTuner. Xing Li 0023, Zeyu Xing 0002, Linping Qu, Hui-Ling Zhen, Yiwu Yao, Wulong Liu, Sinno Jialin Pan, Mingxuan Yuan |
ICML | 1 |
| 2025 | Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM InferenceabstractKey-Value (KV) cache eviction---which retains the KV pairs of the most important tokens while discarding less important ones---is a critical technique for optimizing both memory usage and inference latency in large language models (LLMs).
However, existing approaches often rely on simple heuristics---such as attention weights---to measure token importance, overlooking the spatial relationships between token value states in the vector space.
This often leads to suboptimal token selections and thus performance degradation.
To tackle this problem, we propose a novel method, namely **AnDPro** (**An**chor **D**irection **Pro**jection), which introduces a projection-based scoring function to more accurately measure token importance.
Specifically, AnDPro operates in the space of value vectors and leverages the projections of these vectors onto an *``Anchor Direction''*---the direction of the pre-eviction output---to measure token importance and guide more accurate token selection.
Experiments on $16$ datasets from the LongBench benchmark demonstrate that AnDPro can maintain $96.07\\%$ of the full cache accuracy using only $3.44\\%$ KV cache budget, reducing KV cache budget size by $46.0\\%$ without compromising quality compared to previous state-of-the-arts. Zijie Geng, Jie Wang 0005, Xing Li 0023, Mingxuan Yuan, Jianye Hao, Defu Lian, Enhong Chen, Feng Wu 0001 |
NeurIPS | 6 |
| 2025 | AttentionPredictor: Temporal Patterns Matter for KV Cache CompressionabstractWith the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation.
To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention scores. However, these methods often struggle to accurately determine critical tokens as they neglect the *temporal patterns* in attention scores, resulting in a noticeable degradation in LLM performance.
To address this challenge, we propose **AttentionPredictor**, which is the **first learning-based method to directly predict attention patterns for KV cache compression and critical token identification**.
Specifically, AttentionPredictor learns a lightweight, unified convolution model to dynamically capture spatiotemporal patterns and predict the next-token attention scores. An appealing feature of AttentionPredictor is that it accurately predicts the attention score and shares the unified prediction model, which consumes negligible memory, among all transformer layers. Moreover, we propose a cross-token critical cache prefetching framework that hides the token estimation time overhead to accelerate the decoding stage. By retaining most of the attention information, AttentionPredictor achieves **13$\times$** KV cache compression and **5.6$\times$** speedup in a cache offloading scenario with comparable LLM performance, significantly outperforming the state-of-the-arts. The code is available at https://github.com/MIRALab-USTC/LLM-AttentionPredictor. Qingyue Yang, Jie Wang 0005, Xing Li 0023, Chen Chen 0077, Lei Chen 0031, Xianzhi Yu, Wulong Liu, Jianye Hao, Mingxuan Yuan, Bin Li 0025 |
NeurIPS | 3 |
| 2025 | A Delay-Driven Iterative Technology Mapping FrameworkabstractTechnology mapping is the pivotal synthesis step that translates abstract logical models into technology-dependent implementations using the designated library, e.g., standard cells for ASICs. The efficient solutions heavily rely on the gate selection guided by estimated delay. However, estimating these delays is sophisticated due to the absence of actual interconnect load and transition time during the mapping. In this article, we revisit the difficulties of the delay-driven mapping problem and explore three key insights to address these. Inspired by the insights, we first design a structure-aware load-slew model that integrates input transitions and output loads for gate delay estimations. Benefiting from the model, we propose a delay-iterative framework that progressively reduces the overall circuit delay by further aligning library characteristics with logical network structures. Finally, experiments with 130 nm and 7 nm libraries show its superiority, which averagely reduces circuit delay by 10% with nonlinear delay model, and 6% in delay after P&R, as compared to ABC. Liwei Ni, Lei Chen 0031, Xing Li 0023, Shuai Ma 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | A Unified Parallel Framework for LUT Mapping and Logic OptimizationabstractLookup-table (LUT) mapping has been extensively utilized in logic synthesis, including being an indispensable step in FPGA design, serving as a building block in high-effort synthesis flows, and providing an algorithmic framework for logic optimization. Hence, a fast mapping algorithm is vital to satisfying the demand for synthesizing high-quality, large-scale modern VLSI designs. This article proposes two efficient GPU-parallel algorithms, namely LUT mapping and and-inverter graph (AIG) optimization using a precomputed database, which rely on a common parallel mapping framework that consists of novel fine-grained parallel mapping passes with high degree of parallelism. The mapping pass is enhanced by specifically tailored cut evaluation and memory management methods for GPUs that enable fast mapping of large circuits with limited GPU memory. Parallel timing analysis passes and parallel cut expansion passes are also proposed for constructing a fully GPU-accelerated LUT mapping flow. The core of parallel AIG optimization is a plugin of the mapping framework, which contains a self-adaptive parallel candidate structure evaluation procedure with high time efficiency and low hardware resource usage. Experiments show that on average, GPU LUT mapping and AIG optimization achieve$34.6\times $and$99.9\times $speedup with similar result quality, compared with the high-performance LUT mapper and AIG optimization algorithm with a database implemented in ABC, respectively, on large benchmarks. When combining the two algorithms with other GPU logic optimization algorithms, a GPU-based sequence targeting LUT network synthesis achieves$46.7\times $speedup with 4.7% smaller area and 0.2% smaller delay over ABC. Tianji Liu, Lei Chen 0031, Xing Li 0023, Mingxuan Yuan, Evangeline F. Y. Young |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | FineMap: A Fine-grained GPU-parallel LUT Mapping EngineabstractLookup-table (LUT) mapping is an indispensable step in FPGA design flows, and also serves as a building block in many technology-independent optimization algorithms. Therefore, it is crucial to accelerate LUT mapping in order to satisfy the demand for synthesizing high-quality, large-scale VLSI designs. Previous work on GPU LUT mapping suffers from low speedup due to limited degree of parallelism. In this paper, we propose an ultra-fast GPU-parallel LUT mapping engine named FineMap, which is composed of a novel fine-grained mapping phase with a high degree of parallelism, a parallel cut expansion phase and a parallel timing analysis pass. The mapping phase is enhanced by specifically tailored cut evaluation and memory management algorithms for GPUs that enable fast mapping of large circuits with limited GPU memory. Experiments show that compared with the high-performance mapper implemented in ABC, FineMap achieves 128.7× speedup with better quality in terms of area on large benchmarks. Tianji Liu, Lei Chen 0031, Xing Li 0023, Mingxuan Yuan, Evangeline F. Y. Young |
ASPDAC | 3 |
| 2024 | RTLRewriter: Methodologies for Large Models aided RTL Code OptimizationabstractRegister Transfer Level (RTL) code optimization is crucial for enhancing the efficiency and performance of digital circuits during early synthesis stages. Currently, optimization relies heavily on manual efforts by skilled engineers, often requiring multiple iterations based on synthesis feedback. In contrast, existing compiler-based methods fall short in addressing complex designs. This paper introduces RTLRewriter, an innovative framework that leverages large models to optimize RTL code. A circuit partition pipeline is utilized for fast synthesis and efficient rewriting. A multi-modal program analysis is proposed to incorporate vital visual diagram information as optimization cues. A specialized search engine is designed to identify useful optimization guides, algorithms, and code snippets that enhance the model's ability to generate optimized RTL. Additionally, we introduce a Cost-aware Monte Carlo Tree Search (C-MCTS) algorithm for efficient rewriting, managing diverse retrieved contents and steering the rewriting results. Furthermore, a fast verification pipeline is proposed to reduce verification cost. To cater to the needs of both industry and academia, we propose two benchmarking suites: the long Rewriter benchmark, targeting complex scenarios with extensive circuit partitioning, optimization trade-offs, and verification challenges, and the short Rewriter benchmark, designed for a wider range of scenarios and patterns. Our comparative analysis with established compilers such as Yosys and E-graph demonstrates significant improvements, highlighting the benefits of integrating large models into the early stages of circuit design. We provide our benchmarks at https://github.com/yaoxufeng/RTLRewriter-Bench. Xufeng Yao, Xing Li 0023, Yingzhao Lian, Ran Chen 0001, Lei Chen 0031, Mingxuan Yuan, Hong Xu 0001, Bei Yu 0001 |
ICCAD | 3 |
| 2024 | A Circuit Domain Generalization Framework for Efficient Logic Synthesis in Chip DesignabstractLogic Synthesis (LS) plays a vital role in chip design. A key task in LS is to simplify circuits---modeled by directed acyclic graphs (DAGs)---with functionality-equivalent transformations. To tackle this task, many LS heuristics apply transformations to subgraphs---rooted at each node on an input DAG---sequentially. However, we found that a large number of transformations are ineffective, which makes applying these heuristics highly time-consuming. In particular, we notice that the runtime of the Resub and Mfs2 heuristics often dominates the overall runtime of LS optimization processes. To address this challenge, we propose a novel data-driven LS heuristic paradigm, namely PruneX, to reduce ineffective transformations. The major challenge of developing PruneX is to learn models that well generalize to unseen circuits, i.e., the out-of-distribution (OOD) generalization problem. Thus, the major technical contribution of PruneX is the novel circuit domain generalization framework, which learns domain-invariant representations based on the transformation-invariant domain-knowledge. To the best of our knowledge, PruneX is the first approach to tackle the OOD problem in LS heuristics. We integrate PruneX with the aforementioned Resub and Mfs2 heuristics. Experiments demonstrate that PruneX significantly improves their efficiency while keeping comparable optimization performance on industrial and very large-scale circuits, achieving up to $3.1\times$ faster runtime. Lei Chen 0031, Jie Wang 0005, Yinqi Bai, Xing Li 0023, Xijun Li, Mingxuan Yuan, Jianye Hao, Yongdong Zhang 0001, Feng Wu 0001 |
ICML | 5 |
| 2024 | Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation FrameworkabstractLogic Synthesis (LS) aims to generate an optimized logic circuit satisfying a given functionality, which generally consists of circuit translation and optimization. It is a challenging and fundamental combinatorial optimization problem in integrated circuit design. Traditional LS approaches rely on manually designed heuristics to tackle the LS task, while machine learning recently offers a promising approach towards next-generation logic synthesis by neural circuit generation and optimization. In this paper, we first revisit the application of differentiable neural architecture search (DNAS) methods to circuit generation and found from extensive experiments that existing DNAS methods struggle to exactly generate circuits, scale poorly to large circuits, and exhibit high sensitivity to hyper-parameters. Then we provide three major insights for these challenges from extensive empirical analysis: 1) DNAS tends to overfit to too many skip-connections, consequently wasting a significant portion of the network's expressive capabilities; 2) DNAS suffers from the structure bias between the network architecture and the circuit inherent structure, leading to inefficient search; 3) the learning difficulty of different input-output examples varies significantly, leading to severely imbalanced learning. To address these challenges in a systematic way, we propose a novel regularized triangle-shaped circuit network generation framework, which leverages our key insights for completely accurate and scalable circuit generation. Furthermore, we propose an evolutionary algorithm assisted by reinforcement learning agent restarting technique for efficient and effective neural circuit optimization. Extensive experiments on four different circuit benchmarks demonstrate that our method can precisely generate circuits with up to 1200 nodes. Moreover, our synthesized circuits significantly outperform the state-of-the-art results from several competitive winners in IWLS 2022 and 2023 competitions. Jie Wang 0005, Qingyue Yang, Yinqi Bai, Xing Li 0023, Lei Chen 0031, Jianye Hao, Mingxuan Yuan, Bin Li 0025, Yongdong Zhang 0001, Feng Wu 0001 |
NeurIPS | 5 |
| 2023 | Lightweight Structural Choices Operator for Technology MappingabstractTechnology mapping quality heavily depends on the subject graph structure. To overcome structural biases, operators construct choice nodes to enable mappings with improved node and level counts. Nevertheless, state-of-the-art structural choice operators scale poorly with graph size.We present the lightweight structural choices (LCH) operator that incorporates equivalencies by processing only subparts of the graph. We propose multiple heuristics that rely on specific node extraction orders and subpart sizes to extract non-overlapping components. Compared to state-of-the-art methods on EPFL circuits, LCH is 2.35x faster enduring a small sacrifice in node count (3%) and level reduction (2%). Antoine Grosnit, Matthieu Zimmer, Rasul Tutunov, Xing Li 0023, Lei Chen 0031, Mingxuan Yuan, Haitham Bou-Ammar |
DAC | 4 |
| 2023 | A Database Dependent Framework for K-Input Maximum Fanout-Free Window RewritingabstractRewriting is a widely used logic optimization approach incorporated in most commercial logic synthesis tools. In this paper, we present a new rewriting method based on And-Inverted Graph (AIG). Rather than focusing on cut rewriting, it considers a novel sub-structure called Maximum Fanout-Free Window (MFFW) and rewrites with a more compact implementation. Both exact synthesis and heuristic methods can be adopted to optimize MFFWs. A database dependent framework is proposed to store the optimal sub-structures to accelerate the processing. We further propose the semi-canonicalization to reduce the scale of the database, which could reduce more than 98% of the 4-input MFFW database. Extensive experiments on benchmark datasets demonstrate both the effectiveness and efficiency of our proposed framework. Xuliang Zhu, Ruofei Tang, Lei Chen 0002, Xing Li 0023, Xin Huang 0001, Mingxuan Yuan, Weihua Sheng, Jianliang Xu |
DAC | 4 |
| 2023 | CPP: A Multi-Level Circuit Partitioning Predictor for Hardware Verification SystemsabstractCircuit partitioning is a critical step in hardware-assisted functional verification that involves splitting a circuit into multiple partitions and assigning them to specific hardware. However, partitioning a large circuit can require considerable computation resources and time, especially when complex hardware constraints are involved. Moreover, the path delay after partitioning can have a significant impact on verification efficiency, making early path delay prediction crucial for refining the circuit effectively. In this work, we propose a novel circuit partitioning predictor, named CPP, to rapidly and accurately predict the path delay after partitioning. To achieve this, we use circuit coarsening to develop a multi-level path representation and employ a convolutional neural network (CNN) that can capture both local and global path structures for delay prediction. Through extensive experiments on large industrial circuits, we demonstrate the superiority of our prediction framework. Xinshi Zang, Lei Chen 0031, Xing Li 0023, Wilson W. K. Thong, Weihua Sheng, Evangeline F. Y. Young, Martin D. F. Wong |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | EffiSyn: Efficient Logic Synthesis with Dynamic Scoring and PruningabstractLogic synthesis tools synthesize circuit structures to optimize specific targets given reasonable constraints and runtime using a set of well-defined operators. The efficiency of these operators is critical to achieving better runtime and optimization convergence. However, most synthesis operators are designed heuristically with fixed and sub-optimal traversal orders of nodes, cuts, and candidate subgraphs that are independent of circuit structures and functionalities. This leads to redundant computation and loss of optimization opportunities. Due to the spatial structure similarity, sub-circuits already synthesized in the same circuit contain meaningful information to guide more efficient synthesis for unvisited sub-circuits. Historical evaluation and gain features can learn from conflicts and be utilized to predict synthesis gain and prune invalid sub-circuits. Instead of using high-weight feature extraction and models, we utilize efficient prediction models with lightweight structural and functional features to reduce overhead. We thus propose a generalizable and dynamic scoring and pruning framework EffiSyn to accelerate logic synthesis operators while maintaining synthesis effectiveness. For example, we further improve the highly optimized operator drw by scoring and pruning invalid cuts and precomputed subgraphs. Extensive experiments on 20 public and industrial circuits validate that EffiSyn can accelerate drw by about 35% with negligible effectiveness loss or even with effectiveness improvement. Experiments over diverse circuits and synthesis sequences also validate the generalization of the proposed framework. Xing Li 0023, Lei Chen 0002, Jiantang Zhang, Shuang Wen 0007, Weihua Sheng, Yu Huang 0005, Mingxuan Yuan |
ICCAD | 1 |
| 2023 | EasyMap: Improving Technology Mapping via Exploration-Enhanced Heuristics and Adaptive SequencingabstractTechnology mapping is a crucial step in the logic synthesis in chip design e.g. Field Programmable Gate Arrays (FPGAs) design, where a logic network is transformed into a K-bounded lookup tables (K-LUTs) network. Traditional mapping algorithms converges quickly to a suboptimal result, which limits the exploration capacity for further improvement. In this paper, we propose a new mapping method called Exploration-enhanced heuristics and Adaptive sequencing for Technology Mapping (EasyMap). EasyMap includes a pool of new heuristics and considers the mapping exploration as a conditional sequence optimization problem. During the mapping exploration procedure, heuristic algorithms with specific parameters are selected and applied sequentially. Our EasyMap outperforms the widely used IfMap in ABC by a significant margin. In particular, when optimizing area with a level constraint, EasyMap outperforms IfMap by reducing 9.1% more area on arithmetic circuits of the EPFL benchmark. Moreover, when optimizing area without level constraints at the same time, EasyMap can reduce 19% more area than IfMap on arithmetic circuits. Peiyu Wang, Anqi Lu, Xing Li 0023, Junjie Ye 0002, Lei Chen 0031, Mingxuan Yuan, Jianye Hao, Junchi Yan |
ICCAD | 3 |
| 2023 | AiMap: Learning to Improve Technology Mapping for ASICs via Delay PredictionabstractTechnology mapping is an essential process in the EDA flow which aims to find an optimal implementation of a logic network from a technology library. In ASIC designs, the estimated cell delay w.r.t. the cut has a significant impact on both area and delay of the mapped network. In this work, we first propose formulating cell delay estimation as a regression learning task by incorporating multiple perspective features, such as the structure of logic networks and non-linear cell delays, to guide the mapper search. We design a learning model that incorporates a customized attention mechanism to be aware of the pin delay and jointly learns the hierarchy between the logic network and library, with the help of proposed parameterizable strategies to generate learning labels. Experimental results show that our proposed method noticeably improves area by 12% and delay by 1%, compared with ABC. Liwei Ni, Min Zhou 0006, Lei Chen 0031, Xing Li 0023, Shuai Ma 0001 |
ICCD | 6 |
| 2023 | SGDP: A Stream-Graph Neural Network Based Data PrefetcherabstractData prefetching is important for storage system optimization and access performance improvement. Traditional prefetchers work well for mining access patterns of sequential logical block address (LBA) but cannot handle complex non-sequential patterns that commonly exist in real-world applications. The state-of-the-art (SOTA) learning-based prefetchers cover more LBA accesses. However, they do not adequately consider the spatial interdependencies between LBA deltas, which leads to limited performance and robustness. This paper proposes a novel Stream-Graph neural network-based Data Prefetcher (SGDP). Specifically, SGDP models LBA delta streams using a weighted directed graph structure to represent interactive relations among LBA deltas and further extracts hybrid features by graph neural networks for data prefetching. We conduct extensive experiments on eight real-world datasets. Empirical results verify that SGDP outperforms the SOTA methods in terms of the hit ratio by 6.21%, the effective prefetching ratio by 7.00%, and speeds up inference time by 3.13× on average. Besides, we generalize SGDP to different variants by different stream constructions, further expanding its application scenarios and demonstrating its robustness. SGDP offers a novel data prefetching solution and has been verified in commercial hybrid storage systems in the experimental phase. Our codes and appendix are available at https://github.com/yyysjz1997/SGDP/. Yiyuan Yang, Rongshang Li, Qiquan Shi, Xijun Li, Xing Li 0023, Mingxuan Yuan |
IJCNN | 6 |
| 2022 | HIMap: a heuristic and iterative logic synthesis approachabstractRecently, many models show their superiority in sequence and parameter tuning. However, they usually generate non-deterministic flows and require lots of training data. We thus propose a heuristic and iterative flow, namely HIMap, for deterministic logic synthesis. In which, domain knowledge of the functionality and parameters of synthesis operators and their correlations to netlist PPA is fully utilized to design synthesis templates for various objetives. We also introduce deterministic and effective heuristics to tune the templates with relatively fixed operator combinations and iteratively improve netlist PPA. Two nested iterations with local searching and early stopping can thus generate dynamic sequence for various circuits and reduce runtime. HIMap improves 13 best results of the EPFL combinational benchmarks for delay (5 for area). Especially, for several arithmetic benchmarks, HIMap significantly reduces LUT-6 levels by 11.6 ~ 21.2% and delay after P&R by 5.0 ~ 12.9%. Xing Li 0023, Lei Chen 0031, Mingxuan Yuan, Hongli Yan, Yupeng Wan |
DAC | 1 |
| 2021 | Block Access Pattern Discovery via Compressed Full Tensor TransformerabstractThe discovery and prediction of block access patterns in hybrid storage systems is of crucial importance for effective tier management. Existing methods are usually based on heuristics and unable to handle complex patterns. This work newly introduces transformer to block access pattern prediction. We remark that block accesses in the tier management systems are aggregated temporally and spatially as multivariate time series of block access frequency, so the runtime requirements are relaxed, making complex models applicable for the deployment. Moreover, enormous and rarely accessed blocks in storage systems and the structure of traditional transformer models would result in millions of redundant parameters and make them impractical to be deployed. We incorporate Tensor-Train Decomposition (TTD) with transformer and propose the Compressed Full Tenor Transformer (CFTT), in which all linear layers in the vanilla transformer are replaced with tensor-train layers. Weights of input and output layers are shared to further reduce parameters and reuse knowledge implicitly. CFTT can significantly reduce the model size and computation cost, which is critical to save storage space and inference time. Extensive experiments are conducted on synthetic and real-world datasets. The results demonstrate that transformers achieve state-of-the-art performance stably in terms of top-k hit rates. Moreover, the proposed CFTT compresses transformers 16× to 461× and speeds up inference 5× without sacrificing performance on the whole, which facilitates its applications in tier management in hybrid storage systems. Xing Li 0023, Qiquan Shi, Lei Chen 0031, Yiyuan Yang, Mingxuan Yuan |
CIKM | 1 |