Ye Cai 0001

dblp:27/10805 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0003-6470-0364ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AiLO: A Predictive Framework for Logic Optimization Using Multi-Scale Cross-Attention Transformer
abstract
Logic Optimization (LO) is a critical stage in the chip design process, focused on improving the Quality of Results (QoR) by optimizing circuit designs to minimize area and delay. During logic optimization, evaluating the QoR after each iteration requires completing logic optimization and technology mapping. The evaluation process is highly time-consuming, restricting the number of optimization iterations possible within a given time. To address this, the AI-aided logic optimization framework (AiLO) is developed to explore more optimization operator sequences (recipes). AiLO framework consists of two core components: AI-based metric evaluation and optimization exploration. To achieve accurate evaluation, different prediction models can be integrated. A multi-scale cross-attention Transformer (CrossLO) is introduced to simulate the optimization structure of recipes across circuit at various scales to enhance the prediction accuracy. Moreover, the AI evaluation module can effectively maintain the recipe ranking, even when prediction accuracy is biased. The logic optimization exploration algorithm integrated with CrossLO (AI evaluation) shows an average improvement of 14.75% over the initial version. NSGA-II (optimization module) integrated with CrossLO achieves a significant lead over other algorithms in the same time. In addition, the AiLO framework continues to grow with the performance of the two components, demonstrating strong adaptability and flexibility.
Ye Cai 0001, Rui Wang 0189, Liwei Ni, Xiaoze Lin, Biwei Xie
ACM Trans. Design Autom. Electr. Syst.1
2025 Verilua: An Open Source Versatile Framework for Efficient Hardware Verification and Analysis Using LuaJIT
abstract
The growing complexity of hardware verification highlights limitations in existing frameworks, particularly regarding flexibility and reusability. Current methodologies often require multiple specialized environments for functional verification, waveform analysis, and simulation, leading to toolchain fragmentation and inefficient code reuse. This paper presents Verilua, a unified framework leveraging LuaJIT and the Verilog Procedural Interface (VPI), which integrates three core functional-ities: Lua-based functional verification, a scripting engine for RTL simulation, and waveform analysis. By enabling complete code reuse through a unified Lua codebase, the framework achieves a 12x speedup in RTL simulation compared to cocotb and a 70x improvement in waveform analysis over state-of-the-art solutions. Through consolidating verification tasks into a single platform, Verilua enhances efficiency while reducing tool fragmentation and learning overhead, addressing critical challenges in modern hardware design.
Ye Cai 0001, Chuyu Zheng, Dan Tang 0002
DATE1
2025 BPTCN: a low-latency branch prediction model based on temporal convolutional networks
Ye Cai 0001, Jinghui Zhong, Hao Liao
J. Supercomput.1
2024 iEDA: An Open-source infrastructure of EDA
abstract
By leveraging the power of open-source software, the EDA tool offers a cost-effective and flexible solution for designers, researchers, and hobbyists alike. Open-source EDA promotes collaboration, innovation, and knowledge sharing within the EDA community. It emphasizes the role of the toolchain in accelerating the development of electronic systems, reducing design costs, and improving design quality. This paper presents an open-source EDA project, iEDA, aiming to build a basic infrastructure for EDA technology evolution and closing the industrial-academic gap in the EDA area. As the foundation for developing EDA tools and researching EDA algorithms and technologies, iEDA is mainly composed of file system, database, manager, operator and interface. To demonstrate the effectiveness of iEDA, we implement and tape out four chips of different scales (from 700k to 500M gates) on different process nodes (110nm and 28nm) with iEDA. iEDA is publicly available on the project home page https://github.com/OSCC-Project/iEDA.
Zengrong Huang, Simin Tao, Zhipeng Huang 0009, Chunan Zhuang, Yihang Qiu, Guojie Luo, Huawei Li 0001, Haihua Shen, Mingyu Chen 0001, Dongbo Bu, Wenxing Zhu, Ye Cai 0001, Xiaoming Xiong, Yi Heng, Peng Zhang 0007, Bei Yu 0001, Biwei Xie, Yungang Bao
ASPDAC15
2024 Net Resource Allocation: A Desirable Initial Routing Step
abstract
In modern IC design, routing significantly impacts chip performance, power, area, and design iteration count. Critical challenges in routing include generating a rectilinear Steiner minimum tree (RSMT) for each net and handling routing resources among nets. Due to limited resources and net order, congestion is inevitable in VLSI circuit routing. Most competitive routers address congestion after routing without prior net guidance, leading to difficulty in managing resources among nets. We suggest introducing a net resource allocation step to tackle routing and congestion as a potentially desirable initial routing stage. Firstly, we introduce the net region probability density (NRPD) concept to achieve suitable net resource allocation. Using a prior NRPD, we model the resource allocation problem as linear programming (LP). We solve the LP problem and obtain a posterior NRPD for each net on each grid. Based on the posterior NRPD and congestion map, we introduce a cost scheme to guide net routing. This cost scheme supports a weighted RSMT construction technique for better topological solutions. We propose an iterative method for global routing and track assignment, improving detailed routing quality and optimizing design rule violations. Experimental results show the effectiveness of net resource allocation and demonstrate the superior performance of our router over OpenROAD's router across multiple metrics.
Zhisheng Zeng, Jikang Liu, Zhipeng Huang 0009, Ye Cai 0001, Biwei Xie, Yungang Bao
DAC4
2024 Parallel AIG Refactoring via Conflict Breaking
abstract
Algorithm parallelization to leverage multi-core platforms for improving the efficiency of Electronic Design Automation (EDA) tools plays a significant role in enhancing the scalability of Integrated Circuit (IC) designs. Logic optimization is a key process in the EDA design flow to reduce the area and depth of the circuit graph by finding logically equivalent graphs for substitution, which is typically time-consuming. To address these challenges, in this paper, we first analyze two types of conflicts that need to be handled in the parallelization framework of refactoring And-Inverter Graph (AIG). We then present a fine-grained parallel AIG refactoring method, which strikes a balance between the degree of parallelism and the conflicts encountered during the refactoring operations. Experiment results show that our parallel refactor is 28x averagely faster than the sequential algorithm on large benchmark tests with 64 physical CPU cores, and has comparable optimization quality.
Ye Cai 0001, Liwei Ni, Biwei Xie
ISCAS1
2023 Structured DFT Development Approach for Chisel-Based High Performance RISC-V Processors
abstract
Research has shown that agile language Chisel and related agile design methodology is promising to sustain the scaling computing performance in a more efficient way. Design For Test, as an economical and effective method for chip production testing, must be integrated into the agile development system to meet the requirements of mass production. Due to the lack of support for the traditional EDA toolchain, chip development through Chisel has not been widely accepted in the industry. This paper is based on the research of XiangShan, an open-source project for RISC-V high-performance processor. XiangShan establishes a structured DFT(degisn for test) development approach and Chisel-based DFT agile design flow. XiangShan develops a flexible chisel-based XS-shared bus mbist interface to improve design PPA and proposes the MarchSLD algorithm to enhance defect detection in the FinFet process node. XiangShan’s DFT development Approach is moving towards hierarchical, flexible, and reliable in two generations of processors. Experimental results show that the optimized Chisel-based DFT design flow can support agile development requirements and achieve industry-competitive performance.
Ye Cai 0001, Zhiheng He, Sen Liang
ITC-Asia2
2022 Towards Developing High Performance RISC-V Processors Using Agile Methodology
abstract
While research has shown that the agile chip design methodology is promising to sustain the scaling of computing performance in a more efficient way, it is still of limited usage in actual applications due to two major obstacles: 1) Lack of tool-chain and developing framework supporting agile chip design, especially for large-scale modern processors. 2) The conventional verification methods are less agile and become a major bottleneck of the entire process. To tackle both issues, we propose MINJIE, an open-source platform supporting agile processor development flow. MINJIE integrates a broad set of tools for logic design, functional verification, performance modelling, pre-silicon validation and debugging for better development efficiency of state-of-the-art processor designs. We demonstrate the usage and effectiveness of MINJIE by building two generations of an open-source superscalar out-of-order RISC-V processor code-named XIANGSHAN using agile methodologies. We quantify the performance of XIANGSHAN using SPEC CPU2006 benchmarks and demonstrate that XIANGSHAN achieves industry-competitive performance.
Yinan Xu 0001, Dan Tang 0002, Guokai Chen, Lingrui Gou, Qianruo Li, Zuojun Li, Jiazhan Tan, Huaqiang Wang, Huizhe Wang, Kaifan Wang, Chuanqi Zhang, Fawang Zhang, Linjuan Zhang, Zifei Zhang 0001, Yaoyang Zhou, Yike Zhou, Jiangrui Zou, Ye Cai 0001, Dandan Huan, Zusong Li, Jiye Zhao, Qiyuan Quan, Xingwu Liu, Sa Wang, Kan Shi, Ninghui Sun, Yungang Bao
MICRO26
2019 FPGA-Based Parallel Multi-Core GZIP Compressor in HDFS
abstract
With the development of Big Data, data storage has been exposed to more challenges. Data compression which can save both storage and network bandwidth, is a very important technology to deal with the challenges. In this paper, we present an end-to-end, complete, high-throughput parallel multi-core GZIP compressor in FPGA for HDFS. The GZIP compressor is designed by the scalable architecture, which supports to increase throughput by expanding multiple compression cores based on systolic array architecture. We implemented and evaluated the hardware compressor in Alpha Data Adm-Pcie-KU3 FPGA board, utilizing RIFFA for data transfers over PCI Express. According to the evaluation results, up to 16-cores compressor can be implemented and the peak compression throughput exceeds 1.1 GB/s. It is 70X speedup compared with the software compression solution. When we load the hardware compressor into HDFS, the performance of HDFS is twice as much as that without loading the compressor.
Haoxin Luo, Ye Cai 0001, Qiuming Luo, Rui Mao 0001
PDCAT2
2018 Algorithms designed for compressed-gene-data transformation among gene banks with different references
abstract
BACKGROUND: With the reduction of gene sequencing cost and demand for emerging technologies such as precision medical treatment and deep learning in genome, it is an era of gene data outbreaks today. How to store, transmit and analyze these data has become a hotspot in the current research. Now the compression algorithm based on reference is widely used due to its high compression ratio. There exists a big problem that the data from different gene banks can't merge directly and share information efficiently, because these data are usually compressed with different references. The traditional workflow is decompression-and-recompression, which is too simple and time-consuming. We should improve it and speed it up. RESULTS: In this paper, we focus on this problem and propose a set of transformation algorithms to cope with it. We will 1) analyze some different compression algorithms to find the similarities and the differences among all of them, 2) come up with a naïve method named TDM for data transformation between difference gene banks and finally 3) optimize former method TDM and propose the method named TPI and the method named TGI. A number of experiment result proved that the three algorithms we proposed are an order of magnitude faster than traditional decompression-and-recompression workflow. CONCLUSIONS: Firstly, the three algorithms we proposed all have good performance in terms of time. Secondly, they have their own different advantages faced with different dataset or situations. TDM and TPI are more suitable for small-scale gene data transformation, while TGI is more suitable for large-scale gene data transformation.
Qiuming Luo, Chao Guo 0004, Ye Cai 0001, Gang Liu 0028
BMC Bioinform.4
2014 Optimization of Uncore Data Flow on NUMA Platform
Qiuming Luo, Chang Kong, Ye Cai 0001
NPC5
2013 Analyzing the Characteristics of Memory Subsystem on Two Different 8-Way NUMA Architectures
Qiuming Luo, Chang Kong, Ye Cai 0001, Xiaohui Lin 0001
NPC5
2012 MAP-numa: Access Patterns Used to Characterize the NUMA Memory Access Optimization Techniques and Algorithms
Qiuming Luo, Chengjian Liu, Chang Kong, Ye Cai 0001
NPC4
2012 Quantitatively Measuring the Memory Locality Leakage on NUMA Systems Based on Instruction-Based-Sampling
abstract
Sustaining the memory locality is critical for obtaining high performance in NUMA system. But how to identify a locality leakage problem and how to measure the leakage is still open issue. This paper provides an algorithm to quantitatively measure the locality leakage based on the memory trace produced by IBS (Instruction-Based-Sampling). A """"perfect matrix"""" PM is generated from virtual memory address trace, which represents the highest locality pattern. A """"communication matrix"""" CM is obtained from physical memory address trace to describe the actual memory access pattern. The penalty factors are calculated from PM or CM with considering of the hardware NUMA factor. The leakage is measured by the difference between the penalty factors of PM and the penalty factors of CM, which can be used to estimate the performance decrease and guide the optimization. Some experiment results are show to testify the effectiveness and accuracy of our quantitative measurement.
Qiuming Luo, Chengjian Liu, Chang Kong, Ye Cai 0001
PDCAT4
2011 Performance Evaluation of OpenMP Constructs and Kernel Benchmarks on a Loongson-3A Quad-Core SMP System
abstract
As a competitor and alternative to mainstream general-purpose CPU (Intel/AMD/etc.), Loongson is a family of general-purpose MIPS-compatible CPUs developed at the ICT of CAS in China. The quad-core Loongson 3A is evaluated in this paper. The performance of the basic OpenMP constructs on Loongson-3A quad-core SMP is obtained by applying the EPCC Micro benchmarks. And then the performance of NAS kernel codes is obtained by applying NAS Parallel Benchmarks (NPB). These benchmarking are carried out for three different OpenMP compilers (and the runtime system), which includes GCC, OMPipth (OMPi with pthread library) and OMPi-psth (OMPi with psthread library). The results show that OMPI-pth's performance is the best and OMPi-psth's performance is the worst. Those test results might help to program the OpenMP codes as well as to select the appropriate compiler and its runtime system. And an Intel core i5 quad-core platform is used for comparison purpose, by running NPB, which implies that Loongson 3A's performance is nearly one tenth of i5's. The NPB results can help to defining a Loongson system's scale when replacing an Intel i5 system for a given problem size.
Qiuming Luo, Chang Kong, Ye Cai 0001, Gang Liu 0028
PDCAT3