EDBT 2026 Demo / reviewers in the wild / expert
Jiangfang Yi
dblp:80/10510
· DBLP profile ↗
9ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-4401-3551ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the Gap between Hardware Fuzzing and Industrial VerificationabstractAs hardware design complexity increases, hardware fuzzing emerges as a promising tool for automating the verification process. However, a significant gap still exists before it can be applied in industry. This paper aims to summarize the current progress of hardware fuzzing from an industry-use perspective and propose solutions to bridge the gap between hardware fuzzing and industrial verification. First, we review recent hardware fuzzing methods and analyze their compatibilities with industrial verification. We establish criteria to assess whether a hardware fuzzing approach is compatible. Second, we examine whether current verification tools can efficiently support hardware fuzzing. We identify the bottlenecks in hardware fuzzing performance caused by insufficient support from the industrial environment. To overcome the bottlenecks, we propose a prototype, HwFuzzEnv, providing the necessary support for hardware fuzzing. With this prototype, the previous hardware fuzzing method can achieve a several hundred times speedup in industrial settings. Our work could serve as a reference for EDA companies, encouraging them to enhance their tools to support hardware fuzzing efficiently in industrial verification. Tianhao Wei, Jiaxi Zhang 0001, Jiangfang Yi, Guojie Luo |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | Hyperion: A Highly Effective Page and PC Based Delta PrefetcherabstractHardware prefetching plays an important role in modern processors for hiding memory access latency. Delta prefetchers show great potential at the L1D cache level, as they can impose small storage overhead by recording deltas. Furthermore, local delta prefetchers, such as Berti, have been shown to achieve high L1D accuracy. However, there is still room for improving the L1D coverage of existing delta prefetchers. Our goal is to develop a delta prefetcher capable of achieving both high L1D coverage and accuracy. We explore delta prefetchers trained on various types of contextual information, ranging from coarse-grained to fine-grained, and analyze their L1D coverage and accuracy. Our findings indicate that training deltas based on the access histories of both PCs and memory pages for individual PCs and memory pages can lead to increased L1D coverage alongside high accuracy. Therefore, we introduce Hyperion, a highly efficient Page and PC-based delta prefetcher. In terms of the vital component of recording access histories, we implement three different structures and engage in a detailed discussion about them. Furthermore, Hyperion utilizes micro-architecture information (e.g., L1D hits or misses, PQ occupancy) and real-time L1D accuracy to dynamically adjust its issuing mechanism, further enhancing performance and L1D accuracy. Our results show that Hyperion achieves an L1D accuracy of 92.4% and an L1D coverage of 51.9%, along with an L2C coverage of 63.0% and an LLC coverage of 67.5% across a diverse range of applications, including SPEC CPU2006, SPEC CPU2017, GAP, and PARSEC, with a baseline of no prefetching. Regarding performance, Hyperion achieves a 50.1% performance gain, outperforming the state-of-the-art delta prefetcher Berti by 5.0% over baseline across all memory-intensive traces from the four benchmark suites. Wei Chen 0168, Xu Cheng 0001, Jiangfang Yi |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | FlexPointer: Fast Address Translation Based on Range TLB and Tagged PointersabstractPage-based virtual memory relies on TLBs to accelerate the address translation. Nowadays, the gap between application workloads and the capacity of TLB continues to grow, bringing many costly TLB misses and making the TLB a performance bottleneck. Previous studies seek to narrow the gap by exploiting the contiguity of physical pages. One promising solution is to group pages that are both virtually and physically contiguous into a memory range. Recording range translations can greatly increase the TLB reach, but ranges are also hard to index because they have arbitrary bounds. The processor has to compare against all the boundaries to determine which range an address falls in, which restricts the usage of memory ranges. In this article, we propose a tagged-pointer-based scheme, FlexPointer, to solve the range indexing problem. The core insight of FlexPointer is that large memory objects are rare, so we can create memory ranges based on such objects and assign each of them a unique ID. With the range ID integrated into pointers, we can index the range TLB with IDs and greatly simplify its structure. Moreover, because the ID is stored in the unused bits of a pointer and is not manipulated by the address generation, we can shift the range lookup to an earlier stage, working in parallel with the address generation. According to our trace-based simulation results, FlexPointer can reduce nearly all the L1 TLB misses, and page walks for a variety of memory-intensive workloads. Compared with a 4K-page baseline system, FlexPointer shows a 14% performance improvement on average and up to 2.8x speedup in the best case. For other workloads, FlexPointer shows no performance degradation. Dongwei Chen, Dong Tong 0001, Jiangfang Yi, Xu Cheng 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2022 | FlexPointer: Fast Address Translation Based on Range TLB and Tagged PointersabstractPage-based virtual memory suffers from costly page walks because of the gap between application workload sizes and TLB capacity. In this paper, we propose a tagged-pointer-based system, FlexPointer, to solve this problem. FlexPointer creates memory ranges from objects larger than a certain threshold by allocating contiguous physical pages for them. Virtual addresses within a range share a common translation entry, thus greatly expanding the TLB capacity. Because such objects are rare, we can assign each of them a unique ID and pass it to the hardware in pointer tags. According to our trace-based simulation results, FlexPointer can reduce nearly all the L1 TLB misses and page walks for a variety of memory-intensive workloads, providing a 14% performance improvement on average. Dongwei Chen, Dong Tong 0001, Jiangfang Yi, Xu Cheng 0001 |
PACT | 4 |
| 2013 | VFCC: A verification framework of cache coherence using parallel simulationabstractA cache coherence protocol is a vital component of a multiprocessor to maintain the data consistency. In this paper, we proposed VFCC, which is a simulation framework to validate a cache-coherence protocol implementation of a commercial 64-bit superscalar multiprocessor. It exploits multiple-level parallelism to accelerate validation without overheads among threads. Our experimental results demonstrate VFCC has a 5.0× speedup than a traditional simulator on a conventional 16-core host machine. Qiaoli Xiong, Jiangfang Yi, Tianbao Song, Zichao Xie, Dong Tong 0001 |
ASP-DAC | 2 |
| 2012 | S/DC: A storage and energy efficient data prefetcherabstractEnergy efficiency is becoming a major constraint in processor designs. Every component of the processor should be reconsidered to reduce wasted energy and area. Prefetching is an important technique for tolerating memory latency. Prefetcher designs have important impact on the energy efficiency of the memory hierarchy. Stride prefetchers require little storage, but cannot handle irregular access patterns. Delta correlation (DC) prefetchers can handle complicated access patterns, but waste storage because of storing multiple miss addresses for a stride pattern. Moreover, DC prefetchers waste the bandwidth and energy of the memory hierarchy because they cannot identify whether an address has been prefetched and generate a large number of redundant prefetches. In this paper, we propose a storage and energy efficient data prefetcher called stride/DC (S/DC) to combine the advantages of stride and DC prefetchers. S/DC uses a pattern prediction table (PPT) which stores two recent miss addresses in each entry to capture stride patterns. PPT avoids recording multiple miss addresses for a stride pattern, and thus improves the storage efficiency. When handling stride patterns, each PPT entry maintains a counter for obtaining the last prefetched address to avoid generating redundant prefetches. When handling other patterns, S/DC compares the new predicted address with earlier generated addresses in the prefetch queue and filters the redundant ones. In addition, to expand the filtering scope, S/DC uses a prefetch filter to store addresses evicted from the prefetch queue. In this way, S/DC reduces the bandwidth requirements and energy consumption of prefetching. Experimental results demonstrate that S/DC achieves comparable performance with only 24% of the storage and reduces 11.46% of the L2 cache energy, as compared to the CZone/DC prefetcher. Xianglei Dang, Xiaoyin Wang, Dong Tong 0001, Junlin Lu, Jiangfang Yi |
DATE | 5 |
| 2012 | SOLE: Speculative one-cycle load execution with scalability, high-performance and energy-efficiencyabstractConventional superscalar processors usually contain large CAM-based LSQ (load/store queue) with poor scalability and high energy consumption. Recently proposals only focus on improving the LSQ scalability to increase the in-flight instruction capacity, but with poor performance improvement and energy efficiency. This paper presents a novel speculative store-load forwarding mechanism, named SOLE (speculative one-cycle load execution)1. Firstly, SOLE uses address identifiers to determine the memory disambiguation, rather than the exact memory addresses as the traditional LSQ does. Since the address identifier is just simple hash from the address base and offset, the speculative store-load forwarding could be advanced earlier to reduce the load execution latency and avoid unnecessary energy consumption by filtering unnecessary accesses to the data cache. Secondly, SOLE enlarges the forwarding communication range by using SSN (store sequential number) to determine the age order between stores, which further improves the performance. Finally, the implementation of SOLE all uses set-associative structures that avoid the non-scalable problem of CAM-based LSQ. Experiments show that performance of SOLE outperforms the traditional LSQ by 13.57% in terms of performance, with only 75.2% execution energy consumption of the loads and stores. Zhen-Hao Zhang, Dong Tong 0001, Xiaoyin Wang, Jiangfang Yi |
ICCD | 4 |
| 2012 | Active Store Window: Enabling Far Store-Load Forwarding with Scalability and Complexity-Efficiency
Zhen-Hao Zhang, Xiaoyin Wang, Dong Tong 0001, Jiangfang Yi, Junlin Lu |
J. Comput. Sci. Technol. | 4 |
| 2010 | Research Progress of UniCore CPUs and PKUnity SoCs
Xu Cheng 0001, Xiaoyin Wang, Junlin Lu, Jiangfang Yi, Dong Tong 0001, Xuetao Guan, Xianhua Liu 0001, Yi Feng 0003 |
J. Comput. Sci. Technol. | 4 |