Fangzhou Liu 0004

dblp:57/7824-4 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-0715-7313ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021
YearPublicationVenuePosition
2024 Parallel Loop Locality Analysis for Symbolic Thread Counts
abstract
Data movement limits program performance. This bottleneck is more significant in multi-thread programs but more difficult to analyze, especially for multiple thread counts.
Fangzhou Liu 0004, Yifan Zhu 0003, Shaotong Sun, Chen Ding 0001, Wesley Smith, Kaave Hosseini
PACT1
2023 Cache Programming for Scientific Loops Using Leases
abstract
Cache management is important in exploiting locality and reducing data movement. This article studies a new type of programmable cache called the lease cache. By assigning leases, software exerts the primary control on when and how long data stays in the cache. Previous work has shown an optimal solution for an ideal lease cache. This article develops and evaluates a set of practical solutions for a physical lease cache emulated in FPGA with the full suite of PolyBench benchmarks. Compared to automatic caching, lease programming can further reduce data movement by 10% to over 60% when the data size is 16 times to 3,000 times the cache size, and the techniques in this article realize over 80% of this potential. Moreover, lease programming can reduce data movement by another 0.8% to 20% after polyhedral locality optimization.
Benjamin Reber, Matthew Gould, Alexander H. Kneipp, Fangzhou Liu 0004, Ian Prechtl, Chen Ding 0001, Dorin Patru
ACM Trans. Archit. Code Optim.4
2022 CARL: Compiler Assigned Reference Leasing
abstract
Data movement is a common performance bottleneck, and its chief remedy is caching. Traditional cache management is transparent to the workload: data that should be kept in cache are determined by the recency information only, while the program information, i.e., future data reuses, is not communicated to the cache. This has changed in a new cache design named Lease Cache . The program control is passed to the lease cache by a compiler technique called Compiler Assigned Reference Lease (CARL). This technique collects the reuse interval distribution for each reference and uses it to compute and assign the lease value to each reference. In this article, we prove that CARL is optimal under certain statistical assumptions. Based on this optimality, we prove miss curve convexity, which is useful for optimizing shared cache, and sub-partitioning monotonicity, which simplifies lease compilation. We evaluate the potential using scientific kernels from PolyBench and show that compiler insertions of up to 34 leases in program code achieve similar or better cache utilization (in variable size cache) than the optimal fixed-size caching policy, which has been unattainable with automatic caching but now within the potential of cache programming for all tested programs and most cache sizes.
Chen Ding 0001, Dong Chen 0015, Fangzhou Liu 0004, Benjamin Reber, Wesley Smith
ACM Trans. Archit. Code Optim.3
2021 Uniform lease vs. LRU cache: analysis and evaluation
abstract
Lease caching is a new technique that provides greater control of the cache than what is allowed in conventional caches. The simplest control is uniform lease (UL), which means that all leases are identical in length. The UL cache is prescriptive and based on allocation. In comparison, a conventional cache is reactive and based on replacement. They represent two fundamentally different approaches to cache management.
Dong Chen 0015, Chen Ding 0001, Fangzhou Liu 0004, Benjamin Reber, Wesley Smith, Pengcheng Li 0001
ISMM3
2020 PLUM: static parallel program locality analysis under uniform multiplexing
abstract
Data movement has a significant impact on program performance. For multithread programs, this impact is amplified, since different threads often interfere with each other by competing for shared cache space. However, recent de facto locality metrics consider either sequential execution only, or derive locality for multithread programs in an inefficient way, i.e. exhaustive simulation.
Fangzhou Liu 0004, Dong Chen 0015, Wesley Smith, Chen Ding 0001
PPoPP1
2018 Prediction and bounds on shared cache demand from memory access interleaving
abstract
Cache in multicore machines is often shared, and the cache performance depends on how memory accesses belonging to different programs interleave with one another. The full range of performance possibilities includes all possible interleavings, which are too numerous to be studied by experiments for any mix of non-trivial programs.
Jacob Brock, Chen Ding 0001, Rahman Lavaee, Fangzhou Liu 0004
ISMM4
2018 Locality analysis through static parallel sampling
abstract
Locality analysis is important since accessing memory is much slower than computing. Compile-time locality analysis can provide detailed program-level feedback for compilers or runtime systems faster than trace-based locality analysis.
Dong Chen 0015, Fangzhou Liu 0004, Chen Ding 0001, Sreepathi Pai
PLDI2