EDBT 2026 Demo / reviewers in the wild / expert
Yuhang Liu 0001
dblp:131/6710-1
· DBLP profile ↗
17ranked-venue papers
8as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Planaria: Pattern Directed Cross-page Composite PrefetcherabstractGiven the memory wall, the performance of the memory system significantly influences the user experience of mobile phones. The system cache (SC), located on the memory side, is shared among all CPUs and GPUs within the mobile phone, serving as the last line of defense before resorting to time-consuming off-chip memory access. Managing SC is a challenge due to its large working set and irregular access patterns. Despite occupying a substantial on-chip area, SC's effectiveness in terms of hit rate is relatively low. It has been observed that neither state-of-the-art cache replacement policies nor increasing cache size significantly improve SC performance. Prefetchers designed for higher-level caches cannot be seamlessly applied to SC due to the absence of the required program counter (PC) on the memory side and the violation of stringent power constraints in mobile phones by aggressive prefetch traffic. Yuhang Liu 0001, Mingyu Chen 0001 |
DAC | 1 |
| 2024 | Suppressing the Interference Within a Datacenter: Theorems, Metric and StrategyabstractAs the paradigm of cloud computing, a datacenter accommodates many co-running applications sharing system resources. Although highly concurrent applications improve resource utilization, the resulting resource contention can increase the uncertainty of quality of services (QoS). Previous studies have shown that achieving high resource utilization and high QoS simultaneously is challenging. Moreover, quantifying the intensity of interference across multiple concurrent applications in a datacenter, where applications can be either latency-critical (LC) or best-effort (BE), poses a significant challenge. To address these issues, we propose Ah-Q, which comprises a series of theorems, a quantification theory and a scheduling strategy. Firstly, we present the necessary and sufficient conditions to precisely test whether a datacenter is both QoS guaranteed and high-throughput. We also present and prove a theorem that reveals the relationship between tail latency and throughput. Our theoretical results are insightful and useful for building datacenters that have desirable performance. By applying our theoretical results, datacenter architects can more effectively balance the trade-off between resource utilization and QoS, leading to improved performance for co-running applications. Secondly, we propose the “System Entropy” (E$\rm {_{S}}$) theory to quantitatively and analytically measure interference in a datacenter. Interference arises due to resource scarcity or irrational scheduling, and effective scheduling can alleviate resource scarcity. To assess the effectiveness of a resource scheduling strategy, we introduce the concept of “resource equivalence”. We evaluate various resource scheduling strategies to demonstrate the correctness and effectiveness of the proposed theory. Thirdly, we introduce a new resource scheduling strategy, ARQ, that leverages both isolation and sharing of resources. Our evaluations show that ARQ significantly outperforms state-of-the-art strategies PARTIES and CLITE in reducing the tail latency of LC applications and increasing the IPC of BE applications. Yuhang Liu 0001, Jiapeng Zhou, Mingyu Chen 0001, Yungang Bao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Ah-Q: Quantifying and Handling the Interference within a Datacenter from a System PerspectiveabstractInterference among applications frequently occurs in a datacenter and significantly influences the cost-efficiency and the user experience. However, it is challenging for us to quantify the exact intensity of the interference that occurred in the overall system of a datacenter, because there are many concurrent applications in a datacenter, and their type can be either latency-critical (LC) and best-effort (BE). To address this issue, we present the Ah-Q which includes a theory and a strategy.First, we propose the "system entropy" (ES) theory to holistically and analytically quantify the interference in a datacenter to address this vital issue. The interference is caused by the scarcity of resources or/and the irrationality of scheduling. As more appropriate scheduling can compensate for resource scarcity, we derive the concept of "resource equivalence" to quantify the effectiveness of a resource scheduling strategy. We evaluate different resource scheduling strategies to validate the correctness and effectiveness of the proposed theory.Moreover, using the theory to eliminate interference, we propose a new resource scheduling strategy; i.e., ARQ, which dynamically allocates the isolated resources and the shared resources to simultaneously harvest the benefits of isolation and sharing. Our results show that compared to the state-of-the-art strategies (PARTIES and CLITE), ARQ is more effective to reduce the tail latency of the LC applications and to increase the IPC of the BE applications. Compared with PARTIES and CLITE, ARQ increases the yield (the ratio of satisfactory LC applications) by 25% and 20%, respectively; when the load is low, ARQ increases IPC of BE applications by 63.8% and 37.1%, respectively; ARQ reduces ESby 36.4% and 33.3%, respectively. The effectiveness of ARQ has saved resources significantly to achieve the same satisfactory overall user experience. Yuhang Liu 0001, Jiapeng Zhou, Mingyu Chen 0001, Yungang Bao |
HPCA | 1 |
| 2021 | LSP: Collective Cross-Page Prefetching for NVMabstractAs an emerging technique, non-volatile memory (NVM) provides valuable opportunities for boosting the memory system, which is vital for the computing system performance. However, one challenge preventing NVM from replacing DRAM as the main memory is that NVM row activation's latency is much longer (by approximately 10x) than that of DRAM. To address this issue, we present a collective cross-page prefetching scheme that can accurately open an NVM row in advance and then prefetch the data blocks from the opened row with low overhead. We identify a memory access pattern (referred to as a ladder stream) to facilitate prefetching that can cross page boundary, and propose the ladder stream prefetcher (LSP) for NVM. In LSP, two crucial components have been well designed. Collective Prefetch Table is proposed to reduce the interference with demand requests caused by prefetching through speculatively scheduling the prefetching according to the states of the memory queue. It is implemented with low overhead by using single entry to track multiple prefetches. Memory Mapping Table is proposed to accurately prefetch future pages by maintaining the mapping between physical and virtual addresses. Experimental evaluations show that LSP improves the memory system performance with no prefetching by 66%, and the improvement over the state-of-the-art prefetchers, Access Map Pattern Matching Prefetcher (AMPM), Best-Offset Prefetcher (BOP) and Signature Path Prefetcher (SPP) is 26.6%. 21.7% and 27.4%. respectively. Haiyang Pan, Yuhang Liu 0001, Tianyue Lu, Mingyu Chen 0001 |
DATE | 2 |
| 2020 | IMPULP: A Hardware Approach for In-Process Memory Protection via User-Level Partitioning
Mingyu Chen 0001, Yuhang Liu 0001, Zong-Hao Yang, Zonghui Hong, Yunge Guo |
J. Comput. Sci. Technol. | 3 |
| 2019 | Make Page Coloring more Efficient on Slice-Based Three-Level CacheabstractOn modern multi-core machines, page coloring has been used to alleviate the competition at Last Level Cache (LLC). However, the latest development of CPU architecture has brought new issues to page coloring. Firstly, in the case of three-level cache, previous works about page coloring did not discuss the impact on L2 cache of color allocation and the competition for L2 cache is not considered concurrently under hyper-threading. In addition, as the last level cache structure is changed from shared to slice-based and undocumented hash function is applied, page coloring is more complex and slice information is also not fully utilized. This paper presents solutions to these issues. Firstly, by making small changes to the traditional page coloring, the problem that page coloring may waste L2 cache is alleviated. At the same time, we rethink the vertical allocation of L2 cache and LLC in page coloring under hyper-threading, and discuss the impact of color allocation on programs, especially those with different sensitivity to L2 cache and LLC. Finally, we make full use of slice information and propose Partial Conflict Color (PCC). At the same time, we also propose a fast method to obtain PCC. Experiments show that using PCC can improve system performance when the number of colors is insufficient. Tianyue Lu, Yuhang Liu 0001, Mingyu Chen 0001 |
ICPADS | 3 |
| 2019 | HCMA: Supporting High Concurrency of Memory Accesses with Scratchpad Memory in FPGAsabstractCurrently many researches focus on new methods of accelerating memory accesses between memory controller and memory modules. However, the absence of an accelerator for memory accesses between CPU and memory controller wastes the performance benefits of new methods. Therefore, we propose a coordinated batch method to support high concurrency of memory accesses (HCMA). Compared to the conventional method of holding outstanding memory access requests in miss status handling registers (MSHRs), HCMA method takes advantage of scratchpad memory in FPGAs or SoCs to circumvent the limitation of MSHR entries. The concurrency of requests is only limited by the capacity of scratchpad memory. Moreover, to avoid the higher latency when searching more entries, we design an efficient coordinating mechanism based on circular queues.We evaluate the performance of HCMA method on an MP-SoC FPGA platform. Compared to conventional methods based on MSHRs, HCMA method supports ten times of concurrent memory accesses (from 10 to 128 entries on our evaluation platform). HCMA method achieves up to 2.72× memory bandwidth utilization for applications that access memory with massive fine-grained random requests, and to 3.46× memory bandwidth utilization for stream-based memory accesses. For real applications like CG, our method improves speedup performance by 29.87%. Yuhang Liu 0001, Mingyu Chen 0001 |
NAS | 2 |
| 2019 | LPM: A Systematic Methodology for Concurrent Data Access Pattern Optimization from a Matching PerspectiveabstractAs applications become increasingly data intensive, conventional computing systems become increasingly inefficient due to data access performance bottlenecks. While intensive efforts have been made in developing new memory technologies and in designing special purpose machines, there is a lack of solutions for evaluating and utilizing recent hardware advancements to address the memory-wall problem in a systematic way. In this study, we present the memory Layered Performance Matching (LPM) methodology to provide a systematic approach for data access performance optimization. LPM uniquely presents and utilizes the data access concurrency, in addition to data access locality, in a memory hierarchical system. The LPM methodology consists of models and algorithms, and is supported with a series of analytic results for its correctness. The rationale of LPM is to reduce the overall data access delay through the matching of data request rate and data supply rate at each layer of a memory hierarchy, with a balanced consideration of data locality, data concurrency, and latency hiding of data flow. Extensive experimentations on both physical platforms and software simulators confirm our theoretical findings, and they show that the LPM approach can be applied in diverse computing platforms and can effectively guide performance optimization of memory systems. Yuhang Liu 0001, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | CaL: Extending Data Locality to Consider Concurrency for Performance OptimizationabstractBig data applications demand a better memory performance. Data Locality has been the focus of reducing data access delay. Data access concurrency, however, has become prevalent in modern memory systems in recent years. How to extend existing locality-based performance optimization to consider data concurrency becomes a timely issue facing the researchers and practitioners in the field of computing, especially in the field of big data computing. In this study, we introduce the concept and definition of Concurrency-aware data access Locality (CaL), which, as its name states, extends the concept of locality by considering concurrency. Compared to the conventional concept of locality, CaL accurately reflects the combined impact of data access locality and concurrency in modern memory systems and is very effective for data intensive applications. The value of CaL can be quantitatively measured directly by performance counters in mainstream commercial processors and is practically feasible. Two theoretical results are presented to reveal the relationships between CaL and existing memory system performance metrics of memory accesses per cycle (APC), average memory access time (AMAT), and memory bandwidth (B). In this way, we provide a methodology to use existing locality-based optimization methods directly or in combination with data concurrency optimizations, to improve the value of CaL and to improve the performance of a memory system. To demonstrate the practical value of CaL, we conduct four case studies to illustrate the power of concurrency-aware locality optimization. Compared with the conventional locality based optimization, the CaL-aware design has achieved significant performance improvement. It achieved a 3.12-fold speedup on K-means, which is a widely-used data analytic kernel from the big data benchmarks. Yuhang Liu 0001, Xian-He Sun |
IEEE Trans. Big Data | 1 |
| 2018 | PTAT: An Efficient and Precise Tool for Tracing and Profiling Detailed TLB MissesabstractAs the memory access footprints of applications in areas like data analytics increase, the latency overhead of translation lookaside buffer (TLB) misses increases. Thus, the efficiency of TLB becomes increasingly critical for overall system performance. Analyzing TLB miss traces is useful for hardware architecture design and software application optimization. Utilizing cycle-accurate simulators or instrumentation tools is very time-consuming and/or inaccurate for tracing and profiling TLB misses. In this article, we propose an efficient and precise tool to collect and profile last-level TLB misses. This tool utilizes a novel software method called Page Table Access Tracing (PTAT), storing last-level page table entries of certain workload processes into a reserved uncached memory region. Therefore, each last-level TLB miss incurred by user process corresponds to one uncached page table access to main memory, which can be captured and recorded by a hardware memory bus monitor. The detected information is then dumped into offline storage. In this manner, full TLB miss traces are collected and can be analyzed flexibly. Compared to previous software-based methods, this method achieves higher performance. Experiments show that, compared with a state-of-the-art kernel instrumentation method (BadgerTrap), which lacks complete dumping trace function, the speedup is still up to 3.88-fold for memory-intensive benchmarks. Due to the improved efficiency and completeness of tracing, case studies validate that more flexible profiling can be conducted, which is of great significance for TLB performance optimization. The accuracy of PTAT is verified by both dedicated sequence and performance counters. Jiutian Zhang, Yuhang Liu 0001, Mingyu Chen 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | SMEFF: A scalable memory extension fabric for FPGAabstractIn resource-constrained FPGA systems, off-chip memory plays an important role in both prototype verification and acceleration systems for big data. As the scale of applications become increasingly large and complex, the data to be processed grows exponentially. In contrast, FPGAs provide limited memory capacity and bandwidth, severely limiting the scale and performance of prototype verification systems and acceleration systems. Furthermore, data movement is expected to be a dominant consumer of energy, thus inefficient data movement between different DRAM modules also incurs significant performance and energy penalties. This paper proposes a practical design: A Scalable Memory Extension Fabric for FPGA (SMEFF), which is an asynchronous memory access mechanism and exploits cascaded technology to solve the problem of memory capacity and bandwidth. SMEFF uses two key technologies to achieve memory capacity and bandwidth improvements, and shrink the latency and overhead of data movement-the first is an FPGA-based high-speed serial bus to build a multi-level memory fabric instead of the traditional parallel bus mechanism to solve the signal integrity problem. The second is a module to module (M-To-M) DMA data movement technology, which reduces the latency and overhead of data movement between memory modules. We implement SMEFF on an FPGA-based prototype to demonstrate the feasibility of our approach. Experimental results show that SMEFF provides 5x memory capacity increase and up to 3.6x memory bandwidth improvement compared to state-of-the-art FPGA-based memory systems, and outperforms PCIe-based systems. The data movement of M-TO-M's DMA technology obtains up to 3x latency reduction, and average of 21.1% to 61.1% energy reduction compared to state-of-the-art FPGA-based memory systems. SMEFF thus increases FPGA-based prototype memory systems capacity and bandwidth. More importantly, our architecture provides opportunities for the design of scalable, cost-effective FPGA-based memory subsystems. Yuhang Liu 0001, Mingyu Chen 0001 |
FPT | 3 |
| 2017 | TDV Cache: Organizing Off-Chip DRAM Cache of NVMM from a Fusion PerspectiveabstractEmerging Non-Volatile Memory (NVM) provides both larger memory capacity and higher energy efficiency, but has much longer access latency than traditional DRAM, thus DRAM can be used as an efficient cache to hide the long latency of Non-Volatile Main Memory (NVMM) system. Transparent Off-chip DRAM cache (TOD cache) is a new DRAM cache structure where off-chip DRAM module is used as L4 cache and managed by hardware. The capacity and latency ratio of TOD cache over NVM are both quite different from those of traditional on-chip SRAM or die-stacked DRAM cache over off-chip DRAM memory. All the factors including hit latency, miss latency and hit rate need to be re-considered for TOD cache design. In this study, we first point out that three types of traditional cache schemes cannot be used directly for TOD cache, since set-associative cache suffers from extra tag lookup latency, direct-mapped cache has low hit rate and tag cache is too small to efficiently hold the working sets of tags for DRAM cache. Based on these observations, we propose a novel cache scheme, TDV, that fuses these three different types of cache together to take their advantages. In TDV, a direct-mapped cache is used as the first-level cache to achieve short access latency, a set-associative victim cache is taken as the second-level cache to obtain extra high hit rate, and a SRAM tag cache only serves for the victim cache rather than the whole DRAM cache and thus improves the hit rate of tag cache significantly. The simulation results show that, TDV cache has a performance improvement of 6.3% and 8.3% on average than state-of-the-art direct-mapped (Alloy cache) and set-associative cache (ATCache) with same DRAM and SRAM capacity. Tianyue Lu, Yuhang Liu 0001, Haiyang Pan, Mingyu Chen 0001 |
ICCD | 2 |
| 2017 | PTAT: An efficient and precise tool for collecting detailed TLB miss tracesabstractIt is well known that the TLB performance impacts the memory system performance, which is critical for overall system performance. Similar to multi-level caches, multilevel TLBs have become an important leverage for boosting data access performance. Applications have increasingly large working sets. Servers targeting such applications have thus been built with ever larger main memory capacities, but there has been no commensurate growth in TLB sizes. Designing high performance and energy efficient memory hierarchies require insight into the behavior of current designs: when do they work well, and when do they fall short of expectations. Profiling the TLB misses is the prerequisite for memory system optimization. Both designing efficient TLB architecture and TLB-friendly applications require analysis of TLB miss behavior. Although researchers have extensively studied TLB behavior, current approaches have some issues in either efficiency or precision. Jiutian Zhang, Yuhang Liu 0001, Yuan Ruan, Mingyu Chen 0001 |
ISPASS | 2 |
| 2016 | Efficient design space exploration via statistical sampling and AdaBoost learningabstractDesign space exploration (DSE) has become a notoriously difficult problem due to the exponentially increasing size of design space of microprocessors and time-consuming simulations. To address this issue, machine learning techniques have been widely employed to build predictive models. However, most previous approaches randomly sample the training set leading to considerable simulation cost and low prediction accuracy. In this paper, we propose an efficient and precise DSE methodology by combining statistical sampling and Adaboost learning technique. The proposed method includes three phases. (1) Firstly, orthogonal design based feature selection is employed to prune design space. (2) Sencondly, an orthogonal array based training data sampling method is introduced to select the representative configurations for simulation. (3) Finally, a new active learning approach ActBoost is proposed to build predictive model. Evaluations demonstrate that the proposed framework is more efficient and precise than state-of-art DSE techniques. Shuzhen Yao, Yuhang Liu 0001, Senzhang Wang, Xian-He Sun |
DAC | 3 |
| 2015 | LPM: Concurrency-Driven Layered Performance MatchingabstractData access has become the preeminent performance bottleneck of computing. In this study, a Layered Performance Matching (LPM) model and its associated algorithm are proposed to match the request and reply speed for each layer of a memory hierarchy to improve memory performance. The rationale of LPM is that the performance of each layer of a memory hierarchy should and can be optimized to closely match the request of the layer directly above it. The LPM model simultaneously considers both data access concurrency and locality. It reveals the fact that increasing the effective overlapping between hits and misses of the higher layer will alleviate the performance impact of the lower layer. The terms pure miss and pure miss penalty are introduced to measure the effectiveness of such hit-miss overlapping. By distinguishing between (general) miss and pure miss, we have made LPM optimization practical and feasible. Our evaluation shows the data stall time can be reduced significantly with an optimized hardware configuration. We also have achieved noticeable performance improvement by simply adopting smart LPM scheduling without changing the underlying hardware configurations. Analysis and experimental results show LPM is feasible and effective. It provides a novel and efficient way to cope with the ever-widening memory wall problem, and to optimize the vital memory system design. Yuhang Liu 0001, Xian-He Sun |
ICPP | 1 |
| 2015 | C2-bound: a capacity and concurrency driven analytical model for many-core designabstractIn this paper, we propose C2-Bound, a data-driven analytical model, that incorporates both memory capacity and data access concurrency factors to optimize many-core design. C2-Bound is characterized by combining the newly proposed latency model, concurrent average memory access time (C-AMAT), with the well-known memory-bounded speedup model (Sun-Ni's law) to facilitate computing tasks. Compared to traditional chip designs that lack the notion of memory concurrency and memory capacity, C2-Bound model finds memory bound factors significantly impact the optimal number of cores as well as their optimal silicon area allocations, especially for data-intensive applications with a none parallelizable sequential portion. Therefore, our model is valuable to the design of new generation many-core architectures that target big data processing, where working sets are usually larger than conventional scientific computing. These findings are evidenced by our detailed simulations, which show with C2-Bound the design space can be narrowed down significantly up to four orders of magnitude. C2-Bound analytic results can be either used in reconfigurable hardware environments or, by software designers, applied to scheduling, partitioning, and allocating resources among diverse applications. Yuhang Liu 0001, Xian-He Sun |
SC | 1 |
| 2015 | Reevaluating Data Stall Time with the Consideration of Data Access Concurrency
Yuhang Liu 0001, Xian-He Sun |
J. Comput. Sci. Technol. | 1 |