VLDB 2026 Research / reviewers in the wild / expert
Tianyue Lu
dblp:125/2112
· DBLP profile ↗
16ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0001-8808-0935ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ANIMo: Accelerating Nested Isolation with Monitor-free Domain Transition
Yibin Xu, Tianyi Huang, Tianyue Lu, Mingyu Chen 0001 |
ASP-DAC | 5 |
| 2026 | S-MSHR: A Scalable MSHR Architecture Using Cache Tag Data-Ready Bits and Index Queues
Xu Zhang 0033, Yibin Xu, Tianyue Lu, Mingyu Chen 0001 |
CCGrid | 5 |
| 2026 | LightSerial: Accelerating In-Process Isolation via Implicit Dependency ExposureabstractWith the frequent emergence of in-process security threats, in-process isolation has become essential for software security. While hardware security primitives enable lightweight isolation, an implicit dependency persists between permission-setting instructions and their subsequent checked instructions. To enforce this implicit dependency, modern out-of-order processors resort to enforcing strict serialization by flushing the entire pipeline during permission switches. This approach severely degrades instruction-level parallelism. Moreover, serialization leaves an exploitable security window for side-channel attacks by failing to revert speculative microarchitectural side effects. Yibin Xu, Tianyue Lu, Mingyu Chen 0001 |
CF | 4 |
| 2026 | AExec: Asynchronous Multi-accelerator Execution and Management Mechanism
Xiaokun Pei, Zhuolun Jiang, Mingyu Chen 0001, Songyue Wang, Tianyue Lu |
CF | 5 |
| 2025 | DASICS: Efficient In-Process Protection with Hardware-Assisted Dynamic CompartmentalizationabstractHardware-assisted in-process compartmentalization reduces attack surface at low cost, but existing methods face practical challenges: inefficient dynamic permission management, weak metadata/instruction protection, and limited resource isolation. To tackle these problems, this paper proposes DASICS, a lightweight and efficient design of hardware-assisted in-process compartmentalization. DASICS partitions the process code segments into trusted and untrusted compartments and implements a user-mode protection runtime in the trusted compartment for dynamic permission management. It employs boundary registers to enforce dynamic access-control restrictions on instructions within different untrusted compartments. Additionally, it applies metadata access restriction, control-flow checks, and systemcall filtering for the untrusted compartments to achieve comprehensive protection. We implemented a hardware prototype of DASICS on the RISC-V XiangShan superscalar out-of-order processor and validated its effectiveness on FPGA. Our prototype increases less than 5% LUTs cost, and experimental results show that DASICS isolation incurs an average overhead of${6.02 \%}$on Memcached key-value store and 8.18% on NGINX webserver. DASICS project is publicly available at github.com/DASICS-ICT. Yibin Xu, Tianyi Huang, Tianyue Lu, Mingyu Chen 0001 |
ICCD | 6 |
| 2024 | Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory AccessabstractThe growing memory demands of modern applications have driven the adoption of far memory technologies in data centers to provide cost-effective, high-capacity memory solutions. However, far memory presents new performance challenges because its access latencies are significantly longer and more variable than local DRAM. For applications to achieve acceptable performance on far memory, a high degree of memory-level parallelism (MLP) is needed to tolerate the long access latency. While modern out-of-order processors are capable of exploiting a certain degree of MLP, they are constrained by resource limitations and hardware complexity. The key obstacle is the synchronous memory access semantics of traditional load/store instructions, which occupy critical hardware resources for a long time. The longer far memory latencies exacerbate this limitation. This article proposes a set of Asynchronous Memory Access Instructions (AMI) and its supporting function unit, Asynchronous Memory Access Unit (AMU), inside contemporary Out-of-Order Core. AMI separates memory request issuing from response handling to reduce resource occupation. Additionally, AMU architecture supports up to several hundreds of asynchronous memory requests through re-purposing a portion of L2 Cache as scratchpad memory (SPM) to provide sufficient temporal storage. Together with a coroutine-based programming framework, this scheme can achieve significantly higher MLP for hiding far memory latencies. Evaluation with a cycle-accurate simulation shows AMI achieves 2.42× speedup on average for memory-bound benchmarks with 1μs additional far memory latency. Over 130 outstanding requests are supported with 26.86× speedup for GUPS (random access) with 5 μs latency. These demonstrate how the techniques tackle far memory performance impacts through explicit MLP expression and latency adaptation. Luming Wang, Xu Zhang 0033, Songyue Wang, Zhuolun Jiang, Tianyue Lu, Mingyu Chen 0001, Siwei Luo, Keji Huang |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | Rethinking Design Paradigm of Graph Processing System with a CXL-like Memory Semantic FabricabstractWith the evolution of network fabrics, message-passing clusters have been promising solutions for large-scale graph processing. Alternatively, the shared-memory model is also introduced to avoid redundant copies and extra storage space of graph data. Compared to conventional network fabrics, with the capability of fine-grained, byte-addressable remote memory access, emerging memory semantic interconnects and fabrics, e.g., Intel's Compute Express Link (CXL), are intuitively more appropriate for adoption in shared-memory clusters. However, due to the latency gap between local and remote memory, it is still challenging to take advantage of the shared-memory graph processing with memory semantic fabrics. To tackle this problem, in this paper, we first investigate memory access characterizations of graph vertex propagation based on the shared-memory model. Then we propose GraCXL, a series of design paradigms to address high-frequency and long-latency of remote memory access potentially incurred in CXL-based clusters. For system adaptiveness, we elaborate GraCXL towards the general-purpose CPU cluster and the domain-specific FPGA accelerator array, respectively. We design a custom fabric with the CXL.mem protocol and leverage a couple of ARM SoC-equipped FPGAs to build an evaluation prototype in the absence of commodity CXL hardware and platforms. Experimental results show that the proposed GraCXL CPU and FPGA clusters achieve 1.33x-8.92x and 2.48x-5.01x performance improvement, respectively. Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Zhang 0017, Mingyu Chen 0001 |
CCGrid | 3 |
| 2023 | MARB: Bridge the Semantic Gap between Operating System and Application Memory Access BehaviorabstractThe virtual memory subsystem (VMS) is a long-standing and integral part of an operating system (OS). It plays a vital role in enabling remote memory systems over fast data center networks and is promising in terms of transparency and generality. Specifically, these systems use three VMS mechanisms: demand paging, page swapping, and page prefetching. However, the VMS inherent data path is costly, which takes a huge toll on performance. Despite prior efforts to propose page swapping and prefetching algorithms to minimize the occurrences of the data path, they still fall short due to the semantic gap between the OS and applications - the VMS has limited knowledge of its running applications' memory access behaviors. In this paper, orthogonal to prior efforts, we take a fundamen-tally different approach by building an efficient framework to collect full memory access traces at the local bus, and make them available to the OS through CPU cache. Consequently, the page swapping and page prefetching can use this trace to make better decisions, thereby improving the overall performance of systems. We implement a proof-of-concept prototype on commodity x86 servers using a hardware-based memory tracking tool. To show-case our framework's benefits, we integrate it with a state-of-the-art remote memory system and the default kernel page eviction subsystem. Our evaluation shows promising improvements. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yisong Chang, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
DATE | 5 |
| 2023 | HoPP: Hardware-Software Co-Designed Page Prefetching for Disaggregated MemoryabstractMemory disaggregation is a promising direction to mitigate memory contention in datacenters. To make memory disaggregation practical, prior efforts expose remote memory to applications transparently via virtual memory subsystem’s swapping interface. However, due to the semantic gap between OS and applications – OS cannot know the memory accessing sequences of an application but via page faults. This approach has two limitations. First, it learns little from page faults’ access history, which leads to sub-optimal prefetching predictions. Second, a page fault can still occur even if there is a prefetch-hit which leads to a large kernel overhead.To address such limitations, our key insight is to decouple the address capturing from page faults by collecting full memory access traces in the memory controller. Using this idea, we buildHoPP– a hardware-software co-designed prefetching framework.HoPPadds hardware modules to the memory controller to feed sufficient hot pages to OS in real-time, which has three benefits inHoPP’s software design: 1) it improves existing prefetching algorithms with simple revamps, also offers more insights to build better policies; 2) the prefetch algorithm can run as a separate data path alongside the normal remote data path via page faults, potentially hiding the swap latency from applications, and enabling fine-grained control over prefetching behaviors; 3) the prefetch-hit overhead can be eliminated by early page table entry (PTE) injection, i.e., inject PTE for the prefetched page as soon as it returns. We implemented a proof-of-concept prototype using commodity servers along with a hardware-based memory tracking tool calledHMTTto emulate a modified memory controller. Results show that compared to Fastswap and Leap,HoPP-optimized prefetching algorithm achieves over 90% accuracy and coverage, which leads to up to 59% completion time improvement for various datacenter applications. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
HPCA | 5 |
| 2023 | Morpheus: An Adaptive DRAM Cache with Online Granularity Adjustment for Disaggregated MemoryabstractDisaggregated memory introduces a cost-effective solution for improving the memory utilization rate of data centers, by sharing a distributed memory pool among several individual servers. However, latency penalty in the existing connection between a computing node and the memory pool introduces performance degradation due to frequent far memory accesses. Based on our observation, page caching in the local DRAM, despite its reductions in the number of far memory accesses, still faces severe data over-fetching problem.With a detailed analysis of far memory access traces collected via several representative real-world applications, we argue that exploiting the various page-specific preferences of caching granularity is the key point of solving the data over-fetching problem in the DRAM cache. Consequently, in this paper, we present that it is influential to enable 1) dynamic selection of caching granularity for each page to not only guarantee sufficient spatial localities compared to the conventional fine-grained cache lines but also avoid data over-fetching caused by the coarse-grained pages, as well as 2) adaptive adjustment of cache capacity during execution for each granularity to accommodate the varying proportion of pages with different granularity preferences. Specifically, we propose Morpheus, an adaptive DRAM cache architecture that determines an optimal page-specific caching granularity at run-time and dynamically adjusts capacity occupations of different caching granularities. Based on our modeling and evaluations within the DRAMSim3 simulator, Morpheus exhibits 1.17-1.34x performance speedup for a wide range of workloads against the state-of-the-art DRAM cache design. Xu Zhang 0033, Tianyue Lu, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001 |
ICCD | 2 |
| 2022 | GraFF: A Multi-FPGA System with Memory Semantic Fabric for Scalable Graph ProcessingabstractFPGA has been a promising solution for graph processing in many scenarios. With a rapid growth in graph size, the on/off-chip memory capacity of a single FPGA is insufficient to hold large-scale graphs. To tackle such problem, in this position paper, we introduce GraFF, a Graph processing system with multiple FPGAs interconnected via a custom memory semantic Fabric. In order to efficiently exploit system parallelism, we first split the traversal of graph data into a series of independent fine-grained flits that are concurrently delivered among FPGAs as sheer memory semantic transactions. Then we relax FPGAs' synchronization from strict barrier boundaries between adjacent supersteps to fully parallelize graph traversing and computing. We build a prototype of GraFF with four custom FPGA nodes. Preliminary evaluation result based on the Breadth First Search (BFS) algorithm shows that the peak performance of GraFF reaches up to 6.23 GTEPS. Moreover, GraFF exhibits linear scalability when the number of FPGAs rises from one to four. Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Liu 0004, Ke Zhang 0017, Mingyu Chen 0001 |
FPT | 3 |
| 2021 | LSP: Collective Cross-Page Prefetching for NVMabstractAs an emerging technique, non-volatile memory (NVM) provides valuable opportunities for boosting the memory system, which is vital for the computing system performance. However, one challenge preventing NVM from replacing DRAM as the main memory is that NVM row activation's latency is much longer (by approximately 10x) than that of DRAM. To address this issue, we present a collective cross-page prefetching scheme that can accurately open an NVM row in advance and then prefetch the data blocks from the opened row with low overhead. We identify a memory access pattern (referred to as a ladder stream) to facilitate prefetching that can cross page boundary, and propose the ladder stream prefetcher (LSP) for NVM. In LSP, two crucial components have been well designed. Collective Prefetch Table is proposed to reduce the interference with demand requests caused by prefetching through speculatively scheduling the prefetching according to the states of the memory queue. It is implemented with low overhead by using single entry to track multiple prefetches. Memory Mapping Table is proposed to accurately prefetch future pages by maintaining the mapping between physical and virtual addresses. Experimental evaluations show that LSP improves the memory system performance with no prefetching by 66%, and the improvement over the state-of-the-art prefetchers, Access Map Pattern Matching Prefetcher (AMPM), Best-Offset Prefetcher (BOP) and Signature Path Prefetcher (SPP) is 26.6%. 21.7% and 27.4%. respectively. Haiyang Pan, Yuhang Liu 0001, Tianyue Lu, Mingyu Chen 0001 |
DATE | 3 |
| 2019 | Make Page Coloring more Efficient on Slice-Based Three-Level CacheabstractOn modern multi-core machines, page coloring has been used to alleviate the competition at Last Level Cache (LLC). However, the latest development of CPU architecture has brought new issues to page coloring. Firstly, in the case of three-level cache, previous works about page coloring did not discuss the impact on L2 cache of color allocation and the competition for L2 cache is not considered concurrently under hyper-threading. In addition, as the last level cache structure is changed from shared to slice-based and undocumented hash function is applied, page coloring is more complex and slice information is also not fully utilized. This paper presents solutions to these issues. Firstly, by making small changes to the traditional page coloring, the problem that page coloring may waste L2 cache is alleviated. At the same time, we rethink the vertical allocation of L2 cache and LLC in page coloring under hyper-threading, and discuss the impact of color allocation on programs, especially those with different sensitivity to L2 cache and LLC. Finally, we make full use of slice information and propose Partial Conflict Color (PCC). At the same time, we also propose a fast method to obtain PCC. Experiments show that using PCC can improve system performance when the number of colors is insufficient. Tianyue Lu, Yuhang Liu 0001, Mingyu Chen 0001 |
ICPADS | 2 |
| 2017 | TDV Cache: Organizing Off-Chip DRAM Cache of NVMM from a Fusion PerspectiveabstractEmerging Non-Volatile Memory (NVM) provides both larger memory capacity and higher energy efficiency, but has much longer access latency than traditional DRAM, thus DRAM can be used as an efficient cache to hide the long latency of Non-Volatile Main Memory (NVMM) system. Transparent Off-chip DRAM cache (TOD cache) is a new DRAM cache structure where off-chip DRAM module is used as L4 cache and managed by hardware. The capacity and latency ratio of TOD cache over NVM are both quite different from those of traditional on-chip SRAM or die-stacked DRAM cache over off-chip DRAM memory. All the factors including hit latency, miss latency and hit rate need to be re-considered for TOD cache design. In this study, we first point out that three types of traditional cache schemes cannot be used directly for TOD cache, since set-associative cache suffers from extra tag lookup latency, direct-mapped cache has low hit rate and tag cache is too small to efficiently hold the working sets of tags for DRAM cache. Based on these observations, we propose a novel cache scheme, TDV, that fuses these three different types of cache together to take their advantages. In TDV, a direct-mapped cache is used as the first-level cache to achieve short access latency, a set-associative victim cache is taken as the second-level cache to obtain extra high hit rate, and a SRAM tag cache only serves for the victim cache rather than the whole DRAM cache and thus improves the hit rate of tag cache significantly. The simulation results show that, TDV cache has a performance improvement of 6.3% and 8.3% on average than state-of-the-art direct-mapped (Alloy cache) and set-associative cache (ATCache) with same DRAM and SRAM capacity. Tianyue Lu, Yuhang Liu 0001, Haiyang Pan, Mingyu Chen 0001 |
ICCD | 1 |
| 2014 | Achieving efficient packet-based memory system by exploiting correlation of memory requestsabstractPacket-based interface is a trend for future memory system to alleviate memory capacity and bandwidth bottlenecks. On the other hand fine-grained memory access has been proven to efficiently reduce memory power. However leveraging both these two technologies will result in high packet overhead, because previous implementations of packet-based interface all adopt a simple design that a single packet is dedicated to a single request (SPSR). In this paper, we propose three optimizations to overcome the problem by exploiting correlations of memory requests. First, we propose a novel single packet multiple requests (SPMR) interface that encapsulates multiple requests into a packet to share packet header and tail. Second, we propose an adaptive address compression mechanism within a packet by adopting a base-difference algorithm. Third, we propose a mechanism to merge multiple memory requests with continuous access addresses into a single request before packing. By this way, the granularity constraint of cache line size is broken to enable efficiently row buffer scheduling. The experimental results show that, for certain memory-intensive workloads, the optimizations can effectively reduce packet overhead by about 53.9% and improve system performance by about 63.6% in average. Tianyue Lu, Licheng Chen, Mingyu Chen 0001 |
DATE | 1 |
| 2014 | MIMS: Towards a Message Interface Based Memory System
Licheng Chen, Mingyu Chen 0001, Yuan Ruan, Yongbing Huang, Zehan Cui, Tianyue Lu, Yungang Bao |
J. Comput. Sci. Technol. | 6 |