EDBT 2026 Demo / reviewers in the wild / expert
Xu Zhang 0033
dblp:98/5660-33
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0009-0027-9818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S-MSHR: A Scalable MSHR Architecture Using Cache Tag Data-Ready Bits and Index Queues
Xu Zhang 0033, Yibin Xu, Tianyue Lu, Mingyu Chen 0001 |
CCGrid | 2 |
| 2026 | RaidenSwap: A Multi-Swap Remote System for Multi-core ApplicationsabstractKernel-based remote memory systems are gaining traction in datacenters due to their significant improvement in memory utilization and their ability to transparently provide applications with unlimited memory capacity. However, the high degree of parallelism in contemporary applications leads to a significant demand for remote access throughput, which mismatches with the state-of-the-art kernel swap path due to its inherently limited parallelism. As a result, it cannot scale up the multi-core applications. We dive into the implementation of the swap path and identify the root cause behind it - significant lock contentions and inefficient swap tasks offloading. Kefan Liu, Ke Liu 0004, Xu Zhang 0033, Ning Liu 0031, Sa Wang, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
EuroSys | 3 |
| 2025 | DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack Communication
Xu Zhang 0033, Ke Liu 0004, Yuan Hui 0001, Yisong Chang, Yizhou Shan, Ke Zhang 0017, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
USENIX ATC | 1 |
| 2024 | Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory AccessabstractThe growing memory demands of modern applications have driven the adoption of far memory technologies in data centers to provide cost-effective, high-capacity memory solutions. However, far memory presents new performance challenges because its access latencies are significantly longer and more variable than local DRAM. For applications to achieve acceptable performance on far memory, a high degree of memory-level parallelism (MLP) is needed to tolerate the long access latency. While modern out-of-order processors are capable of exploiting a certain degree of MLP, they are constrained by resource limitations and hardware complexity. The key obstacle is the synchronous memory access semantics of traditional load/store instructions, which occupy critical hardware resources for a long time. The longer far memory latencies exacerbate this limitation. This article proposes a set of Asynchronous Memory Access Instructions (AMI) and its supporting function unit, Asynchronous Memory Access Unit (AMU), inside contemporary Out-of-Order Core. AMI separates memory request issuing from response handling to reduce resource occupation. Additionally, AMU architecture supports up to several hundreds of asynchronous memory requests through re-purposing a portion of L2 Cache as scratchpad memory (SPM) to provide sufficient temporal storage. Together with a coroutine-based programming framework, this scheme can achieve significantly higher MLP for hiding far memory latencies. Evaluation with a cycle-accurate simulation shows AMI achieves 2.42× speedup on average for memory-bound benchmarks with 1μs additional far memory latency. Over 130 outstanding requests are supported with 26.86× speedup for GUPS (random access) with 5 μs latency. These demonstrate how the techniques tackle far memory performance impacts through explicit MLP expression and latency adaptation. Luming Wang, Xu Zhang 0033, Songyue Wang, Zhuolun Jiang, Tianyue Lu, Mingyu Chen 0001, Siwei Luo, Keji Huang |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Rethinking Design Paradigm of Graph Processing System with a CXL-like Memory Semantic FabricabstractWith the evolution of network fabrics, message-passing clusters have been promising solutions for large-scale graph processing. Alternatively, the shared-memory model is also introduced to avoid redundant copies and extra storage space of graph data. Compared to conventional network fabrics, with the capability of fine-grained, byte-addressable remote memory access, emerging memory semantic interconnects and fabrics, e.g., Intel's Compute Express Link (CXL), are intuitively more appropriate for adoption in shared-memory clusters. However, due to the latency gap between local and remote memory, it is still challenging to take advantage of the shared-memory graph processing with memory semantic fabrics. To tackle this problem, in this paper, we first investigate memory access characterizations of graph vertex propagation based on the shared-memory model. Then we propose GraCXL, a series of design paradigms to address high-frequency and long-latency of remote memory access potentially incurred in CXL-based clusters. For system adaptiveness, we elaborate GraCXL towards the general-purpose CPU cluster and the domain-specific FPGA accelerator array, respectively. We design a custom fabric with the CXL.mem protocol and leverage a couple of ARM SoC-equipped FPGAs to build an evaluation prototype in the absence of commodity CXL hardware and platforms. Experimental results show that the proposed GraCXL CPU and FPGA clusters achieve 1.33x-8.92x and 2.48x-5.01x performance improvement, respectively. Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Zhang 0017, Mingyu Chen 0001 |
CCGrid | 1 |
| 2023 | Morpheus: An Adaptive DRAM Cache with Online Granularity Adjustment for Disaggregated MemoryabstractDisaggregated memory introduces a cost-effective solution for improving the memory utilization rate of data centers, by sharing a distributed memory pool among several individual servers. However, latency penalty in the existing connection between a computing node and the memory pool introduces performance degradation due to frequent far memory accesses. Based on our observation, page caching in the local DRAM, despite its reductions in the number of far memory accesses, still faces severe data over-fetching problem.With a detailed analysis of far memory access traces collected via several representative real-world applications, we argue that exploiting the various page-specific preferences of caching granularity is the key point of solving the data over-fetching problem in the DRAM cache. Consequently, in this paper, we present that it is influential to enable 1) dynamic selection of caching granularity for each page to not only guarantee sufficient spatial localities compared to the conventional fine-grained cache lines but also avoid data over-fetching caused by the coarse-grained pages, as well as 2) adaptive adjustment of cache capacity during execution for each granularity to accommodate the varying proportion of pages with different granularity preferences. Specifically, we propose Morpheus, an adaptive DRAM cache architecture that determines an optimal page-specific caching granularity at run-time and dynamically adjusts capacity occupations of different caching granularities. Based on our modeling and evaluations within the DRAMSim3 simulator, Morpheus exhibits 1.17-1.34x performance speedup for a wide range of workloads against the state-of-the-art DRAM cache design. Xu Zhang 0033, Tianyue Lu, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001 |
ICCD | 1 |
| 2022 | GraFF: A Multi-FPGA System with Memory Semantic Fabric for Scalable Graph ProcessingabstractFPGA has been a promising solution for graph processing in many scenarios. With a rapid growth in graph size, the on/off-chip memory capacity of a single FPGA is insufficient to hold large-scale graphs. To tackle such problem, in this position paper, we introduce GraFF, a Graph processing system with multiple FPGAs interconnected via a custom memory semantic Fabric. In order to efficiently exploit system parallelism, we first split the traversal of graph data into a series of independent fine-grained flits that are concurrently delivered among FPGAs as sheer memory semantic transactions. Then we relax FPGAs' synchronization from strict barrier boundaries between adjacent supersteps to fully parallelize graph traversing and computing. We build a prototype of GraFF with four custom FPGA nodes. Preliminary evaluation result based on the Breadth First Search (BFS) algorithm shows that the peak performance of GraFF reaches up to 6.23 GTEPS. Moreover, GraFF exhibits linear scalability when the number of FPGAs rises from one to four. Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Liu 0004, Ke Zhang 0017, Mingyu Chen 0001 |
FPT | 1 |