Yuanchao Xu 0001

dblp:78/10373-1 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0003-4165-9138ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 6 first-author · 20 since 2021Software engineering, systems software and programming languages · 12 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CPU-Oblivious Offloading of Failure-Atomic Transactions for Disaggregated Memory
abstract
Memory disaggregation introduces new challenges for application reliability, as compute server or interconnection failures can interrupt execution and lead to data inconsistency in the memory server. This paper presents Fanmem, a novel failure-atomic transaction system designed specifically for disaggregated memory architectures. Fanmem ensures data consistency in the presence of failures, drawing inspiration from persistent memory transactions while tailored for memory disaggregation. The key innovations of Fanmem include an asynchronous transaction model and the integration of a processing unit within the switch, enabling the offloading of time-consuming log persistency operations to the switch processing unit and significantly reducing the overhead on the compute servers. Evaluation confirms the effectiveness of Fanmem on two representative memory-disaggregated architectures. Compared to the state-of-the-art persistent memory transaction system, Fanmem achieves an average performance improvement of 1.2X and 1.7X on the respective architectures.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001, Yan Solihin
ASPLOS (2)3
2026 PIPM: Partial and Incremental Page Migration for Multi-host CXL Disaggregated Shared Memory
abstract
The emerging Compute Express Link (CXL) interconnect supports multi-host cache-coherent disaggregated shared memory (CXL-DSM). However, existing page migration approaches, designed primarily for single-host systems, are inefficient in multi-host CXL-DSM scenarios. To address this, we propose Partial and Incremental Page Migration (PIPM), a hardware-based solution that transparently leverages host-side local memory. PIPM is co-designed with the CXL multi-host coherence protocol, enabling coherent access to data residing in local DRAM. To overcome limitations of existing migration methods, PIPM supports fine-grained data migration and integrates hardware-based monitoring and decision-making mechanisms to optimize data placement. Evaluation results demonstrate that PIPM delivers performance improvements of up to 2.54× (1.86× on average) over the default multi-host CXL-DSM configuration.
Gangqi Huang, Heiner Litz, Yuanchao Xu 0001
ASPLOS (2)3
2026 Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory Management
Xiangyue Huang 0001, Yanan Guo 0002, Yuanchao Xu 0001
ISCA3
2026 LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page Prefetcher
Xiangyue Huang 0001, Yanan Guo 0002, Yuanchao Xu 0001
ISCA3
2026 Fenc2: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment Encoding
Zhaoting Gong, Nuo Xu 0013, Yuanchao Xu 0001, Fan Yao 0001, Wujie Wen
ISCA4
2026 Recurrent Neural Networks Meet Context-Free Grammar: Two Birds with One Stone
abstract
This work addresses a key challenge in the effective adoption of Recurrent Neural Networks (RNNs) by reducing inference time and expanding the scope of a prediction. It introduces compressed learning, a novel approach that integrates Context-Free Grammar (CFG) and online tokenization into the training and inference of RNNs for streaming inputs. Through a hierarchical compression algorithm, it compresses an input sequence to a CFG and makes predictions based on the compressed sequence. Its algorithm design employs a set of techniques to overcome the issues from the myopic nature of online tokenization, the tension between inference accuracy and compression rate, and other complexities in sequence compression and prediction. Its effectiveness is theoretically analyzed and empirically validated on 16 real-world sequences, including program function calls, memory traces, and system logs. Empirical results demonstrate that compressed learning can successfully recognize and leverage repetitive patterns in input sequences, and effectively translate them into dramatic (1–1,762 \(\times\) ) inference speedups as well as much (1–7,830 \(\times\) ) expanded prediction scope, while keeping the inference accuracy satisfactory.
Hui Guan 0001, Umang Chaudhary, Yuanchao Xu 0001, Lin Ning 0001, Lijun Zhang 0005, Xipeng Shen
ACM Trans. Knowl. Discov. Data3
2025 ${MC}^{3}$: Memory Contention-Based Covert Channel Communication on Shared DRAM System-on-Chips
abstract
Shared memory system-on-chips (SM-SoCs) are ubiquitously employed by a wide range of computing platforms, including edge/IoT devices, autonomous systems, and smartphones. In SM-SoCs, system-wide shared memory enables a convenient and cost-effective mechanism for making data accessible across dozens of processing units (PUs), such as CPU cores and domain-specific accelerators. Due to the diverse computational characteristics of the PUs they embed, SM-SoCs often do not employ a shared last-level cache (LLC). Although covert channel attacks have been widely studied in shared memory systems, high-throughput communication has previously been feasible only by relying on an LLC or by possessing privileged or physical access to the shared memory subsystem. In this study, we introduce a new memory-contention-based covert communication attack,$\boldsymbol{MC}^{3}$, which specifically targets shared system memory in mobile SoCs. Unlike existing attacks, our approach achieves high-throughput communication without the need for an LLC or elevated access to the system. We explore the effectiveness of our methodology by demonstrating the tradeoff between the channel transmission rate and the robustness of the communication. We evaluate$\boldsymbol{MC}^{3}$on NVIDIA Orin AGX, NX, and Nano platforms and achieve transmission rates up to 6.4 Kbps with less than 1% error rate.
Ismet Dagli, James Crea, Soner Seçkiner, Yuanchao Xu 0001, Selçuk Köse, Mehmet Esat Belviranli
DATE4
2025 FluidFaaS: A Dynamic Pipelined Solution for Serverless Computing with Strong Isolation-based GPU Sharing
abstract
Prompted by the rise of artificial intelligence (AI) or machine learning (ML), more serverless workloads demand efficient GPU support. Recent years have witnessed a shift of interest from weak isolation-based methods, such as Multi-Process Service (MPS), to strong isolation-based methods, such as Multi-Instance GPU (MIG), for GPU support on serverless platforms, thanks to concerns about performance interference and security. The current MIG-based solution for serverless computing is, however, subject to severe GPU resource fragmentation and under-utilization. This paper identifies the reason as the gap between current MIG supports in serverless computing and the rigid constraints in MIG (re)configurations. It proposes FluidFaaS, a solution that enables flexible MIG management for serverless computing. Through a novel programming system support, FluidFaaS enables fine-grained resource assignment to the components within a serverless function, based on which, it equips the invokers with runtime support that constructs pipelines on MIGs on the fly for a serverless function. The innovations, along with a hotness-aware eviction-based time sharing of MIG slices, significantly reduce GPU resource fragmentation and enhance system throughput. Evaluations demonstrate that FluidFaaS outperforms the state-of-the-art solutions by 25%-75% in throughput while achieving up to 90% higher SLO hit rates in various workloads.
Xinning Hui, Yuanchao Xu 0001, Xipeng Shen
HPDC2
2025 Efficient Security Support for CXL Memory through Adaptive Incremental Offloaded (Re-)Encryption
abstract
Current DRAM technologies face critical scaling limitations, significantly impacting the expansion of memory bandwidth and capacity required by modern data-intensive applications.Compute eXpress Link (CXL) emerges as a promising technology to address these limitations, enabling efficient cache-coherent memory expansion through direct connections between processors and CXL memory devices.Despite its potential, broad adoption of CXL memory in public cloud computing introduces substantial security challenges.Trusted Execution Environments (TEEs), such as Intel SGX/TDX and AMD SEV, provide robust protection for data integrity and confidentiality in cloud environments, complemented by CXL Integrity and Data Encryption (CXL IDE), which employs XTS encryption and Galois/Counter Mode (GCM) for secure message transmission.However, this approach incurs significant performance overhead due to the latency of XTS encryption on memory-intensive workloads.To mitigate this, we propose Adaptive Incremental Offloaded (Re-)Encryption (AIORE), an adaptive security framework combining Counter (CTR) and XTS encryption.AIORE dynamically selects encryption schemes based on page access frequency, implements incremental and offloaded re-encryption strategies, and leverages memory node computation to reduce overheads.Evaluation with Gem5 across diverse benchmarks reveals that AIORE significantly reduces security overhead by 62.8% on average and maintains overhead within 3.7% relative to an insecure baseline.
Chuanhan Li, Jishen Zhao, Yuanchao Xu 0001
MICRO3
2025 Security and Performance Implications of GPU Cache Eviction Priority Hints
abstract
NVIDIA provides cache eviction priority hints such as evict_first and evict_last on recent GPUs.These hints allow users to specify the eviction priority that should be used for individual cache lines to improve cache utilization.However, NVIDIA does not disclose the microarchitectural details of these hints or cache eviction behaviors when using them, which makes their security and performance implications unclear.In this paper, we first reverse engineer the detailed comprehensive behaviors of these eviction priority hints.Then based on our findings, we analyze their impact on system security and performance.First, we found that these priority hints introduce new security problems.Specifically, we develop a new covert channel using the evict_first priority hint, which is more efficient than existing GPU covert channels.We also demonstrate a performance degradation attack using the evict_last priority hint, which is more stealthy compared to the known methods.Second, from the performance perspective, we show that marking a cache line as evict_last does not always keep it in the cache.In fact, if more than 12/16 (or 3/16, depending on the driver version) of the L2 cache size worth of data are marked as evict_last, cache thrashing can occur, which leads to performance degradation for real GPU workloads.
Qizhong Wang, Xiangyue Huang 0001, Yanan Guo 0002, Yuanchao Xu 0001
MICRO4
2024 Data Enclave: A Data-Centric Trusted Execution Environment
abstract
Trusted Execution Environments (TEEs) protect sensitive applications in the cloud with the minimal trust in the cloud provider. Existing TEEs with integrity protection however lack support for data management primitives, causing data sharing between enclaves either insecure or cumbersome. This paper proposes a new data abstraction for TEEs, data enclave. As a data-centric abstraction, data enclave is decoupled from an enclave's existence, is equipped with flexible secure permission controls, and crytographically isolated. It eliminates the hurdles for enclaves to cooperate efficiently, and at the same time, enables dynamic shrinking of the height of integrity tree for performance. This paper presents this new abstraction, its properties, and the architecture support. Experiments on synthetic benchmarks and three real-world applications all show that data enclave can help improve the efficiency of enclaves and inter-enclave cooperations significantly while enhancing the security protection.
Yuanchao Xu 0001, James Pangia, Chencheng Ye 0001, Yan Solihin, Xipeng Shen
HPCA1
2024 ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
abstract
Recent years have witnessed increasing interest in machine learning inferences on serverless computing for its auto-scaling and cost effective properties. Existing serverless computing, however, lacks effective job scheduling methods to handle the schedule space dramatically expanded by GPU sharing, task batching, and intertask relations. Prior solutions have dodged the issue by neglecting some important factors, leaving some large performance potential locked. This paper presents ESG, a new scheduling algorithm that directly addresses the difficulties. ESG treats sharable GPU as a first-order factor in scheduling. It employs an optimality-guided adaptive method by combining A*-search and a novel dual-blade pruning to dramatically prune the scheduling space without compromising the quality. It further introduces a novel method, dominator-based SLO distribution, to ensure the scalability of the scheduler. The results show that ESG can significantly improve the SLO hit rates (61%-80%) while saving 47%-187% costs over prior work.
Xinning Hui, Yuanchao Xu 0001, Zhishan Guo, Xipeng Shen
HPDC2
2024 Outback: Fast and Communication-efficient Index for Key-Value Store on Disaggregated Memory
abstract
Disaggregated memory systems achieve resource utilization efficiency and system scalability by distributing computation and memory resources into distinct pools of nodes. RDMA is an attractive solution to support high-throughput communication between different disaggregated resource pools. However, existing RDMA solutions face a dilemma: one-sided RDMA completely bypasses computation at memory nodes, but its communication takes multiple round trips; two-sided RDMA achieves one-round-trip communication but requires non-trivial computation for index lookups at memory nodes, which violates the principle of disaggregated memory. This work presents Outback, a novel indexing solution for key-value stores with a one-round-trip RDMA-based network that does not incur computation-heavy tasks at memory nodes. Outback is the first to utilize dynamic minimal perfect hashing and separates its index into two components: one memory-efficient and compute-heavy component at compute nodes and the other memory-heavy and compute-efficient component at memory nodes. We implement a prototype of Outback and evaluate its performance in a public cloud. The experimental results show that Outback achieves higher throughput than both the state-of-the-art one-sided RDMA and two-sided RDMA-based in-memory KVS by 1.06--5.03×, due to the unique strength of applying a separated perfect hashing index.
Yi Liu 0115, Minghao Xie, Shouqian Shi, Yuanchao Xu 0001, Heiner Litz, Chen Qian 0001
Proc. VLDB Endow.4
2023 SpecPMT: Speculative Logging for Resolving Crash Consistency Overhead of Persistent Memory
abstract
Crash consistency overhead is a long-standing barrier to the adoption of byte-addressable persistent memory in practice. Despite continuous progress, persistent transactions for crash consistency still incur a 5.6X slowdown, making persistent memory prohibitively costly in practical settings. This paper introduces speculative logging, a new method that forgoes most memory fences and reduces data persistence overhead by logging data values early. This technique enables a novel persistent transaction model, speculatively persistent memory transactions (SpecPMT). Our evaluation shows that SpecPMT reduces the execution time overheads of persistent transactions substantially to just 10%.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Yan Sha, Xiaofei Liao, Hai Jin 0001, Yan Solihin
ASPLOS (2)2
2023 Reconciling Selective Logging and Hardware Persistent Memory Transaction
abstract
Log creation, maintenance, and its persist ordering are known to be performance bottlenecks for durable transactions on persistent memory. Existing hardware persistent memory transactions overlook an important opportunity for improving performance: some persistent data is algorithmically redundant such that it can be recovered from other data, removing the need for logging such data. The paper presents an ISA extension that enables selective logging for hardware persistent memory transactions for the first time. The ISA extension features two novel components: fine-grain logging and lazy persistency. Fine-grain logging allows hardware to log updates on data in the granularity of words without lengthening the critical path of data accesses. Lazy persistency allows updated data to remain in the cache after the transaction commits. Together, the new hardware persistent memory transaction outperforms the state-of-the-art hardware counterpart by 1.8× on average.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Yan Sha, Xiaofei Liao, Hai Jin 0001, Yan Solihin
HPCA2
2022 Temporal Exposure Reduction Protection for Persistent Memory
abstract
The long-living nature and byte-addressability of persistent memory (PM) amplifies the importance of strong memory protections. This paper develops temporal exposure reduction protection (TERP) as a framework for enforcing memory safety. Aiming to minimize the time when a PM region is accessible, TERP offers a complementary dimension of memory protection. The paper gives a formal definition of TERP, explores the semantics space of TERP constructs, and the relations with security and composability in both sequential and parallel executions. It proposes programming system and architecture solutions for the key challenges for the adoption of TERP, which draws on novel supports in both compilers and hardware to efficiently meet the exposure time target. Experiments validate the efficacy of the proposed support of TERP, in both efficiency and exposure time minimization.
Yuanchao Xu 0001, Chencheng Ye 0001, Xipeng Shen, Yan Solihin
HPCA1
2022 FFCCD: fence-free crash-consistent concurrent defragmentation for persistent memory
abstract
Persistent Memory (PM) is increasingly supplementing or substituting DRAM as main memory. Prior work have focused on reusability and memory leaks of persistent memory but have not addressed a problem amplified by persistence, persistent memory fragmentation, which refers to the continuous worsening of fragmentation of persistent memory throughout its usage. This paper reveals the challenges and proposes the first systematic crash-consistent solution, Fence-Free Crash-consistent Concurrent Defragmentation (FFCCD). FFCCD resues persistent pointer format, root nodes and typed allocation provided by persistent memory programming model to enable concurrent defragmentation on PM. FFCCD introduces architecture support for concurrent defragmentation that enables a fence-free design and fast read barrier, reducing two major overheads of defragmenting persistent memory. The techniques is effective (28--73% fragmentation reduction) and fast (4.1% execution time overhead).
Yuanchao Xu 0001, Chencheng Ye 0001, Yan Solihin, Xipeng Shen
ISCA1
2022 Brief Industry Paper: Enabling Level-4 Autonomous Driving on a Single $1k Off-the-Shelf Card
abstract
In the past few years we have developed hardware computing systems for commercial autonomous vehicles, but inevitably the high development cost and long turn-around time have been major roadblocks for commercial deployment. Hence we also explored the potential of software optimization. This paper, for the first-time, shows that it is feasible to enable full leve1-4 autonomous driving workloads on a single off-the-shelf card (Jetson AGX Xavier) for less than ${\$}1\mathrm{k}$, an order of magnitude less than the state-of-the-art systems, while meeting all the requirements of latency. The success comes from the resolution of some important issues shared by existing practices through a series of measures and innovations.
Hsin-Hsuan Sung, Yuanchao Xu 0001, Jiexiong Guan, Wei Niu 0002, Bin Ren 0002, Yanzhi Wang 0001, Shaoshan Liu, Xipeng Shen
RTAS2
2022 Preserving Addressability Upon GC-Triggered Data Movements on Non-Volatile Memory
abstract
This article points out an important threat that application-level Garbage Collection (GC) creates to the use of non-volatile memory (NVM). Data movements incurred by GC may invalidate the pointers to objects on NVM and, hence, harm the reusability of persistent data across executions. The article proposes the concept of movement-oblivious addressing (MOA), and develops and compares three novel solutions to materialize the concept for solving the addressability problem. It evaluates the designs on five benchmarks and a real-world application. The results demonstrate the promise of the proposed solutions, especially hardware-supported Multi-Level GPointer, in addressing the problem in a space- and time-efficient manner.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Hai Jin 0001, Xiaofei Liao, Yan Solihin
ACM Trans. Archit. Code Optim.2
2021 Hardware-Based Address-Centric Acceleration of Key-Value Store
abstract
Efficiently retrieving data is essential for key-value store applications. A major part of the retrieving time is on data addressing, that is, finding the location of the value in memory that corresponds to a key. This paper introduces an address-centric approach to speed up the addressing by creating a shortcut for the translation of a key to the physical address of the value. The new technique is materialized with a novel in-memory table, STLT, a virtual-physical address buffer, and two new instructions. It creates a fast path for data addressing and meanwhile opens up opportunities for the use of simpler and faster hash tables to strike a better tradeoff between hashing conflicts and hashing overhead. Together, the new technique brings up to 1.4× speedups on key-value store application Redis and up to 13× speedups on some widely used indexing data structures, consistently outperforming prior solutions significantly.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Xiaofei Liao, Hai Jin 0001, Yan Solihin
HPCA2
2021 Recurrent Neural Networks Meet Context-Free Grammar: Two Birds with One Stone
abstract
Recurrent Neural Networks (RNN) are widely used for various prediction tasks on sequences such as text, speed signals, program traces, and system logs. Due to RNNs’ inherently sequential behavior, one key challenge for the effective adoption of RNNs is to reduce the time spent on RNN inference and to increase the scope of a prediction. This work introduces CFG-guided compressed learning, an approach that creatively integrates Context-Free Grammar (CFG) and online tokenization into RNN learning and inference for streaming inputs. Through a hierarchical compression algorithm, it compresses an input sequence to a CFG and makes predictions based on the compressed sequence. Its algorithm design employs a set of techniques to overcome the issues from the myopic nature of online tokenization, the tension between inference accuracy and compression rate, and other complexities. Experiments on 16 real-world sequences of various types validate that the proposed compressed learning can successfully recognize and leverage repetitive patterns in input sequences, and effectively translate them into dramatic (1-1762×) inference speedups as well as much (1-7830×) expanded prediction scope, while keeping the inference accuracy satisfactory.
Hui Guan 0001, Umana Chaudhary, Yuanchao Xu 0001, Lin Ning 0001, Lijun Zhang 0005, Xipeng Shen
ICDM3
2021 Supporting Legacy Libraries on Non-Volatile Memory: A User-Transparent Approach
abstract
As mainstream computing is poised to embrace the advent of byte-addressable non-volatile memory (NVM), an important roadblock has remained largely unnoticed, support of legacy libraries on NVM. Libraries underpin modern software everywhere. As current NVM programming interfaces all designate special types and constructs for NVM objects and references, legacy libraries, being incompatible with these data types, will face major obstacles for working with future applications written for NVM. This paper introduces a simple approach to mitigating the issue. The novel approach centers around user-transparent persistent reference, a new concept that allows programmers to reference a persistent object in the same way as reference a normal (volatile) object. The paper presents the implementation of the concept, carefully examines its soundness, and describes compiler and simple architecture support for keeping performance overheads very low.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Xiaofei Liao, Hai Jin 0001, Yan Solihin
ISCA2
2021 PCCS: Processor-Centric Contention-aware Slowdown Model for Heterogeneous System-on-Chips
abstract
Many slowdown models have been proposed to characterize memory interference of workloads co-running on heterogeneous System-on-Chips (SoCs). But they are mostly for post-silicon usage. How to effectively consider memory interference in the SoC design stage remains an open problem. This paper presents a new approach to this problem, consisting of a novel processor-centric slowdown modeling methodology and a new three-region interference-conscious slowdown model. The modeling process needs no measurement of co-running of various combinations of applications, but the produced slowdown models can be used to estimate the co-run slowdowns of arbitrary workloads on various SoC designs that embed a newer generation of accelerators, such as deep learning accelerators (DLA), in addition to CPUs and GPUs. The new method reduces average prediction errors of the state-of-art model from 30.3% to 8.7% on GPU, from 13.4% to 3.7% on CPU, from 20.6% to 5.6% on DLA and demonstrates much improved efficacy in guiding SoC designs.
Yuanchao Xu 0001, Mehmet Esat Belviranli, Xipeng Shen, Jeffrey S. Vetter
MICRO1
2021 UDF to SQL translation through compositional lazy inductive synthesis
abstract
Many data processing systems allow SQL queries that call user-defined functions (UDFs) written in conventional programming languages. While such SQL extensions provide convenience and flexibility to users, queries involving UDFs are not as efficient as their pure SQL counterparts that invoke SQL’s highly-optimized built-in functions. Motivated by this problem, we propose a new technique for translating SQL queries with UDFs to pure SQL expressions. Unlike prior work in this space, our method is not based on syntactic rewrite rules and can handle a much more general class of UDFs. At a high-level, our method is based on counterexample-guided inductive synthesis (CEGIS) but employs a novel compositional strategy that decomposes the synthesis task into simpler sub-problems. However, because there is no universal decomposition strategy that works for all UDFs, we propose a novel lazy inductive synthesis approach that generates a sequence of decompositions that correspond to increasingly harder inductive synthesis problems. Because most realistic UDF-to-SQL translation tasks are amenable to a fine-grained decomposition strategy, our lazy inductive synthesis method scales significantly better than traditional CEGIS. We have implemented our proposed technique in a tool called CLIS for optimizing Spark SQL programs containing Scala UDFs. To evaluate CLIS, we manually study 100 randomly selected UDFs and find that 63 of them can be expressed in pure SQL. Our evaluation on these 63 UDFs shows that CLIS can automatically synthesize equivalent SQL expressions in 92% of the cases and that it can solve 2.4× more benchmarks compared to a baseline that does not use our compositional approach. We also show that CLIS yields an average speed-up of 3.5× for individual UDFs and 1.3× to 3.1× in terms of end-to-end application performance.
Yuanchao Xu 0001, Xipeng Shen, Isil Dillig
Proc. ACM Program. Lang.2
2020 MERR: Improving Security of Persistent Memory Objects via Efficient Memory Exposure Reduction and Randomization
abstract
This paper proposes a new defensive technique for memory, especially useful for long-living objects on Non-Volatile Memory (NVM), or called Persistent Memory objects (PMOs). The method takes a distinctive perspective, trying to reduce memory exposure time by largely shortening the overhead in attaching and detaching PMOs into the memory space. It does it through a novel idea, embedding page table subtrees inside PMOs. The paper discusses the complexities the technique brings, to permission controls and hardware implementations, and provides solutions. Experimental results show that the new technique reduces memory exposure time by 60% with a 5% time overhead (70% with 10.9% overhead). It allows much more frequent address randomizations (shortening the period from seconds to less than 41.4us), offering significant potential for enhancing memory security.
Yuanchao Xu 0001, Yan Solihin, Xipeng Shen
ASPLOS1
2020 Hardware-Based Domain Virtualization for Intra-Process Isolation of Persistent Memory Objects
abstract
Persistent memory has appealing properties in serving as main memory. While file access is protected by system calls, an attached persistent memory object (PMO) is one load/store away from accidental (or malicious) reads or writes, which may arise from use of just one buggy library. The recent progress in intra-process isolation could potentially protect PMO by enabling a process to partition sensitive data and code into isolated components. However, the existing intra-process isolations (e.g., Intel MPK) support isolation of only up to 16 domains, forming a major barrier for PMO protections. Although there is some recent effort trying to virtualize MPK to circumvent the limit, it suffers large overhead. This paper presents two novel architecture supports, which provide 11 - 52 × higher efficiency while offering the first known domain-based protection for PMOs.
Yuanchao Xu 0001, Chencheng Ye 0001, Yan Solihin, Xipeng Shen
ISCA1
2018 Taming the "Monster": Overcoming Program Optimization Challenges on SW26010 Through Precise Performance Modeling
abstract
This paper presents an effort for overcoming the complexities of program optimizations on SW26010, the heterogeneous many-core processor that powers Sunway TaihuLight, the world top one supercomputer. The solution centers around a precise, static performance model for modern many-core processor. Through a careful design that leverages the special properties of SW26010 and an effective treatment to massive parallelism, the model achieves a high accuracy, showing less than 5% average errors in estimating program execution performance. The precise performance model opens many opportunities for analyzing and guiding code optimizations. The paper demonstrates the usefulness by revealing a series of insights on the effects of some important code optimizations on SW26010. Moreover, it demonstrates that with such a precise performance model, it is feasible to replace empirical auto-tuning with static auto-tuning for optimizing regular loops on heterogeneous many-core systems. Such a replacement speeds up the tuning process by as much as a factor of 43 while keeping the tuning quality loss below 6%.
Shizhen Xu, Yuanchao Xu 0001, Wei Xue 0003, Xipeng Shen, Fang Zheng 0015, Xiaomeng Huang, Guangwen Yang 0002
IPDPS2
2016 Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer
abstract
This paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day.
Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002
SC14