VLDB 2026 Research / reviewers in the wild / expert
Shoumeng Yan
dblp:08/6611
· DBLP profile ↗
45ranked-venue papers
1as first author
38since 2021 · last 2026
0009-0007-9580-5395ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 1 first-author · 23 since 2021Security and privacy · 12 · 12 since 2021Software engineering, systems software and programming languages · 10 · 7 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Framework for Developing and Optimizing Fully Homomorphic Encryption Programs on GPUsabstractIn sensitive domains such as healthcare and finance, machine learning increasingly employs Fully Homomorphic Encryption (FHE) to secure both user data and models. Although FHE's intrinsic parallelism naturally aligns with GPU architectures, optimizing GPU kernels alone remains insufficient for efficient end-to-end FHE application development. The inherent complexity of FHE schemes and intricate GPU-specific details impede developers from focusing on high-level program logic. Additionally, FHE's high memory requirements, fine-grained memory operations, and redundant computations introduce further optimization challenges, resulting in inefficiencies even when GPU kernels are individually optimized. This paper introduces EasyFHE, a framework designed to simplify the development and optimization of GPU-accelerated FHE applications. Similar to PyTorch, EasyFHE provides high-level interfaces for defining computational logic while automatically handling low-level tasks, such as implementation selection and memory management. Furthermore, it incorporates an optimization framework that systematically addresses performance bottlenecks by applying tailored optimization passes during the lowering from high-level FHE programs to GPU kernels. Compared to state-of-the-art open-source GPU FHE libraries, EasyFHE uniquely supports FHE programs with memory requirements exceeding typical GPU capacities, achieving an average speedup of 2.88× with a peak of 4.39×. Jianyu Zhao 0004, Xueyu Wu 0001, Guang Fan 0001, Mingzhe Zhang 0005, Shoumeng Yan, Lei Ju 0001, Zhuoran Ji |
ASPLOS (2) | 5 |
| 2026 | Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption AcceleratorabstractFully homomorphic encryption (FHE) enables computation on encrypted data without compromising privacy, positioning it as a promising solution for secure cloud computing. However, its substantial computational overhead impedes practical deployment, prompting the development of dedicated hardware accelerators. In practice, when deploying cryptographic algorithm optimizations on FHE accelerators, hardware constraints typically such as limited memory capacity, often lead to a disparity between theoretical algorithmic advantage and achievable hardware efficiency. Liang Kong 0005, Xianglong Deng, Guang Fan 0001, Shengyu Fan, Yilan Zhu, Geng Yang 0001, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005 |
ASPLOS (2) | 9 |
| 2026 | Pyramid: A Secure, Resource-Efficient, and Pluggable Kubernetes for Multi-TenancyabstractThis work aims to achieve the best of both worlds with two prominent techniques adopted in cloud computing systems: hardware trusted execution environments (TEEs) for data processing security, and Kubernetes (k8s) for efficient container orchestration and resource management for multi-tenancy. A secure, resource-efficient, and pluggable container orchestration system, called Pyramid, is proposed, which incurs minimal intrusive modifications to the commercial k8s. Pyramid puts a separate trusted k8s on top of the original k8s cluster and carefully cooperates between the two layers. The workflow within each layer is maximally preserved without significant changes. The untrusted layer manages resource scheduling across different tenants to improve utilization and passes the resource information to the trusted layer to launch actual computations secured by TEEs, with the help of carefully designed interface and protection mechanisms. Evaluation results show that Pyramid achieves 1.4X higher throughput on the data plane, with comparable control-plane performance to previous work. Xiang Li 0156, Weijie Liu 0004, Fabing Li, Hongliang Tian, Zheli Liu, Shoumeng Yan, Mingyu Gao 0001 |
EuroSys | 6 |
| 2026 | MlsDisk: Trusted Block Storage for TEEs Based on Layered Secure Logging
Erci Xu, Lujia Yin, Xinyuan Luo, Shaowei Song, Qingsong Chen, Shoumeng Yan, Jiwu Shu, Hongliang Tian, Yiming Zhang 0003 |
FAST | 7 |
| 2026 | An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design AutomationabstractFully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs. Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005 |
HPCA | 13 |
| 2026 | HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Geng Yang 0001, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng, Meng Li 0004, Yisong Chang, Shoumeng Yan, Mingzhe Zhang 0005 |
ISCA | 15 |
| 2026 | SoK: Analysis of Accelerator TEE Designs
Chenxu Wang 0005, Yujun Liang, Xuanyao Peng, Yuqun Zhang, Fengwei Zhang, Jiannong Cao 0001, Rui Hou 0001, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 10 |
| 2026 | DCS3: A Dual-Layer Co-Aware Scheduler With Stealing Balance and Synchronized Priority in Virtualization EnvironmentsabstractVirtualization environments (e.g., containers and hypervisors) achieve isolation of multiple runtime entities but result in two mutually isolated guest and host layers. Such cross-ayer isolation could cause high latency and low throughput of the system. Previous aware scheduling and double scheduling fail to achieve bidirectional coordination between the guest and host layers. To address this challenge, we develop DCS3, a Dual-layer Co-aware Scheduler that combines stealing balance and synchronized priority. Stealing balancing migrates tasks between virtual CPU (vCPU) queues for load balance based on the workloads of physical CPUs (pCPUs). Synchronized priority dynamically adjusts the thread priorities running on the pCPUs according to the current vCPU workloads. The vCPUs and pC-PUs belong to the guest and host layers, respectively. Compared with aware scheduling, double scheduling, and DCS2 (i.e., DCS3 without synchronized priority), DCS3 has the following obvious advantages: 1) Requests Per Second (RPS) increases by up to 52%, 55%, and 2%, respectively; 2) request latency decreases by up to 72%, 71%, and 20%, respectively. Chenglai Xiong, Guoqi Xie, Zhongjia Wang, Zhenli He, Shaowen Yao 0001, Jianfeng Tan, Tiwei Bie, Shoumeng Yan |
IEEE Trans. Computers | 11 |
| 2026 | Keystone-Vault: Hardware-Assisted Efficient Intra-Enclave Isolation for RISC-V TEEs
Tianming Yan, Kun Yang 0012, Hongliang Tian, Shoumeng Yan, Kui Ren 0001 |
IEEE Trans. Computers | 4 |
| 2026 | Exploration of Karatsuba Algorithm for Efficient Barrett Modular Multiplication
Bo Zhang 0098, Mingzhe Zhang 0005, Shoumeng Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | Building Confidential Accelerator Computing Environment for Arm CCA
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2026 | Complementing Confidential Computing Environment for Applications on Arm CCA
Yiming Zhang 0030, Zhenyu Ning, Fengwei Zhang, Xiapu Luo, Haoyang Huang, Shoumeng Yan, Zhengyu He |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | CROSS-TEE: A Distributed Trusted Execution Environment Architecture for Cross-Module Automotive Security
Kun Yang 0012, Hongliang Tian, Shoumeng Yan, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | The Future of Fully Homomorphic Encryption System: From a Storage I/O Perspective
Erci Xu, Shengyu Fan, Xianglong Deng, Guiming Shi, Guang Fan 0001, Liang Kong 0005, Yilan Zhu, Shoumeng Yan, Mingzhe Zhang 0005 |
APPT | 10 |
| 2025 | ALLMod: Exploring Area-Efficiency of LUT-based Large Number Modular Reduction via Hybrid WorkloadsabstractModular arithmetic, particularly modular reduction, is widely used in cryptographic applications such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). High-bit-width operations are crucial for enhancing security; however, they are computationally intensive due to the large number of modular operations required. The lookup-table-based (LUT-based) approach, a “space-for-time” technique, reduces computational load by segmenting the input number into smaller bit groups, pre-computing modular reduction results for each segment, and storing these results in LUTs. While effective, this method incurs significant hardware overhead due to extensive LUT usage. In this paper, we introduce ALLMod, a novel approach that improves the area efficiency of LUT-based largenumber modular reduction by employing hybrid workloads. Inspired by the iterative method, ALLMod splits the bit groups into two distinct workloads, achieving lower area costs without compromising throughput. We first develop a template to facilitate workload splitting and ensure balanced distribution. Then, we conduct design space exploration to evaluate the optimal timing for fusing workload results, enabling us to identify the most efficient design under specific constraints. Extensive evaluations show that ALLMod achieves up to $\lt sup\gt1\lt/sup\gt|.65 \times$ and $3 \times$ improvements in area efficiency over conventional LUT-based methods for bit-widths of 128 and 8,192, respectively. Fangxin Liu, Haomin Li 0002, Zongwu Wang, Bo Zhang 0098, Mingzhe Zhang 0005, Shoumeng Yan, Li Jiang 0002, Haibing Guan |
DAC | 6 |
| 2025 | AtomicDisk: A Secure Virtual Disk for TEEs against Eviction Attacks
Hongliang Tian, Shaowei Song, Qingsong Chen, Weijie Liu 0004, Erci Xu, Shoumeng Yan, Yiming Zhang 0003 |
FAST | 9 |
| 2025 | WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresabstractThe application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by $\mathbf{7 3 \%}$ and pipeline stalls by $\mathbf{8 6 \%}$ compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to $2.12 \times$ improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of $13.4 \times$ and $3.5 \times$, respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves $2.8 \times$ the performance of TensorFHE. Guang Fan 0001, Mingzhe Zhang 0005, Fangyu Zheng, Shengyu Fan, Xianglong Deng, Wenxu Tang, Liang Kong 0005, Shoumeng Yan |
HPCA | 10 |
| 2025 | ccAI: A Compatible and Confidential System for AI ComputingabstractConfidential xPU computing has emerged as a prominent technique for effectively securing users' AI computing workloads on heterogeneous systems equipped with xPUs.Although the industry adopts this technology in cutting-edge hardware (e.g.NVIDIA H100 GPU) to safeguard high-performance AI computing, most clouds still rely on legacy xPUs and suffer from data leakage problems. Chenxu Wang 0005, Danqing Tang, Changxu Ci, Yankai Xu, Fengwei Zhang, Jiannong Cao 0001, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
MICRO | 9 |
| 2025 | HAWK: Fully Homomorphic Encryption Accelerator with Fixed-Word Key Decomposition Switching
Liang Kong 0005, Shengyu Fan, Xianglong Deng, Guang Fan 0001, Guiming Shi, Yilan Zhu, Geng Yang 0001, Shoumeng Yan, Mingzhe Zhang 0005 |
MICRO | 9 |
| 2025 | The Road to Trust: Building Enclaves within Confidential VMs
Wenhao Wang 0001, Linke Song, Benshan Mei, Shijun Zhao, Shoumeng Yan, XiaoFeng Wang 0001, Dan Meng 0002, Rui Hou 0001 |
NDSS | 6 |
| 2025 | SCRUTINIZER: Towards Secure Forensics on Compromised TrustZone
Yiming Zhang 0030, Fengwei Zhang, Xiapu Luo, Rui Hou 0001, Xuhua Ding, Zhenkai Liang, Shoumeng Yan, Tao Wei 0002, Zhengyu He |
NDSS | 7 |
| 2025 | CortenMM: Efficient Memory Management with Strong Correctness GuaranteesabstractModern memory management systems suffer from poor performance and subtle concurrency bugs, slowing down applications while introducing security vulnerabilities. We observe that both issues stem from the conventional design of memory management systems with two levels of abstraction: a software-level abstraction (e.g., VMA trees in Linux) and a hardware-level abstraction (typically, page tables). This design increases portability but requires correctly and efficiently synchronizing two drastically different and complex data structures, which is generally challenging. Junyang Zhang 0003, Xiangcan Xu, Yonghao Zou, Xinyi Wan 0001, Siyuan Wang 0026, Di Wang 0017, Hao Chen 0023, Lin Huang 0005, Shoumeng Yan, Yuval Tamir, Yingwei Luo, Xiaolin Wang 0001, Huashan Yu, Zhenlin Wang 0003, Hongliang Tian, Diyu Zhou |
SOSP | 12 |
| 2025 | ASTERINAS: A Linux ABI-Compatible, Rust-Based Framekernel OS with a Small and Sound TCB
Yuke Peng, Hongliang Tian, Junyang Zhang 0003, Jinyi Xian, Xiaolin Wang 0001, Chenren Xu, Diyu Zhou, Yingwei Luo, Shoumeng Yan, Yinqian Zhang |
USENIX ATC | 12 |
| 2025 | DAHE: Parameter-Adaptive and Memory Efficient FPGA Acceleration of Homomorphic EncryptionabstractWhile homomorphic encryption (HE) has been well-recognized as a promising data privacy protection technique, there are many challenges to the real-world deployment of HE applications. In this work, we propose a design flow for parameter-adaptive and memory-efficient FPGA acceleration of homomorphic encryption. In the framework, we explore the correlations between HE parameter selection to meet various design objectives and the huge design space due to underlying FPGA hardware resource allocation. Particularly, we demonstrate that adaptive management of the FPGA memory hierarchy is crucial to supporting diverse cryptosystem parameter selection for application-level security, accuracy, and performance requirements. We propose a resource-efficient and flexible micro-architectural design for HE operations, where data access patterns in various pipeline execution stages are optimized for high memory bandwidth utilization. Furthermore, a memory-aware performance model is built for automatic design space exploration for cryptosystem parameter selection and hardware resource provisioning. Experimental results show 1.50X and 1.16X speedup for the NTT and Rotation operations w.r.t. the state-of-the-art FPGA implementation. Meanwhile, the proposed framework generates flexible and high-performance accelerator code for real HE application kernels with different cryptosystem parameters on a wide range of FPGA devices. Yilan Zhu, Honghui You, Wei Zhang 0173, Jiming Xu, Qian Lou, Shoumeng Yan, Lei Ju 0001 |
IEEE Trans. Computers | 6 |
| 2025 | GIF-FHE: A Comprehensive Implementation and Evaluation of GPU-Accelerated FHE With Integer and Floating-Point Computing PowerabstractFully Homomorphic Encryption (FHE) allows computations on encrypted data without revealing the plaintext, garnering significant interest from both academic and industrial communities. However, its broader adoption has been hindered by performance limitations. Consequently, researchers have turned to GPUs for efficient FHE implementation. Nevertheless, most have predominantly favored integer units due to their ease of use, overlooking the considerable computational potential of floating-point units in GPUs. Recognizing this untapped floating-point computational power, our paper introducesGIF-FHE, an extensive exploration and implementation of FHE, leveraging GPUs' integer and floating-point instructions for FHE acceleration. We develop a comprehensive suite of low-level and middle-level FHE primitives, offering multiple implementation variants with support for three word size configurations ($64/52/32$-bit). Particularly, we make innovative use of floating-point implementations, employing a novel methodology to efficiently leverage the floating-point unit's fused multiply-add (FMA) instructions. This represents the pioneering integration of floating-point units into FHE acceleration. To bridge our highly-optimized FHE primitives with practical applications, this paper also provides a high-level FHE implementation and interfaces that can be directly applied by upper-level applications such as neural network inference. Finally, we undertake a comprehensive experiment evaluation and comparison involving three types of arithmetic: FP64/INT64/INT32 with varying word size configurations and computation units. Notably, our fundamental function implementations consistently outperform counterparts on the same platform, achieving speedups ranging from$2.0\times$to$4.2\times$. In the context of CKKS FHE schemes, our homomorphic operation implementation surpasses the state-of-the-art GPU-based solution with a speedup of up to$3.8\times$, and exceeds the performance of the widely adopted CPU-based library, SEAL, with a remarkable speedup of over$300\times$. Fangyu Zheng, Guang Fan 0001, Wenxu Tang, Yuan Zhao 0015, Jiankuo Dong, Jingqiang Lin 0001, Shoumeng Yan, Jiwu Jing |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2024 | Verifying Rust Implementation of Page Tables in a Software Enclave HypervisorabstractAs trusted execution environments (TEE) have become the corner stone for secure cloud computing, it is critical that they are reliable and enforce proper isolation, of which a key ingredient is spatial isolation. Many TEEs are implemented in software such as hypervisors for flexibility, and in a memory-safe language, namely Rust to alleviate potential memory bugs. Still, even if memory bugs are absent from the TEE, it may contain semantic errors such as mis-configurations in its memory subsystem which breaks spatial isolation. Zhenyang Dai, Vilhelm Sjöberg, Xupeng Li, Yu Chen 0004, Wenhao Wang 0001, Yuekai Jia, Sean Noble Anderson, Laila Elbeheiry, Shubham Sondhi, Yu Zhang 0313, Zhaozhong Ni, Shoumeng Yan, Ronghui Gu, Zhengyu He |
ASPLOS (2) | 13 |
| 2024 | HMNTT: A Highly Efficient MDC-NTT Architecture for Privacy-preserving ApplicationsabstractIn privacy-preserving applications like Post-Quantum Cryptography (PQC) and Fully Homomorphic Encryption (FHE), polynomial multiplication is common, and the Number Theoretic Transform (NTT) is a key algorithm for reducing its complexity. In this paper, we present HMNTT, a highly efficient MDC-NTT architecture. Utilizing the four-step NTT algorithm and a pipelined transpose module, HMNTT offers a highly efficient and scalable architecture for handling NTT with large degrees. We optimize the processing element (PE) to alleviate backpressure and data conflicts in data flow. Leveraging FPGA characteristics, we construct a modular multiplication module to reduce resource usage and improve operating frequency. Evaluation results indicate that HMNTT achieves an average of 2.34 × and 1.26 × reduction in Area-Time Product compared to the latest pipelined NTT architectures. Changxu Liu, Danqing Tang, Hao Zhou 0015, Shoumeng Yan, Fan Yang 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | A Compiler-Like Framework for Optimizing Cryptographic Big Integer Multiplication on GPUsabstractWith the growth of digital data and rising security concerns, techniques for privacy-preserving computation have become increasingly essential. Big integer multiplication, pivotal for these applications, is compute-intensive but poses challenges for GPU acceleration due to its complexity and the need for application-specific tailored implementations. This paper presents IMCompiler, a compiler-like framework that automatically gen-erates optimized GPU kernels for integer multiplications used in cryptosystems. It features a frontend-IR-backend structure, where the Intermediate Representation (IR) employs a segmented integer multiplication algorithm to decouple architecture-specific optimizations from high-level parameters. The frontend can then easily translate integer multiplication with various high-level parameters into the IR, while the backend focuses on fine-tuning a single GPU kernel for each device, enabling automatic code generation. Moreover, we introduce a computation diagram to facilitate the analysis of parallelization strategies, inspiring many optimizations, including two-dimensional parallelization, tailored caching strategy, index transposing, and lazy carrying. Experiments show that IMCompiler achieves a 4.47× speedup compared to the widely used baseline and 1.42 × over Nvidia's official library. The speedup will be even higher for larger integers and higher-capacity GPUs. Zhuoran Ji, Jianyu Zhao 0004, Jiming Xu, Shoumeng Yan, Lei Ju 0001 |
MICRO | 5 |
| 2024 | CAGE: Complementing Arm CCA with GPU Extensions
Chenxu Wang 0005, Fengwei Zhang, Yunjie Deng 0001, Kevin Leach, Jiannong Cao 0001, Zhenyu Ning, Shoumeng Yan, Zhengyu He |
NDSS | 7 |
| 2024 | Area-Efficient Barrett Modular Multiplication With Optimized Karatsuba AlgorithmabstractThis article presents an area-efficient Barrett modular multiplication (BMM) algorithm, facilitating the development of cryptosystems like fully homomorphic encryption. Instead of implementing three normal multiplications required by classic BMM, our proposed BMM introduces optimizations for multiplication AB, truncated multiplication$\lfloor AB/2^{f} \rfloor $, and modular multiplication (MM)$AB ~\text {mod}~2^{f}$. Taking the 4-term Karatsuba algorithm as an example, an N-bit multiplication AB can be decomposed into$9~(N/4)$-bit multiplications. Our optimized approaches for truncated multiplication and MM require an area equivalent to only$6.5~(N/4)$-bit multiplications when$f\approx N$. Furthermore, our optimized Karatsuba multiplications introduce efficient (E, I) matrix pairs, circumventing area overhead from complex I matrices and sign extension in multiplication. We also employ encode algorithm to eliminate many additions needed in BMM and inside multiplications, significantly shortening critical path. Experimental results demonstrate the advantages of our proposed BMM in terms of throughput and area efficiency. Bo Zhang 0098, Shoumeng Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Building a Lightweight Trusted Execution Environment for Arm GPUsabstractA wide range of Arm endpoints leverage integrated and discrete GPUs to accelerate computation. However, Arm GPU security has not been explored by the community. Existing work has used Trusted Execution Environments (TEEs) to address GPU security concerns on Intel-based platforms, but there are numerous architectural differences that lead to novel technical challenges in deploying TEEs for Arm GPUs. There is a need for generalizable and efficient Arm-based GPU security mechanisms. To address these problems, we presentStrongBox, the first GPU TEE for secured general computation on Arm endpoints.StrongBoxprovides an isolated execution environment by ensuring exclusive access to GPU. Our approach is based in part on a dynamic, fine-grained memory protection policy as Arm-based GPUs typically share a unified memory with the CPU. Furthermore,StrongBoxreduces runtime overhead from the redundant security introspection operations. We also design an effective defense mechanism withinsecure worldto protect the confidential GPU computation. Our design leverages the widely-deployed Arm TrustZone and generic Arm features, without hardware modification or architectural changes. We prototypeStrongBoxusing an off-the-shelf Arm Mali GPU and perform an extensive evaluation. Results show thatStrongBoxsuccessfully ensures GPU computation security with a low (4.70%–15.26%) overhead. Chenxu Wang 0005, Yunjie Deng 0001, Zhenyu Ning, Kevin Leach, Jin Li 0002, Shoumeng Yan, Zhengyu He, Jiannong Cao 0001, Fengwei Zhang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Lost along the Way: Understanding and Mitigating Path-Misresolution Threats to Container IsolationabstractFilesystem isolation enforced by today's container technology has been found to be less effective in the presence of host-container interactions increasingly utilized by container tools. This weakened isolation has led to a type of path misresolution (Pamir) vulnerabilities, which have been considered to be highly risky and continuously reported over the years. In this paper, we present the first systematic study on the Pamir risk and the existing fixes to related vulnerabilities. Our research reveals that in spite of significant efforts being made to patch vulnerable container tools and address the risk, the Pamir vulnerabilities continue to be discovered, including a new vulnerability (CVE-2023-0778) we rediscovered from patched software. A key insight of our study is that the Pamir risk is inherently hard to prevent at the level of container tools, due to their heavy reliance on third-party components. While security inspections should be applied to all components to mediate host-container interactions, third-party component developers tend to believe that container tools should perform security checks before invoking their components, and are therefore reluctant to patch their code with the container-specific protection. Moreover, due to the large number of components today's container tools depend on, re-implementing all of them is impractical. Zhi Li 0048, Weijie Liu 0004, XiaoFeng Wang 0001, Bin Yuan 0002, Hongliang Tian, Hai Jin 0001, Shoumeng Yan |
CCS | 7 |
| 2023 | Mind Your Enclave Pointers! Detecting Privacy Leaks for SGX Apps via Sparse Taint AnalysisabstractIntel SGX is a promising TEE technique that can protect programs running in user space from being maliciously accessed by the host operating system. Although it provides hardware access control and memory encryption, improper implementation of a code snippet running inside an enclave can still leak private data. While existing research mainly detects bugs of enclave code, this paper serves as a first attempt to study the privacy leakage issues caused by pointer misuse. In particular, we focus on explicit pointer declarations that may incur data copy from trusted to untrusted spaces, and we summarize five common patterns of such leakage code. Further, we propose a novel approach to detect these patterns based on static sparse taint analysis. Our approach starts from suspicious pointers of predefined patterns and performs forward analysis to recognize all taint sinks. It then backward analyzes the values being leaked through these sinks and checks if their identifiers are sensitive. We have implemented a prototype, namely STELLA, and conducted real-world experiments with dozens of open-source SGX programs. Results show that STELLA found 80 new leakage issues previously unknown in 13 projects. We hope our work can remind SGX developers of pointer misuse issues and inspire better designs toward mitigating the problem. Shoumeng Yan, Hui Xu 0009 |
ISSRE | 3 |
| 2023 | OOM-Guard: Towards Improving the Ergonomics of Rust OOM Handling via a Reservation-Based ApproachabstractOut of memory (OOM) is an exceptional system state where any further memory allocation requests may fail. Such allocation failures would crash the process or system if not handled properly, and they may also lead to an inconsistent program state that cannot be recovered easily. Current mechanisms for preventing such hazards highly rely on the manual effort of the programmers themselves. This paper studies the OOM issues of Rust, which is an emerging system programming language that stresses the importance of memory safety but still lacks handy mechanisms to handle OOM well. Even worse, Rust employs an infallible mode of memory allocations by default. As a result, the program written by Rust would simply abort itself when OOM occurs. Such crashes would lead to critical robustness issues for services or modules of operating systems. We propose OOM-Guard, a handy approach for Rust programmers to handle OOM. OOM-Guard is by nature a reservation-based approach that aims to convert the handlings for many possible failed memory allocations into handlings for a smaller number of reservations. In order to achieve efficient reservation, OOM-Guard incorporates a subtle cost analysis algorithm based on static analysis and a proxy allocator. We then apply OOM-Guard to two well-known Rust projects, Bento and rCore. Results show that OOM-Guard can largely reduce developers' efforts for handling OOM and incurs trivial overhead in both memory space and execution time. Hongliang Tian, Shoumeng Yan, Hui Xu 0009 |
ESEC/SIGSOFT FSE | 4 |
| 2023 | CipherH: Automated Detection of Ciphertext Side-channel Vulnerabilities in Cryptographic Implementations
Mengyuan Li 0004, Yining Tang, Shuai Wang 0011, Shoumeng Yan, Yinqian Zhang |
USENIX Security Symposium | 5 |
| 2023 | SHELTER: Extending Arm CCA with Isolation in User Space
Yiming Zhang 0030, Zhenyu Ning, Fengwei Zhang, Xiapu Luo, Haoyang Huang, Shoumeng Yan, Zhengyu He |
USENIX Security Symposium | 7 |
| 2022 | StrongBox: A GPU TEE on Arm EndpointsabstractA wide range of Arm endpoints leverage integrated and discrete GPUs to accelerate computation such as image processing and numerical processing applications. However, in spite of these important use cases, Arm GPU security has yet to be scrutinized by the community. By exploiting vulnerabilities in the kernel, attackers can directly access sensitive data used during GPU computing, such as personally-identifiable image data in computer vision tasks. Existing work has used Trusted Execution Environments (TEEs) to address GPU security concerns on Intel-based platforms, while there are numerous architectural differences that lead to novel technical challenges in deploying TEEs for Arm GPUs. In addition, extant Arm-based GPU defenses are intended for secure machine learning, and lack generality. There is a need for generalizable and efficient Arm-based GPU security mechanisms. Yunjie Deng 0001, Chenxu Wang 0005, Shunchang Yu, Shiqing Liu, Zhenyu Ning, Kevin Leach, Jin Li 0002, Shoumeng Yan, Zhengyu He, Jiannong Cao 0001, Fengwei Zhang |
CCS | 8 |
| 2022 | HyperEnclave: An Open and Cross-platform Trusted Execution Environment
Yuekai Jia, Wenhao Wang 0001, Yu Chen 0004, Zhengde Zhai, Shoumeng Yan, Zhengyu He |
USENIX ATC | 6 |
| 2020 | Occlum: Secure and Efficient Multitasking Inside a Single Enclave of Intel SGXabstractIntel Software Guard Extensions (SGX) enables user-level code to create private memory regions called enclaves, whose code and data are protected by the CPU from software and hardware attacks outside the enclaves. Recent work introduces library operating systems (LibOSes) to SGX so that legacy applications can run inside enclaves with few or even no modifications. As virtually any non-trivial application demands multiple processes, it is essential for LibOSes to support multitasking. However, none of the existing SGX LibOSes support multitasking both securely and efficiently. Youren Shen, Hongliang Tian, Yu Chen 0004, Kang Chen 0001, Runji Wang, Yubin Xia, Shoumeng Yan |
ASPLOS | 8 |
| 2020 | DroidCloud: Scalable High Density AndroidTM Cloud RenderingabstractCloud rendering is an emerging technology in which rendering-heavy applications run on the cloud server and then stream the rendered contents to the end-user device. High density and high scalability of the cloud rendering services are crucial to support millions of users concurrently and cost-effectively. However, it is still challenging to run Android OS in cloud smoothly with high density and high scalability without compromising user experience. This paper presents DroidCloud, the first open-source Android\footnoteAndroid is a trademark of Google LLC. cloud rendering solution focusing on the scalable design and density aspect optimization to the best of our knowledge. To cloudify Android OS, DroidCloud utilizes thevHAL technology in order to support remote devices and keep transparent to Android applications. And aFlexible rendering scheduling policy is introduced to break the boundary of GPU physical locations. Thus, both remote GPUs and local GPUs can accommodate render tasks by forwarding rendering tasks and making it possible to support multiple Android OSes with GPU acceleration. Besides, to further improve the density, DroidCloud optimizes the resource cost both in a single instance and across instances. We show that DroidCloud can run hundreds of Android OSes on a single Intel Xeon server with GPU acceleration simultaneously, increasing the density at the scale of one order of magnitude compared to current cloud gaming systems. Further experimental results demonstrate that DroidCloud can transparently run Android applications at native speed with lower CPU, memory, and storage utilization. Linsheng Li, Cathy Bao, Randy Xu, Mohammad R. Haghighat, Jerry W. Hu, Shoumeng Yan, Zhengwei Qi |
ACM Multimedia | 9 |
| 2019 | Towards Standardization of AV Safety: C++ Library for Responsibility Sensitive SafetyabstractThe need for safety in Automated Driving (AD) is becoming increasingly critical with the accelerating deployment of this technology. Beyond functional safety, industry must guarantee the operational safety of automated vehicles. Towards that end, Mobileye introduced the Responsibility Sensitive Safety (RSS), a model-based approach to Safety [1]. In this paper we expand upon this work introducing the C++ Library for Responsibility Sensitive Safety, an open source executable that implements a subset of RSS. We provide architectural details to integrate the C++ Library for Responsibility Sensitive Safety with AD Software pipelines as safety module overseeing decision making of driving policies. We illustrate this application with an example integration with the Baidu Apollo AD stack and simulator, [2] and [3], that provides safety validation of the planning module. Furthermore, we show how the C++ Library for Responsibility Sensitive Safety can be used to explore the usefulness of the RSS model through parameter exploration and analysis on minimum safe longitudinal distance, (dmin), considering different weather conditions. We also compare these results with half-of-speed rule followed in some parts of the world. We expect that the C++ Library for Responsibility Sensitive Safety becomes a critical component of future tools for formal verification, testing and validation of AD safety and that it helps bootstrap the AD research efforts towards standardization of safety. Bernd Gaßmann, Fabian Oboril, Cornelius Bürkle, Shoumeng Yan, Maria Soledad Elli, Ignacio J. Alvarez, Naveen Aerrabotu, Suhel Jaber, Peter van Beek, Darshan Iyer, Jack Weast |
IV | 5 |
| 2017 | Learning Efficient Convolutional Networks through Network SlimmingabstractThe deployment of deep convolutional neural networks (CNNs) in many real world applications is largely hindered by their high computational cost. In this paper, we propose a novel learning scheme for CNNs to simultaneously 1) reduce the model size; 2) decrease the run-time memory footprint; and 3) lower the number of computing operations, without compromising accuracy. This is achieved by enforcing channel-level sparsity in the network in a simple but effective way. Different from many existing approaches, the proposed method directly applies to modern CNN architectures, introduces minimum overhead to the training process, and requires no special software/hardware accelerators for the resulting models. We call our approach network slimming, which takes wide and large networks as input models, but during training insignificant channels are automatically identified and pruned afterwards, yielding thin and compact models with comparable accuracy. We empirically demonstrate the effectiveness of our approach with several state-of-the-art CNN models, including VGGNet, ResNet and DenseNet, on various image classification datasets. For VGGNet, a multi-pass version of network slimming gives a 20× reduction in model size and a 5× reduction in computing operations. Zhuang Liu 0003, Gao Huang 0001, Shoumeng Yan, Changshui Zhang |
ICCV | 5 |
| 2016 | Relationship-aware code search for JavaScript frameworksabstractJavaScript frameworks, such as jQuery, are widely used for developing web applications. To facilitate using these JavaScript frameworks to implement a feature (e.g., functionality), a large number of programmers often search for code snippets that implement the same or similar feature. However, existing code search approaches tend to be ineffective, without taking into account the fact that JavaScript code snippets often implement a feature based on various relationships (e.g., sequencing, condition, and callback relationships) among the invoked framework API methods. To address this issue, we present a novel Relationship-Aware Code Search (RACS) approach for finding code snippets that use JavaScript frameworks to implement a specific feature. In advance, RACS collects a large number of code snippets that use some JavaScript frameworks, mines API usage patterns from the collected code snippets, and represents the mined patterns with method call relationship (MCR) graphs, which capture framework API methods’ signatures and their relationships. Given a natural language (NL) search query issued by a programmer, RACS conducts NL processing to automatically extract an action relationship (AR) graph, which consists of actions and their relationships inferred from the query. In this way, RACS reduces code search to the problem of graph search: finding similar MCR graphs for a given AR graph. We conduct evaluations against representative real-world jQuery questions posted on Stack Overflow, based on 308,294 code snippets collected from over 81,540 files on the Internet. The evaluation results show the effectiveness of RACS: the top 1 snippet produced by RACS matches the target code snippet for 46% questions, compared to only 4% achieved by a relationship-oblivious approach. Zerui Wang, Qianxiang Wang, Shoumeng Yan, Tao Xie 0001, Hong Mei 0001 |
SIGSOFT FSE | 4 |
| 2009 | Terascale chip multiprocessor memory hierarchy and programming modelabstractSmall scale chip multiprocessors are being shipped in volume by all microprocessor vendors. Many of these vendors are also investigating large scale chip multiprocessors targeted towards highly parallel workloads in media, graphics, and others. One of the most challenging aspects of architecting terascale processors is the design of a scalable memory hierarchy. Current proposals for providing coherent shared memory in terascale systems require a sophisticated coherence protocol and memory hierarchy. In this paper we propose an alternate memory configuration along with a programming model that significantly simplifies the terascale memory hierarchy. Our proposal still provides fully coherent shared memory but eliminates the hardware coherence protocol. Our programming model enables the programmer to better express the memory characteristic of terascale workloads. Finally, our proposed memory hierarchy performs better and is more scalable than conventional designs. Shoumeng Yan, Xiaocheng Zhou, Sai Luo, Peinan Zhang, Naveen Cherukuri, Ronny Ronen, Bratin Saha |
HiPC | 1 |
| 2009 | Programming model for a heterogeneous x86 platformabstractThe client computing platform is moving towards a heterogeneous architecture consisting of a combination of cores focused on scalar performance, and a set of throughput-oriented cores. The throughput oriented cores (e.g. a GPU) may be connected over both coherent and non-coherent interconnects, and have different ISAs. This paper describes a programming model for such heterogeneous platforms. We discuss the language constructs, runtime implementation, and the memory model for such a programming environment. We implemented this programming environment in a x86 heterogeneous platform simulator. We ported a number of workloads to our programming environment, and present the performance of our programming environment on these workloads. Bratin Saha, Xiaocheng Zhou, Shoumeng Yan, Mohan Rajagopalan, Jesse Fang, Peinan Zhang, Ronny Ronen, Avi Mendelson |
PLDI | 5 |