Mengmei Ye

dblp:193/4299 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-3434-1968ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Privacy-Preserving Multimedia Mobile Cloud Computing Using Cost-Effective Protective Perturbation
abstract
Mobile cloud computing has been adopted in many multimedia applications, where resource-constrained mobile devices send multimedia data (e.g., images) to remote cloud servers to request computation intensive multimedia services (e.g., image recognition). Despite the performance improvement, the cloud-based mechanism often causes privacy concerns as the user data is offloaded to untrusted cloud servers. Existing solutions require computation-intensive perturbation generation on resource-constrained mobile devices. Also, the protected images are not compliant with standard image compression algorithms, leading to significant bandwidth consumption. We develop a novel privacy-preserving multimedia mobile cloud computing framework, namely PMC2, to address the resource and bandwidth challenges. PMC2 employs confidential computing on an edge server to deploy the perturbation generator, which addresses the on-device resource challenge. Also, we develop a neural compressor for the protected images to address the bandwidth challenge. Our evaluations of PMC2 demonstrate superior latency, power efficiency, and bandwidth consumption while maintaining high accuracy in the target multimedia service.
Zhongze Tang, Mengmei Ye, Yao Liu 0001, Sheng Wei 0001
NOSSDAV3
2025 Guest Editorial Special Issue on Emerging Hardware Security and Trust Technologies - AsianHOST 2023
abstract
If no abstract provided do not include one in the JATS XML
Xinmiao Zhang 0001, Chongyan Gu, Mengmei Ye, Reza Azarderakhsh, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Securing AI Inference in the Cloud: Is CPU-GPU Confidential Computing Ready?
abstract
Many applications have been offloaded onto cloud environments to achieve higher agility, access to more powerful computational resources, and obtain better infrastructure management. Although cloud environments provide solid security solutions, users with highly sensitive data or regulatory compliance requirements, such as HIPAA (Health Insurance Portability and Accountability Act) and GDPR (General Data Protection Regulation), still hesitate to move such application domains to the cloud. To address these concerns, cloud service providers have started to offer solutions to protect data confidentiality and integrity through trusted execution environments (TEEs). While so far these were limited to CPU TEEs only, NVIDIA's Hopper architecture has shifted the landscape by enabling confidential computing features essential to protecting confidentiality and integrity for real-world applications offloaded to GPUs, such as large language models (LLMs). However, there lacks a sufficient study on how much performance overhead confidential computing introduces in a TEE comprised of a CPU-GPU configuration. In this paper we evaluate a confidential computing environment comprised of an Intel TDX system and NVIDIA H100 GPUs through various micro benchmarks and real workloads including BERT, LLaMA, and Granite large language models and provide discussions on the overhead incurred by confidential computing when GPUs are utilized. We show that while LLMs are sensitive to the model types and batch sizes, when larger models with pipelined processing are deployed, the performance of LLM inference in CPU-GPU TEEs can be close to par with their non-confidential setups.
Apoorve Mohan, Mengmei Ye, Hubertus Franke, Mudhakar Srivatsa, Nelson Mimura Gonzalez
CLOUD2
2024 S2TAR: Shared Secure Trusted Accelerators with Reconfiguration for Machine Learning in the Cloud
abstract
The demand for hardware accelerators such as Tensor Processing Units (TPUs) and Graphics Processing Units (GPUs) is rapidly increasing due to growing Machine Learning (ML) workloads. As with any shared computing resources, there is a growing need to dynamically adjust and scale accelerator services while ensuring data privacy and confidentiality, especially in cloud environments. We propose a secure and reconfigurable TPU design with confidential computing support, achieved through a Trusted Execution Environment (TEE) framework tailored for reconfigurable TPU in a multi-tenant cloud. Our contributions include a novel TPU design based on switchbox-enabled systolic arrays to support rapid dynamic partitioning. We evaluate our TPU design with TEEs in shared environments, achieving up to 42.1 % higher performance for realistic ML inference workloads. Our remote attestation protocol extends to sub-device partitions, providing trustworthiness on a fine-grained level and decouples host and accelerator TEEs into separate attestation reports without degrading security guarantees. Our work presents a new TEE framework for secure and reconfigurable ML accelerators in a multi-tenant cloud environment.
Sandhya Koteshwara, Mengmei Ye, Hubertus Franke, Deming Chen
CLOUD3
2023 Free the Turtles: Removing Nested Virtualization for Performance and Confidentiality in the Cloud
abstract
With the growing popularity of containers running in cloud environments, one might believe techniques like virtual machines (VMs) and nested virtualization are no longer essential for cloud computing, but this is not true. First, they are foundational for the Infrastructure as a Service (IaaS) base that most clouds use for their Kubernetes (K8s) offerings - nodes are deployed in VMs instead of on bare-metal machines to improve flexibility and resource utilization. Second, state-of-theart technology like Kata Containers or KubeVirt deploys VMs inside K8s, to improve container isolation or to offer a cloudnative way to deploy VMs. When such techniques are coupled with VM-based K8s worker node deployments, this requires the use of nested virtualization. However, nested virtualization introduces a larger trusted computing base (TCB) and therefore raises security concerns. In addition, it introduces an important performance overhead and cannot be easily applied to confidential computing, an emerging technology to allow the execution of sensitive workloads in the cloud. In this paper, we propose the secondary-VM (secVM) framework, an alternative to nested virtualization which addresses these challenges by flattening the nested hierarchy, but maintains the resource isolation features of nested virtualization. By removing the emulation layer required for nested virtualization, the secVM framework reduces the TCB, is compatible with any confidential computing techniques, and removes the overhead associated with nested virtualization. Therefore, in reference to the Turtles Project that introduced the concept of nested virtualization in the open-source virtualization technology - Kernel-based Virtual Machine (KVM), we argue it is time to free the “turtles”.
Mengmei Ye, Angelo Ruocco, Daniele Buono, James Bottomley, Hubertus Franke
CLOUD1
2023 Remote attestation of confidential VMs using ephemeral vTPMs
abstract
Trying to address the security challenges of a cloud-centric software deployment paradigm, silicon and cloud vendors are introducing confidential computing – an umbrella term aimed at providing hardware and software mechanisms for protecting cloud workloads from the cloud provider and its software stack. Today, Intel Software Guard Extensions (SGX), AMD secure encrypted virtualization (SEV), Intel trust domain extensions (TDX), etc., provide a way to shield cloud applications from the cloud provider through encryption of the application’s memory below the hardware boundary of the CPU, hence requiring trust only in the CPU vendor. Unfortunately, existing hardware mechanisms do not automatically enable the guarantee that a protected system was not tampered with during configuration and boot time. Such a guarantee relies on a hardware root of trust, i.e., an integrity-protected location that can store measurements in a trustworthy manner, extend them, and authenticate the measurement logs to the user (remote attestation).
Vikram Narayanan, Cláudio Carvalho, Angelo Ruocco, Gheorghe Almási 0001, James Bottomley, Mengmei Ye, Tobin Feldman-Fitzthum, Daniele Buono, Hubertus Franke, Anton Burtsev
ACSAC6
2023 AccShield: a New Trusted Execution Environment with Machine-Learning Accelerators
abstract
Machine learning accelerators such as the Tensor Processing Unit (TPU) are already being deployed in the hybrid cloud, and we foresee such accelerators proliferating in the future. In such scenarios, secure access to the acceleration service and trustworthiness of the underlying accelerators become a concern. In this work, we present AccShield, a new method to extend trusted execution environments (TEEs) to cloud accelerators which takes both isolation and multi-tenancy into security consideration. We demonstrate the feasibility of accelerator TEEs by a proof of concept on an FPGA board. Experiments with our prototype implementation also provide concrete results and insights for different design choices related to link encryption, isolation using partitioning and memory encryption.
William Kozlowski, Sandhya Koteshwara, Mengmei Ye, Hubertus Franke, Deming Chen
DAC4
2022 Visual privacy protection in mobile image recognition using protective perturbation
abstract
Deep neural networks (DNNs) have been widely adopted in mobile image recognition applications. Considering intellectual property and computation resources, the image recognition model is often deployed at the service provider end, which takes input images from the user's mobile device and accomplishes the recognition task. However, from the user's perspective, the input images could contain sensitive information that is subject to visual privacy concerns, and the user must protect the privacy while offloading them to the service provider. To address the visual privacy issue, we develop a protective perturbation generator at the user end, which adds perturbations to the input images to prevent privacy leakage. Meanwhile, the image recognition model still runs at the service provider end to recognize the protected images without the need of being re-trained. Our evaluations using the CIFAR-10 dataset and 8 image recognition models demonstrate effective visual privacy protection while maintaining high recognition accuracy. Also, the protective perturbation generator achieves premium timing performance suitable for real-time image recognition applications.
Mengmei Ye, Zhongze Tang, Huy Phan, Yi Xie 0001, Bo Yuan 0001, Sheng Wei 0001
MMSys1
2021 Runtime Fault Injection Detection for FPGA-based DNN Execution Using Siamese Path Verification
abstract
Deep neural networks (DNNs) have been deployed on FPGAs to achieve improved performance, power efficiency, and design flexibility. However, the FPGA-based DNNs are vulnerable to fault injection attacks that aim to compromise the original functionality. The existing defense methods either duplicate the models and check the consistency of the results at runtime, or strengthen the robustness of the models by adding additional neurons. However, these existing methods could introduce huge overhead or require retraining the models. In this paper, we develop a runtime verification method, namely Siamese path verification (SPV), to detect fault injection attacks for FPGA-based DNN execution. By leveraging the computing features of the DNN and designing the weight parameters, SPV adds neurons to check the integrity of the model without impacting the original functionality and, therefore, model retraining is not required. We evaluate the proposed SPV approach on Xilinx Virtex-7 FPGA using the MNIST dataset. The evaluation results show that SPV achieves the security goal with low overhead.
Xianglong Feng, Mengmei Ye, Ke Xia, Sheng Wei 0001
DATE2
2021 Fake Gradient: A Security and Privacy Protection Framework for DNN-based Image Classification
abstract
Deep neural networks (DNNs) have demonstrated phenomenal success in image classification applications and are widely adopted in multimedia internet of things (IoT) use cases, such as smart home systems. To compensate for the limited resources on the IoT devices, the computation-intensive image classification tasks are often offloaded to remote cloud services. However, the offloading-based image classification could pose significant security and privacy concerns to the user data and the DNN model, leading to effective adversarial attacks that compromise the classification accuracy. The existing defense methods either impact the original functionality or result in high computation or model re-training overhead. In this paper, we develop a novel defense approach, namely Fake Gradient, to protect the privacy of the data and defend against adversarial attacks based on encryption of the output. Fake Gradient can hide the real output information by generating fake classes and further mislead the adversarial perturbation generation based on fake gradient knowledge, which helps maintain a high classification accuracy on the perturbed data. Our evaluations using ImageNet and 7 popular DNN models indicate that Fake Gradient is effective in protecting the privacy and defending against adversarial attacks targeting image classification applications.
Xianglong Feng, Yi Xie 0001, Mengmei Ye, Zhongze Tang, Bo Yuan 0001, Sheng Wei 0001
ACM Multimedia3
2018 HISA: hardware isolation-based secure architecture for CPU-FPGA embedded systems
abstract
Heterogeneous CPU-FPGA systems have been shown to achieve significant performance gains in domain-specific computing. However, contrary to the huge efforts invested on the performance acceleration, the community has not yet investigated the security consequences due to incorporating FPGA into the traditional CPU-based architecture. In fact, the interplay between CPU and FPGA in such a heterogeneous system may introduce brand new attack surfaces if not well controlled. We propose a hardware isolation-based secure architecture, namely HISA, to mitigate the identified new threats. HISA extends the CPU-based hardware isolation primitive to the heterogeneous FPGA components and achieves security guarantees by enforcing two types of security policies in the isolated secure environment, namely the access control policy and the output verification policy. We evaluate HISA using four reference FPGA IP cores together with a variety of reference security policies targeting representative CPU-FPGA attacks. Our implementation and experiments on real hardware prove that HISA is an effective security complement to the existing CPU-only and FPGA-only secure architectures.
Mengmei Ye, Xianglong Feng, Sheng Wei 0001
ICCAD1
2018 EvoIsolator: Evolving Program Slices for Hardware Isolation Based Security
abstract
To provide strong security support for today’s applications, microprocessor manufacturers have introduced hardware isolation, an on-chip mechanism that provides secure accesses to sensitive data. Currently, hardware isolation is still difficult to use by software developers because the process to identify access points to sensitive data is error-prone and can lead to under and over protection of sensitive data. Under protection can lead to security vulnerabilities. Over protection can lead to an increased attack surface and excessive communication overhead. In this paper we describe EvoIsolator , a search-based framework to (i) automatically generate executable minimal slices that include all access points to a set of specified sensitive data; and (ii) automatically optimize (for small code block size and low communication overhead) the code modules for hardware isolation. We demonstrate, through a small feasibility study, the potential impact of our proposed code optimizer.
Mengmei Ye, Myra B. Cohen, Witawas Srisa-an, Sheng Wei 0001
SSBSE1
2016 Two-way real time multimedia stream authentication using physical unclonable functions
abstract
Multimedia authentication is an integral part of multimedia signal processing in many real-time and security sensitive applications, such as video surveillance. In such applications, a full-fledged video digital rights management (DRM) mechanism is not applicable due to the real time requirement and the difficulties in incorporating complicated license/key management strategies. This paper investigates the potential of multimedia authentication from a brand new angle by employing hardware-based security primitives, such as physical unclonable functions (PUFs). We show that the hardware security approach is not only capable of accomplishing the authentication for both the hardware device and the multimedia stream but, more importantly, introduce minimum performance, resource, and power overhead. We justify our approach using a prototype PUF implementation on Xilinx FPGA boards. Our experimental results on the real hardware demonstrate the high security and low overhead in multimedia authentication obtained by using hardware security approaches.
Mehrdad Zaker Shahrak, Mengmei Ye, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001
MMSP2