EDBT 2026 Demo / reviewers in the wild / expert
Sandhya Koteshwara
dblp:178/2406
· DBLP profile ↗
10ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0003-3182-219XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vela: A Virtualized LLM Training System with GPU Direct RoCEabstractVela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark. Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam |
ASPLOS (2) | 17 |
| 2024 | S2TAR: Shared Secure Trusted Accelerators with Reconfiguration for Machine Learning in the CloudabstractThe demand for hardware accelerators such as Tensor Processing Units (TPUs) and Graphics Processing Units (GPUs) is rapidly increasing due to growing Machine Learning (ML) workloads. As with any shared computing resources, there is a growing need to dynamically adjust and scale accelerator services while ensuring data privacy and confidentiality, especially in cloud environments. We propose a secure and reconfigurable TPU design with confidential computing support, achieved through a Trusted Execution Environment (TEE) framework tailored for reconfigurable TPU in a multi-tenant cloud. Our contributions include a novel TPU design based on switchbox-enabled systolic arrays to support rapid dynamic partitioning. We evaluate our TPU design with TEEs in shared environments, achieving up to 42.1 % higher performance for realistic ML inference workloads. Our remote attestation protocol extends to sub-device partitions, providing trustworthiness on a fine-grained level and decouples host and accelerator TEEs into separate attestation reports without degrading security guarantees. Our work presents a new TEE framework for secure and reconfigurable ML accelerators in a multi-tenant cloud environment. Sandhya Koteshwara, Mengmei Ye, Hubertus Franke, Deming Chen |
CLOUD | 2 |
| 2024 | STRonG: System Topology Risk Analysis on GraphsabstractProduction systems run complex stacks comprising constantly-evolving hardware and software components. Vulnerabilities in such stacks continuously pose security risks to both the service provider and customers, thus calling for a solution to analyze and quantify security risks. STRonG is a framework that leverages a layered graph-based approach to model, analyze, and quantify security risks in complex software and hardware stacks of systems. We propose using adjustable templates/stencils for relatively tractable and consistent modeling and allow user-defined scoring methods to be applied. STRonG quantitatively assesses how structure, components, or attribute modifications impact the security risk of critical parts of a system stack during the design or early stages of the development process. The framework’s efficacy is demonstrated by applying STRonG to the control stack of OpenStack cloud infrastructure and performing risk assessment before and after introducing a novel security layer, Secure Hypervisor Channel (SHC). We also demonstrate how introducing SHC can quantitatively reduce system risk. Lars Schneidenbach, Sandhya Koteshwara, Martin Ohmacht, Apoorve Mohan |
CCGrid | 2 |
| 2023 | AccShield: a New Trusted Execution Environment with Machine-Learning AcceleratorsabstractMachine learning accelerators such as the Tensor Processing Unit (TPU) are already being deployed in the hybrid cloud, and we foresee such accelerators proliferating in the future. In such scenarios, secure access to the acceleration service and trustworthiness of the underlying accelerators become a concern. In this work, we present AccShield, a new method to extend trusted execution environments (TEEs) to cloud accelerators which takes both isolation and multi-tenancy into security consideration. We demonstrate the feasibility of accelerator TEEs by a proof of concept on an FPGA board. Experiments with our prototype implementation also provide concrete results and insights for different design choices related to link encryption, isolation using partitioning and memory encryption. William Kozlowski, Sandhya Koteshwara, Mengmei Ye, Hubertus Franke, Deming Chen |
DAC | 3 |
| 2020 | Performance Optimization of Lattice Post-Quantum Cryptographic Algorithms on Many-Core ProcessorsabstractCurrent public-key cryptography systems are vulnerable to quantum computing based attacks. Post-quantum cryptographic (PQC) schemes, based on mathematical paradigms such as lattice-based hard problems, are under consideration by NIST as quantum-safe alternatives. Profiling of several latticebased cryptography algorithms reveals that polynomial multiplication and random number generation are the most time consuming components. The nature of these computations and challenges in vectorizing them are discussed in this paper. Vectorization of the identified time-consuming primitives results in 52% and 83% improvement in performance for the CRYSTALS-Kyber KEM SHA3 variant and AES variant, respectively. Sandhya Koteshwara, Manoj Kumar 0006, Pratap Pattnaik |
ISPASS | 1 |
| 2019 | Architecture Optimization and Performance Comparison of Nonce-Misuse-Resistant Authenticated Encryption AlgorithmsabstractThis paper presents a performance comparison of new authenticated encryption (AE) algorithms which are aimed at providing better security and resource efficiency compared to existing standards. Specifically, these algorithms improve the security of existing AE standards by providing a critical property termed nonce-misuse resistance. This paper addresses algorithm to architectural mappings of several candidates from the ongoing Competition for AE: Security, Applicability, and Robustness as well as a submission from the Crypto Forum Research Group. Implementations of the architectures on both field-programmable gate arrays and application-specific integrated circuits platforms are provided and compared with the architecture of a popular standard: Advanced Encryption Standard in Galois Counter mode (AES-GCM). Optimizations that are applicable to AE, in general, and nonce-misuse-resistant architectures, in particular, are presented. A hardware-software codesign approach to optimization is also discussed. The implementations via proposed optimizations demonstrate that new AE algorithms can provide comparable performance as standard AES-GCM while enhancing security and resource utilization for specific use-case scenarios. Sandhya Koteshwara, Amitabh Das, Keshab K. Parhi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Low-Energy Architectures of Linear Classifiers for IoT Applications using Incremental Precision and Multi-Level ClassificationabstractThis paper presents a novel incremental-precision classification approach that leads to a reduction in energy consumption of linear classifiers for IoT applications. Features are first input to a low-precision classifier. If the classifier successfully classifies the sample, then the process terminates. Otherwise, the classification performance is incrementally improved by using a classifier of higher precision. This process is repeated until the classification is complete. The argument is that many samples can be classified using the low-precision classifier, leading to a reduction in energy. To achieve incremental-precision, a novel data-path decomposition is proposed to design of fixed-width adders and multipliers. These components improve the precision without recalculating the outputs, thus reducing energy. Using a linear classification example, it is shown that the proposed incremental-precision based multi-level classifier approach can reduce energy by about 41% while achieving comparable accuracies as that of a full-precision system. Sandhya Koteshwara, Keshab K. Parhi |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Key-Based Dynamic Functional Obfuscation of Integrated Circuits Using Sequentially Triggered Mode-Based DesignabstractThis paper proposes a novel technique for hardware obfuscation termed dynamic functional obfuscation. Hardware obfuscation refers to a set of countermeasures used against IC counterfeiting and illegal overproduction. Traditionally, obfuscation encrypts semiconductor circuits using key inputs which must be set to a correct value to operate the circuit correctly. By keeping the key values secret during the manufacturing process, any attempt by unauthorized parties to overproduce chips or pirate designs is thwarted. The proposed dynamic technique differs from existing fixed obfuscation schemes as the obfuscating signals change over time. This results in inconsistent circuit behavior upon input of incorrect key, where the chip operates correctly sometimes and fails sometimes. The advantage of dynamic obfuscation is that it results in stronger obfuscation by increasing the time complexity of deciphering the correct key using brute-force attack, even with shorter keys. Moreover, the dynamic nature of these circuits also makes them resistant to reverse engineering and SAT solver-based attacks. To achieve dynamic obfuscation, ideas from hardware Trojan literature and sequentially triggered counters are utilized. A demonstration of obfuscation on sequential circuits implementing fast Fourier transform (FFT) algorithm and Ethernet IP shows low overall area and power overheads of less than 1%. Security in terms of time to attack for the FFT circuit (for a key size of 30 bits and a system operating at 100 MHz) is increased to 1021,055 years using dynamic obfuscation compared with only 5.36 s using fixed obfuscation schemes. For the Ethernet IP core, time to attack of dynamic obfuscation with a key size of 32 bits is 1046,423,135 years compared with 21.47s with fixed obfuscation. It is also shown that for a key size of K bits, the lower bound for time to attack using brute-force is proportional to K2Kand K22Kfor the proposed design using one and two random number generators, respectively. Sandhya Koteshwara, Chris H. Kim, Keshab K. Parhi |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | FPGA implementation and comparison of AES-GCM and Deoxys authenticated encryption schemesabstractAuthenticated Encryption (AE) schemes are key-based cryptographic algorithms that provide both goals of confidentiality of message and authenticity of the sender, simultaneously. Traditionally, Advanced Encryption Standard (AES) in Galois Counter Mode (AES-GCM), among several other approaches, has been employed for Authenticated Encryption. However, several lightweight cryptographic applications such as those used in sensor networks or RFID security can benefit from new AE schemes which can be constructed more efficiently. In this paper we provide evaluations for Deoxys, a third round candidate from the ongoing Competition for Authenticated Encryption: Security, Applicability, and Robustness (CAESAR). We describe simplified flow diagrams and a detailed summary on the timing performance, area, memory and energy requirements of AES-GCM and Deoxys, using our own implementations on Altera Cyclone V FPGAs. Our analysis shows that Deoxys requires 10% less energy per bit and 25% less LUTs as compared to AES-GCM. Sandhya Koteshwara, Amitabh Das, Keshab K. Parhi |
ISCAS | 1 |
| 2017 | Hierarchical functional obfuscation of integratec circuits using a mode-based approachabstractHardware obfuscation has been proposed as a hardware security measure against reverse engineering, intellectual property (IP) piracy and integrated circuits (IC) overbuilding. In this paper, we present a novel method of obfuscation using a hierarchical approach. In the design flow, IP vendors obfuscate their designs using a set of keys and provide these keys to the design house. The design house then integrates all the IPs and adds its own keys to create a complete obfuscated system. This prevents both misuse of IPs and illegal use of ICs since only secure parties have access to the correct keys. The obfuscation at each level is performed using a mode-based approach in which the design can operate in meaningful and non-meaningful modes. The design is functionally correct in only one mode. An attacker needs to work through different levels of the design to correctly decipher its operation and correct working mode. Since each of the IPs can work in multiple meaningful modes, the attack becomes more difficult as the number of IPs increases. These ideas are demonstrated using a convolution architecture with fast Fourier transform (FFT) blocks. With only about 13% area and 15% power overhead over an unobfuscated design, it is shown that the proposed design has the flexibility to be obfuscated with different key sizes and overheads depending on the level of security. Sandhya Koteshwara, Chris H. Kim, Keshab K. Parhi |
ISCAS | 1 |