VLDB 2026 Research / reviewers in the wild / expert
Muhammad Husni Santriaji
dblp:177/5255 · also Muhammad Santriaji
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-2308-0002ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fully Homomorphic Encryption Inference of Neural Networks Using CKKS-TFHE Scheme Switching and Accelerated Linear LayersabstractFully homomorphic encryption (FHE) allows neural network inference to be performed directly on encrypted data and models, preserving end-to-end privacy. One of the most widely used FHE schemes, CKKS, supports approximate arithmetic over real numbers and is well-suited for computing polynomial operations. However, CKKS does not natively support conditional branching, making it unsuitable for non-linear functions such as ReLU, which behaves differently depending on whether the input is negative or nonnegative. To address this, ReLU is typically approximated using polynomials. Unfortunately, this approach introduces two key drawbacks: (1) the approximation is only valid within a certain input range, and (2) identifying this range typically requires prior exposure of cleartext data to perform profiling, thus partially compromising privacy. This requirement is problematic in practice, as small profiling errors can lead to substantial accuracy degradation. On the other hand, the TFHE scheme supports binary gate-level computation, enabling precise implementation of conditional operations such as ReLU. Yet TFHE lacks the SIMD capabilities of CKKS, leading to significant computational latency. This paper explores the use of scheme switching for encrypted neural network inference: leveraging CKKS for polynomial-compatible layers and switching to TFHE for layers that require conditional logic, such as ReLU. We demonstrate that this hybrid approach substantially improves the numerical fidelity of encrypted neural network inference: while polynomial approximations suffer from numerical divergence, scheme switching more closely matches non-encrypted inference. Anas Banta Seutia, Muhammad Zaky Firdaus, Muhammad Alfi Ramadhan, Kabul Kurniawan, Muhammad Husni Santriaji, Alfian Amrizal, Reza Pulungan, Hiroyuki Takizawa |
AsiaCCS | 5 |
| 2026 | Coral: Covariance-Guided Resource Adaptive Learning for Efficient Edge Inference
Ahmad N. L. Nabhaan, Zaki Sukma, Rakandhiya D. Rachmanto, Muhammad Husni Santriaji, Byungjin Cho, Arief Setyanto, In Kee Kim |
ICFEC | 4 |
| 2026 | WASL: Harmonizing Uncoordinated Adaptive Modules in Multi-Tenant Cloud SystemsabstractModern cloud applications increasingly rely on adaptive control modules, such as dynamic resource tuning or system reconfiguration, to meet strict quality-of-service (QoS) objectives. However, when multiple independently developed adaptation modules are colocated on a shared infrastructure, their uncoordinated behavior causes interference leading to QoS violations. Existing approaches require centralized control or inter-module communication, violating modularity and limiting adoption in multi-tenant environments. Ahsan Pervaiz, Anwesha Das 0001, Vedant Kodagi, Muhammad Husni Santriaji, Henry Hoffmann |
ICPE | 4 |
| 2025 | NAPER: Fault Protection for Real-Time Resource-Constrained Deep Neural NetworksabstractFault tolerance in Deep Neural Networks (DNNs) deployed on resource-constrained systems presents unique challenges for high-accuracy applications with strict timing requirements. Memory bit-flips can severely degrade DNN accuracy, while traditional protection approaches like Triple Modular Redundancy (TMR) often sacrifice accuracy to maintain reliability, creating a three-way dilemma between reliability, accuracy, and timeliness. We introduce NAPER, a novel protection approach that addresses this challenge through ensemble learning. Unlike conventional redundancy methods, NAPER employs heterogeneous model redundancy, where diverse models collectively achieve higher accuracy than any individual model. This is complemented by an efficient fault detection mechanism and a real-time scheduler that prioritizes meeting deadlines by intelligently scheduling recovery operations without interrupting inference. Our evaluations demonstrate NAPER's superiority: 40% faster inference in both normal and fault conditions, maintained accuracy 4.2% higher than TMR-based strategies, and guaranteed uninterrupted operation even during fault recovery. NAPER effectively balances the competing demands of accuracy, reliability, and timeliness in real-time DNN applications. Rian Adam Rajagede, Muhammad Husni Santriaji, Muhammad Arya Fikriansyah, Hilal Hudan Nuha, Yanjie Fu, Yan Solihin |
IOLTS | 2 |
| 2025 | DataSeal: Ensuring the Verifiability of Private Computation on Encrypted DataabstractFully Homomorphic Encryption (FHE) allows computations to be performed directly on encrypted data without needing to decrypt it first. This “encryption-in-use” feature is crucial for securely outsourcing computations in privacy-sensitive areas such as healthcare and finance. Nevertheless, in the context of FHE-based cloud computing, clients often worry about the integrity and accuracy of the outcomes. This concern arises from the potential for a malicious server or server-side vulnerabilities that could result in tampering with the data, computations, and results. Ensuring integrity and verifiability with low overhead remains an open problem, as prior attempts have not yet achieved this goal. To tackle this challenge and ensure the verification of FHE's private computations on encrypted data, we introduce DataSeal, which combines the low overhead of the algorithm-based fault tolerance (ABFT) technique with the confidentiality of FHE, offering high efficiency and verification capability. Through thorough testing in diverse contexts, we demonstrate that DataSeal achieves much lower overheads for providing computation verifiability for FHE than other techniques that include MAC, ZKP, and TEE. DataSeal's space and computation overheads decrease to nearly negligible as the problem size increases. Muhammad Husni Santriaji, Yancheng Zhang, Qian Lou, Yan Solihin |
SP | 1 |
| 2023 | TrojBits: A Hardware Aware Inference-Time Attack on Transformer-Based Language ModelsabstractTransformer-based language models demonstrate exceptional performance in Natural Language Processing (NLP) tasks but remain susceptible to backdoor attacks involving hidden input triggers. Trojan injection via hardware bitflips presents a significant challenge for contemporary language models. However, previous research overlooks practical hardware considerations, such as DRAM and cache memory structures, resulting in unrealistic attacks that demand the manipulation of an excessive number of parameters and bits. In this paper, we present TrojBits, a novel approach requiring minimal bit-flips to effectively insert Trojans into real-world Transformer language model systems. This is achieved through a three-module framework designed to efficiently target Transformer-based language models, consisting of Vulnerable Parameters Ranking (VPR), Hardware-aware Attack Optimization (HAO), and Vulnerable Bits Pruning (VBP). Within the VPR module, we are the first to employ Gradient-guided Fisher information to identify the most susceptible Transformer parameters, specifically in the word embedding layer. The HAO module then redistributes these parameters across multiple triggers, conforming to hardware constraints by incorporating a regularization term in the trojan optimization methodology. Finally, the VBP module aims to reduce the number of bit-flips by discarding less significant bits. We evaluate TrojBits on two representative NLP models, BERT and XLNE, on three classification tasks (SST2, OffensEval, and AG’s News). Our results demonstrate that our TrojBits successfully achieves the inference-time attack with only 64 parameters out of 116 million and 90-bit flips while maintaining the model performance. Mansour Al Ghanim, Muhammad Husni Santriaji, Qian Lou, Yan Solihin |
ECAI | 2 |
| 2020 | ALERT: Accurate Learning for Energy and Timeliness
Chengcheng Wan 0001, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann, Michael Maire, Shan Lu 0001 |
USENIX ATC | 2 |
| 2018 | MERLOT: Architectural Support for Energy-Efficient Real-Time Processing in GPUsabstractCorrect functioning of embedded systems requires strict timing guarantees. Traditionally, enforcing timing guarantees is the operating system's responsibility. The OS scheduler assigns sufficient resources to an application task to ensure it meets its deadline. Meeting hard real-time deadlines requires the scheduler to be conservative and allocate for the worst case timing; when behavior is not worst case, extra resources are allocated and energy is wasted. Some software schedulers reduce this energy waste by recognizing when an application is ahead of a worst case schedule and reclaiming unneeded resources, but they are fundamentally limited by (1) overhead and (2) a lack of visibility into low-level resource usage. Therefore, this paper advocates hardware assistance for energy management of hard real-time tasks. Specifically, we propose MERLOT, a hardware-based resource manager for GPUs that enforces software-specified timing guarantees with minimal energy. We implement MERLOT in VHDL and find that its performance, power, and area overheads are minuscule. We implement MERLOT in GPGPU-Sim to test timing and energy consumption and compare to two software-only approaches: one that always allocates for worst case timing and an intelligent approach that reduces resource usage when it recognizes better than worst case behavior. Compared to the approach that always allocates for worst case, MERLOT reduces energy by 16.73% on average, with 16:43% in 1.5X and 17:03% in 2.0X WCET. Compared to the intelligent software-only approach, MERLOT produces a 15:6% energy savings. MERLOT uses less energy than softwareonly approaches because it recognizes better than worst case behavior earlier, by monitoring hardware events that are not visible to software and quickly adjusting resource usage. Muhammad Husni Santriaji, Henry Hoffmann |
RTAS | 1 |
| 2016 | GRAPE: Minimizing energy for GPU applications with performance requirementsabstractMany applications have performance requirements (e.g., real-time deadlines or quality-of-service goals) and we can save tremendous energy by tailoring resource usage so the application just meets its performance using the minimal resources. This problem is a classic constrained optimization: the performance goal is the constraint and energy consumption is the objective to be optimized. While several existing hardware approaches solve unconstrained optimizations (i.e., maximizing performance or minimizing energy), we are not aware of a hardware approach that minimizes GPU energy under an externally defined performance constraint. Therefore, we propose GRAPE, a hardware control system for GPUs that coordinates core usage, wavefront/warp action, core speed, and memory speed to deliver user-specified performance while minimizing energy. We implement GRAPE in VHDL (to demonstrate feasibility) and as an extension to GPGPU-Sim (for performance and power measurement). We find that GRAPE can be implemented with very low hardware overhead; however, compared to the no-overhead approach of race-to-idle, GRAPE reduces energy by 9-26% (depending on the performance goal), while meeting performance goals with an average error of 0.75%. Muhammad Husni Santriaji, Henry Hoffmann |
MICRO | 1 |