VLDB 2026 Research / reviewers in the wild / expert
Songqiao Cui
dblp:348/7604
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0001-9407-1050ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extending and Accelerating Inner Product Masking with Fault Detection via Instruction Set ExtensionabstractInner product masking is a well-studied masking countermeasure against side-channel attacks. IPM-FD further extends the IPM scheme with fault detection capabilities. However, implementing IPM-FD in software especially on embedded devices results in high computational overhead. Therefore, in this work we perform a detailed analysis of all building blocks for IPM-FD scheme and propose a Masked Processing Unit to accelerate all operations, for example multiplication and IPM-FD specific Homogenization. We can then offload these computational extensive operations with dedicated hardware support. With only 4.05% and 4.01% increase in Look-Up Tables and Flip-Flops (Random Number Generator excluded), respectively, compared with baseline cv32e40p RISC-V core, we can achieve up to 16.55× speed-up factor with optimal configuration. We then practically evaluate the side-channel security via uni- and bivariate Test Vector Leakage Assessment which exhibits no leakage. Finally, we use two different methods to simulate the injected fault and confirm the fault detection capability of up to k−1 faults, with k being the replication factor. Songqiao Cui, Geng Luo, Junhan Bao, Josep Balasch, Ingrid Verbauwhede |
DATE | 1 |
| 2024 | Configurable Loop Shuffling via Instruction Set ExtensionsabstractHiding is a popular countermeasure against side-channel attacks. In software contexts, it typically involves adding time domain randomizations by shuffling the execution order of operations and/or by inserting dummy instructions. The combination of such countermeasures demands the presence of a random permutation algorithm, in addition to extra program flow control steps, leading to significant overheads in the software implementation. In this work, we improve the performance of such hiding countermeasures by designing a hardware engine capable of shuffling the execution order of software loops. Our engine can be easily integrated into a processor architecture and configured by means of custom instruction set extensions. We demonstrate this by prototyping and evaluating it on the CV32E40P RISC-V core (formerly RI5CY). For this particular platform, we additionally combine our engine with its native hardware loop feature and propose instructions capable of permuting memory access addresses at runtime. We validate the functional correctness of our design by targeting two algorithms explored in related works: the popular AES block cipher and the Number Theoretic Transform (NTT) found in several post-quantum cryptographic algorithms. For both designs, we benchmark the performance overheads of the resulting implementations and validate their increased security by means of practical experiments on an FPGA. Songqiao Cui, Josep Balasch |
ASAP | 1 |
| 2023 | Efficient Software Masking of AES through Instruction Set ExtensionsabstractMasking is a well-studied countermeasure to protect software implementations against side-channel attacks. For the case of AES, incorporating masking often requires to implement internal transformations using finite field arithmetic. This results in significant performance overheads, mostly due to finite field multiplications, which are even worsened when no lookup tables are used. In this work, we extend a RISC-V core with custom instructions to accelerate AES finite field arithmetic. With a 3.3 % area increase, we measure 7.2x and S.4x speed up over software-only implementations of first-order Boolean Masking and Inner Product Masking, respectively. We also investigate vectorized instructions capable of exploiting the intra-block and inter-block parallelism in the implementation. Our implementations avoid the use of lookup tables, run in constant time, and show no evidence of first-order leakage when evaluated on an FPG A. Songqiao Cui, Josep Balasch |
DATE | 1 |