EDBT 2026 Demo / reviewers in the wild / expert
Yao Liu 0006
dblp:64/424-6
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-2719-4055ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Low-Cost Secure Branch Predictor to Mitigate the Speculative Attacks by Disrupting Setup PhaseabstractMany types of speculative attacks that exploit branch prediction bear a malicious training process on the branch predictor (BP) in the setup phase. Currently, defense mechanisms on the BP have been rarely studied, and the existing works are restricted to typical scenarios. In this article, we propose a low-cost secure BP to mitigate speculative attacks by monitoring suspicious branch prediction behaviors in all circumstances. We propose a secure mechanism to evaluate the risk level for every branch in the pattern history table. Additionally, we utilize a random number generator to randomly invert the prediction result so as to disrupt the training process, according to the risk level and the generated random number. The maximum inversion probability can be real-time configured during operation. Typically, we implement it on the BI-MODE BP in XuanTie-C910 RISC-V core with a minimal system on chip on field-programmable gate array (FPGA). The realistic hardware evaluation under SPEC2017 with Linux shows that, under the optimal tradeoff between security and performance with the maximum inversion probability of 20%, the hardware and performance overhead are less than 1%, and the speculative attacks including Spectre v1.x and Meltdown can be mitigated with the rates of 44%–88%. Runye Ding, Yao Liu 0006, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Unified Lossless-Throughput Architecture for AES and SM4 Encryption with Changeable KeysabstractNetwork devices targeting to implement data-intensive applications often require the outstanding performance of symmetric encryption, when dealing with multiple concurrent requests from multiple users. Despite the numerous works on high-performance implementation of AES and SM4, hybrid architectures with lossless throughput when the key changes have not been proposed. In this paper, we propose a unified fully-pipelined architecture of AES and SM4 targeting high-performance Galois/Counter Mode application scenarios. The architecture is able to maintain the consistent throughput of input and output datastreams with changeable keys. Compared with state-of-the-art works implemented with the TSMC 65nm process, our design can reduce the area by 26.83% by using a shared composite S-box. With the one-hot S-box, our design can reduce power consumption by 30.93% and increase throughput by 34.21%. Zhishuo Huang, Haosong Zhao, Donald Donglong Chen, Shuyan Zhu, Yinjin Fu, Nong Xiao 0001, Yao Liu 0006 |
ISCAS | 8 |
| 2024 | REALISE-IoT: RISC-V-Based Efficient and Lightweight Public-Key System for IoT ApplicationsabstractLoRa is a promising choice for deploying an IoT network due to its lightweight feature and the extensive support by LoRa Alliance. However, as a fundamental part of LoRa, the typical LoRaWAN protocol confronts severe security challenges because it insecurely utilizes AES-128 to support the low-cost feature. In this paper, we propose a systematic solution that is compatible with LoRaWAN for IoT applications. We extend the standard LoRaWAN protocol with public-key infrastructures. Public-key features like Key exchange and authentication are supported by lightweight hardware implementations of SHA-2, ECDH, EdDSA, and TRNG. A lightweight RISC-V processor with a security coprocessor is implemented and verified using FPGA technology. The security protocol and the prototype hardware system are validated and evaluated on practical applications from our industrial partner. The prototyped development board consumes a static power of 0.116Wand a dynamic power of 0.206 W. The proposed system can achieve a 5.6x-144.7x speed up and reduce memory usage by 2.4x-12.3x for security computations. Gaoyu Mao, Yao Liu 0006, Wangchen Dai, Guangyan Li, Alan H. F. Lam, Ray C. C. Cheung |
IEEE Internet Things J. | 2 |
| 2023 | In-Network Aggregation with Transport Transparency for Distributed TrainingabstractRecent In-Network Aggregation (INA) solutions offload the all-reduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation. Shuo Liu 0002, Qiaoling Wang, Junyi Zhang 0005, Wenfei Wu, Qinliang Lin, Yao Liu 0006, Marco Canini, Ray C. C. Cheung, Jianfei He |
ASPLOS (3) | 6 |
| 2021 | Scalable Fully Pipelined Hardware Architecture for In-Network Aggregated AllReduce CommunicationabstractThe Ring-AllReduce framework is currently the most popular solution to deploy industry-level distributed machine learning tasks. However, only about half of the maximum bandwidth can be achieved in the optimal condition. In recent years, several in-network aggregation frameworks have been proposed to overcome the drawback, but limited hardware information have been disclosed. In this paper, we propose a scalable fully-pipelined architecture that handles tasks like forwarding, aggregation and retransmission with no bandwidth loss. The architecture is implemented on a Xilinx Ultrascale FPGA that connects to 8 working servers with 10 Gb/s network adapters, and it is able to scale to more complicated scenarios involving more workers. Compared with Ring-AllReduce, using AllReduce-Switch improves the efficient bandwidth of AllReduce communication with a ratio of$1.75\times $. In image training tasks, the proposed hardware architecture helps to achieve up to$1.67\times $speedup to the training process. For computing-intensive models, the speedup from communication may be partially hidden by computing. In particular, for ResNet-50, AllReduce-Switch improves the training process with MPI and NCCL by$1.30\times $and$1.04\times $respectively. Yao Liu 0006, Junyi Zhang 0005, Shuo Liu 0002, Qiaoling Wang, Wangchen Dai, Ray C. C. Cheung |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2018 | Lightweight Secure Processor Prototype on FPGAabstractLightweight devices usually come with an additional cryptographic co-processor for enabling the secrecy, in contrast, the master processor is typically a commercial processor where the required protection mechanism is missing. In this paper, an on-going effort in secured architecture named S-RISC-V based on RISC-V core is introduced. The mechanism of key generation used for memory protection is supported together with the joint efforts in the following perspectives, including ISA extension, compiler improvement, and hardware implementation. The architecture has been verified on Zedboard running at 25MHz, driven by the host ARM core. The area overhead is less than 10%, compared with the original RISC-V core. Yao Liu 0006, Ray C. C. Cheung, Hei Wong |
FPL | 1 |