EDBT 2026 Demo / reviewers in the wild / expert
Tian Ye 0002
dblp:230/9617
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-9298-8906ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Model-Hardware Co-design Framework for Robust and Efficient CNN-Based SAR ATRabstractConvolutional Neural Networks (CNNs) have achieved state-of-the-art accuracy in Synthetic Aperture Radar (SAR) Automatic Target Recognition (ATR), but their high computational cost, latency, and memory usage make it challenging on resource-constrained platforms such as small satellites. While adversarial robustness is critical for real-world SAR ATR, it is often overlooked in system-level optimizations. Addressing both challenges requires more than model compression or accelerator design alone, demanding joint optimization in a unified framework. Sachini Wickramasinghe, Tian Ye 0002, Cauligi S. Raghavendra, Viktor Prasanna 0001 |
FPGA | 2 |
| 2025 | FAST: FPGA Acceleration of Fully Homomorphic Encryption with Efficient BootstrappingabstractBootstrapping is a critical operation in Fully Homomorphic Encryption (FHE) for privacy-preserving computation. Due to its significant computational overhead, accelerating bootstrapping is crucial for practical FHE applications involving deep evaluation circuits. In this paper, we introduce FAST, an FPGA-based accelerator for efficient FHE bootstrapping. We propose novel datapath optimizations for two key operations in bootstrapping: homomorphic linear transformation (HLT) and polynomial evaluation. Our memory-efficient datapath designed for HLT significantly reduces off-chip ciphertext access. We also speed up the polynomial evaluation process by reducing the number of required HE operations. We conduct an in-depth analysis of the Advanced Bootstrapping Algorithm (ABA) and highlight its computational advantages. FAST is the first accelerator to support ABA, demonstrating significant speedup for bootstrapping. In addition, we develop a novel versatile permutation circuit to handle diverse permutation patterns in FHE, achieving high throughput and efficient resource utilization. Compared with the state-of-the-art (SOTA) GPU and FPGA designs, FAST achieves 8.84× and 5.89× speedups for bootstrapping, respectively. As illustrative examples of deep FHE applications, we show that FAST delivers over 20× speedup for logistic regression training compared with the SOTA GPU implementation and outperforms the SOTA FPGA design by 1.43× for ResNet-20 inference. Zhihan Xu, Tian Ye 0002, Rajgopal Kannan, Viktor Prasanna 0001 |
FPGA | 2 |
| 2022 | End-to-End Acceleration of Homomorphic Encrypted CNN Inference on FPGAsabstractHomomorphic Encryption is a promising approach to perform secure inference on Machine Learning models such as CNNs by allowing cloud servers to perform computations on encrypted data directly. However, CNN inference over encrypted images has high computational complexity. Prior works propose accelerators for individual HE primitives on FPGAs. In this work, we focus on an integrated design for end-to-end acceleration of inference on encrypted data. We develop parameterized IP cores for HE primitives and CNN layers. To understand the tradeoffs between various parameters such as hardware resources and performance and optimize the overall performance, we develop a parameterized performance model to evaluate the resource consumption and latency of the complete design. The performance model allows design space exploration to identify the optimal architectural parameters of the accelerator for a given FPGA, CNN model, and security requirements. We implement our design on a Xilinx VU13P FPGA and compare its performance with software implementation on a state-of-the-art server with a multi-core CPU. Our implementation for a widely studied 8-layer CNN inference for a batch of 8K images achieves average inference time of 38.8 ms per image, which is 4.1× improvement over the software baseline on the state-of-the-art server. Tian Ye 0002, Rajgopal Kannan, Viktor Prasanna 0001 |
FPGA | 1 |
| 2022 | Estimating the Impact of Communication Schemes for Distributed Graph ProcessingabstractExtreme scale graph analytics is imperative for several real-world Big Data applications with the underlying graph structure containing millions or billions of vertices and edges. Since such huge graphs cannot fit into the memory of a single computer, distributed processing of the graph is required. Several frameworks have been developed for performing graph processing on distributed systems. The frameworks focus primarily on choosing the right computation model and the partitioning scheme under the assumption that such design choices will automatically reduce the communication overheads. For any computational model and partitioning scheme, communication schemes — the data to be communicated and the virtual interconnection network among the nodes — have significant impact on the performance. To analyze this impact, in this work, we identify widely used communication schemes and estimate their performance. Analyzing the trade-offs between the number of compute nodes and communication costs of various schemes on a distributed platform by brute force experimentation can be prohibitively expensive. Thus, our performance estimation models provide an economic way to perform the analyses given the partitions and the communication scheme as input. We validate our model on a local HPC cluster as well as the cloud hosted NSF Chameleon cluster. Using our estimates as well as the actual measurements, we compare the communication schemes and provide conditions under which one scheme should be preferred over the others. Tian Ye 0002, Sanmukh R. Kuppannagari, César A. F. De Rose, Sasindu Wijeratne, Rajgopal Kannan, Viktor Prasanna 0001 |
ISPDC | 1 |
| 2021 | Performance Modeling and FPGA Acceleration of Homomorphic Encrypted ConvolutionabstractPrivacy of data is a critical concern when applying Machine Learning (ML) techniques to domains with sensitive data. Homomorphic Encryption (HE), by enabling computations on encrypted data, has emerged as a promising approach to perform inference on ML models such as Convolution Neural Network (CNN) in a privacy preserving manner. A significant portion of the total inference latency is in performing convolution over homomorphic encrypted data (HE-Convolution). For performing convolution over plaintext data, low latency accelerator designs have been proposed using algorithms such as im2col, frequency domain convolution, etc. However, developing accelerators for the HE versions of these algorithms is non-trivial. In this work, we develop a unified FPGA design that enables low latency execution of both im2col and frequency domain HE-Convolution. To enable selection of the efficient algorithm for each convolution layer of a CNN, we develop a performance model that takes the parameters about the encryption and convolution layer as input and outputs the computation and resource requirements of the two algorithms for that layer. We use the performance model to select convolution algorithm for each layer of ResNet-50 and obtain the first low latency batch-1 inference accelerator for CNN inference with HE-Convolution targeting FPGAs using HLS. We compare our design against prior techniques on CPUs and show that our accelerator achieves speedups in the range of $3.4\times\sim 6.7\times$ in latency. Tian Ye 0002, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FPL | 1 |