EDBT 2026 Demo / reviewers in the wild / expert
Nazim Altar Koca
dblp:352/9876
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accuracy-preserving Layer Normalization Approximations for Efficient Transformer Hardware AcceleratorsabstractTransformers have emerged as leading models for solving natural language processing (NLP) problems. Layer normalization (LN) has been found to be a throughput and latency bottleneck for Transformer networks. The LN datapath involves sequential processing that is dependent on the input data, and GPU and CPU unfriendly nonlinear square root and reciprocal operations. These make pipelining and hardware simplification challenging without compromising the accuracy of pretrained models. In this paper, we propose a hardware-efficient and accuracy-preserving design for LN approximation. The LN core can be directly plugged into pretrained Transformer networks to accelerate NLP tasks without fine tuning. The FPGA implementation of our design outperforms the state-of-the-art FPGA implementation of LN, with 3.5 × higher throughput per LUT and 1.3 × higher throughput per DSP. Our approximation introduces negligible average and worst-case accuracy drops of only 0.28% and 1.09%, respectively on the main language benchmark tasks. Additionally, a better pipelined design with pairwise variance calculation is proposed to reduce the number of iterations for LN, resulting in a further 27% reduction in latency with a slight increase in hardware resource requirement. Nazim Altar Koca, Anh-Tuan Do, Chip-Hong Chang |
ISCAS | 1 |
| 2024 | Efficient Fast Additive Homomorphic Encryption Cryptoprocessor for Privacy-Preserving Federated Learning AggregationabstractPrivacy leakage is a critical concern of collaboratively training a large-scale deep learning model from multiple clients. To protect the local data, homomorphic encryption (e.g., Paillier) could be utilized for data aggregation on the central server. Nevertheless, even with CPU-optimized libraries or FPGA-based accelerators, the computing power and throughput limitations remain a stumbling block for practical deployment of Paillier scheme. In this paper, we present an efficient and high-throughput cryptoprocessor based on a recently introduced Fast Additive Homomorphic Encryption (FAHE) algorithm. For encryption, we incorporate the asymmetric decomposition, time multiplexing resource reuse and hard-macro based wide-bus logic operations to efficiently map the large (>40 kbits) integer multiplications for low latency FPGA implementation. For decryption, we propose a table lookup method for rapid modular reduction by leveraging the relative short modulus size of FAHE. The single large precomputed lookup table is carefully partitioned into multiple subtables and deployed in dual-port RAMs to enable resource-efficient parallel computation. The FAHE cryptoprocessor is implemented on a Xilinx ZCU102 FPGA board for performance evaluation and comparison. The results show that the throughput of our design is 354 × to 404 × higher than the state-of-the-art Paillier accelerators. Compared to the FAHE software implementation, the latency of our proposed design is 14.95 × and 11.42 × lower for encryption and decryption, respectively. Wenye Liu, Nazim Altar Koca, Chip-Hong Chang |
DATE | 2 |
| 2024 | Exploring Error Correction Circuits on RISC-V based Systems for Space ApplicationsabstractRISC-V systems are becoming increasingly adopted in space applications. SRAM (Static Random-Access Memory) data memory is a critical component that occupies a large portion of the processor peripheral system. SRAM is vulnerable to single-event upsets (SEUs). Existing studies mainly considered Hamming error correction codes (ECCs) for memory protection in RISC-V processor. In this paper, we explore different ECCs as well as the triple modular redundancy (TMR) as solutions to mitigate SEU effects on SRAMs for RISC-V system. To overcome the area and power overhead of TMR with higher fault tolerance than ECCs, we propose a dual modular redundancy (DMR) with ECC memory protection scheme. We conduct a comprehensive error analysis to evaluate different fault-tolerant designs under various radiation attack scenarios, utilizing real-world data to devise the fault-injection campaign. An efficient scheduling for the self-refresh operation of SRAMs is proposed to prevent the error accumulation. The proposed DMR with ECC design reduces the power and area overhead of TMR by 28% and 11% respectively and improve the error resilience of the SRAM significantly compared with Hsiao and Hamming ECC schemes. Nazim Altar Koca, Chip-Hong Chang, Anh-Tuan Do, Vishnu P. Nambiar |
ISCAS | 1 |
| 2023 | Hardware-efficient Softmax Approximation for Self-Attention NetworksabstractSelf-attention networks such as Transformer have become state-of-the-art models for natural language processing (NLP) problems. Softmax function, which serves as a normalizer to produce attention scores, turns out to be a severe throughput and latency bottleneck of a Transformer network. Softmax datapath consists of data-dependent sequential nonlinear exponentiation and division operations, which are not amenable to pipelining and parallelism, nor can they be directly linearized for pretrained models without substantial accuracy drop. In this paper, we proposed a hardware efficient Softmax approximation which can be used as a direct plug-in substitution into pretrained transformer network to accelerate NLP tasks without compromising its accuracy. Experiment results on FPGA implementation show that our design outperforms vanilla Softmax designed using Xilinx IPs with 15x less LUTs, 55x less registers and 23x lower latency at similar clock frequency and less than 1% accuracy drop on main language benchmark tasks. We also propose a pruning method to reduce the input entropy of Softmax for NLP problems with high number of inputs. It was validated on CoLA task to achieve a further 25% reduction of latency. Nazim Altar Koca, Anh-Tuan Do, Chip-Hong Chang |
ISCAS | 1 |