EDBT 2026 Demo / reviewers in the wild / expert
Infall Syafalni
dblp:119/3703
· DBLP profile ↗
12ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-9922-5688ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Base Conversion RNS using Hybrid Barret-Karatsuba-Based Modulus Prime Cores for Homomorphic MultiplicationabstractResidue Number Systems (RNS) offer a promising approach to enhancing the computational efficiency of Fully Homomorphic Encryption (FHE). This paper proposes an efficient base conversion method for RNS, leveraging parallel Processing Unit (PU) cores, where each PU core operates under different moduli, along with a Karatsuba-based Barrett modular multiplier to accelerate modular arithmetic in homomorphic encryption schemes. Our approach addresses the computational bottleneck of base conversions in FHE by utilizing optimized parallelism and efficient multiplication methods. Experimental results show that, with nearly the same latency for the computing unit, the proposed design can achieve higher throughput for implementing base conversion (BConv). The performance is further evaluated with different parallel core configurations (2, 3, and 6 cores), and these scenarios are compared against a single-core implementation. The design with 6 parallel cores achieves more than 5Gbps, demonstrating the significant benefit of increased parallelism for BConv in FHE operations. Muhammad Ogin Hasanuddin, Rafael Aditya Cahyo W, Infall Syafalni, Nana Sutisna, Hanho Lee, Trio Adiono |
ISCAS | 3 |
| 2025 | Reconfigurable Radix-2/4/8 of Unified 2N Points NTT-INTT for Homomorphic EncryptionabstractFully homomorphic encryption is an encryption that allows direct mathematical operation on its ciphertext. One of the challenges of implementing FHE is the costly operation of polynomial multiplication over an integer ring. One the of candidates to solve this problem is using Number Theoretic Transform (NTT) which is an FFT evaluated on integer ring. In this paper, we propose a unified reconfigurable radix for the NTT-INTT operation to balance between the flexibility of radix-2 to process 2Npoints of polynomial and low computational cycle of higher radix. Experimental results indicate a significant improvement up to 9.06× in throughput per slice when compared to recent alternative designs handling large polynomial inputs (≥ 12), although it remains inferior to architectures optimized for smaller inputs. Nonetheless, this reconfigurable design provides a viable pathway to enhance overall throughput by capitalizing on parallel processing, albeit at the cost of an increased number of slices. Infall Syafalni, Nicholas Teffandi, Fauzan Ibrahim, Indira Pramudita, Nana Sutisna, Trio Adiono |
ISCAS | 1 |
| 2025 | Design and implementation of a real-time SDR-FPGA-based versatile radio frequency OFDM systemabstractOrthogonal frequency-division multiplexing (OFDM) is widely used in modern wireless, telecommunications, and broadcasting standards due to its multipath tolerance and high spectral efficiency. It enables high-data-rate transmission while mitigating inter-symbol interference by extending symbol duration. Many OFDM system designs have been developed, but most of them have put a halt on MATLAB simulation or register-transfer level (RTL) implementation only, lacking real-time testing and TCP/IP integration -- a feature that is needed in many wireless communication devices nowadays. This paper presents a full design integration and implementation of an OFDM transceiver system on the ADRV9361-Z7035 Software-Defined Radio (SDR) FPGA development board, which integrates the AD9361 RF Agile Transceiver with the Xilinx Z7035 Zynq-7000 SoC. The design process includes creating MATLAB models for each OFDM block, developing system objects for data transmission, and verifying the system using an SDR MATLAB model. The implementation involves constructing an RTL model and developing Linux kernel drivers for local processing, independent of MATLAB and Simulink. Transmitter power performance measurements across various modulation and coding schemes indicate optimal performance within the 700 MHz to 900 MHz range, with a maximum power of 13.6 dBm using BPSK FEC 1/2. Receiver sensitivity measurements show that most modulation and coding schemes (MCS) have met their standard sensitivity levels. Throughput performances for each MCS were compared to theoretical values and benchmarked against the Silvus Streamcaster 4200, highlighting strong performance across the MCS, with minor areas for improvement in 16-QAM configurations. The integrated OFDM system has been able to be used for multimedia applications, such as local video streaming and internet browsing. Trio Adiono, Michael Jonathan, Erwin Setiawan, Syifaul Fuada, Nana Sutisna, Rahmat Mulyawan, Infall Syafalni |
Integr. | 7 |
| 2025 | Parallel and Pipeline Multicore OFDM Baseband Processor for Wideband UOWC SystemsabstractOptical wireless communication (OWC) has been considered a promising solution for overcoming the spectrum limitations of RF in wireless communication. However, studies on OWC systems, particularly on actual real-time prototyping, are limited and only focus on simulation and experiments using waveform generators and oscilloscopes. Meanwhile, the data rate requirements for optical communication systems are rapidly increasing, approaching the Gbps range. Therefore, the development of real-time prototypes with high data speeds is essential and requires super sample rates (SSR), necessitating signal processing blocks capable of processing multiple samples in parallel within a single clock cycle while operating at very high clock frequencies. To address these requirements, this work proposes a parallel implementation of an orthogonal frequency division multiplexing (OFDM) baseband processor on an FPGA for OWC systems. The proposed system operates with a bandwidth of 320 MHz, achieving a maximum raw data rate of 720 Mbps at 16-QAM 3/4 modulation. This performance is comparable to that of the IEEE 802.11be WiFi 7 standard. The proposed system is applicable not only to underwater OWC but also for indoor OWC (e.g. Li-Fi), broadening its potential use cases in next-generation wireless communication networks. Trio Adiono, Erwin Setiawan, Michael Jonathan, Nana Sutisna, Rahmat Mulyawan, Infall Syafalni, Wasiu O. Popoola |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | FPGA Implementation of SFO for OFDM-based Network Enabled Li-Fi SystemabstractLight Fidelity (Li-Fi) is the complementary technology to the RF based communication in 6G. It uses optical spectrum which is completely free and unregulated. Sampling frequency offset (SFO) is an issue related to the mismatch between TX and RX clock oscillators. However, this issue is often assumed to be ideal in many Li-Fi research. In this work, we aim to address this issue by proposing an FPGA implementation of SFO compensator module for OFDM system. We optimize an SFO compensation method so that it can be implemented in FPGA efficiently in terms of resource usage. We evaluate the module by integrating it with the OFDM processor, Linux TCP/IP stack, and tested with TCP data. Real-time experiment results demonstrate that with our SFO compensator, the EVM of our Li-Fi system can be decreased up to 6.7% and 1.4% for QPSK and 16-QAM, respectively. Frame lost is decreased by 22.42% and TCP data rate is increased by 8.25x. Trio Adiono, Erwin Setiawan, Michael Jonathan, Rahmat Mulyawan, Nana Sutisna, Infall Syafalni, Wasiu O. Popoola |
ISCAS | 6 |
| 2024 | Low-Complexity and High-Throughput Number Theoretic Transform Architecture for Polynomial Multiplication in Homomorphic EncryptionabstractThe computationally intensive of polynomial ring multiplication in homomorphic encryption (HE) schemes demand an optimized hardware accelerator design, specifically targeted for edge devices. The state-of-the-art algorithm for calculating polynomial ring multiplication is the Number Theoretic Transform (NTT) which is capable of reducing the number of operations needed from the school book multiplication algorithm. In this work, we propose a novel NTT hardware accelerator design that is suitable for use in implementing a partially homomorphic encryption scheme with a 192-bit level of security. The main novelty presented in this paper is the realization of an area-efficient higher radix NTT and inverse NTT (INTT) accelerator which is achieved via mathematical optimization and a point-based approach to calculating NTT. Compared to previous works with similar throughput per slice metric, the design is able to deliver 5.16× higher throughput for a 2.46× increase in slice count. Nana Sutisna, Elkhan J. Brillianshah, Infall Syafalni, Muhammad Ogin Hasanuddin, Trio Adiono, Tutun Juhana |
ISCAS | 3 |
| 2024 | Adaptive Traffic Light Controller with Reinforcement Learning for Reducing Traffic CongestionabstractReinforcement Learning is a machine learning method that has been adopted in smart traffic lights to reduce traffic congestion effectively. The algorithm is employed as an agent, like a traffic police officer, that can manage traffic jams with the intelligent capability. In this paper, we revisit the basic Q- Learning method applied to adaptive traffic lights by proposing systematic state-action definition, formulating reward policy, and evaluating the performance. The proposed method basically aims to adjust the duration of green lights at each intersection according to the level of congestion at that intersection dynamically, hence, can adapt with traffic condition. Simulation results confirm that the propose method works effectively in managing traffic congestion by maintaining vehicle queue length almost constant at around 40 vehicles queue, offering an improvement from the conventional method that give growth in queue more than 60 vehicles. In the future, we plan to implement the proposed work in an embedded hardware systems to address the real-time performance requirement. Anisa Dwizarah, Jalu Reswara, Infall Syafalni, Nana Sutisna, Trio Adiono |
TENCON | 3 |
| 2023 | Fast and Scalable Multicore YOLOv3-Tiny Accelerator Using Input Stationary Systolic ArchitectureabstractThis article proposes a scalable accelerator for deep learning (DL) implementation on edge computing, which is often limited by power, storage, and computation speed. The accelerator is based on systolic array cores with 126 processing elements (PEs) and optimized for YOLOv3-Tiny with$448\times448$input images. Two multicast (MC) network architectures, feature map multicasting and weight multicasting, are introduced to control data stream distribution within the multicores. Results show that the proposed weight multicast (W-MC) systems outperformed the feature map multicast (FMAP-MC) systems in multicore scenarios, with up to$2.23\times $frame rates per second (FPS). The 4-core W-MC system achieved the best efficiency with an overall frame rate of 13.73 FPS/W and an overall throughput of 35.83 GOPS/W. The 8-core W-MC system delivered the best performance, with a frame rate of 38.50 FPS after normalization to the standard YOLOv3-Tiny network. The proposed accelerator offers better computational efficiency and greater accelerator utilization in real-world inference scenarios, compared to previous state-of-the-art works. Trio Adiono, Rhesa Muhammad Ramadhan, Nana Sutisna, Infall Syafalni, Rahmat Mulyawan, Chang Hong Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | A Method to Detect Bit Flips in a Soft-Error Resilient TCAMabstractTernary content addressable memories (TCAMs) are special memories which are widely used in high-speed network applications such as routers, firewalls, and network address translators. In high-reliability network applications such as aerospace and defense systems, soft-error tolerant TCAMs are indispensable to prevent data corruption or faults caused by radiation. This paper shows a novel way in generating keys to cover the correct match. It proposes a novel soft-error tolerant TCAM for multiple-bit-flip errors using partial don't-care keys (X-keys). First, this paper observes the case of single-bit-flip errors by X-TCAM. Second, it extends the X-TCAM to the case of multiple-bit-flip errors called KX-TCAM, where K stands for the maximum number of errors k. KX-TCAM corrects up to k-bit-flip errors and enhances the tolerance of the TCAM against soft errors, where k is the maximum number of bit flips in a word of a TCAM. KX-TCAM consists of a TCAM, a preprocessed don't-care-bit index look-up memory (X look-up), and a backup error checking and correction (ECC)-SRAM. First, KX-TCAM randomly selects a search key. After that, KX-TCAM detects multiple-bit-flip errors by the generated X-keys using the X look-up. If the keys match the different locations, then a soft error is suspected and KX-TCAM refreshes the TCAM words by using the backup ECC-SRAM. Experimental results show that the soft-error tolerance capability of KX-TCAM significantly outperforms existing state-of-the-art schemes. Moreover, the hardware overhead of KX-TCAM is small due to the use of a single TCAM. KX-TCAM can be easily implemented and is useful for fault-tolerant packet classifiers. Infall Syafalni, Tsutomu Sasao, Xiaoqing Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Assertion-based verification of industrial WLAN systemabstractFor the last decades, the advancement of system on chips complexity and size are driving the verification process of the digital designs to be much more complicated and time consuming. Moreover, nowadays, the verification process takes up to 80% of overall design development time. This paper emphasizes the importance of automatic testbench generator for industrial wireless local area network (iWLAN) for factory automation system (FA). The proposed generated testbench is written in SystemVerilog to represent the functions and to reduce the code size compared to traditional Verilog testbench. It is also supported by assertion-based method to simplify the code and to improve the observability in SoC verification. Experimental results show that our proposed method produces smaller code size and can improve the efficiency of the verification of SoC designs. Moreover, we shows the verified iWLAN system data transfer. Infall Syafalni, Nico Surantha, Duc Khai Lam, Nana Sutisna, Yuhei Nagao, Katsuhiko Wakasugi, Yang Tongxin, Hiroshi Ochi, Taadaki Tsuchiya |
ISCAS | 1 |
| 2015 | A soft-error tolerant TCAM using partial don't-care keysabstractThis paper proposes a novel soft-error tolerant TCAM using partial don't-care keys (X-keys), namely TX, which significantly enhances the tolerance of the TCAM against soft errors. Experimental results show that the soft-error tolerance of the TX outperforms existing schemes. Moreover, the overhead of the TX is very small. Infall Syafalni, Tsutomu Sasao, Xiaoqing Wen, Stefan Holst, Kohei Miyase |
ETS | 1 |
| 2013 | A TCAM generator for packet classificationabstractIn the internet, packets are classified by source and destination addresses and ports, as well as protocol type. Ternary content addressable memories (TCAMs) are often used to perform this operation. This paper shows a method to reduce the number of words in TCAM for multi-field classification functions. We use head-tail expressions to represent a multi-field classification rule. Furthermore, we present an O(r2)-algorithm, called MFHT, to generate simplified TCAMs for two-field classification functions, where r is the number of rules. Experimental results show that MFHT achieves a 58% reduction of words for random rules and a 52% reduction of words for ACL and FW rules. Moreover, MFHT is fast and useful for simplifying TCAM for packet classification. Infall Syafalni, Tsutomu Sasao |
ICCD | 1 |