EDBT 2026 Demo / reviewers in the wild / expert
Fasih Ud Din Farrukh
dblp:244/7533
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-0178-8130ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HyperGS: Efficient Real-Time 3D Gaussian Rendering Processor Through Hierarchical SortingabstractThis paper proposes HyperGS, the first complete hardware accelerator implementation for 3D Gaussian Splatting with algorithmic optimization. Through an in-depth analysis of the Gaussian rendering pipeline characteristics, we designed an efficient and practical hardware architecture with optimized resource utilization. Specifically, we proposed a parallel large-scale sorting unit to enable parallel processing of depth sorting and rasterization, saving 22% of computing time. We also designed a fully pipelined preprocessing calculation unit with low resource usage for model parameter preprocessing, generating key-value pairs. We propose an efficient rasterization unit based on factorization. The rendering pipeline has been improved by reducing preprocessing parameters and decreasing the number of external memory accesses. Experimental results show that HyperGS reduces power consumption by a factor of 84 and improves performance by a factor of 44 compared with existing NVIDIA Jetson GPUs. Using only 19.2GB/s of bandwidth, we achieved rendering speeds ranging from 20.76 to 43.38 frames per second. Our work provides a feasible solution for efficient real-time 3D Gaussian rendering on resource-constrained edge computing devices. Cheng Nian, Xiaorui Mo, Jiaying Peng, Weiyi Zhang 0002, Fasih Ud Din Farrukh, Chun Zhang 0001 |
ISCAS | 5 |
| 2025 | An 197-μJ/Frame Single-Frame Bundle Adjustment Hardware Accelerator for Mobile Visual OdometryabstractThis article presents an energy-efficient hardware accelerator for optimized bundle adjustment (BA) for mobile high-frame-rate visual odometry (VO). BA uses graph optimization techniques to optimize poses and landmarks and the applications are robot navigation, virtual reality (VR), and augmented reality (AR). Existing software implementations of BA optimization involve complex computational flows, numerical calculations, Lie group, and Lie algebra conversions. This poses challenges of slow computational speeds and high power consumption. A two-level reuse hardware architecture is proposed and implemented that efficiently updates the Jacobian matrix while reducing the field-programmable gate array (FPGA) hardware resources by 25%. A set of methodologies is proposed to quantify the errors caused by fixed-point systems during optimization. A fully pipelined architecture is implemented to increase computational speed while reducing hardware resources by 29%. This design features a parallel equation solver that improves processing speed by$2\times $compared to conventional approaches. This article employs a single-frame local BA VO on the KITTI dataset and EuRoC dataset, achieving an average translational error of 0.75% and a rotational error of$0.0028~^{\circ } $/m. The proposed hardware achieves a performance ranging from 188 to 345 frames/s in optimizing two main feature extraction methods with a maximum of 512 extracted feature points. Compared to state-of-the-art implementations, the accelerator achieved a minimum energy efficiency ratio of 11.6 mJ and$191~\mu $J on the FPGA platform and application-specific integrated circuits (ASICs) platform, respectively. These improvements underscore the potential of FPGAs to enhance VO systems’ adaptability and efficiency in complex environments. Cheng Nian, Xiaorui Mo, Weiyi Zhang 0002, Fasih Ud Din Farrukh, Yushi Guo, Chun Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | A 77.79 GOPs/W Retentive Network FPGA Inference Accelerator with Optimized WorkloadabstractRetention Network (RetNet) is a novel neural network model, with a inference complexity of O(1), considered to be the successor of the Transformer model. This work presents an inference accelerator supporting RetNet and DNN algorithms focus on dataflow optimizing, model quantization for improved parallelism and reusability, and addressing K-V cache issues in Transformers. To meet model inference storage bandwidth requirements, we used output block stationary (OBS) dataflow to explore data reusability in time and space. A hardware accelerator for the position encoding module is designed, reducing storage bandwidth and on-chip storage during Extrapolatable Position Embedding (XPOS). Compared to Transformer models of the same size architecture, this work has better inference accuracy, achieving energy efficiency ratio of 77.79 GOPs/W, a 12.5x speedup, and 2.13x energy efficiency compared to GPU implementations, which is 1.33x better than the transformer work baseline. Meanwhile, the inference process of Muti-Scale-Retention(MSR) is completed with the smallest on-chip RAM of 350KB, thereby enabling a more efficient ASIC implementation of hardware accelerators. Cheng Nian, Weiyi Zhang 0002, Fasih Ud Din Farrukh, Liting Niu, Dapeng Jiang, Chun Zhang 0001 |
IECON | 3 |
| 2023 | Hardware-Software Co-Design of Matrix-Solving for Non-Linear Optimization in SLAM SystemsabstractSimultaneous Localization and Mapping (SLAM) is one of the most important techniques for autonomous robots that enables the robot aware of its current position and the surrounding environment. There is a significant improvement in the accuracy with the advancement in SLAM algorithms. However, the computation complexity increases accordingly and the embedded processors of autonomous robots struggle to support heavy calculation. Matrix-solving contributes a major portion of calculation time and considering sub-tasks such as bundle adjustment takes over 40% of total time. Therefore, it is significant to optimize the calculations required for matrix-solving. However, previous works for matrix-solving accelerators are generalized and the specific matrix form in SLAM problems is not fully considered. This work concentrates on the dedicated software and hardware codesign of the matrix-solving task in SLAM systems and provides three solutions for different scales of matrix-solving problems in SLAM. The proposed FSFI-Cholesky and FI-Iterative method have achieved up to 120.2x speed improvement over the non-optimized Cholesky algorithm. Moreover, this work also reduces the execution time by more than 7.0x compared to the state-of-the-art design with fewer DSPs used for both dense and sparse matrices. Liting Niu, Weiyi Zhang 0002, Cheng Nian, Fei Shao, Fasih Ud Din Farrukh, Chun Zhang 0001 |
IECON | 5 |
| 2019 | A Solution to Optimize Multi-Operand Adders in CNN Architecture on FPGAabstractConvolutional Neural Network (CNN) is a very popular method in recent times to solve many computer vision tasks. However, CNN is becoming computationally intensive as time is progressing which requires a dedicated hardware for real time implementation. Graphics Processing Unit (GPU) and Field Programmable Gate Array (FPGA) are two hot choices to execute and accelerate the CNN network. FPGA has an advantage over GPU due to its flexible architecture and it can also provide high performance per unit watt of power. These benefits make FPGA a suitable candidate for CNN acceleration. However, optimization is required for FPGA based accelerator design to accommodate more computations. One of the challenges in accelerator design is to perform addition of intermediate results generated in a process of convolution. Therefore, Multi-Operand Adders (MOAs) are necessary in the accelerator design of CNN on FPGA but consume most of the area. Optimization strategy based on WALLACE tree architecture is proposed in this article to replace the typical binary adder tree in CNN accelerator design. Experimental results show an improvement in terms of area optimization and performance in comparison with the previous implementation. Fasih Ud Din Farrukh, Tuo Xie, Chun Zhang 0001, Zhihua Wang 0001 |
ISCAS | 1 |