Syed Mohsin Abbas

dblp:149/0336 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0002-5719-3654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Headsail: One-Year Tape-Out of a 25-mm2 Linux-Capable RISC-V MPSoC
abstract
The Internet-of-Things (IoT) devices feature a broad range of power, memory, and performance requirements. Ultralow-power, low-performance controllers are at one end of the spectrum, while high-performance, power-intensive systems-on-chip (SoCs) are at the other. Heterogeneous and specialized multiprocessor SoC (MPSoC) architectures have emerged as the most effective paradigm for delivering high performance and energy efficiency across a wide range of application workloads. This work introducesHeadsail, an MPSoC application-specific integrated circuit (ASIC) designed by SoC Hub at Tampere University, Finland.Headsailfeatures a 512-KiB primary data buffer, 128 KiB of shared on-chip SRAM, seven CPU cores (including four CVA6 64-bit RISC-V processors), a low power-DDR2 (LP-DDR2) memory controller, two unique chip-to-chip (C2C) interfaces, and shared peripherals.Headsailhas been successfully implemented using a TSMC 22-nm low-power CMOS technology. Testing results show that the samples can support a maximum operating frequency of 1 GHz and achieve a peak performance of 1100 giga operations per second (GOPS), with an implementation area of 25 mm2and a power-consumption range of 64 mW–1.5 W.
Matti Käyrä, Thomas Szymkowiak, Antti Rautakoura, Antti Nurmi, Kari Hepola, Henri Lunnikivi, Toni Jääskeläinen, Abdesattar Kalache, Petteri Toivanen, Roope Keskinen, Andreas Stergiopoulos, Väinö-Waltteri Granat, Arto Oinonen, Joonas Multanen, Pekka Jääskeläinen, Karri Palovuori, Timo Hämäläinen 0001, Syed Mohsin Abbas
IEEE Trans. Very Large Scale Integr. Syst.19
2025 Reconfigurable Image Acquisition and Processing Subsystem for MPSoCs
abstract
Modern commercial Multi-Processor Systems-Onchips (MPSoCs), targeted for smartphones as well as embedded Artificial Intelligence (AI) and Internet-Of-Things (IoT) applications, require camera interfaces that support the CSI-2 standard and are reconfigurable to support various image data types and throughput requirements. However, existing State-Of-the-Art (SoA) implementations of reconfigurable camera interfaces are proprietary and closed-source.This work proposes, to the best of our knowledge, the first open-source implementation of a reconfigurable image acquisition and processing platform for MPSoC integration. The proposed design is implemented and verified on both FPGA (Zynq ZCU104 FPGA) and ASIC (TSMC 22nm) platforms. The hardware implementation results depict that the proposed design can support a throughput of 9.6 Gbps for streaming 4K resolution images.
Antti Nurmi, Syed Mohsin Abbas
ISCAS3
2025 Improved Step-GRAND: Low-Latency Soft-Input Guessing Random Additive Noise Decoding
abstract
The ultrareliable low-latency communication (URLLC) application scenario requires the adoption of short linear block codes to satisfy the low-latency requirements. Guessing random additive noise decoding (GRAND) is a prominent universal decoding solution for short linear block codes that lends itself to efficient hardware implementations. GRAND-based hardware implementations generally offer reduced average decoding latency but their high worst-case (W.C.) latency renders them unsuitable for deployment in mission-critical applications. This article presents an improved version of step-GRAND, a soft-input variant of GRAND that features a novel test error pattern (TEP) generating approach. A novel very large-scale integration (VLSI) architecture is developed for the execution of the improved step-GRAND algorithm with reduced W.C. decoding latency. Application specific integrated circuit (ASIC) implementation results, employing low-power (LP) TSMC 65-nm CMOS technology, demonstrate that the proposed improved step-GRAND can achieve an average decoding latency as low as 10 ns for decoding a$(128,105)$linear block code at a target frame error rate (FER) of$10^{-7}$, while the W.C. decoding latency can reach$300~\text {ns}\sim 1~\mu \text { s}$depending on the parametric settings. Compared with the previously proposed baseline soft-input ordered reliability bits GRAND (ORBGRAND) hardware implementation with similar decoding performance at target FER of$10^{-7}$, the improved step-GRAND hardware achieves$7 \times \sim 17\times $reduction in W.C. latency,$7\times $reduction in power consumption, and$37 \times \sim 66\times $higher area efficiency in the W.C. scenario. Furthermore, the proposed hardware can achieve an average throughput of up to 10.5 Gb/s and a W.C. throughput of$102\sim 350$Mb/s.
Syed Mohsin Abbas, Marwan Jalaleddine, Chi-Ying Tsui, Warren J. Gross
IEEE Trans. Very Large Scale Integr. Syst.1
2023 List-GRAND: A Practical Way to Achieve Maximum Likelihood Decoding
abstract
Guessing random additive noise decoding (GRAND) is a recently proposed universal maximum likelihood (ML) decoder for short-length and high-rate linear block codes. Soft-GRAND (SGRAND) is a prominent soft-input GRAND variant, outperforming the other GRAND variants in decoding performance; nevertheless, SGRAND is not suitable for parallel hardware implementation. Ordered Reliability Bits-GRAND (ORBGRAND) is another soft-input GRAND variant that is suitable for parallel hardware implementation; however, it has lower decoding performance than SGRAND. In this article, we propose List-GRAND (LGRAND), a technique for enhancing the decoding performance of ORBGRAND to match the ML decoding performance of SGRAND. Numerical simulation results show that LGRAND enhances ORBGRAND’s decoding performance by 0.5–0.75 dB for channel codes of various classes at a target frame error rate (FER) of 10−7. For linear block codes of length 127/128 and different code rates, LGRAND’s VLSI implementation can achieve an average information throughput of 47.27–51.36 Gb/s. In comparison to ORBGRAND’s VLSI implementation, the proposed LGRAND hardware has a 4.84% area overhead.
Syed Mohsin Abbas, Marwan Jalaleddine, Warren J. Gross
IEEE Trans. Very Large Scale Integr. Syst.1
2022 High-Throughput and Energy-Efficient VLSI Architecture for Ordered Reliability Bits GRAND
abstract
Ultrareliable low-latency communication (URLLC), a major 5G new-radio (NR) use case, is the key enabler for applications with strict reliability and latency requirements. These applications necessitate the use of short-length and high-rate channel codes. Guessing random additive noise decoding (GRAND) is a recently proposed maximum likelihood (ML) decoding technique for these short-length and high-rate codes. Rather than decoding the received vector, GRAND tries to infer the noise that corrupted the transmitted codeword during transmission through the communication channel. As a result, GRAND can decode any code, structured or unstructured. GRAND has hard-input as well as soft-input variants. Among these variants, ordered reliability bits GRAND (ORBGRAND) is a soft-input variant that outperforms hard-input GRAND and is suitable for parallel hardware implementation. This work reports the first hardware architecture for ORBGRAND, which achieves an average throughput of up to 42.5 Gb/s for a code length of 128 at a target frame error rate (FER) of 10−7. Furthermore, the proposed hardware can be used to decode any code as long as the length and rate constraints are met. In comparison to the GRAND with ABandonment (GRANDAB), a hard-input variant of GRAND, the proposed architecture enhances decoding performance by at least 2 dB. When compared to the state-of-the-art fast dynamic successive cancellation flip decoder (Fast-DSCF) using a 5G polar code (PC) (128, 105), the proposed ORBGRAND VLSI implementation has$49\times $higher average throughput,$32\times $times more energy efficiency, and$5\times $more area efficiency while maintaining similar decoding performance.
Syed Mohsin Abbas, Thibaud Tonnellier, Furkan Ercan, Marwan Jalaleddine, Warren J. Gross
IEEE Trans. Very Large Scale Integr. Syst.1
2021 High-Throughput VLSI Architecture for Soft-Decision Decoding with ORBGRAND
abstract
Guessing Random Additive Noise Decoding (GRAND) is a recently proposed approximate Maximum Likelihood (ML) decoding technique that can decode any linear error-correcting block code. Ordered Reliability Bits GRAND (ORBGRAND) is a powerful variant of GRAND, which outperforms the original GRAND technique by generating error patterns in a specific order. Moreover, their simplicity at the algorithm level renders GRAND family a desirable candidate for applications that demand very high throughput. This work reports the first-ever hardware architecture for ORBGRAND, which achieves an average throughput of up to 42.5 Gbps for a code length of 128 at an SNR of 10 dB. Moreover, the proposed hardware can be used to decode any code provided the length and rate constraints. Compared to the state-of-the-art fast dynamic successive cancellation flip decoder (Fast-DSCF) using a 5G polar (128,105) code, the proposed VLSI implementation has 49× more average throughput while maintaining similar decoding performance.
Syed Mohsin Abbas, Thibaud Tonnellier, Furkan Ercan, Marwan Jalaleddine, Warren J. Gross
ICASSP1
2017 Concatenated LDPC-polar codes decoding through belief propagation
abstract
Owing to their capacity-achieving performance and low encoding and decoding complexity, polar codes have drawn much research interests recently. Successive cancellation decoding (SCD) and belief propagation decoding (BPD) are two common approaches for decoding polar codes. SCD is sequential in nature while BPD can run in parallel. Thus BPD is more attractive for low latency applications. However BPD has some performance degradation at higher SNR when compared with SCD. Concatenating LDPC with Polar codes is one popular approach to enhance the performance of BPD, where a short LDPC code is used as an outer code and Polar code is used as an inner code. In this work we propose a new way to construct concatenated LDPC-Polar code, which not only outperforms conventional BPD and existing concatenated LDPC-Polar code but also shows a performance improvement of 0.5 dB at higher SNR regime when compared with SCD.
Syed Mohsin Abbas, YouZhe Fan, Chi-Ying Tsui
ISCAS1
2017 High-Throughput and Energy-Efficient Belief Propagation Polar Code Decoder
abstract
Owing to their capacity-achieving performance and low encoding and decoding complexity, polar codes have received significant attention recently. Successive cancellation decoding (SCD) and belief propagation decoding (BPD) are two popular approaches for decoding polar codes. SCD, despite having less computational complexity when compared with BPD, suffers from long latency due to the serial nature of the SC algorithm. BPD, on the other hand, is parallel in nature and is more attractive for low-latency applications. However, due to the iterative nature of BPD, the required latency and energy dissipation increase linearly with the number of iterations. In this paper, we propose a novel scheme based on subfactor-graph freezing to reduce the average number of computations as well as the average number of iterations required by BPD, which directly translates into lower latency and energy dissipation. Simulation results show that the proposed scheme has no performance degradation and achieves significant reduction in computation complexity over the existing methods. Moreover, the hardware architecture for the proposed scheme is developed and compared with the state-of-the-art BPD implementations for (1024, 512) polar codes. A decoding throughput of 13.9 Gb/s is achieved along with a 60%-73% improvement in energy reduction and two times increase in hardware efficiency when compared with the existing BPD implementations.
Syed Mohsin Abbas, YouZhe Fan, Chi-Ying Tsui
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Low-latency approximate matrix inversion for high-throughput linear pre-coders in massive MIMO
abstract
This work presents a high-throughput and low-latency matrix inversion design for a linear pre-coder for massive MIMO systems. Because of the large number of Base Station (BS) antennas as well as the multiple User Terminals (UTs) served in a massive MIMO system, the channel matrix dimensions become larger. Inversions of such large matrices using direct inversion methods, such as used in linear pre-coders like Zero Forcing (ZF), would entail prohibitive complexity. For avoiding such complexity, Neumann series based approximate inversion has been suggested for linear pre-coders in massive MIMO systems. However the performance, complexity and convergence speed of the Neumann series approach depends very much on the initial approximation of the inverse used as a starting point. In this work, we present a novel initial approximation for the Neumann series which facilitates the parallel computation of the inverse and hence results in lower latency for inversion as well as better accuracy when compared to the previous approaches. A VLSI architecture of the proposed method is implemented for the inversion of a 16 × 16 matrix, in TSMC 65nm technology. A throughput of 0.54M to 15M matrix inversion per sec is achieved at a clock frequency of 460MHz with a 117K gate count.
Syed Mohsin Abbas, Chi-Ying Tsui
VLSI-SoC1
2015 FSNoC: A Flit-Level Speedup Scheme for Network on-Chips Using Self-Reconfigurable Bidirectional Channels
abstract
In this paper, we explore optimizing the bandwidth utilization of the network-on-chips (NoCs). We propose a flit-level speedup scheme to improve the NoC performance using self-reconfigurable bidirectional channels. For the NoC intrarouter bandwidth, in addition to allowing flits from different packets to use the idle internal bandwidth of the crossbar, our proposed flit-level speedup scheme also allows flits within the same packet to be transmitted simultaneously. For interrouter channels, a distributed channel configuration scheme is developed to dynamically change the link directions. In this way, the effective bandwidth between two routers can change adaptively depending on the run time network traffic. We present the implementation of the proposed flit-level speedup NoC on a 2-D mesh. An input buffer architecture, which supports reading and writing two flits from the same virtual channel at the same time, is proposed. The switch allocator is also designed to support flit-level parallel arbitration. Extensive simulations on both the synthetic traffic and real applications show performance improvement in throughput and latency over the existing architectures using bidirectional channels.
Zhiliang Qian, Syed Mohsin Abbas, Chi-Ying Tsui
IEEE Trans. Very Large Scale Integr. Syst.2
2014 An Efficient Multiple Cell Upsets Tolerant Content-Addressable Memory
abstract
Multiple cell upsets (MCUs) become more and more problematic as the size of technology reaches or goes below 65 nm. The percentage of MCUs is reported significantly larger than that of single cell upsets (SCUs) in 20 nm technology. In SRAM and DRAM, MCUs are tackled by incorporating single-error correcting double-error detecting (SEC-DED) code and interleaved data columns. However, in content-addressable memory (CAM), column interleaving is not practically possible. A novel error correction code (ECC) scheme is proposed in this paper that will cater for ever-increasing MCUs. This work demonstrated that m parity bits are sufficient to cater for up to m-bit MCUs, with an understanding of the physical grouping of MCUs. The results showed that the proposed scheme requires 85% fewer parity bits compared to traditional Hamming distance based schemes.
Syed Mohsin Abbas, Soonyoung Lee, Sanghyeon Baeg, Sungju Park
IEEE Trans. Computers1