Sanchita Saha Ray

dblp:177/2124 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0001-7921-1916ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2023 ${O(N)}$O(N) Memory-Free Hardware Architecture for Burrows-Wheeler Transform
abstract
A novel hardware architecture for Burrows-Wheeler Transform (BWT) scheme is presented. The core idea is to have a memory-free strategy that does not involve any software overhead during BWT operation. This is achieved by introducing a register-file concept and utilizing basic digital logic circuits to perform the entire BWT operation. Additionally, this is a kind of transformation scheme that does not utilize any kind of matrix during transformation, and thereby, it is free from run-time memory consumption. It efficiently handles the string terminating mechanism in the proposed design without involving any extra terminating symbol. This string terminator-free architecture eventually reduces additional operation and storage space to maintain the string, and thereby, the architecture does not necessitate any register read and write operations. This architecture exhibits efficient transformation without involving any indexing method or sorting mechanism during an inverse transformation operation. This architecture achieves$O(N)$time complexity compared to$O(N^{2})$and$O(N \log N)$as experienced by the existing state-of-the-art approaches.
Surajeet Ghosh, Sanchita Saha Ray
IEEE Trans. Computers2
2023 k-Degree Parallel Comparison-Free Hardware Sorter for Complete Sorting
abstract
This article presents a novel parallel comparison-free hardware sorting architecture that sorts$N$,$n$-bit elements consuming linear worst-case sorting latency of$\mathit {O(N)}$clock-cycles utilizing$k$-parallel clusters. A parallel cluster contains a number of identical blocks utilizing a few fundamental logic components. The architecture achieves a speed-up of${}({n}/[{\lceil {({n}/{k})}\rceil +k}])$over nonparallel architectures. The proposed sorting architecture identifies the largest element in every cycle with smaller clock periods compared to state-of-the-art comparison-free approaches and does not require any kind of preprocessing of the input data set, unlike most of the existing hardware sorters. The experimental result shows the degree of parallelism$(k)$for 16, 32, and 64-bit elements are, respectively, {4}, {4, 5, 6, 7, 8}, and {8} to achieve maximum system throughput. This sorter achieves a throughput of ≈194 Million Elements per second (MEps) for$k=6$and, it is ≈142 MEps for$k=8$to sort 256 elements of 32 and 64-bit, respectively, whereas it is ≈159 MEps for$k = 6$to sort 512 elements of 40 bit.
Sanchita Saha Ray, Surajeet Ghosh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Memory Efficient Hash-Based Longest Prefix Matching Architecture With Zero False +ve and Nearly Zero False -ve Rate for IP Processing
abstract
A novel hash-based hardware architecture for longest prefix match (LPM) scheme has been presented for IP processing.The main idea is to have zero false positive (+ve) rate with negligible false negative (-ve) rate. This has been achieved by implementing hardware-based simple hash functions for faster generation of routing table addresses. Prefix table has been designed using multiple read-port memory modules to deploy concurrent access of multiple memory words. Moreover, in this architecture, memory requirement has been reduced by maintaining next-hop address pointers instead of keeping actual next-hop address using global next-hop address memory. Prefix search time is fixed and restricts an LPM search operation within a single clock cycle irrespective of IPv4 and IPv6 address suits. This architecture can accommodate increased prefix growth trend of 40% – 100% with unchanged memory capacity. After rigorous testing, false -ve rate is found mostly zero and reasonably negligible (0.00512%) in worst-cases, however, false +ve rate is always zero. Lower bounds in memory requirements for four numerous schemes of the proposed LPM architecture with their probable failure-rates are analysed. System failure-rates are observed against incremental prefix growth with variable number of hash functions ranging from 6 to 8 to ensure future sustainability of the scheme.
Sanchita Saha Ray, Surajeet Ghosh, Bhaskar Sardar
IEEE Trans. Computers1
2022 Worst Case O(N) Comparison-Free Hardware Sorting Engine
abstract
This article proposes a novel comparison-free hardware sorting engine that sorts$N$unique$n$-bit elements (irrespective of signed and unsigned) consuming linear sorting latency of$O(N)$clock cycles. It can even efficiently sort$N$data elements with a nonzero duplicity rate in less than$O(N)$clock cycles. This sorting engine is designed using$n$-symmetric cascaded blocks utilizing few fundamental logic components. The entire design is synthesized for several data sets from pseudorandomly generated data elements to unique elements, and also from random to completely sorted elements with various duplicity rates. The architecture appears impartial with respect to ordering of elements. Synthesis results indicate that the proposed approach consumes reasonably lower field programmable gate array resources than existing approaches. The architecture takes per-element sorting latency in sorting 512 unique signed elements as 22.56 ns (48 bit) and takes 26.80 ns (64 bit) to sort 256 unique signed elements. The engine achieves sorting throughput rates as$\approx 117$-to-142 Million-Elements-per-second (MEps) (16 bit), 79-to-97 MEps (24 bit) for sorting 256-to-1K, whereas 66-to-73 MEps (32 bit) and 44-to-49 MEps (48 bit) for sorting 256-to-512 elements. However, it is 37 MEps (64 bit) in sorting 256 signed elements. This architecture consumes$\approx 1.52~\mu \text{W}$for the unique signed numbers (SNs) as per-byte processing power and$\approx 1.55~\mu \text{W}$for the SNs with nonzero duplicity rates.
Sanchita Saha Ray, Dulal Adak, Surajeet Ghosh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1