EDBT 2026 Demo / reviewers in the wild / expert
Souradip Sarkar
dblp:38/7993
· DBLP profile ↗
8ranked-venue papers
4as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 62% Interconnection networks and networks-on-chip · 38% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interconnection networks and networks-on-chip › network-on-chip design
application-specific noc |
0.1 | 1 | 2010 | Network-on-Chip Hardware Accelerators for Biological Sequence Alignment · IEEE Trans. Computers 2010 |
Hardware accelerators and domain-specific architectures
bioinformatics accelerator |
0.1 | 1 | 2010 | Network-on-Chip Hardware Accelerators for Biological Sequence Alignment · IEEE Trans. Computers 2010 |
Interconnection networks and networks-on-chip
network-on-chip design |
0.1 | 1 | 2010 | Network-on-Chip Hardware Accelerators for Biological Sequence Alignment · IEEE Trans. Computers 2010 |
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
sequence alignment accelerator |
0.1 | 1 | 2010 | Network-on-Chip Hardware Accelerators for Biological Sequence Alignment · IEEE Trans. Computers 2010 |
Bioinformatics and computational biology
phylogenetics |
0.0 | 1 | 2012 | NoC-Based Hardware Accelerator for Breakpoint Phylogeny · IEEE Trans. Computers 2012 |
Bioinformatics and computational biology
sequence alignment |
0.0 | 1 | 2010 | Network-on-Chip Hardware Accelerators for Biological Sequence Alignment · IEEE Trans. Computers 2010 |
Methods — techniques the papers use, named apart from their topics
traveling salesman problem heuristics · 0.3branch-and-bound · 0.3comparative performance and energy evaluation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Agile Design-Space Exploration of Dynamic Layer-Skipping in Neural ReceiversabstractDynamic Neural Networks (DyNN) adapt their structure during runtime for improved performance and lower power consumption. DyNNs benefit neural network-based wireless receivers, or in short neural receivers, that adapt their performance under varying channel conditions. However, DyNN architectures are not yet explored sufficiently in the literature as this would require a framework for simultaneously evaluating system-level performance and accurate hardware Power-Performance-Area (PPA) under different dynamic scenarios. This paper presents two main contributions: (1) An automated framework that bridges a system-level performance evaluation model of a wireless system and a High-Level Synthesis (HLS) tool for agile Design-Space Exploration (DSE) of DyNNs. (2) A novel DyNN architecture for neural receivers in wireless systems that skips a variable number of layers according to varying channel conditions. Our proposed neural receiver architecture with layer-skipping for a Single Input Multi Output (SIMO) wireless system when implemented in 22 nm FD-SOI technology shows power consumption savings of up to 59.2% at the cost of 200% increase in area compared to a static network. Bram Van Bolderik, Souradip Sarkar, Vlado Menkovski, Sonia M. Heemstra de Groot, Manil Dev Gomony |
DSD | 2 |
| 2019 | A Reconfigurable Architecture for Posit ArithmeticabstractPhysical layer of modern communication systems involves several floating point computations that require higher numerical fidelity and dynamic range for achieving the maximum data rate. Posit number format is a promising alternative for floating point computation as it provides a better dynamic range and numerical fidelity for the same number of bits used for representation. However, the Posit number format has not yet been thoroughly studied in the context of signal processing algorithms in the physical layer. In addition, a configurable Posit arithmetic hardware is essential for adapting to the ever changing communication standards and to configure the dynamic range according to algorithmic demands. The main contributions in this paper are: 1) Performance analysis of common signal processing algorithms using the Posit number format in comparison with the IEEE Standard for Single Precision Floating Point arithmetic (IEEE 754). 2) A novel reconfigurable hardware accelerator for Posit arithmetic operations and comparison with the state-of-the-art Posit and IEEE Floating Point arithmetic architectures. Although our proposed architecture consumes over 3× energy and area compared to IEEE Single Precision arithmetic, our results show that using the Posit number format for physical layer algorithms results in significant performance gain. We achieved over 15 dB and 25 dB gain for FFT and matrix multiplication algorithms. In addition, compared to state-of-the-art Posit arithmetic architecture, our proposed architecture resulted in over 2× speedup in operating frequency and 35% savings in energy consumption. Souradip Sarkar, Purushotham Murugappa, Manil Dev Gomony |
DSD | 1 |
| 2018 | Quater-imaginary base for complex number arithmetic circuitsabstractArithmetic operations involving complex numbers are widely used in the signal processing functions in the physical layer of modern wireless and wireline communication systems, electronic instrumentation and control systems. With the ever increasing throughput requirements of such systems, the power consumption of the hardware realization is increasing beyond the allowed budget. Arithmetic circuits based on binary numeral system that have been optimized rigorously over the past few decades are currently being used for the computation involving complex numbers. In this paper, we present the potential of arithmetic circuits for complex number computations based on the Quater-imaginary (QI) base numeral system to reduce power consumption. We show that for a simple multiplier implementation in the QI base, the savings in power and area consumption could be up to 40% when synthesized in 28nm TSMC standard cell technology node. Souradip Sarkar, Manil Dev Gomony |
DATE | 1 |
| 2012 | Power-aware multi-core simulation for early design stage hardware/software co-optimizationabstractStringent performance targets and power constraints push designers towards building specialized workload-optimized systems across a broad spectrum of the computing arena, including supercomputing applications as exemplified by the IBM BlueGene and Intel MIC architectures. In this paper, we make the case for hardware/software co-design during early design stages of processors for scientific computing applications. Considering an important scientific kernel, namely stencil computation, we demonstrate that performance and energy-efficiency can be improved by a factor of 1.66X and 1.25X, respectively, by co-optimizing hardware and software. Wim Heirman, Souradip Sarkar, Trevor E. Carlson, Ibrahim Hur, Lieven Eeckhout |
PACT | 2 |
| 2012 | NoC-Based Hardware Accelerator for Breakpoint PhylogenyabstractMaximum Parsimony phylogenetic tree reconstruction is based on finding the breakpoint median, given a set of species, and is represented by a bounded edge-weight graph model. This reduces the breakpoint median problem to one of solving multiple instances of the Traveling Salesman Problem (TSP), which is a classical NP-complete problem in graph theory. Exponential time algorithms that apply efficient runtime heuristics, such as branch-and-bound, to dynamically prune the search space are used to solve TSP. In this paper, we present the design and performance evaluation of a network-on-chip (NoC)-based implementation for solving TSP under the bounded edge-weight model, as used in the computation of breakpoint phylogeny. Our approach takes advantage of fine-grain parallelism from the multiple processing elements (PEs) and uses efficient NoC architecture for inter-PE communication. To accelerate the application on hardware, our PE design optimizes a particular lower bound calculation operation which typically tends to be the serial bottleneck in computation of a TSP solution. We also explore two representative NoC architectures-mesh and quad-tree-and show that the latter is more energy-efficient for this application domain. Experimental results show that this new implementation is able to achieve speedups of up to three orders of magnitude over state-of-the-art multithreaded software implementations. Turbo Majumder, Souradip Sarkar, Partha Pratim Pande, Anantharaman Kalyanaraman |
IEEE Trans. Computers | 2 |
| 2010 | An optimized NoC architecture for accelerating TSP kernels in breakpoint median problemabstractTraveling Salesman Problem (TSP) is a classical NP-complete problem in graph theory. It aims at finding a least-cost Hamiltonian cycle that traverses all vertices of an input edge-weighted graph. One application of TSP is in breakpoint median-based Maximum Parsimony phylogenetic tree reconstruction, wherein a bounded edge-weight model is used. Exponential algorithms that apply efficient heuristics, such as branch-and-bound, to dynamically prune the search space are used. We adopted this approach in an NoC-based implementation for solving TSP targeted towards phylogenetics taking advantage of the fine-grained parallelism and efficient communication network. The largest fraction of the solution time for TSP is accounted for by a particular lower bound calculation operation that uses the graph's adjacency matrix. In this paper, we present the design and implementation of the processing elements with a highly optimized lower bound computation kernel and evaluate its performance. Additionally, we explore two major NoC architectures -mesh and quad-tree - and show that the latter is more suitable for this application domain. Turbo Majumder, Souradip Sarkar, Partha Pratim Pande, Anantharaman Kalyanaraman |
ASAP | 2 |
| 2010 | Hardware accelerators for biocomputing: A surveyabstractComputing research has become a vital cog in the machinery required to drive biological discovery. Computing has made possible significant achievements over the last decade, especially in the genomics sector. An emerging area is the investigation of hardware accelerators for speeding up the massive scale of computation needed in large-scale biocomputing applications. Various hardware platforms, such as FPGA, Graphics Processing Unit (GPU), the Cell Broadband Engine (CBE) and multi-core processors are being explored. In this paper, we present a survey of hardware accelerators for biocomputing by choosing a representative set of each. Souradip Sarkar, Turbo Majumder, Anantharaman Kalyanaraman, Partha Pratim Pande |
ISCAS | 1 |
| 2010 | Network-on-Chip Hardware Accelerators for Biological Sequence AlignmentabstractThe most pervasive compute operation carried out in almost all bioinformatics applications is pairwise sequence homology detection (or sequence alignment). Due to exponentially growing sequence databases, computing this operation at a large-scale is becoming expensive. An effective approach to speed up this operation is to integrate a very high number of processing elements in a single chip so that the massive scales of fine-grain parallelism inherent in several bioinformatics applications can be exploited efficiently. Network-on-chip (NoC) is a very efficient method to achieve such large-scale integration. In this work, we propose to bridge the gap between data generation and processing in bioinformatics applications by designing NoC architectures for the sequence alignment operation. Specifically, we 1) propose optimized NoC architectures for different sequence alignment algorithms that were originally designed for distributed memory parallel computers and 2) provide a thorough comparative evaluation of their respective performance and energy dissipation. While accelerators using other hardware architectures such as FPGA, general purpose graphics processing unit (GPU), and the cell broadband engine (CBE) have been previously designed for sequence alignment, the NoC paradigm enables integration of a much larger number of processing elements on a single chip and also offers a higher degree of flexibility in placing them along the die to suit the underlying algorithm. The results show that our NoC-based implementations can provide above 102-103-fold speedup over other hardware accelerators and above 104-fold speedup over traditional CPU architectures. This is significant because it will drastically reduce the time required to perform the millions of alignment operations that are typical in large-scale bioinformatics projects. To the best of our knowledge, this work embodies the first attempt to accelerate a bioinformatics application using NoC. Souradip Sarkar, Gaurav Ramesh Kulkarni, Partha Pratim Pande, Anantharaman Kalyanaraman |
IEEE Trans. Computers | 1 |