Pankaj Bhowmik

dblp:201/5879 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-9241-1071ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 62% Reconfigurable computing and FPGAs · 19% Energy-efficient computing · 19%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
dynamic power reduction
0.112019
Visual Cortex Inspired Pixel-Level Re-configurable Processors for Smart Image Sensors · DAC 2019
Reconfigurable computing and FPGAs
reconfigurable architecture
0.112019
Visual Cortex Inspired Pixel-Level Re-configurable Processors for Smart Image Sensors · DAC 2019

Methods — techniques the papers use, named apart from their topics

predictive coding · 0.4ASIC implementation · 0.4
YearPublicationVenuePosition
2021 ESCA: Event-Based Split-CNN Architecture with Data-Level Parallelism on UltraScale+ FPGA
abstract
This paper presents an event-based split-CNN architecture (ESCA) for running time-critical vision applications with comparatively less memory footprint while consuming low power. ESCA has a dedicated hardware architecture and scheduling of on-chip memory buffering using a split-CNN that reduces memory requirements by splitting the feature maps into small patches and independently executes them. The model emulates the concepts of the biological vision system to obtain possible events from each patch. We save energy and time by processing only the patches with possible events. The data-level parallelism is employed with a deep pipeline strategy that accelerates the system. We implement the design in the Virtex UltraScale+ FPGA at 320 MHz. Simulation results show that the architecture obtains significant speedup while power-saving depends on each image patch's features.
Pankaj Bhowmik, Md Jubaer Hossain Pantho, Joel Mandebi, Christophe Bobda
FCCM1
2020 Architecture Support for FPGA Multi-tenancy in the Cloud
abstract
Cloud deployments now increasingly provision FPGA accelerators as part of virtual instances. While FPGAs are still essentially single-tenant, the growing demand for hardware acceleration will inevitably lead to the need for methods and architectures supporting FPGA multi-tenancy. In this paper, we propose an architecture supporting space-sharing of FPGA devices among multiple tenants in the cloud. The proposed architecture implements a network-on-chip (NoC) designed for fast data movement and low hardware footprint. Prototyping the proposed architecture on a Xilinx Virtex Ultrascale + demonstrated near specification maximum frequency for on-chip data movement and high throughput in virtual instance access to hardware accelerators. We demonstrate similar performance compared to single-tenant deployment while increasing FPGA utilization (we achieved $6 \times$ higher FPGA utilization with our case study), which is one of the major goals of virtualization. Overall, our NoC interconnect achieved about $2 \times$ higher maximum frequency than the state-of-the-art and a bandwidth of 25.6 Gbps.
Joel Mandebi, Alex Shuping, Pankaj Bhowmik, Christophe Bobda
ASAP3
2020 Near-Sensor Inference Architecture with Region Aware Processing
abstract
Convolutional Neural Networks have been adopted in a wide range of vision-based application domains in recent times, due to their success in enabling ubiquitous machine vision and intelligent decisions. However, the overwhelming computation demand of the convolution operations has somewhat limited their use in resource-constrained embedded platforms. This paper presents a pixel processing architecture to facilitate CNN inference near the image sensor. The architecture exploits the high bandwidth available at the sensor interface and incorporates an array of pixel processors to perform inference directly on the sensor device. The proposed design addresses problems related to the mapping of computations onto an array of pixel processors and introduces a suitable network structure for communication. The pixel processors are highly optimized to provide low latency and power for CNN applications. While designing the pixel processors, we focused on reducing redundancies using the concepts of biological vision systems. We prototype the model in a Virtex UltraScale FPGA and implement it in ASIC using the TSMC 90nm technology library. The results suggest that the proposed architecture significantly reduces dynamic power consumption and achieves high-speed up surpassing the computational capabilities of existing embedded processors.
Md Jubaer Hossain Pantho, Pankaj Bhowmik, Christophe Bobda
ICCD2
2019 Event-Based Re-configurable Hierarchical Processors for Smart Image Sensors
abstract
This paper presents a reconfigurable hierarchical hardware architecture at the pixel and region level for smart image sensors to accelerate machine vision applications. This architecture maintains hierarchical processing that begins at the pixel level. It reduces the computational burden on the sequential processor and accelerates the overall processing. There are three hierarchical layers, and each layer passes only the relevant information to the next layer by removing redundant information. This method facilitates the final processor to respond if there is an event in the scene. Identifying relevant information is complex, and we adapted different bio-inspired algorithms in the hierarchical layers to meet this purpose. Besides, these processors are made runtime reconfigurable to different applications, and it adds flexibility after fabrication. This hierarchical processing breaks the traditional sequential image processing and introduces parallelism for the machine vision applications. We evaluate the design in FPGA and achieve the GDSII file in ASIC platform at 800MHz. Simulation results show that the area overhead and power penalty for adding reconfiguration feature stay in an acceptable range. Besides, removing redundant information 84.01% and 96.91% dynamic power can be saved at the pixel-level and region-level, respectively.
Pankaj Bhowmik, Md Jubaer Hossain Pantho, Christophe Bobda
ASAP1
2019 Visual Cortex Inspired Pixel-Level Re-configurable Processors for Smart Image Sensors
abstract
This paper presents a reconfigurable hardware architecture of smart image sensors to speed up low-level image processing applications at the pixel level. For each pixel in the sensor plane, the design includes an activation module and a processor. The processor has a basic structure which is common to all applications and reconfigurable segments for specific applications. Visual cortex inspired computing, like, Predictive Coding in time is implemented in the activation module to remove temporal redundancy. The ASIC implementation shows the design saves up to 84.01% dynamic power and achieves 9x speedup at 800 MHz by accurate prediction.
Pankaj Bhowmik, Md Jubaer Hossain Pantho, Christophe Bobda
DAC1