Sreetama Sarkar

dblp:218/8481 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0009-2182-6200ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
abstract
Despite their remarkable progress in multimodal understanding tasks, large vision language models (LVLMs) often suffer from "hallucination", generating texts misaligned with the visual context.Existing methods aimed at reducing hallucinations through inference time intervention incur a significant increase in latency.To mitigate this, we present SPIN, a task-agnostic attention-guided head suppression strategy that can be seamlessly integrated during inference without incurring any significant compute or latency overhead.We investigate whether hallucination in LVLMs can be linked to specific model components.Our analysis suggests that hallucinations can be attributed to a dynamic subset of attention heads in each layer.Leveraging this insight, for each text query token, we selectively suppress attention heads that exhibit low attention to image tokens, keeping the top-k attention heads intact.Extensive evaluations on visual question answering and image description tasks demonstrate the efficacy of SPIN in reducing hallucination scores up to 2.7× while maintaining F1, and improving throughput by 1.8× compared to existing alternatives.Code is available here.
Sreetama Sarkar, Yue Che, Alex Gavin, Peter A. Beerel, Souvik Kundu 0002
EMNLP1
2025 Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
abstract
Vision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge.
Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi
ICCAD4
2025 STAR-UNet: Sinogram Transformation using Attention based Residual U-Net
abstract
In computed tomography (CT), sinograms acquired under low photon counts or limited angles often suffer from severe noise and substantial data gaps, resulting in pronounced artifacts in reconstructed images. To address these challenges, we propose STAR-UNet, a novel Attention based Residual U-Net for simultaneous sinogram denoising and inpainting. Our design integrates windowed self-attention blocks, multi-scale residual connections, and an attention-guided dual upsampling strategy to effectively capture both local details and global dependencies. We train and evaluate our model on a challenging sinogram dataset derived from the LIDC/IDRI collection, which combines sparse-view sampling with high Poisson noise. Experimental results show that STAR-UNet reduces normalized root mean square error (NRMSE) by up to 31.4% under different sparsity levels and noise conditions. In addition, it improves Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) up to 5.9% and 4.1% respectively, compared to widely used convolutional encoder–decoder baselines when reconstructing CT images via filtered back projection. Moreover, STAR-UNet reduces FLOPs by 1.4 × and parameter count by 1.12 ×. These findings highlight STAR-UNet’s efficiency, robustness, and strong generalization in restoring highly undersampled, noisy sinograms, providing a promising strategy for low-dose CT applications.
Sreetama Sarkar
IJCNN2
2025 MaskVD: Region Masking for Efficient Video Object Detection
abstract
Video tasks are compute-heavy and thus pose a challenge when deploying in real-time applications, particularly for tasks that require state-of-the-art Vision Trans-formers (ViTs). Several research efforts have tried to address this challenge by leveraging the fact that large portions of the video undergo very little change across frames leading to redundant computations in frame-based video processing. In particular, some works leverage pixel or semantic differences across frames, however, this yields limited latency benefits with significantly increased memory overhead. This paper, in contrast, presents a strategy for masking regions in video frames that leverages the semantic information in images and the temporal correlation between frames to significantly reduce FLOPs and latency with little to no penalty in performance over base-line models. In particular, we demonstrate that by lever-aging extracted features from previous frames, ViT back-bones directly benefit from region masking, skipping up to 80% of input regions, improving FLOPs and latency by 3.14x and 1.5x. We improve memory and latency over the state-of-the-art (SOTA) by 2.3x and 1.14x, while maintaining similar detection performance. Additionally, our approach demonstrates promising results on convolutional neural networks (CNNs) and provides latency improvements over the SOTA up to 1.3x using specialized computational kernels. Code is available at: https://github.com/sreetamasarkar/MaskVD
Sreetama Sarkar, Gourav Datta, Souvik Kundu 0002, Chirayata Bhattacharyya, Peter A. Beerel
WACV1
2024 FixPix: Fixing Bad Pixels using Deep Learning
Sreetama Sarkar, Xinan Ye, Gourav Datta, Peter A. Beerel
ICPR (3)1
2023 Technology-Circuit-Algorithm Tri-Design for Processing-in-Pixel-in-Memory (P2M)
abstract
The massive amounts of data generated by camera sensors motivate data processing inside pixel arrays, i.e., at the extreme-edge. Several critical developments have fueled recent interest in the processing-in-pixel-in-memory paradigm for a wide range of visual machine intelligence tasks, including (1) advances in 3D integration technology to enable complex processing inside each pixel in a 3D integrated manner while maintaining pixel density, (2) analog processing circuit techniques for massively parallel low-energy in-pixel computations, and (3) algorithmic techniques to mitigate non-idealities associated with analog processing through hardware-aware training schemes. This article presents a comprehensive technology-circuit-algorithm landscape that connects technology capabilities, circuit design strategies, and algorithmic optimizations to power, performance, area, bandwidth reduction, and application-level accuracy metrics. We present our results using a comprehensive co-design framework incorporating hardware and algorithmic optimizations for various complex real-life visual intelligence tasks mapped onto our P2M paradigm.
Md. Abdullah-Al Kaiser, Gourav Datta, Sreetama Sarkar, Souvik Kundu 0002, Zihan Yin, Manas Garg, Ajey P. Jacob, Peter A. Beerel, Akhilesh Jaiswal 0001
ACM Great Lakes Symposium on VLSI3
2022 Accelerating and pruning CNNs for semantic segmentation on FPGA
abstract
Semantic segmentation is one of the popular tasks in computer vision, providing pixel-wise annotations for scene understanding. However, segmentation-based convolutional neural networks require tremendous computational power. In this work, a fully-pipelined hardware accelerator with support for dilated convolution is introduced, which cuts down the redundant zero multiplications. Furthermore, we propose a genetic algorithm based automated channel pruning technique to jointly optimize computational complexity and model accuracy. Finally, hardware heuristics and an accurate model of the custom accelerator design enable a hardware-aware pruning framework. We achieve 2.44X lower latency with minimal degradation in semantic prediction quality (−1.98 pp lower mean intersection over union) compared to the baseline DeepLabV3+ model, evaluated on an Arria-10 FPGA. The binary files of the FPGA design, baseline and pruned models can be found in github.com/pierpaolomori/SemanticSegmentationFPGA
Pierpaolo Morì, Manoj Rohit Vemparala, Nael Fasfous, Saptarshi Mitra, Sreetama Sarkar, Alexander Frickenstein, Lukas Frickenstein, Domenik Helms, Naveen Shankar Nagaraja, Walter Stechele, Claudio Passerone
DAC5