EDBT 2026 Demo / reviewers in the wild / expert
Tamoghno Das
dblp:329/6680
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-7379-0969ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable AttentionabstractDeformable Transformers achieve state-of-the-art object detection, but deformable attention maps poorly to hardware due to irregular memory access and low arithmetic intensity. We present QUILL, a schedule-aware accelerator that makes MSDeformAttn cache-local and single-pass. QUILL’s Distance-based Out-of-Order Querying (DOOQ) reorders queries by spatial proximity, enabling a look-ahead, double-buffered prefetch that overlaps memory and compute. QUILL also uses a fused MSDeformAttn pipeline that performs interpolation, Softmax, aggregation, and output projection in one pass, avoiding intermediate spills and keeping small tensors on-chip. Implemented in RTL and evaluated end-to-end, QUILL achieves up to 7.29× higher throughput and 47.3× better energy efficiency than an RTX 4090, and improves throughput/energy efficiency over prior accelerators by 3.26–9.82× / 2.01–6.07×. With mixed precision, accuracy stays within ≤ 0.9 AP of FP32 across Deformable and Sparse DETR variants. By converting sparsity into locality and locality into utilization, QUILL delivers consistent end-to-end gains. Hyunwoo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 0001, Wenjun Huang 0001, Tamoghno Das, Suyeon Jang, Mohsen Imani |
DATE | 6 |
| 2026 | T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU ReorganizationabstractRecent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. While ternary quantization enables significant resource savings, existing CPU solutions rely heavily on memory-based lookup tables (LUTs) which limit scalability, and FPGA or GPU accelerators remain impractical for edge use. This paper presents T-SAR, the first framework to achieve scalable ternary LLM inference on CPUs by repurposing the SIMD register file for dynamic, in-register LUT generation with minimal hardware modifications. T-SAR eliminates memory bottlenecks and maximizes data-level parallelism, delivering 5.6–24.5× and 1.1–86.2× improvements in GEMM latency and GEMV throughput, respectively, with only 3.2% power and 1.4% area overheads in SIMD units. T-SAR achieves up to 2.5–4.9× the energy efficiency of an NVIDIA Jetson AGX Orin, establishing a practical approach for efficient LLM inference on edge platforms. Hyunwoo Oh, KyungIn Nam, Rajat Bhattacharjya, Hanning Chen, Tamoghno Das, Sanggeon Yun, Suyeon Jang, Andrew Ding, Nikil Dutt, Mohsen Imani |
DATE | 5 |
| 2025 | iTaskSense: Task-Oriented Object Detection in Resource-Constrained EnvironmentsabstractTask-oriented object detection is increasingly essential for intelligent sensing applications, enabling AI systems to operate autonomously in complex, real-world environments such as autonomous driving, healthcare, and industrial automation. Conventional models often struggle with generalization, requiring vast datasets to accurately detect objects within diverse contexts. In this work, we introduce iTask, a taskoriented object detection framework that leverages large language models (LLMs) to generalize efficiently from limited samples by generating an abstract knowledge graph. This graph encapsulates essential task attributes, allowing iTask to identify objects based on high-level characteristics rather than extensive data, making it possible to adapt to complex mission requirements with minimal samples. iTask addresses the challenges of high computational cost and resource limitations in vision-language models by offering two configuration models: a distilled, task-specific vision transformer optimized for high accuracy in defined tasks, and a quantized version of the model for broader applicability across multiple tasks. Additionally, we designed a hardware acceleration circuit to support real-time processing, essential for edge devices that require low latency and efficient task execution. Our evaluations show that the task-specific configuration achieves a 15% higher accuracy over the quantized configuration in specific scenarios, while the quantized model provides robust multi-task performance. The hardware-accelerated iTask system achieves a $3.5 x$ speedup and a 40% reduction in energy consumption compared to GPU-based implementations. These results demonstrate that iTask’s dual-configuration approach and situational adaptability offer a scalable solution for task-specific object detection, providing robust and efficient performance in resourceconstrained environments. Sungheon Jeong 0001, Hamza Errahmouni Barkam, Hyunwoo Oh, Hanning Chen, Tamoghno Das, Mohsen Imani |
DAC | 5 |
| 2025 | Robust Reasoning and Learning with Brain-Inspired Representations under Hardware-Induced NonlinearitiesabstractTraditional machine learning depends on high-precision arithmetic and near-ideal hardware assumptions, which is increasingly challenged by variability in aggressively scaled semiconductor devices. Compute-in-memory (CIM) architectures alleviate data-movement bottlenecks and improve energy efficiency yet introduce nonlinear distortions and reliability concerns. We address these issues with a hardware-aware optimization framework based on Hyperdimensional Computing (HDC), systematically compensating for non-ideal similarity computations in CIM. Our approach formulates encoding as an optimization problem, minimizing the Frobenius norm between an ideal kernel and its hardware-constrained counterpart, and employs a joint optimization strategy for end-to-end calibration of hypervector representations. Experimental results demonstrate that our method when applied to QuantHD achieves 84\% accuracy under severe hardware-induced perturbations, a 48\% increase over naive QuantHD under the same conditions. Additionally, our optimization is vital for graph-based HDC reliant on precise variable-binding for interpretable reasoning. Our framework preserves the accuracy of RelHD on the Cora dataset, achieving a 5.4$\times$ accuracy improvement over naive RelHD under nonlinear environments. By preserving HDC's robustness and symbolic properties, our solution enables scalable, energy-efficient intelligent systems capable of classification and reasoning on emerging CIM hardware. William Youngwoo Chung, Hamza Errahmouni Barkam, Tamoghno Das, Mohsen Imani |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Revisiting Reconfigurable Acceleration of Vision Transformer with Patch PruningabstractVision Transformers (ViTs) have become the backbone of numerous cutting-edge vision applications. The attention modules within ViTs play a crucial role in modeling spatial relationships between pixels. Although these attention modules enhance the accuracy of ViT models, they also increase computational demands, limiting the deployment of ViTs in edge computing environments. To address this issue, prior research has focused on optimizing ViTs from both software and hardware perspectives. A notable software optimization technique is reducing the image patches involved in attention computations. Two common methods to achieve this are window attention and patch pruning. However, they introduce new challenges for existing hardware platforms regarding attention computation. Therefore, it is essential to develop new hardware modules to simultaneously support pruned attention computations and efficient window shifts. In this study, we introduce an FPGA-based token reduction vision transformer accelerator called TRFPA. Experiments conducted on the Xilinx ZCU104 and Alveo U50 demonstrate that TRFPA outperforms previous FPGA-based ViT accelerators, achieving a 7× speedup and a 3× improvement in energy efficiency. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Hyunwoo Oh, Tamoghno Das, Fei Wen 0003, Mohsen Imani |
ISLPED | 5 |
| 2025 | LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning SegmentationabstractLarge Vision Language Models (LVLMs) have been widely adopted to guide vision foundation models in performing reasoning segmentation tasks, achieving impressive performance. However, the substantial computational overhead associated with LVLMs presents a new challenge. The primary source of this computational cost arises from processing hundreds of image tokens. Therefore, an effective strategy to mitigate such overhead is to reduce the number of image tokens-a process known as image token pruning. Previous studies on image token pruning for LVLMs have primarily focused on high-level visual understanding tasks, such as visual question answering and image captioning. In contrast, guiding vision foundation models to generate accurate visual masks based on textual queries demands precise semantic and spatial reasoning capabilities. Consequently, pruning methods must carefully control individual image tokens throughout the LVLM reasoning process. Our empirical analysis reveals that existing methods struggle to adequately balance reductions in computational overhead with the necessity to maintain high segmentation accuracy. In this work, we propose LVLM_CSP, a novel training-free visual token pruning method specifically designed for LVLM-based reasoning segmentation tasks. LVLM_CSP consists of three stages: clustering, scattering, and pruning. Initially, the LVLM performs coarse-grained visual reasoning using a subset of selected image tokens. Next, fine-grained reasoning is conducted, and finally, most visual tokens are pruned in the last stage. Extensive experiments demonstrate that LVLM_CSP achieves a 65% reduction in image token inference FLOPs with virtually no accuracy degradation, and a 70% reduction with only a minor 1% drop in accuracy on the 7B LVLM. Hanning Chen, Yang Ni 0001, Wenjun Huang 0001, Hyunwoo Oh, Yezi Liu, Tamoghno Das, Mohsen Imani |
ACM Multimedia | 6 |
| 2024 | Bayesian-Informed Hyperdimensional Learning for Intelligent and Efficient Data ProcessingabstractIn machine learning (ML), near-sensor AI is transforming edge computing by reducing response times and data transmission, ultimately saving energy and bandwidth. Despite challenges like limited computational resources and the need for transparent decision-making, this approach aims to enhance the intelligence and autonomy of edge devices. Our research presents a novel framework that adds a layer of abstract intelligence to sensors, boosting system efficiency and accuracy through transparent, interpretable sub-symbolic AI. We combine Bayesian algorithms with hyperdimensional computing (HDC), inspired by the human brain's operational efficiency, to deliver an energy-efficient solution matching the accuracy of traditional cloud systems without constant server dependence. This framework uses a binary classifier with Bayesian insights to choose the best data processing location---locally or in the cloud---adapting to data environments. Our method ensures cloud-level performance while significantly reducing energy consumption, improving the sustainability of sensor-based systems. It also enables continual adaptation and learning directly at the sensor level, enriching cloud models with fresh edge insights. Our results have shown to bridge the gap from around 38% quality loss between the standalone near-sensor HDC model and the SOTA cloud-based model to improve the quality loss to only 9% while simultaneously saving 45.34% of energy by not using the cloud. This framework paves the way for more sustainable, efficient, and accurate edge computing in the ML landscape by bridging the gap between simple near-sensor models and their advanced cloud-based counterparts. Hamza Errahmouni Barkam, Tamoghno Das, Prathyush Poduval, Sungheon Jeong 0001, Calvin Yeung 0002, Mostafa A. Solitan, Mohsen Imani |
ICCAD | 2 |