Manan Suri

dblp:97/11214 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0003-1417-3570ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 9 since 2021Systems, architecture and hardware · 12 · 3 since 2021
YearPublicationVenuePosition
2025 ChartLens: Fine-grained Visual Attribution in Charts
abstract
Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Dinesh Manocha. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan Rossi, Dinesh Manocha
ACL (1)1
2025 Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
abstract
Flowcharts are a critical tool for visualizing decision-making processes.However, their non-linear structure and complex visual-textual relationships make it challenging to interpret them using LLMs, as vision-language models frequently hallucinate nonexistent connections and decision paths when analyzing these diagrams.This leads to compromised reliability for automated flowchart processing in critical domains such as logistics, health, and engineering.We introduce the task of Fine-grained Flowchart Attribution, which traces specific components grounding a flowchart referring LLM response.Flowchart Attribution ensures the verifiability of LLM predictions and improves explainability by linking generated responses to the flowchart's structure.We propose FlowPathAgent, a neurosymbolic agent that performs fine-grained post hoc attribution through graph-based reasoning.It first segments the flowchart, then converts it into a structured symbolic graph, and then employs an agentic approach to dynamically interact with the graph, to generate attribution paths.Additionally, we present FlowExplainBench, a novel benchmark for evaluating flowchart attributions across diverse styles, domains, and question types.Experimental results show that FlowPathAgent mitigates visual hallucinations in LLM answers over flowchart QA, outperforming strong baselines by 10-14% on our proposed FlowExplainBench dataset.
Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan Rossi, Vivek Gupta 0001, Dinesh Manocha
EMNLP1
2025 VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
abstract
Manan Suri, Puneet Mathur, Franck Dernoncourt, Kanika Goswami, Ryan A. Rossi, Dinesh Manocha. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Manan Suri, Puneet Mathur, Franck Dernoncourt, Kanika Goswami, Ryan Rossi, Dinesh Manocha
NAACL (Long Papers)1
2024 DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
abstract
Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I. Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha
EMNLP1
2024 VPU-CIM: A 130nm, 33.98 TOPS/W RRAM based Compute-In-Memory Vector Co-Processor
abstract
Deep Learning inference on edge devices requires reduced memory load/store latency and low bit-precision computations. To address these challenges, we present VPU-CIM: a novel RRAM-based Compute-In-Memory (CIM) variable bit-precision vector co-processor. We introduce vector extensions to the RISC-V ISA and implement it as an in-memory compute unit with a unique data mapping strategy. The design is implemented using open-source Skywater 130nm PDK, with area estimates provided for TSMC 28nm and ASAP 7nm PDKs. Our design achieves an energy efficiency of 33.98 TOPS/W for a 4,4 (I, W) precision configuration. The results demonstrate the potential of RRAM-based vector computations in memory.
J. Chithambara Moorthii, Vinay Rayapati, Nanditha Rao, Manan Suri
ISCAS4
2024 POEM: Performance Optimization and Endurance Management for Non-volatile Caches
abstract
Non-volatile memories (NVMs), with their high storage density and ultra-low leakage power, offer promising potential for redesigning the memory hierarchy in next-generation Multi-Processor Systems-on-Chip (MPSoCs). However, the adoption of NVMs in cache designs introduces challenges such as NVM write overheads and limited NVM endurance. The shared NVM cache in an MPSoC experiences requests from different processor cores and responses from the off-chip memory when the requested data is not present in the cache. Besides, upon evictions of dirty data from higher-level caches, the shared NVM cache experiences another source of write operations, known as writebacks . These sources of write operations—writebacks and responses—further exacerbate the contention for the shared bandwidth of the NVM cache and create significant performance bottlenecks. Uncontrolled write operations can also affect the endurance of the NVM cache, posing a threat to cache lifetime and system reliability. Existing strategies often address either performance or cache endurance individually, leaving a gap for a holistic solution. This study introduces the Performance Optimization and Endurance Management (POEM) methodology, a novel approach that aggressively bypasses cache writebacks and responses to alleviate the NVM cache contention. Contrary to the existing bypass policies that do not pay adequate attention to the shared NVM cache contention and focus too much on cache data reuse, POEM’s aggressive bypass significantly improves the overall system performance, even at the expense of data reuse. POEM also employs effective wear leveling to enhance the NVM cache endurance by careful redistribution of write operations across different cache lines. Across diverse workloads, POEM yields an average speedup of 34% over a naïve baseline and 28.8% over a state-of-the-art NVM cache bypass technique while enhancing the cache endurance by 15% over the baseline. POEM also explores diverse design choices by exploiting a key policy parameter that assigns varying priorities to the two system-level objectives.
Aritra Bagchi, Dharamjeet, Ohm Rishabh, Manan Suri, Preeti Ranjan Panda
ACM Trans. Design Autom. Electr. Syst.4
2023 ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER
abstract
Sreyan Ghosh, Utkarsh Tyagi, Manan Suri, Sonal Kumar, Ramaneswaran S, Dinesh Manocha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Sreyan Ghosh, Utkarsh Tyagi, Manan Suri, Sonal Kumar, Ramaneswaran S., Dinesh Manocha
ACL (1)3
2023 CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network
abstract
The tremendous growth of social media users interacting in online conversations has led to significant growth in hate speech affecting people from various demographics.Most of the prior works focus on detecting explicit hate speech, which is overt and leverages hateful phrases, with very little work focusing on detecting hate speech that is implicit or denotes hatred through indirect or coded language.In this paper, we present CoSyn, a context synergized neural network that explicitly incorporates user-and conversational-context for detecting implicit hate speech in online conversations.CoSyn introduces novel ways to encode these external contexts and employs a novel context interaction mechanism that clearly captures the interplay between them, making independent assessments of the amounts of information to be retrieved from these noisy contexts.Additionally, it carries out all these operations in the hyperbolic space to account for the scalefree dynamics of social media.We demonstrate the effectiveness of CoSyn on 6 hate speech datasets and show that CoSyn outperforms all our baselines in detecting implicit hate speech with absolute improvements in the range of 1.24% -57.8%.We make our code available 1 .
Sreyan Ghosh, Manan Suri, Purva Chiniya, Utkarsh Tyagi, Sonal Kumar, Dinesh Manocha
EMNLP2
2023 Low-Power Lossless Image Compression on Small Satellite Edge using Spiking Neural Network
abstract
The emerging trend of small satellites for earth observation missions has enabled commercial organisations to exploit the horizon for various business applications related to weather forecasting/monitoring, Land Use Land Cover (LULC) classifications, disaster (such as oil spill or forest fire) monitoring etc. However, the limited power and computational capacity of these small satellites arising out of the size and weight restrictions have posed newer challenges, primarily related to low-power on-board data processing and transmission. One possible approach is to harness the capabilities of the evolving neuromorphic computing paradigm for such low-power computing requirements. One possible application can be lossless compression of high-resolution earth observation images before sending those downstream to ground stations for further analysis. In this paper, we propose a novel method of lossless image compression based on classical Arithmetic Encoding that exploits the low power computing capability of Spiking Neural Networks and neuromorphic platforms. We experimentally prove that our SNN approach achieves compression ratio at-par with state of the art ANN methods with an estimated 2.5× power efficiency and 50% lower latency with a much smaller model - thereby enabling on-board image compression and at the same time, saving on a corresponding amount of energy during transmission.
Sayan Kahali, Sounak Dey, Chetan Kadway, Arijit Mukherjee, Arpan Pal 0001, Manan Suri
IJCNN6
2023 Neuromorphic Recurrent Spiking Neural Networks for EMG Gesture Classification and Low Power Implementation on Loihi
abstract
In this work, we show an efficient Electromyograph (EMG) gesture recognition using Double Exponential Adaptive Threshold (DEXAT) neuron based Recurrent Spiking Neural Network (RSNN). Our network achieves a classification accuracy of 90% while using lesser number of neurons compared to the best reported prior art on Roshambo EMG dataset. Further, to illustrate the benefits of dedicated neuromorphic hardware, we show hardware implementation of DEXAT neuron using multicompartment methodology on Intel's neuromorphic Loihi chip. RSNN implementation on Loihi (Nahuku 32) achieves significant energy/latency benefits of ~983X/19X compared to GPU for batch size = 50.
Ahmed Shaban, Sai Sukruth Bezugam, Manan Suri
ISCAS3
2023 Online time-series forecasting using spiking reservoir
Arun M. George, Sounak Dey, Dighanchal Banerjee, Arijit Mukherjee, Manan Suri
Neurocomputing5
2021 A Hybrid CMOS-Memristive Approach to Designing Deep Generative Models
abstract
Deep learning and its applications have gained tremendous interest recently in both academia and industry. Restricted Boltzmann machines (RBMs) offer a key methodology to implement deep learning paradigms. This brief presents a novel approach for realizing hybrid CMOS-memristive-based deep generative models (DGMs). In our proposed DGM architecture, HfOx-based (filamentary-type switching) memristive devices are extensively used for realizing both computational as well as storage functions, such as: 1) synapses (weights); 2) internal neuron-state storage; 3) stochastic neuron activation; and 4) programmable signal normalization. To validate the proposed scheme, we have simulated two different architectures: 1) deep belief network (DBN) for classification and 2) stacked denoising autoencoder for the reconstruction of handwritten digits from the MNIST data set. The maximum test accuracy achieved by pretraining of the proposed DBN was 92.6%, whereas the best case mean squared error (mse) achieved by pretraining of the proposed SDA network was 0.046. When the proposed model-based weights are used for weight initialization, they offer a significant advantage in terms of learning performance in comparison with randomized initialization.
Vivek Parmar, Manan Suri
IEEE Trans. Neural Networks Learn. Syst.2
2020 Unified Characterization Platform for Emerging NVM Technology: Neural Network Application Benchmarking using off-the-Shelf NVM Chips
abstract
In this paper, we present a unified FPGA based electrical test-bench for characterizing different emerging NonVolatile Memory (NVM) chips. In particular, we present detailed electrical characterization and benchmarking of multiple commercially available, off-the-shelf, NVM chips viz.: MRAM, FeRAM, CBRAM, and ReRAM. We investigate important NVM parameters such as: (i) current consumption patterns, (ii) endurance, and (iii) error characterization. The proposed FPGA based testbench is then utilized for a Proof-of-Concept (PoC) Neural Network (NN) image classification application. Four emerging NVM chips are benchmarked against standard SRAM and Flash technology for the AI application as active weight memory during inference mode.
Supriya Chakraborty, Manan Suri
ISCAS3
2020 Methodology for Realizing VMM with Binary RRAM Arrays: Experimental Demonstration of Binarized-ADALINE using OxRAM Crossbar
abstract
In this paper, we present an efficient hardware mapping methodology for realizing vector matrix multiplication (VMM) on resistive memory (RRAM) arrays. Using the proposed VMM computation technique, we experimentally demonstrate a binarized-ADALINE (Adaptive Linear) classifier on an OxRAM crossbar. An 8×8 OxRAM crossbar with Ni/3-nm HfO2/7 nm Al-doped-TiO2/TiN device stack is used. Weight training for the binarized-ADALINE classifier is performed ex-situ on UCI cancer dataset. Post weight generation the OxRAM array is carefully programmed to binary weight-states using the proposed weight mapping technique on a custom-built testbench. Our VMM powered binarized-ADALINE network achieves a classification accuracy of 78% in simulation and 67% in experiments. Experimental accuracy was found to drop mainly due to crossbar inherent sneak-path issues and RRAM device programming variability.
Sandeep Kaur Kingra, Vivek Parmar, Shubham Negi, Sufyan Khan, Boris Hudec, Tuo-Hung Hou, Manan Suri
ISCAS7
2019 NV-BNN: An Accurate Deep Convolutional Neural Network Based on Binary STT-MRAM for Adaptive AI Edge
abstract
Binary STT-MRAM is a highly anticipated embedded nonvolatile memory technology in advanced logic nodes < 28 nm. How to enable its in-memory computing (IMC) capability is critical for enhancing AI Edge. Based on the soon-available STT-MRAM, we report the first binary deep convolutional neural network (NV-BNN) capable of both local and remote learning. Exploiting intrinsic cumulative switching probability, accurate online training of CIFAR-10 color images (~ 90%) is realized using a relaxed endurance spec (switching ≤ 20 times) and hybrid digital/IMC design. For offline training, the accuracy loss due to imprecise weight placement can be mitigated using a rapid non-iterative training-with-noise and fine-tuning scheme.
Chih-Cheng Chang, Ming-Hung Wu, Jia-Wei Lin, Chun-Hsien Li, Vivek Parmar, Heng-Yuan Lee, Jeng-Hua Wei, Shyh-Shyuan Sheu, Manan Suri, Tian-Sheuan Chang, Tuo-Hung Hou
DAC9
2019 Hyperspectral Image Classification for Remote Sensing Using Low-Power Neuromorphic Hardware
abstract
In this paper, we present a novel feature extraction algorithm based approach for performing Hyperspectral Image Classification using a low-power Neuromorphic hardware. The application of interest for this study is HSI image classification for remote sensing. We demonstrate energy-efficient data processing pipeline optimized to use with on-edge neuromorphic hardware. The dataset used for the study is Salinas-A. We use the Brilliant USB stick with 4 NM500 chips for prototyping the application. Achieved recognition time is 18.4 μs and energy consumption is ~10 μJ with an accuracy of ~ 97%.
Vivek Parmar, Jung-Ho Ahn 0003, Manan Suri
IJCNN3
2019 Optimized Implementation of Neuromorphic HATS Algorithm on FPGA
abstract
In this paper, we present first-ever optimized hardware implementation of a state-of-the-art neuromorphic approach Histogram of Averaged Time Surfaces (HATS) algorithm to event-based object classification in FPGA for asynchronous time-based image sensors (ATIS). Our Implementation achieves latency of 3.3 ms for the N-CARS dataset samples and is capable of processing 2.94 Mevts/s. Speed-up is achieved by using parallelism in the design and multiple Processing Elements can be added. As development platform, Zynq-7000 SoC from Xilinx is used. The tradeoff between Average Absolute Error and Resource Utilization for fixed precision implementation is analyzed and presented. The proposed FPGA implementation is ~ 32 × power efficient compared to software implementation.
Khushal Sethi, Manan Suri
ISCAS2
2018 Design Exploration of IoT centric Neural Inference Accelerators
abstract
Neural networks have been successfully deployed in a variety of fields like computer vision, natural language processing, pattern recognition, etc. However most of their current deployments are suitable for cloud-based high-performance computing systems. As the computation of neural networks is not suited to traditional Von-Neumann CPU architectures, many novel hardware accelerator designs have been proposed in literature. In this paper we present the design of a novel, simplified and extensible neural inference engine for IoT systems. We present a detailed analysis on the impact of various design choices like technology node, computation block size, etc on overall performance of the neural inference engine. The paper demonstrates the first design instance of a power-optimized ELM neural network using ReLU activation. Comparison between learning performance of simulated hardware against the software model of the neural network shows a variation of ~ 1% in testing accuracy due to quantization. The accelerator compute blocks manage to achieve a performance per Watt of ~ 290 MSPS/W (Million samples per second per Watt) with a network structure of size: 8 x 32 x 2. Minimum energy of 40 pJ is acheived per sample processed for a block size of 16. Further, we show through simulations that an added power-saving of ~ 30 % can be acheived if SRAM based main memory is replaced with emerging STT-MRAM technology.
Vivek Parmar, Manan Suri
ACM Great Lakes Symposium on VLSI2
2018 Efficient Low-Power Material Analysis using Neuromorphic Hardware: A spectral case study
abstract
In this paper, we propose an accurate and fast technique to perform qualitative spectral analysis. We train a bio-inspired dedicated ASIC with the spectral signatures of the constituents and test the network performance on real NMR, IR and artificially synthesised FT-IR binary and ternary mixtures. Constituents are detected with a best case accuracy of ~95%, ~79% and ~78% in order of their respective proportions. Ability to detect the presence of samples even when present in low proportions coupled with advantages gained in throughput and power enables it to be deployed for real time intelligent detection in portable hand-held spectrometers.
Narayani Bhatia, Manan Suri
IJCNN2
2018 MASTISK: Simulation Framework For Design Exploration Of Neuromorphic Hardware
abstract
In this paper, we present MASTISK (MAchine-learning and Synaptic-plasticity Technology Integrated Simulation frameworK). MASTISK is an open-source versatile and flexible tool developed in MATLAB for design exploration of dedicated neuromorphic hardware using nanodevices and hybrid CMOS-nanodevice circuits. MASTISK has a hierarchical organization capturing details at the level of devices, circuits (i.e., neurons or activation functions, synapses or weights) and architectures (i.e., topology, learning-rules, algorithms). In the current version, MASTISK provides user-friendly interface for design and simulation of spiking neural networks (SNN) powered by spatio-temporal learning rules such as Spike-Timing Dependent Plasticity (STDP). Users may provide network definition as a simple input parameter file and the framework is capable of performing automated learning/inference simulations. To validate the working of MASTISK, we present 2 case-studies: (i) RRAM based synapses, and (ii) PCM based neurons. The proposed framework offers new functionalities, compared to similar simulation tools in literature, such as: (i) arbitrary synaptic circuit modeling capability with both identical and non-identical stimuli, (ii) arbitrary spike modeling, and (iii) nanodevice based neuron emulation. The code of MASTISK is available on request at: https: //gitlab.commVMhome.
Tinish Bhattacharya, Vivek Parmar, Manan Suri
IJCNN3
2018 OxRAM Resistive Switching for DR Improvement
abstract
On-chip presence of emerging nonvolatile resistive memory devices provide opportunities to perform different types of computing and storage operations with several advantages. A unique application of oxide based resistive memory (OxRAM) devices in CMOS Image Sensors (CIS) for improvement of pixel dynamic range (DR) is proposed in this paper. In CIS, DR signifies the detectable range of light intensity. A modified 3T-APS (Active Pixel Sensor) circuit that incorporates OxRAM in 1T-1R configuration is proposed for DR improvement in case of low exposure situations. We show relative pixel DR improvement through OxRAM resistive switching during exposure and study the impact of different preconditioning states. Best case, simulated relative pixel DR improvement of ~34 dB was obtained for the proposed design. To study the system level impact of the proposed pixel a simplified behavioral model based simulation of an array with 580 × 950 pixels was performed.
Mukul Sarkar, Manan Suri
ISCAS3
2018 Current Optimized Coset Coding for Efficient RRAM Programming
Supriya Chakraborty, Tinish Bhattacharya, Manan Suri
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Stochastic CBRAM-Based Neuromorphic Time Series Prediction System
abstract
In this research, we present a Conductive-Bridge RAM (CBRAM)-based neuromorphic system which efficiently addresses time series prediction. We propose a new (i) voltage-mode, stochastic, multiweight synapse circuit based on experimental bi-stable CBRAM devices, (ii) a voltage-mode neuron circuit based on the concept of charge sharing, and (iii) an optimized training methodology powered by a stochastic implementation of the Least-Mean-Squares (SLMS) training rule. To validate the proposed design, we use time series prediction for short-term electrical load forecasting in smart grids. Our system is able to forecast hourly electrical loads with a mean accuracy of 96%, an estimated power dissipation of 15 μW, and area of 14.5 μm 2 at 65 nm CMOS technology.
Cory E. Merkel, Dhireesha Kudithipudi, Manan Suri, Bryant T. Wysocki
ACM J. Emerg. Technol. Comput. Syst.3
2015 OXRAM based ELM architecture for multi-class classification applications
abstract
In this paper, we show how metal-oxide (OxRAM) based nanoscale memory devices can be exploited to design low-power Extreme Learning Machine (ELM) architectures. In particular we fabricated HfO2and TiO2based OxRAM devices, and exploited their intrinsic resistance spread characteristics to realize ELM hidden layer weights and neuron biases. To validate our proposed OxRAM-ELM architecture, full-scale learning and multi-class classification simulations were performed for two complex datasets: (i) Land Satellite images and (ii) Image segmentation. Dependence of classification performance on neuron gain parameter and OxRAM device properties was studied in detail.
Manan Suri, Vivek Parmar, Gilbert Sassine, Fabien Alibart
IJCNN1
2015 NoC router using STT-MRAM based hybrid buffers with error correction and limited flit retransmission
abstract
In this paper, we present a unique methodology to implement deep IO buffers for Network-on-Chip (NoC) platform, based on a hybrid design involving conventional SRAM and emerging Spin-Transfer-Torque Magnetic Random Access Memory (STT-MRAM) technology. We focus on the system-level impact of probabilistic switching of STT-MRAM devices, arising when write latency of STT-MRAM is reduced through conservative programming and aggressive scaling. We incorporate STT-MRAM specific error detection and correction schemes at the input buffers, and propose a new limited flit retransmission scheme to reduce flit errors due to the probabilistic switching. Our hybrid STT-MRAM buffers along with additional logic consume less than 80% of the area of SRAM-only FIFOs of the same depth. We demonstrate optimum NoC throughput at moderate injection rates on a mesh NoC.
Turbo Majumder, Manan Suri, Vinay Shekhar
ISCAS2
2011 Phase change memory for synaptic plasticity application in neuromorphic systems
abstract
In this paper, we show that Phase Change Memory (PCM) can be used to emulate specific functions of a biological synapse similar to Long Term Potentiation (LTP) and Long Term Depression (LTD) plasticity effects. The dependence of synaptic weight on programming pulse width and pulse amplitude is shown experimentally for the PCM devices. Different combinations of consecutive LTD and LTP events have been experimentally demonstrated and analyzed for the PCM synapse.
Manan Suri, Veronique Sousa, Luca Perniola, Dominique Vuillaume, Barbara De Salvo
IJCNN1