Deepak Mathaikutty

dblp:m/DeepakMathaikutty · also Deepak A. Mathaikutty · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
5since 2021 · last 2026
0009-0004-3768-1337ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 9 · 7 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Algorithm-Hardware Co-Design of Digital Compute-in-Memory Architecture Supporting Flexible and Temporal N:M Sparsity
abstract
Structured pruning with fixed N:M sparsity ratios in large language models (LLMs) significantly constrains model expressivity, often leading to suboptimal accuracy. While supporting multiple N:M configurations can enhance representational flexibility, such a capability typically introduces substantial hardware complexity and overhead. To overcome these limitations, we first present FLOW, a flexible, layer-wise, outlier-density-aware N:M sparsity selection framework. FLOW adaptively determines the optimal N and M values per layer within a specified range by jointly considering the magnitude and distribution of outliers, thereby improving sparsity allocation and model fidelity. Extending this idea, we introduce FLOW++, which generalizes flexible N:M sparsity to the temporal domain for reasoning LLMs. FLOW++ enables temporal pruning by dynamically adapting sparsity patterns across reasoning thought types. To support efficient deployment of models with such dynamically varying sparsity patterns, we propose FlexCiM, a flexible, low-overhead, digital compute-in-memory (DCiM) architecture. FlexCiM partitions the DCiM macro into smaller submacros, which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. We conduct experiments across different LLM families and state-space models and conclusively demonstrate that the proposed algorithm-hardware co-design framework achieves up to 36% higher accuracy,$1.75\times $faster inference, and$1.5\times $lower energy consumption compared to existing alternatives, establishing an effective balance between flexibility and hardware efficiency in sparse LLM inference. Code is available at:https://github.com/FLOW-open-project/FLOW
Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak Mathaikutty, Tushar Krishna
IEEE Trans. Very Large Scale Integr. Syst.5
2025 GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
abstract
Graph Neural Networks (GNNs) are crucial for learning and reasoning over graph-structured data, with applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices, such as client PCs and laptops, enables real-time processing, enhances privacy, and reduces cloud dependency. For instance, GNNs can augment Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparse graphs, and dynamic structures lead to high latency and energy consumption on resource-constrained devices. Modern edge processors combine CPUs, GPUs, and NPUs, where NPUs excel at data-parallel tasks but face challenges with irregular GNN computations. To address these gaps, we present GraNNite, the first hardware-aware framework tailored to optimize GNN deployment on commercial-off-the-shelf (COTS) state-of-the-art (SOTA) DNN accelerators using a systematic three-step methodology: (1) enabling GNN execution on NPUs, (2) optimizing performance, and (3) trading accuracy for further performance and energy efficiency gains. Towards that end, the first category includes techniques such as GraphSplit for workload distribution and StaGr for static graph aggregation, while GrAd and NodePad handle real-time updates for dynamic graphs. Next, performance improvement is acquired through techniques such as EffOp for control-heavy operations and GraSp for sparsity exploitation. For Graph Convolution layers, PreG, SymG, and CacheG reduce redundancy and memory transfers. The final class of techniques deals with quality vs efficiency tradeoffs – QuantGr applies INT8 quantization to lower memory usage and computation time, while GrAx1, GrAx2, and GrAx3 optimize graph attention, broadcast-add, and sample-and-aggregate (SAGE)-max aggregation for higher throughput with minimal quality loss. Experimental evaluations on Intel® Core™ Ultra Series 1 and 2 AI PCs demonstrate that GraNNite achieves speedups of 2.6× to 7.6× over default NPU mappings, with energy efficiency improvements up to 8.6× compared to CPUs and GPUs. Across various GNN models, GraNNite delivers up to 10.8× and 6.7× higher performance than CPUs and GPUs, respectively. Our code implementation is available at this link.
Arghadip Das, Shamik Kundu, Arnab Raha, Soumendu Kumar Ghosh, Deepak Mathaikutty, Vijay Raghunathan
IJCNN5
2025 OpenAssert: Towards Secure Assertion Generation using Large Language Models
abstract
Assertions are critical components used in hardware verification, ensuring robust functionality, fortifying design security, and providing essential verification features. Traditional hardware assertion methods are not automated, complicate security audits, and require effort, causing prolonged development cycles. Recent studies have highlighted the potential of commercial Large Language Models (LLMs) to generate security-focused assertions by leveraging textual data from design specifications. However, reliance on proprietary models like GPT-4 severely jeopardizes IP privacy and data confidentiality, undermining transparency and accountability in data handling practices. In this paper, we address secure hardware assertion generation by proposing a practical approach to significantly enhance the feasibility of open-source LLMs. Our proposed method, OpenAssert, involves fine-tuning existing models to be utilized locally at the user’s end without compromising confidentiality. Additionally, we employ Retrieval Augmentation Generation to refine these models, mitigating hallucinations and security-related errors. OpenAssert demonstrates improvements, achieving up to a 44% increase in rouge-1 score, a 49% improvement in cosine similarity, and a 43.4% reduction in word error rate for security-critical designs compared to open-source models.
Anand Menon, Samit Shahnawaz Miftah, Amisha Srivastava, Shamik Kundu, Shovik Kundu, Arnab Raha, Suvadeep Banerjee, Deepak Mathaikutty, Kanad Basu
VTS8
2024 SwiSS: Switchable Single-Sided Sparsity-based DNN Accelerators
abstract
Deep Neural Networks (DNNs) exhibit sparsity in both activation and weight tensors, but certain layers have higher weight sparsity, while others have higher activation sparsity. This challenges the conventional approach of fixing sparsity acceleration to either weights or activations alone. Conversely, harnessing both-sided sparsity necessitates complex design logic for identifying participating non-zero weights and activation pairs during a multiply-accumulate operation, leading to a significant impact on energy efficiency and area overhead in the edge accelerator. In this paper, we, for the first time, exploit the unbalanced sparsity in DNNs to propose the concept of dynamically Switchable Single-sided Sparsity, SwiSS, to improve energy efficiency in edge DNN accelerators. Through a novel self-adaptive dynamic sparsity selection algorithm, SwiSS can determine whether to enable one-sided weight or one-sided activation sparsity for a sparsity-enabled DNN accelerator. This capability allows SwiSS to dynamically exploit both sides of sparsity while maximizing the associated power and area benefits in the accelerator. Evaluation on state-of-the-art network-dataset configurations conducted on FlexNN [12] accelerator architecture demonstrates that SwiSS yields up to 30.76% and 8.29% improvements in power and area overheads, respectively (which translates to 1.42X and 1.08X improvement in TOPS/W and TOPS/mm2, respectively), compared to a combined two-sided sparsity scenario, with a negligible drop in sparsity acceleration.
Shamik Kundu, Soumendu Kumar Ghosh, Arnab Raha, Deepak Mathaikutty
ISLPED4
2021 Special Session: Approximate TinyML Systems: Full System Approximations for Extreme Energy-Efficiency in Intelligent Edge Devices
abstract
Approximate computing (AxC) has advanced from being an emerging design paradigm to becoming one of the most popular and effective methods of energy optimization for applications in the domains of computer vision, image/video processing, data mining, analytics, and search. The simultaneous rise of artificial intelligence (AI) has provided an additional thrust to the adoption of various AxC techniques in intelligent edge platforms where energy-efficiency is not only desirable but necessary. In spite of the big rise in interest for AxC, the adoption of approximate hardware has mostly been limited to only one component of the system (usually the processing subsystem) which often contributes only a fraction of the overall system-level power. A full system approach to AxC enables us to extend approximations to other subsystems, such as the memory, sensor, and communications subsystems. This paper presents the foundational concepts of an approximate TinyML system that applies approximations synergistically to multiple subsystems in an edge inference device. These approximations are applied intelligently to significantly reduce energy while incurring a negligible loss in application-level quality. We demonstrate multiple versions of an approximate smart camera system that can execute state-of-the-art deep neural networks (DNNs) while consuming only a fraction of the total energy in a typical system.
Arnab Raha, Soumendu Kumar Ghosh, Debabrata Mohapatra, Deepak Mathaikutty, Raymond Sung, Cormac Brick, Vijay Raghunathan
ICCD4
2008 Formal Transformation of a KPN Specification to a GALS Implementation
abstract
Kahn process networks (KPNs) provide a model of computation for streaming audio, video and various multimedia applications. However, the KPN model consists of unbounded FIFOs between these communicating processes which need to be realized by other means. Application of a design transformation process to a KPN style specification towards a Globally asynchronous locally synchronous (GALS) implementation is one way of achieving this. Furthermore, this transformation process needs to preserve the Kahn principle. In this paper, our main contribution is the presentation of one such refinement based design transformation that preserves the Kahn principle. We present correctness preserving transformation towards a lookup-based architecture where the communication between processes is facilitated by a shared on-chip lookup storage structure. This refinement methodology is generic, and various alternate schemes of GALS implementation can be derived.
Syed Suhaib, Bijoy Antony Jose, Sandeep K. Shukla, Deepak Mathaikutty
FDL4
2008 MMV: A Metamodeling Based Microprocessor Validation Environment
abstract
With increasing levels of integration of multiple processing cores and new features to support software functionality, recent generations of microprocessors face difficult validation challenges. The systematic validation approach starts with defining the correct behaviors of the hardware and software components and their interactions. This requires new modeling paradigms that support multiple levels of abstraction. Mutual consistency of models at adjacent levels of abstraction is crucial for manual refinement of models from the full chip level to production register transfer level, which is likely to remain the dominant design methodology of complex microprocessors in the near future. In this paper, we present microprocessor modeling and validation environment (MMV), a validation environment based on metamodeling, that can be used to create models at various abstraction levels and to generate most of the important validation collaterals, viz., simulators, checkers, coverage, and test generation tools. We illustrate the functionalities in MMV by modeling a 32-bit reduced instruction set computer processor at the system, instruction set architecture, and microarchitecture levels. We show by examples how consistency across levels is enforced during modeling and also how to generate constraints for automatic test generation.
Deepak Mathaikutty, Sreekumar V. Kodakara, Ajit Dingankar, Sandeep K. Shukla, David J. Lilja
IEEE Trans. Very Large Scale Integr. Syst.1
2008 MCF: A Metamodeling-Based Component Composition Framework - Composing SystemC IPs for Executable System Models
abstract
Reusing Intellectual Property (IP)-cores accompanied by automated generation of the glue-logic, and automated composability checks can help designers to create efficient system-level models quickly and correctly for fast design space exploration. Furthermore, with the rise of multiple transaction level and register-transfer level abstractions, constructing models with mixed abstraction levels is also important. A framework that allows designers to: 1) describe the structure of components, their interfaces, and their interactions, with a semantically rich visual frontend; 2) automatically select IPs from a component library-based on sound-type theoretic principles; and 3) perform constraint based checks for composability, is highly desirable in this context. A metamodel based framework brings forth further advantages. It helps in: 1) providing rigorous semantics to the visual models; 2) imposing restrictions on the model and on interactions between components through constraints expressed in a constraint language; and 3) enabling type-checking and inference-based facilities. Furthermore, using XML-based schemas to store and process meta-information about the IPs as well as the schematic visual model, allows for an IP selection and integration methodology using existing XML processing tools. With these in mind, we present MCF, a metamodeling-based component composition framework for SystemC-based IP core composition at multiple and mixed abstraction levels, with all the advantages stated above.
Deepak Mathaikutty, Sandeep K. Shukla
IEEE Trans. Very Large Scale Integr. Syst.1
2008 A Trace-Based Framework for Verifiable GALS Composition of IPs
abstract
Composing intellectual property (IP) blocks running at different clock speeds over asynchronous communication links for a system-on-chip (SoC) design is a challenging task, especially for ensuring the functional correctness of the overall design. In this paper, we propose a trace-based framework that helps in identifying a class of IPs that can be composed to ldquocorrect-by-constructionrdquo globally asynchronous locally synchronous (GALS) designs, and their correctness is maintained with respect to their synchronous compositions. Our notion of correctness is latency equivalence. Latency equivalence means that the order of valid values is same on the corresponding signals in the synchronous as well as asynchronous compositions. We also provide a description of the protocol to be inserted between the IPs to obtain this equivalence.
Syed Suhaib, Deepak Mathaikutty, Sandeep K. Shukla
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Design fault directed test generation for microprocessor validation
abstract
Functional validation of modern microprocessors is an important and complex problem. One of the problems in functional validation is the generation of test cases that has higher potential to find faults in the design. We propose a model based test generation framework that generates tests for design fault classes inspired from software validation. There are two main contributions in this paper. Firstly, we propose a microprocessor modeling and test generation framework that generates test suites to satisfy modified condition decision coverage (MCDC), a structural coverage metric that detects most of the classified design faults as well as the remaining faults not covered by MCDC. Secondly, we show that there exists good correlation between types of design faults proposed by software validation and the errors/bugs reported in case studies on microprocessor validation. We demonstrate the framework by modeling and generating tests for the microarchitecture of VESPA, a 32-bit microprocessor. In the results section, we show that the tests generated using our framework's coverage directed approach detects the fault classes with 100% coverage, when compared to model-random test generation
Deepak Mathaikutty, Sandeep K. Shukla, Sreekumar V. Kodakara, David J. Lilja, Ajit Dingankar
DATE1
2007 A Metamodeling based Framework for Architectural Modeling and Simulator Generation
Deepak Mathaikutty, Ajit Dingankar, Sandeep K. Shukla
FDL1
2007 Type Inference for IP Composition
abstract
Type inference and type matching algorithms in the context of a component composition framework are described in this paper. These algorithms facilitate automatic construction of system models from existing SystemC IPs. The approach uses a component composition language to describe an architecture for the system under design and then through automated selection of IPs from an IP library instantiate the architecture. This approach gives rise to many typing problems, and our efficient solutions produce an effective IP-reuse based system modeling and architectural exploration tool that provides productivity gain.
Deepak Mathaikutty, Sandeep K. Shukla
MEMOCODE1
2007 EWD: A metamodeling driven customizable multi-MoC system modeling framework
abstract
We present the EWD design environment and methodology, a modeling and simulation framework suited for complex and heterogeneous embedded systems with varying degrees of expressibility and modeling fidelity. This environment promotes the use of multiple models of computation (MoCs) to support heterogeneity and metamodeling for conformance tests of syntactic and static semantics during the process of modeling. Therefore, EWD is a multiple MoC modeling and simulation framework that ensures conformance of the MoC formalisms during model construction using a metamodeling approach. In addition, EWD provides a suite of translation tools that generate executable models for two simulation frameworks to demonstrate its language-independent modeling framework. The EWD methodology uses the Generic Modeling Environment for customization of the MoC-specific modeling syntax into a visual representation. To embed the execution semantics of the MoCs into the models, we have built parsing and translation tools that leverage an XML-based interoperability language. This interoperability language is then translated into executable Standard ML or Haskell models that can also be analyzed by existing simulation frameworks such as SML-Sys or ForSyDe. In summary, EWD is a metamodeling driven multitarget design environment with multi-MoC modeling capability.
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla, Axel Jantsch
ACM Trans. Design Autom. Electr. Syst.1
2006 Mining Metadata for Composability of IPs from SystemC IP Library
Deepak Mathaikutty, Sandeep K. Shukla
FDL1
2006 MCF: A Metamodeling-based Visual Component Composition Framework
Deepak Mathaikutty, Sandeep K. Shukla
FDL1
2006 Validating Families of Latency Insensitive Protocols
abstract
With increasing clock frequencies, the signal delay on some interconnects in a System on Chip (SoC) often exceeds the clock period, which necessitates latency insensitive protocols (LIPs). The correctness of a system composed of synchronous blocks communicating via LIPs is established by showing latency equivalence between a completely synchronous composition of the blocks, and the LIP-based composition. Every time a new LIP is conceived, it needs to be debugged and then proven correct. Mathematical theorems to establish correctness, though elegant, are error prone, and tedious to create for every new variant of LIPs. In this work, we present validation frameworks for families of LIPs, both for dynamic validation, useful for early debug cycles, and formal verification for formal proof of correctness. This can be a useful framework in the hands of designers trying to create new LIPs or to optimize existing ones for design convergence.
Syed Suhaib, Deepak Mathaikutty, David Berner, Sandeep K. Shukla
IEEE Trans. Computers2
2006 CARH: service-oriented architecture for validating system-level designs
abstract
Existing system-level design languages (SLDLs) and frameworks mainly provide a modeling and a simulation framework. However, there is an increasing demand for supporting tools to aid designers in quick and faster design space and architectural exploration. As a result, numerous tools such as integrated development environments (IDEs) and others that help in debugging, visualization, validation, and verification are commonly employed by designers. As with most tools, they are targeted for a specific purpose, making it difficult for designers to possess all desired features from one particular tool. Only public-domain tools can be easily extended or interfaced with other existing tools, which a lot of the existing commercial tools do not promote. Having an extendable framework allows designers to implement their own desirable features and incorporate them into their framework. However, for technology reuse and transfer, it is important to have a tidy infrastructure for interfacing the extension with the framework, such that the added solution is not highly coupled with the environment, making distribution and deployment to other frameworks difficult, if not impossible. This requires a plug-and-play framework where features can be easily integrated. These issues of extendibility, deployment, and the inadequacies in SLDLs and frameworks are tackled by presenting a service-oriented architecture for validating SLDs for SystemC, called CARH, We code name our software systems after famous computer scientists. CARH which uses a variety of open-source technologies such as Doxygen, Apache's Xerces extensible markup language parsers, SystemC, and the adaptive communication environment (ACE) object request broker.
Hiren D. Patel, Deepak Mathaikutty, David Berner, Sandeep K. Shukla
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 SystemCXML: An Exstensible SystemC Front end Using XML
David Berner, Jean-Pierre Talpin, Hiren D. Patel, Deepak Mathaikutty, Sandeep K. Shukla
FDL4
2005 Modelling Environment for Heterogeneous Systems based on MoCs
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla, Axel Jantsch
FDL1
2005 XFM: An incremental methodology for developing formal models
abstract
We present an agile formal methodology named eXtreme Formal Modeling (XFM), based on Extreme Programming (XP) concepts to construct abstract models from natural language specifications of complex systems. In particular, we focus on Prescriptive Formal Models (PFMs) that capture the specification of the system under design in a mathematically precise manner. Such models can be used as golden reference models for formal verification, test generation, coverage monitor generation, etc. This methodology for incrementally building PFMs works by adding user stories expressed as LTL formulae gleaned from the natural language specifications, one by one, into the model. XFM builds the models, retaining correctness with respect to incrementally added properties by regressively model-checking all the LTL properties captured theretofore in the model. We illustrate XFM with a graded set of examples consisting of a traffic light controller and a DLX pipeline. To make the regressive model-checking steps feasible with current model-checking tools, we need to control the model size increments at each subsequent step in the process. We therefore analyze the effects of ordering the LTL properties in XFM on the statespace growth rate of the model. We compare three different property-ordering methodologies: ad hoc ordering, property-based ordering, and predicate-based ordering. We experiment on the models of the ISA bus monitor and the arbitration phase of the Pentium Pro bus. We experimentally show and mathematically reason that the predicate-based ordering is the best among these orderings. Finally, we present a GUI-based toolbox that we implemented to build PFMs using XFM.
Syed Suhaib, Deepak Mathaikutty, Sandeep K. Shukla, David Berner
ACM Trans. Design Autom. Electr. Syst.2
2004 A Functional Programming Framework of Heterogeneous Model of Computation for System Design
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla
FDL1