Dinesh Pamunuwa

dblp:67/3424 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
5since 2021 · last 2026
0000-0002-4838-7932ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Clock-Independent, Time-Domain Rapid Calibration Method for Memristor-based Analog Computing AI Processors
abstract
Memristor-based analog computing shows promising energy efficiency compared to digital counterparts. High- performance memristor-based AI processor has multi-level memristors, combined with the effect of device retention and ageing, the memristor readout system has tight readout margin to accurately determining the state of the memristor. Because the accuracy of the readout comparator determines the maximum accuracy the system can achieve. To address this, we proposed a high precision, clock-independent offset calibration scheme. The calibration scheme uses edge-triggered pulse-based calibration signal that makes the calibration accuracy independent of clock speed. The coarse-fine calibration steps are used to rapidly converge to minimum offset with high precision. A built-in digital calibration flag indicates the calibration status, that enables the synchronisation across arrays of comparators to facilitate scaling. The comparator is designed, laid out and simulated using 12 nm FinFET process. The comparator achieves average input voltage offset of 120 μV with 50 μV standard deviation after calibration, while the conventional method has the offset of 1.35 mV with standard deviation of 390 μV.
Matthew Schormans, Zhiqiang Que, Dinesh Pamunuwa, Jiayang Li 0002
ISCAS4
2025 Nanoelectromechanical Binary Comparator for Edge-Computing Applications
abstract
Bitwise comparison is a fundamental operation in many digital arithmetic functions and is ubiquitous in both datapath and control elements; for example, many machine learning algorithms depend on binary comparison. This work proposes a new class of binary comparator circuit using 4-terminal nanoelectromechanical (NEM) relays that use just 6 devices compared to 9 transistors in CMOS implementations. Moreover, NEM implementations are capable of withstanding much higher temperatures, up to$300^{\circ}\mathrm{C}$, and radiation levels, well over 1 Mrad absorbed dose, conditions which are common across many industrial edge applications, with near zero standby power. A 1-bit magnitude and equality comparators comprising two in-plane silicon 4-terminal relays each were fabricated on a silicon-on-insulator substrate and electrically characterized for proof of concept, the first such demonstration. Using the 1-bit comparators as building blocks, a scalable tree-based topology is proposed to implement higher-order comparators, resulting in ≈47% reduction in device count over a CMOS implementation for a 64-bit comparator. Circuit level simulations of the comparators using accurate device models show that a single operation consumes at most 21 fJ a 9-fold reduction over the best CMOS offering in an equivalent process node.
Victor Marot, Manu Bala Krishnan, Mukesh Kumar Kulsreshath, Elliott Worsey, Roshan Weerasekera, Dinesh Pamunuwa
DATE6
2025 Nanomechanical Relay-based Switchblock For FPGA Interconnect
abstract
Reprogrammable non-volatile nanomechanical (NEM) switches have single pole double throw functionality alongside zero standby power that allows the creation of four-way switch nodes made of four NEM switches as opposed to 36 or higher for a SRAM-based implementation. The reduced device count per switch node and accompanying reductions in dynamic power, in turn, allows the creation of island-style FPGA switchblocks with more flexibility than traditional switchblocks by enabling switch matrices fully populated with switch nodes. Simulations for a matrix switchblock within a Stratix IV architecture showed an average reduction of over 40% in the channel width and total critical path wire length, as well as a reduction of nearly 50% in total wire length, over a Wilton topology switchblock. The improved routing efficiency results in more efficient mapping, reduced critical path latency, and an overall reduction in the chip size to map a given function. These improvements, in combination with near total elimination of leakage power, makes a compelling case for NEM-CMOS FPGAs, especially for edge computing applications where energy and footprint constraints are far more stringent.
Victor Marot, Dinesh Pamunuwa
ISCAS2
2025 Measurement and Analysis of Dynamic Energy Consumption in Microelectromechanical Relays
abstract
Electrostatically operated micro and nanoelectromechanical (MEM/NEM) relays have been proposed as digital switches to replace transistors due to their sharp turn-on/off transient, zero leakage current between drain and source in the off state, and capability to operate at far higher temperatures and radiation levels than CMOS. However, the dynamic energy associated with charging the gate capacitance have not been investigated or verified to date. Here, we present a detailed analysis starting from first principle formulations and derive a new closed form formula for the dynamic energy consumption of MEM/NEM relays as a function of the pull-in voltage, and open-state and closed-state gate capacitances. We compare against measurements carried out on prototype silicon MEM relays to verify our analysis. We also use the derived analytic model for energy consumption to carry out a scaling study, and compare against CMOS for room temperature and high temperature operation, and show that significant energy savings are possible. The models, analyses and measurement methodologies presented here constitute a set of essential techniques for accurate estimation of the energy consumption of MEM relays in ultra-low power circuit applications.
Elliott Worsey, Mukesh Kumar Kulsreshath, Simon J. Bleiker, Harold M. H. Chong, Frank Niklaus, Dinesh Pamunuwa
ISCAS9
2024 Nanoelectromechanical analog-to-digital converter for low power and harsh environments
abstract
State of the art analog to digital conversion at the edge prioritises either low power consumption, or environmental resilience. Nanoelectromechanical relays show promise as a solution to both with low leakage current and operation within elevated temperatures and radiation levels. This paper introduces an analog to digital flash converter design constructed from four-terminal nanoelectromechanical relays, each selectively biased to partition the voltage domain. Supported by three-terminal relays forming a binary encoder, 2-bit and 3-bit implementations of the design are simulated in Cadence to give estimates for sampling rates (444kHz,278kHz) and power consumption (15.58µW,32.95µW).
Elliott Worsey, Manu Bala Krishnan, Mukesh Kumar Kulsreshath, Dinesh Pamunuwa
ISCAS5
2015 Design Methodologies, Models and Tools for Very-Large-Scale Integration of NEM Relay-Based Circuits
abstract
Integrated circuits based on nano-electro-mechanical (NEM) relays are a promising alternative to conventional CMOS technology in ultra-low energy applications due to their (near) zero stand-by energy consumption. Here we describe the details of an overarching design framework for NEM relays, including automated synthesis from design entry in RTL to layout, based on commercially available EDA tools and engines. Critical differences between relays and FETs manifest in fundamentally different timing characteristics, which significantly affect static timing analysis and the requisite timing models. The adaptation of existing EDA methods, models, tools and platforms for logic and physical synthesis to account for these differences are described, providing insight into large-scale design of NEM relay-based digital processors. A historically well-known processor, the Intel 4004, and a modern MIPS32 compatible processor are synthesized based on a NEM relay-based standard cell library to demonstrate the customized synthesis methodology. An energy study is carried out using the proposed design framework on benchmark circuits implemented in existing CMOS nodes and NEM node, to better understand the energy saving potential of NEM technology.
Sunil Rana, Dinesh Pamunuwa
ICCAD3
2013 Modelling NEM relays for digital circuit applications
abstract
A reduced-order model for NEM relays is presented that combines electro-mechanical beam actuation and landing of beam tip on the surface electrode. This model shows a deviation of less than 2%, for the DC as well as the transient response for beam actuation in a circuit simulation, when compared to a finite-element simulation. It also shows an excellent match for the energy. The model allows accurate circuit simulation to aid in NEM-relay based logic design, and facilitates the quantification of key gate-level metrics.
Sunil Rana, Dinesh Pamunuwa, Daniel Grogg, Michel Despont, Yu Pu, Christoph Hagleitner
ISCAS3
2011 Modeling the computational efficiency of 2-D and 3-D silicon processors for early-chip planning
abstract
Hierarchical models from physical to system-level are proposed for architectural exploration of high-performance silicon systems to quantify the performance and cost trade offs for 2-D and 3-D IC implementations. We show that 3-D systems can reduce interconnect delay and energy by up to an order of magnitude over 2-D, with an increase of 20-30% in performance-per-watt for every doubling of stack height. Contrary to previous analysis, the improved energy efficiency is achievable at a favorable cost. The models are packaged as a standalone tool and can provide fast estimation of coarse-grain performance and cost limitations for a variety of processing systems to be used at the early chip-planning phase of the design cycle.
Matt Grange, Axel Jantsch, Roshan Weerasekera, Dinesh Pamunuwa
ICCAD4
2011 Optimal network architectures for minimizing average distance in k-ary n-dimensional mesh networks
abstract
A general expression for the average distance for meshes of any dimension and radix, including unequal radices in different dimensions, valid for any traffic pattern under zero-load condition is formulated rigorously to allow its calculation without network-level simulations. The average distance expression is solved analytically for uniform random traffic and for a set of local random traffic patterns. Hot spot traffic patterns are also considered and the formula is empirically validated by cycle true simulations for uniform random, local, and hot spot traffic. Moreover, a methodology to attain closed-form solutions for other traffic patterns is detailed. Furthermore, the model is applied to guide design decisions. Specifically, we show that the model can predict the optimal 3-D topology for uniform and local traffic patterns. It can also predict the optimal placement of hot spots in the network. The fidelity of the approach in suggesting the correct design choices even for loaded and congested networks is surprising. For those cases we studied empirically it is 100%.
Matt Grange, Roshan Weerasekera, Dinesh Pamunuwa, Axel Jantsch, Awet Yemane Weldezion
NOCS3
2011 3-D integration and the limits of silicon computation
abstract
The intrinsic computational efficiency (ICE) of silicon defines the upper limit of the amount of computation within a given technology and power envelope. The effective computational efficiency (ECE) and the effective computational density (ECD) of silicon, by taking computation, memory and communication into account, offer a more realistic upper bound for computation of a given technology. Among other factors, they consider how distributed the memory is, how much area is occupied by computation, memory and interconnect, and the geometric properties of 3-D stacked technology with through silicon vias (TSV) as vertical links. We use the ECE and ECD to study the limits of performance under different memory distribution, power, thermal and cost constraints for various 2-D and 3-D topologies, in current and future technology nodes.
Dinesh Pamunuwa, Matt Grange, Roshan Weerasekera, Axel Jantsch
VLSI-SoC1
2010 On signalling over Through-Silicon Via (TSV) interconnects in 3-D Integrated Circuits
abstract
This paper discusses signal integrity (SI) issues and signalling techniques for Through Silicon Via (TSV) interconnects in 3-D Integrated Circuits (ICs). Field-solver extracted parasitics of TSVs have been employed in Spice simulations to investigate the effect of each parasitic component on performance metrics such as delay and crosstalk and identify a reduced-order electrical model that captures all relevant effects. We show that in dense TSV structures voltage-mode (VM) signalling does not lend itself to achieving high data-rates, and that current-mode (CM) signalling is more effective for high throughput signalling as well as jitter reduction. Data rates, energy consumption and coupled noise for the different signalling modes are extracted.
Roshan Weerasekera, Matt Grange, Dinesh Pamunuwa, Hannu Tenhunen
DATE3
2009 Design of Robust Molecular Electronic Circuits
abstract
A compact model for molecular electronic devices that considers the accumulation of charge on the molecule by an equivalent capacitive charging process which is suitable for transient analyses is presented. The model is used to examine the viability of molecular devices in future electronics applications. Further Monte Carlo simulations are carried out to examine digital behaviour in the face of a statistical spread in device characterising parameters, revealing relatively robust circuit behaviour and showcasing the model's capability to relate fundamental physical constants governing the quantum mechanical behavior of the device to abstract figures of merit.
Ci Lei, Dinesh Pamunuwa, Steven Bailey, Colin Lambert
ISCAS2
2009 Scalability of network-on-chip communication architecture for 3-D meshes
abstract
Design constraints imposed by global interconnect delays as well as limitations in integration of disparate technologies make 3D chip stacks an enticing technology solution for massively integrated electronic systems. The scarcity of vertical interconnects however imposes special constraints on the design of the communication architecture. This article examines the performance and scalability of different communication topologies for 3D network-on-chips (NoC) using through-silicon-vias (TSV) for inter-die connectivity. Cycle accurate RTL-level simulations are conducted for two communication schemes based on a 7-port switch and a centrally arbitrated vertical bus using different traffic patterns. The scalability of the 3D NoC is examined under both communication architectures and compared to 2D NoC structures in terms of throughput and latency in order to quantify the variation of network performance with the number of nodes and derive key design guidelines.
Awet Yemane Weldezion, Matt Grange, Dinesh Pamunuwa, Zhonghai Lu, Axel Jantsch, Roshan Weerasekera, Hannu Tenhunen
NOCS3
2009 Two-Dimensional and Three-Dimensional Integration of Heterogeneous Electronic Systems Under Cost, Performance, and Technological Constraints
abstract
Present day market demand for high-performance high-density portable hand-held applications has shifted the focus from 2-D planar system-on-a-chip-type single-chip solutions to alternatives such as tiled silicon and single-level embedded modules as well as 3-D die stacks. Among the various choices, finding an optimal solution for system implementation deals usually with cost, performance, power, thermal, and technological tradeoff analyses at the system conceptual level. It has been estimated that decisions made in the first 20% of the design cycle influence up to 80% of the final product cost. In this paper, we discuss realistic metrics appropriate for performance and cost tradeoff analyses both at the system conceptual level in the early stages of the design cycle and in the implementation phase, for verification. In order to validate the proposed metrics and methodology, two ubiquitous electronic systems are analyzed under various implementation schemes and the performance tradeoffs discussed. This case study is used to highlight the importance of a cost and performance tradeoff analysis early in the design flow.
Roshan Weerasekera, Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 Memory Technology for Extended Large-Scale Integration in Future Electronics Applications
abstract
Extending 2-D planar topologies in integrated circuits (ICs) to a 3-D implementation has the obvious benefits of reducing the overall footprint and average interconnection length, with associated improvements in cost, and delay and energy consumption, while also providing an opportunity to integrate disparate technologies. Such advances are very much technology driven, and early research into 3-D integration has now crystallised into commercially viable options that are being pursued by many companies. Being able to position memory in closer proximity to processing elements in a NoC architecture as afforded by a 3-D physical architecture has the potential to improve the memory bandwidth and mitigate the general nature of delay constrained performance in IC design. Understanding the nature of the opportunities and constraints provided in such a 3-D physical architecture is crucial in realising the true benefits of 3-D integration in future applications.
Dinesh Pamunuwa
DATE1
2008 Minimal-Power, Delay-Balanced Smart Repeaters for Global Interconnects in the Nanometer Regime
abstract
A smart repeater is proposed for driving capacitively-coupled, global-length on-chip interconnects that alters its drive strength dynamically to match the relative bit pattern on the wires and thus the effective capacitive load. This is achieved by partitioning the driver into main and assistant drivers; for a higher effective load capacitance both drivers switch, while for a lower effective capacitance the assistant driver is quiet. In a UMC 0.18-mum technology the potential energy saving is around 10% and the reduction in jitter 20%, in comparison to a traditional repeater for typical global wire lengths. It is also shown that the average energy saving for nanometer technologies is in the range of 20% to 25%. The driver architecture exploits the fact that as feature sizes decrease, the capacitive load per transistor shrinks, whereas global wire loads remain relatively unchanged. Hence, the smaller the technology, the greater the potential saving.
Roshan Weerasekera, Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Extending systems-on-chip to the third dimension: performance, cost and technological tradeoffs
abstract
Because of the today’s market demand for high- performance, high-density portable hand-held applications, elec- tronic system design technology has shifted the focus from 2-D planar SoC single-chip solutions to different alternative options as tiled silicon and single-level embedded modules as well as 3- D integration. Among the various choices, finding an optimal solution for system implementation dealt usually with cost, performance and other technological trade-off analysis at the system conceptual level. It has been identified that the decisions made within the first 20% of the total design cycle time will ultimately result upto 80% of the final product cost. In this paper, we discuss appropriate and realistic metric for performance and cost trade-off analysis both at system conceptual level (up-front in the design phase) and at implementation phase for verification in the three-dimensional integration. In order to validate the methodology, two ubiquitous electronic systems are analyzed under various implementation schemes and discuss the pros and cons of each of them.
Roshan Weerasekera, Lirong Zheng 0001, Dinesh Pamunuwa, Hannu Tenhunen
ICCAD3
2005 Modeling delay and noise in arbitrarily coupled RC trees
abstract
Closed-form equations for second-order transfer functions of general arbitrarily coupled resistance-capacitance (RC) trees with multiple drivers are reported. The models allow precise delay and noise calculations for systems of coupled interconnects with guaranteed stability and represent the minimum complexity associated with this class of circuits. Their accuracy is extensively compared against other relevant models and is found to be better or comparable to more expensive models. All results are derived from a theoretical approach, and their physical basis is examined. The simplicity, accuracy, and generality of the models make them suitable for use in early signal integrity analyses of complex systems and incremental physical optimization.
Dinesh Pamunuwa, Shauki Elassaad, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 A study on the implementation of 2-D mesh-based networks-on-chip in the nanometre regime
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch, Hannu Tenhunen
Integr.1
2003 Analytic Modeling of Interconnects for Deep Sub-Micron Circuits
Dinesh Pamunuwa, Shauki Elassaad, Hannu Tenhunen
ICCAD1
2003 Layout, Performance and Power Trade-Offs in Mesh-Based Network-on-Chip Architectures
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch
VLSI-SOC1
2003 Maximizing throughput over parallel wire structures in the deep submicrometer regime
abstract
In a parallel multiwire structure, the exact spacing and size of the wires determine both the resistance and the distribution of the capacitance between the ground plane and the adjacent signal carrying conductors, and have a direct effect on the delay. Using closed-form equations that map the geometry to the wire parasitics and empirical switch factor based delay models that show how repeaters can be optimized to compensate for dynamic effects, we devise a method of analysis for optimizing throughput over a given metal area. This analysis is used to show that there is a clear optimum configuration for the wires which maximizes the total bandwidth. Additionally, closed form equations are derived, the roots of which give close to optimal solutions. It is shown that for wide buses, the optimal wire width and spacing are independent of the total width of the bus, allowing easy optimization of on-chip buses. Our analysis and results are valid for lossy interconnects as are typical of wires in submicron technologies.
Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.1
2000 Combating digital noise in high speed ULSI circuits using binary BCH encoding
abstract
Increased integration in deep submicron (DSM) technologies has caused very high increases in the RLC parasitics which affect the coupling of noise to signals propagating over interconnect. Error free transmission on-chip will no longer be guaranteed, This paper examines the issue of high speed signaling in DSM and proposes the use of particular BCH codes to improve the bit error rate in the face of noise. We conclude from our results that it is possible to achieve a considerable coding gain by choosing the code properly.
Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
ISCAS1