Dwaipayan Biswas

dblp:133/2215 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 4 first-author · 15 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Thermal Insights of 3-D BS-PDN in Cloud Server SoC Using TCAD Modeling
abstract
In this brief, the thermal performance of a large-scale cloud server system-on-chip (SoC) with the backside power delivery network (BS-PDN) and 3-D integration in memory-on-logic (MoL)/logic-on-memory (LoM) configuration with 2.5-D packaging is analyzed in advanced A10 nanosheet technology node using Sentaurus TCAD platform. The results show a 45.6% (~20.3 K) thermal penalty for the 80-core SoC in MoL with BS-PDN compared with the 2-D-baseline frontside PDN (FS-PDN), using a heatsink with forced cooling. A nonuniform power map further aggravates thermal concerns, which can be mitigated using an LoM configuration with BS-PDN, reducing the penalty to 22% (~15 K). Extending the study to a 320-core SoC, in conjunction with an advanced cooling system, LoM with BS-PDN shows 45.3% (~29 K) lower temperature than conventional MoL BS-PDN. The modeling results provide valuable insights and motivate future research into packaging and cooling techniques for BS-PDN integration.
Subrat Mishra, Herman Oprins, James Myers, Julien Ryckaert, Pieter Woltgens, Dwaipayan Biswas
IEEE Trans. Very Large Scale Integr. Syst.7
2025 TAXI: Traveling Salesman Problem Accelerator with X-bar-based Ising Macros Powered by SOT-MRAMs and Hierarchical Clustering
abstract
Ising solvers with hierarchical clustering have shown promise for large-scale Traveling Salesman Problems (TSPs), in terms of latency and energy. However, most of these methods still face unacceptable quality degradation as the problem size increases beyond a certain extent. Additionally, their hardwareagnostic adoptions limit their ability to fully exploit available hardware resources. In this work, we introduce TAXI – an inmemory computing-based TSP accelerator with crossbar(Xbar)-based Ising macros. Each macro independently solves a TSP subproblem, obtained by hierarchical clustering, without the need for any off-macro data movement, leading to massive parallelism. Within the macro, Spin-Orbit-Torque (SOT) devices serve as compact energy-efficient random number generators enabling rapid “natural annealing”. By leveraging hardware-algorithm co-design, TAXI offers improvements in solution quality, speed, and energy-efficiency on TSPs up to $\mathbf{8 5, 9 0 0}$ cities (the largest TSPLIB instance). TAXI produces solutions that are only $22 \%$ and $20 \%$ longer than the Concorde solver’s exact solution on $\mathbf{3 3, 8 1 0}$ and $\mathbf{8 5, 9 0 0}$ city TSPs, respectively. TAXI outperforms a current state-of-the-art clustering-based Ising solver, being $8 \times$ faster on average across 20 benchmark problems from TSPLib.
Sangmin Yoo, Amod Holla, Sourav Sanyal, Dong Eun Kim, Francesca Iacopi, Dwaipayan Biswas, James Myers, Kaushik Roy 0001
DAC6
2025 Late Breaking Results: Thermal Feasibility of Backside Integrated LDOs in 2.5D/3D System-in-Package Using Nanosheet Technology
abstract
Digital Low Dropout Regulators (LDOs) are an excellent candidate for area-efficient fine-grain power management in heterogeneous systems, leveraging integrated power switches. Relocating the power switches to the backside of the wafer in conjunction with the Backside Power Delivery Network (BSPDN) layer is envisaged as a System Technology Co-Optimization (STCO) booster for finer grain power management and reduced area/cost. We perform a detailed thermal analysis using power-switch-based LDOs enabling per-core DVFS for a high-performance server 3D computing chiplet in a Nanosheet CMOS (A10) technology node with BSPDN. While BSPDN introduces thermal penalties due to a lack of lateral heat spreading, our high-resolution thermal simulations explore the feasibility of moving LDOs to the backside. Increasing the LDO area from 5% to 50% of the backside die area effectively lowers the 2.5/3D System-in-Package (SiP) peak temperature, confirming that thermal concerns do not impede backside LDO integration. This study supports the cost-effective design of next-generation SiPs by demonstrating no adverse thermal impact for relocating power switches to the wafer backside in the nanosheet era.
Yukai Chen, Subrat Mishra, Julien Ryckaert, Dwaipayan Biswas, James Myers
DATE4
2025 Framework for Augmenting Main Memory with CXL-connected Emerging Memory Alternatives
abstract
The rapid evolution of memory technologies and the advent of Compute Express Link (CXL) have opened up new possibilities for scaling main memory by enabling hybrid memory systems with pooled and shared content. System-level evaluation of new memory systems during the early development stage is important for the enablement and further integration of new memory and interconnect technologies. However, existing solutions do not offer a framework neither for emerging memory protocols nor for novel memory technologies. This paper introduces CXL-HMEM to evaluate emerging CXL-based hybrid main memory architectures by applying the System Technology Co-Optimization (STCO) technique. The framework provides flexible performance metrics, workload simulation, and memory traffic analysis to assess system performance under various hybrid memory configurations, including DRAM and tiered memory hierarchies. Key features include support for memory technologies such as IGZO-based DRAM (IGZO) and FeRAM, workload scalability, and an integrated model of the CXL behavior. CXL-HMEM shows that emerging memories can improve system bandwidth and energy consumption by 7%, while having potential to further mitigate particular bottlenecks. CXL-based hybrid main memory can speed up the memory access time by >2× compared to conventional approaches of main memory extension.
Khakim Akhunov, Dwaipayan Biswas, Emil Karimov, Arvind Sharma, Hyungrock Oh, Maarten Rosmeulen, Julien Ryckaert, James Myers
ISCAS2
2025 3D SRAM Disaggregation in Advanced CMOS Nodes using Hybrid Bonding Technology
abstract
This paper studies the potential of hybrid bonded Array-under-CMOS (AuC) technology to partition logic and high-performance L1 cache in advanced technology nodes. By decoupling the SRAM bitcells from the logic tier, we achieve independent optimization of both SRAM and logic devices, as well as back-end-of-line (BEOL) interconnects. Heterogeneous integration and BEOL aspect ratio optimization is implemented with different technology nodes to mitigate the delay penalty due to hybrid bond pad staggering. Addressing the performance degradation associated with scaled technology nodes, we investigate the impact of word-line (WL) and bit-line (BL) resistance on SRAM performance. Leveraging the flexibility of decoupled SRAM BEOL and within the AuC technology framework, we explore the sensitivity to WL and BL metal aspect ratios, comparing their performance against a 2D baseline. Our results demonstrate substantial performance improvements in AuC integration through two key approaches: (1) 5% enhancement via Back-End-of-Line (BEOL) optimization, and (2) 25% improvement enabled by heterogeneous integration, achieved by decoupling memory and logic tiers.
Bhawana Kumari, Anurag Swarnkar, Dawit Burusie Abdi, Fernando García-Redondo, James Myers, Julien Ryckaert, Jaydeep P. Kulkarni, Dwaipayan Biswas
ISCAS8
2025 3D IGZO Charge-Coupled Memory DTCO & STCO Analysis for Compute-near-Memory Applications
abstract
The demand for high-capacity and energy-efficient memory solutions has surged in the era of data-centric computing, particularly for Artificial Intelligence (AI) and Machine Learning (ML) workloads. This paper introduces a novel memory architecture leveraging Charge-Coupled Device (CCD) technology, engineered in a sequential-access block memory configuration, to enhance Compute-near-Memory (CnM) systems. We propose an optimized 3D IGZO CCD block memory as an on-chip weight buffer for high-capacity CnM systems. Our approach achieves 2.95−131.26× improvement in area efficiency and 1.32−4.33× improvement in energy efficiency compared to SRAM solutions.
Khakim Akhunov, Hyungrock Oh, Fernando García-Redondo, Yukai Chen, Arvind Sharma, Jiacong Sun, Sahan Gamage, Maarten Rosmeulen, Swaraj Bandhu Mahato, Rishabh Kishore, Subhali Subhechha, Jaydeep P. Kulkarni, Marian Verhelst, Dwaipayan Biswas, Marie Garcia Bardon, Wim Dehaene, Julien Ryckaert
ISCAS15
2025 SideDRAM: Integrating SoftSIMD Datapaths near DRAM Banks for Energy-Efficient Variable Precision Computation
abstract
By interfacing computing logic directly to the DRAM banks, bank-level Compute-near-Memory (CnM) architectures promise to mitigate the bottleneck at the memory interconnect. While this computation paradigm heavily reduces the energy requirements for data movement across the system, current solutions fail to co-optimize hardware and software to further increase efficiency. Instead, in this manuscript, we present SideDRAM , a co-designed bank-level CnM architecture to enable massively parallel and energy-efficient computations near DRAM. In contrast with past solutions, we support flexible data typing and heterogeneous quantization, relying on the robustness of workloads to employ small bitwidths, and enable a row-wide access to the banks to exploit parallelism and spatial locality. As a result, SideDRAM integrates (1) software-defined SIMD (SoftSIMD) datapaths, supporting low-energy computing with flexible precision, (2) an interface to the banks based on very wide registers (VWRs), enabling asymmetric data access to both utilize the full DRAM bank bandwidth and leverage data locality at the datapath, and (3) a low-overhead distributed control plane, allowing the efficient handling of variable data typing. We benchmark SideDRAM as a near-DRAM solution by analyzing the area, performance, and energy consumption of an HBM2 CnM channel executing heterogeneously quantized machine learning models. The results show that, compared to the state-of-the-art FIMDRAM design, energy improvements of up to 67% are achieved when a DeiT-S inference is executed with a batch size of 16 under the same area constraints, resulting in energy-delay-area product (EDAP) savings that reach 83%. When comparing to a massively parallel mixed-signal CnM solution, SideDRAM consistently obtains similar performance and better energy efficiency results (geomean of 15× improvement across workloads) at a lower area overhead.
Rafael Medina 0001, Pengbo Yu, Alexandre Levisse, Dwaipayan Biswas, Marina Zapater, Giovanni Ansaloni, Francky Catthoor, David Atienza 0001
ACM Trans. Embed. Comput. Syst.4
2025 Bandwidth-Latency-Thermal Co-Optimization of Interconnect-Dominated Many-Core 3D-IC
abstract
The ongoing integration of advanced functionalities in contemporary system-on-chips (SoCs) poses significant challenges related to memory bandwidth, capacity, and thermal stability. These challenges are further amplified with the advancement of artificial intelligence (AI), necessitating enhanced memory and interconnect bandwidth and latency. This article presents a comprehensive study encompassing architectural modifications of an interconnect-dominated many-core SoC targeting the significant increase of intermediate, on-chip cache memory bandwidth and access latency tuning. The proposed SoC has been implemented in 3-D using A10 nanosheet technology and early thermal analysis has been performed. Our workload simulations reveal, respectively, up to 12- and 2.5-fold acceleration in the 64-core and 16-core versions of the SoC. Such speed-up comes at 40% increase in die-area and a 60% rise in power dissipation when implemented in 2-D. In contrast, the 3-D counterpart not only minimizes the footprint but also yields 20% power savings, attributable to a 40% reduction in wirelength. The article further highlights the importance of pipeline restructuring to leverage the potential of 3-D technology for achieving lower latency and more efficient memory access. Finally, we discuss the thermal implications of various 3-D partitioning schemes in High Performance Computing (HPC) and mobile applications. Our analysis reveals that, unlike high-power density HPC cases, 3-D mobile case increases$T_{\max }$only by$2~^{\circ } $C–$3~^{\circ } $C compared to 2-D, while the HPC scenario analysis requires multiconstrained efficient partitioning for 3-D implementations.
Sudipta Das, Samuel Riedel, Mohamed Naeim, Moritz Brunion, Marco Bertuletti, Luca Benini, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic
IEEE Trans. Very Large Scale Integr. Syst.9
2024 3D Partitioning with Pipeline Optimization for Low-Latency Memory Access in Many-Core SoCs
abstract
This paper presents an investigation of System-on-Chip (SoC) communication latency optimization for 3D system integration and highlights the role of architectural modifications to maximize the Power, Performance, & Area (PPA) benefits. An instance of a highly configurable RISC-V SoC is implemented using ∼2nm nanosheet technology and different 3D stacking options using design flow from sign-off tools. The proposed implementation targets performance optimization for different 3D partitioning scenarios: Memory-on-Logic (MoL) & Logic-on-Logic (LoL). We target 2-die 3D Integrated Circuits (3D-IC) with high density 3D interconnect using Face-to-Face (F2F) hybrid bonding (∼1µm), and 3-die stack, as Face-to-Back (F2B) on top of F2F. Our analysis of the 16-core SoC instance shows that the proposed architectural optimizations bring a significant reduction of 4 pipeline stages in the design hierarchy at a marginal cost of 9% effective frequency loss when implemented in 3D in comparison to the baseline 2D architecture. Further, going from 2D to 3D allows more than 40% total system wire-length reduction & 10% less cell area, resulting in 20% power savings. These findings hold promise for further explorations on many-core SoC instances (256 & more) facing system interconnect challenges.
Sudipta Das, Samuel Riedel, Marco Bertuletti, Luca Benini, Moritz Brunion, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic
ISCAS8
2024 Bank on Compute-Near-Memory: Design Space Exploration of Processing-Near-Bank Architectures
abstract
Near-DRAM computing strategies advocate for providing computational capabilities close to where data is stored. Although this paradigm can effectively address the memory-to-processor communication bottleneck, it also presents new challenges: The strict resource constraints in the memory periphery demand careful tailoring of architectural elements. We herein propose a novel framework and methodology to explore compute-near-memory designs that interface to DRAM memory banks, demonstrating the area, energy, and performance tradeoffs subject to the architectural configuration. We exemplify this methodology by conducting two studies on compute-near-bank designs: 1) analyzing the interaction between control and data resources, and 2) exploring the integration of processing units with different DRAM standards. According to our study, the optimal size ratios between instruction and data capacity vary from$2\times $to$4\times $across benchmarks from representative application domains. The retrieved Pareto-optimal solutions from our framework improve state-of-the-art designs, e.g., achieving a 50% performance increase on matrix operations with 15% energy overhead relative to the FIMDRAM design. In addition, the exploration of DRAM shows the interplay between available internal bandwidth, performance, and area overhead. For example, a threefold increase in bandwidth rises performance by 47% across workloads at a 34% extra area cost.
Rafael Medina 0001, Giovanni Ansaloni, Marina Zapater, Alexandre Levisse, Saeideh Alinezhad Chamazcoti, Timon Evenblij, Dwaipayan Biswas, Francky Catthoor, David Atienza 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 Multidie 3-D Stacking of Memory Dominated Neuromorphic Architectures
abstract
Event-driven neuromorphic processors for artificial intelligence (AI) inference on edge/IoT devices require largeon-chip memory capacity, for efficient execution of spiking neural networks (NNs). In this work, we evaluate 3-D stacking benefits on SENECA, a digital neuromorphic accelerator core, sweeping itson-chip memory capacity from 2 up to 32 Mb in both legacy planar and advanced nanosheet CMOS logic nodes. In a planar CMOS node (GF-22 nm), two-die memory-on-logic (MoL) partitioning enables$8\times $moreon-chip memory, and it boosts operating frequency by 7% with 26% less power than the 2-D. Moving to an advanced nanosheet technology (imec A10), multidie (up to 7 dies) MoL stacking enables a performance increase of up to 29% and power savings up to 31%. Furthermore, a core folding (CF) partitioning in A10 shows up to 16% performance improvement with 12% total power savings with respect to the 2-D implementation on the same technology. We also demonstrate no thermal overhead for multidie stacking at advanced nodes for designs exhibiting low power density. These physical design explorations lay the foundation for system technology co-optimization studies for edge devices.
Leandro M. G. Rocha, Refik Bilgic, Mohamed Naeim, Sudipta Das, Herman Oprins, Amirreza Yousefzadeh, Mario Konijnenburg, Dragomir Milojevic, James Myers, Julien Ryckaert, Dwaipayan Biswas
IEEE Trans. Very Large Scale Integr. Syst.11
2024 An Energy Efficient Soft SIMD Microarchitecture and Its Application on Quantized CNNs
abstract
The ever-increasing computational complexity and energy consumption of today’s applications, such as machine learning (ML) algorithms, not only strain the capabilities of the underlying hardware but also significantly restrict their wide deployment at the edge. Addressing these challenges, novel architecture solutions are required by leveraging opportunities exposed by algorithms, e.g., robustness to small-bitwidth operand quantization and high intrinsic data-level parallelism. However, traditional hardware single instruction multiple data (Hard SIMD) architectures only support a small set of operand bitwidths, limiting performance improvement. To fill the gap, this manuscript introduces a novel pipelined processor microarchitecture for arithmetic computing based on the software-defined SIMD (Soft SIMD) paradigm that can define arbitrary SIMD modes through control instructions at run-time. This microarchitecture is optimized for parallel fine-grained fixed-point arithmetic, such as shift/add. It can also efficiently execute sequential shift-add-based multiplication over SIMD subwords, thanks to zero-skipping and canonical signed digit (CSD) coding. A lightweight repacking unit allows changing subword bitwidth dynamically. These features are implemented within a tight energy and area budget. An energy consumption model is established through post-synthesis for performance assessment. We select heterogeneously quantized (HQ) convolutional neural networks (CNNs) from the ML domain as the benchmark and map it onto our microarchitecture. Experimental results showcase that our approach dramatically outperforms traditional Hard SIMD Multiplier-Adder regarding area and energy requirements. In particular, our microarchitecture occupies up to 59.9% less area than a Hard SIMD that supports fewer SIMD bitwidths, while consuming up to 50.1% less energy on average to execute HQ CNNs.
Pengbo Yu, Flavio Ponzina, Alexandre Levisse, Mohit Gupta 0004, Dwaipayan Biswas, Giovanni Ansaloni, David Atienza 0001, Francky Catthoor
IEEE Trans. Very Large Scale Integr. Syst.5
2023 Learning-Oriented Reliability Improvement of Computing Systems From Transistor to Application Level
abstract
Due to technology scaling in modern computing platforms, the safety and reliability issues have increased tremendously, which often accelerate aging, lead to permanent faults, and cause unreliable execution of applications. Failure in some computing systems like avionics may cause catastrophic consequences. Therefore, managing reliability under all circumstances of stress and environmental changes is crucial in all abstraction layers, from application to transistor levels. Machine learning techniques are recently being employed for dynamic reliability estimation and optimization. They can adapt to varying workloads and system conditions. This paper presents reliability improvement approaches from multiple perspectives-from transistor-level to application-level-and discusses their effectiveness and limitations as well as open challenges.
Behnaz Ranjbar, Florian Klemme, Paul R. Genssler, Hussam Amrouch, Jinhyo Jung, Shail Dave, Hwisoo So, Kyongwoo Lee, Aviral Shrivastava, Ji-Yung Lin, Pieter Weckx, Subrat Mishra, Francky Catthoor, Dwaipayan Biswas, Akash Kumar 0001
DATE14
2023 Design Technology co-optimization of 1D-1VCMA to improve read performance for SCM applications
abstract
1-diode 1-Voltage controlled magnetic anisotropy (1D-1VCMA) can be an option for Storage Class Memory (SCM) to bridge the latency gap between DRAM and flash memory. It has low sneak current, high non-linearity and low IR drop. This paper presents the Design Technology Co-optimization (DTCO) study of 1D-1VCMA stack to improve the performance and energy. Thanks to precessional switching of VCMA, the write operation is very fast, but the read determines overall latency as read before write is needed to ensure reliable write operations. The read performance of 1D-1VCMA is penalized due to high VCMA MTJ resistance, hence impacting the overall performance. To improve the read performance, this paper explores two solutions: 1) reducing the VCMA RA product, and 2) improving the read circuit. These solutions improve the read performance by 36% and 260%, respectively.
Mohit Gupta 0004, Manu Perumkunnil Komalan, Dwaipayan Biswas, Saeideh Alinezhad Chamazcoti, Gouri Sankar Kar, Arnaud Furnémont, Julien Ryckaert
ISCAS3
2023 Impact of 3-D Integration on Thermal Performance of RISC-V MemPool Multicore SOC
abstract
Due to the rise in the number of cores in modern multicore architectures, 3-D integration (i.e., vertical stacking of chips) of system-on-a-chip (SOC) promises better performance due to a drastic reduction in global interconnect lengths and die footprint compared with 2-D counterparts. However, thermal issues are predominant in 3-D-SOCs due to the vertical stacking nature of chips which multiplies the transistor power density by the number of dies within the stack. Also, the reduced lateral heat spreading with aggressive die thinning degrades the ON-chip thermal performances. In this article, we investigate the thermal performance analysis of 3-D-SOC and compare the results with the 2-D-SOC designs for a MemPool multicore SOC with shared L1 scratchpad memory (SPM). Simulation results reveal that the 3-D-SOC using memory-on-logic (MOL) configuration increases the ON-chip maximum temperature by more than 20% compared with the baseline 2-D-SOC and the logic die temperature is relatively higher (3.6%) than the memory die. We also explore the impact of architectural floor-planning effects and 3-D functional partitioning on thermal performance of the MemPool instances in the 3-D-SOC with memory capacity ranging from 1 to 8 MiB and benchmarked the thermal performance with the 2-D-SOC designs. We observe that the junction-to-ambient temperature ($T_{\max }$) increases by 44% and is predominant for the SPM capacity of 8 MiB. Further investigations on various 3-D stacking configurations reveal there is an improvement in thermal performance for MOL over logic-on-memory (LOM) for L1 SPM capacity of 1, 2, and 4 MiB, and LOM over the MOL configuration for L1 SPM capacity of 8 MiB.
Sankatali Venkateswarlu, Subrat Mishra, Herman Oprins, Bjorn Vermeersch, Moritz Brunion, Jun-Han Han, Mircea R. Stan, Dwaipayan Biswas, Pieter Weckx, Francky Catthoor
IEEE Trans. Very Large Scale Integr. Syst.8
2018 CORDIC Framework for Quaternion-based Joint Angle Computation to Classify Arm Movements
abstract
We present a novel architecture for arm movement classification based on kinematic properties (joint angle and position), computed from MARG sensors, using a quaternion-based gradient-descent method and a 2-link model of the upper limb. The design based on Coordinate Rotation Digital Computer framework was validated on stroke survivors and healthy subjects performing three elementary arm movements (reach and retrieve, lift arm, rotate arm), involved in `making-a-cup-of-tea' an archetypal daily activity, achieved an overall accuracy of 78% and 85% respectively. The design coded in System Verilog, was synthesized using STMicroelectronics 130 nm technology, occupies 340K NAND2 equivalent area and consumes 292 nW @ 150 Hz, besides being functionally verified up to 25 MHz making it suitable for real-time high speed operations. The orientation, arm position and the joint angle, are computed on-the-fly, with the classification performed at the end of movement duration.
Dwaipayan Biswas, Zixuan Ye, Evangelos B. Mazomenos, Michael Jöbges, Koushik Maharatna
ISCAS1
2018 BiometricNet: Deep Learning based Biometric Identification using Wrist-Worn PPG
abstract
Rapid advances in semiconductor fabrication technology have enabled the proliferation of miniaturized body-worn sensors capable of long term pervasive biomedical signal monitoring. In this paper, we present a novel deep learning-based framework (BiometricNET) on biometric identification using data collected from wrist-worn Photoplethysmography (PPG) signals in ambulatory environments. We have formulated a completely personalized data-driven approach, using a four-layer deep neural network - employing two convolution neural network (CNN) layers in conjunction with two long short-term memory (LSTM) layers, followed by a dense output layer for modelling the temporal sequence inherent within the pulsatile signal representative of cardiac activity. The proposed network configuration was evaluated on the TROIKA dataset collected from 12 subjects involved in physical activity, achieved an average five-fold cross-validation accuracy of 96%.
Luke R. Everson, Dwaipayan Biswas, Madhuri Panwar, Dimitrios Rodopoulos, Amit Acharyya, Chris H. Kim, Chris Van Hoof, Mario Konijnenburg, Nick Van Helleputte
ISCAS2
2018 Modified Huffman based compression methodology for Deep Neural Network Implementation on Resource Constrained Mobile Platforms
abstract
Modern Deep Neural Network (DNN) architectures produce high accuracy across applications, however incur high computational complexity and memory requirements, making it challenging for execution on resource constrained mobile platforms. Driven by application requirements, there has been a shift in execution paradigm of Deep Nets from cloud based computation to sensor/mobile platforms. The limited memory available onboard a mobile platform, necessitates an effective mechanism for storage of network parameters (viz. weights) generated offline post-training. Hence, we propose a modified Huffman encoding-decoding technique, with dynamic usage of net layers, executed on-the-fly in parallel, which can be applied on a memory constrained multicore environment. To the best of our knowledge, this is the first study on applying compression based on multiple bit pattern sequences, to achieve a maximum compression rate of 64 percent and a single module decompression time of about 0.33 seconds without trading-off accuracy.
Chandrajit Pal, Sunil Pankaj, Wasim Akram, Amit Acharyya, Dwaipayan Biswas
ISCAS5
2017 Machine learning for run-time energy optimisation in many-core systems
abstract
In recent years, the focus of computing has moved away from performance-centric serial computation to energy-efficient parallel computation. This necessitates run-time optimisation techniques to address the dynamic resource requirements of different applications on many-core architectures. In this paper, we report on intelligent run-time algorithms which have been experimentally validated for managing energy and application performance in many-core embedded system. The algorithms are underpinned by a cross-layer system approach where the hardware, system software and application layers work together to optimise the energy-performance trade-off. Algorithm development is motivated by the biological process of how a human brain (acting as an agent) interacts with the external environment (system) changing their respective states over time. This leads to a pay-off for the action taken, and the agent eventually learns to take the optimal/best decisions in future. In particular, our online approach uses a model-free reinforcement learning algorithm that suitably selects the appropriate voltage-frequency scaling based on workload prediction to meet the applications' performance requirements and achieve energy savings of up to 16% in comparison to state-of-the-art-techniques, when tested on four ARM A15 cores of an ODROID-XU3 platform.
Dwaipayan Biswas, Vibishna Balagopal, Rishad A. Shafik, Bashir M. Al-Hashimi, Geoff V. Merrett
DATE1
2017 Architecture for complex network measures of brain connectivity
abstract
Cognitive and motor disorders are growing socio-economic concerns where drug treatments although being the first line of action, are not always effective in restoring cognitive and motor functionality. Research has shown that functional brain connectivity, signifying information exchange among different brain regions, is correlated with efficient execution of cognitive and motor tasks. Hence, to analyze the connectivity parameters in real-time for automated disease prognosis and control, an optimized accelerator/hardware design is required which can be integrated within the sensing device. Here we have designed and implemented an optimized hardware architecture of the graph theoretic parameters (computed concurrently) for the clinically significant functional connectivity measure (Phase Lag Index) of human brain network. To the best of our knowledge, this is a first study on the implementation of the complex network topology parameters of brain connectivity measure which has been synthesized at 25 Mhz, using STMicroelectronics 130-nm technology library and having a dynamic power consumption of 10 nW, making it amenable for real-time high speed operations.
Chandrajit Pal, Dwaipayan Biswas, Koushik Maharatna, Amlan Chakrabarti
ISCAS2
2017 Coordinate Rotation-Based Low Complexity K-Means Clustering Architecture
abstract
In this brief, we propose a low-complexity architectural implementation of the K-means-based clustering algorithm used widely in mobile health monitoring applications for unsupervised and supervised learning. The iterative nature of the algorithm computing the distance of each data point from a respective centroid for a successful cluster formation until convergence presents a significant challenge to map it onto a low-power architecture. This has been addressed by the use of a 2-D Coordinate Rotation Digital Computer-based low-complexity engine for computing the n-dimensional Euclidean distance involved during clustering. The proposed clustering engine was synthesized using the TSMC 130-nm technology library, and a place and route was performed following which the core area and power were estimated as 0.36 mm2and 9.21 mW at 100 MHz, respectively, making the design applicable for low-power real-time operations within a sensor node.
Bhagyaraja Adapa, Dwaipayan Biswas, Swati Bhardwaj 0001, Shashank Raghuraman, Amit Acharyya, Koushik Maharatna
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Low-Complexity Framework for Movement Classification Using Body-Worn Sensors
abstract
We present a low-complexity framework for classifying elementary arm movements (reach retrieve, lift cup to mouth, and rotate arm) using wrist-worn inertial sensors. We propose that this methodology could be used as a clinical tool to assess rehabilitation progress in neurodegenerative pathologies tracking occurrence of specific movements performed by patients with their paretic arm. Movements performed in a controlled training phase are processed to form unique clusters in a multidimensional feature space. Subsequent movements performed in an uncontrolled testing phase are associated with the proximal cluster using a minimum distance classifier (MDC). The framework involves performing the compute-intensive clustering on the training data set offline (MATLAB), whereas the computation of selected features on the testing data set and the minimum distance (Euclidean) from precomputed cluster centroids are done in hardware with an aim of low-power execution on sensor nodes. The architecture for feature extraction and MDC are realized using coordinate rotation digital computer-based design that classifies a movement in (9n + 31) clock cycles, n being number of data samples. The design synthesized in STMicroelectronics 130-nm technology consumed 5.3 nW at 50 Hz, besides being functionally verified up to 20 MHz, making it applicable for real-time high-speed operations. Our experimental results show that the system can recognize all three arm movements with average accuracies of 86% and 72% for four healthy subjects using accelerometer and gyroscope data, respectively, whereas for stroke survivors, the average accuracies were 67% and 60%. The framework was further demonstrated as a field-programmable gate array-based real-time system, interfacing with a streaming sensor unit.
Dwaipayan Biswas, Koushik Maharatna, Goran Panic, Evangelos B. Mazomenos, Josy Achner, Jasmin Klemke, Michael Jöbges, Steffen Ortmann
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Detecting Elementary Arm Movements by Tracking Upper Limb Joint Angles With MARG Sensors
abstract
This paper reports an algorithm for the detection of three elementary upper limb movements, i.e., reach and retrieve, bend the arm at the elbow and rotation of the arm about the long axis. We employ two MARG sensors, attached at the elbow and wrist, from which the kinematic properties (joint angles, position) of the upper arm and forearm are calculated through data fusion using a quaternion-based gradient-descent method and a two-link model of the upper limb. By studying the kinematic patterns of the three movements on a small dataset, we derive discriminative features that are indicative of each movement; these are then used to formulate the proposed detection algorithm. Our novel approach of employing the joint angles and position to discriminate the three fundamental movements was evaluated in a series of experiments with 22 volunteers who participated in the study: 18 healthy subjects and four stroke survivors. In a controlled experiment, each volunteer was instructed to perform each movement a number of times. This was complimented by a seminaturalistic experiment where the volunteers performed the same movements as subtasks of an activity that emulated the preparation of a cup of tea. In the stroke survivors group, the overall detection accuracy for all three movements was 93.75% and 83.00%, for the controlled and seminaturalistic experiment, respectively. The performance was higher in the healthy group where 96.85% of the tasks in the controlled experiment and 89.69% in the seminaturalistic were detected correctly. Finally, the detection ratio remains close ( ±6%) to the average value, for different task durations further attesting to the algorithms robustness.
Evangelos B. Mazomenos, Dwaipayan Biswas, Andy Cranny, Amal Rajan, Koushik Maharatna, Josy Achner, Jasmin Klemke, Michael Jöbges, Steffen Ortmann, Peter Langendörfer
IEEE J. Biomed. Health Informatics2
2015 Real-time arm movement recognition using FPGA
abstract
In this paper we present a FPGA-based system to detect three elementary arm movements in real-time (reach and retrieve, lift cup to mouth, rotation of the arm) using data from a wrist-worn accelerometer. Recognition is carried out by accurately mapping transitions of predefined, standard orientations of an accelerometer to the corresponding arm movements. The algorithm is coded in HDL and synthesized on the Altera DE2-115 FPGA board. For real-time operation, interfacing between the streaming sensor unit, host PC and the FPGA was achieved through a combination of Bluetooth, RS232 and an application software developed in C# using the .NET framework to facilitate serial port controls. The synthesized design used 1804 logic elements and recognised the performed arm movement in 41.2 μs, @50 MHz clock on the FPGA. Our experimental results show that the system can recognise all three arm movements with accuracies ranging 85%-96% for healthy subjects and 63%-75% for stroke survivors involved in `making-a-cup-of-tea', typical of an activity of daily living (ADL).
Dwaipayan Biswas, Gerry Juans Ajiwibawa, Koushik Maharatna, Andy Cranny, Josy Achner, Jasmin Klemke, Michael Jöbges
ISCAS1
2013 A Low-Complexity ECG Feature Extraction Algorithm for Mobile Healthcare Applications
abstract
This paper introduces a low-complexity algorithm for the extraction of the fiducial points from the Electrocardiogram (ECG). The application area we consider is that of remote cardiovascular monitoring, where continuous sensing and processing takes place in low-power, computationally constrained devices, thus the power consumption and complexity of the processing algorithms should remain at a minimum level. Under this context, we choose to employ the Discrete Wavelet Transform (DWT) with the Haar function being the mother wavelet, as our principal analysis method. From the modulus-maxima analysis on the DWT coefficients, an approximation of the ECG fiducial points is extracted. These initial findings are complimented with a refinement stage, based on the time-domain morphological properties of the ECG, which alleviates the decreased temporal resolution of the DWT. The resulting algorithm is a hybrid scheme of time and frequency domain signal processing. Feature extraction results from 27 ECG signals from QTDB, were tested against manual annotations and used to compare our approach against the state-of-the art ECG delineators. In addition, 450 signals from the 15-lead PTBDB are used to evaluate the obtained performance against the CSE tolerance limits. Our findings indicate that all but one CSE limits are satisfied. This level of performance combined with a complexity analysis, where the upper bound of the proposed algorithm, in terms of arithmetic operations, is calculated as 2:423N + 214 additions and 1:093N + 12 multiplications for N 861 or 2:553N + 102 additions and 1:093N +10 multiplications for N > 861 (N being the number of input samples), reveals that the proposed method achieves an ideal trade-off between computational complexity and performance, a key requirement in remote CVD monitoring systems.
Evangelos B. Mazomenos, Dwaipayan Biswas, Amit Acharyya, Taihai Chen, Koushik Maharatna, James Rosengarten, John M. Morgan, Nick Curzen
IEEE J. Biomed. Health Informatics2