EDBT 2026 Demo / reviewers in the wild / expert
Kolin Paul
dblp:51/544
· DBLP profile ↗
46ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0001-6970-5509ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Detection of Coronary Artery Disease Using PCG Signals for Wearable Devices: LSTM-Based ApproachabstractIt is acknowledged that coronary artery disease (CAD) is the most common type of heart disease. Continuous monitoring to identify the onset of CAD, even before symptoms are identified, can be helpful. This work attempts to efficiently identify CAD with high accuracy using wearable devices that enable continuous cardiovascular system monitoring. The proposed method utilizes single-channel PCG recordings for binary classification using Long Short-Term Memory. For this purpose, we have utilized the open-source dataset from the Physionet/Computing in Cardiology 2016 Challenge. We then provided a comparison of achieved results with state-of-the-art. The method that is being presented stands out because it does not require feature extraction or denoising, which are typically necessary for other similar works. We implemented five-fold cross-validation and achieved an average accuracy of 98.36±1.6 %. F1-score, precision, specificity, and sensitivity are 97.38±2.32 %, 94±3.97%, 96.3±2.32%, and 95.56±3%, respectively. For computational complexity measurement, we have used floating point operations, i.e., FLOPs (965784) and trainable parameters (1885). Additionally, we have developed an inference mechanism in a C++ program utilizing the weights and biases taken from the Keras model. The inference delay was calculated to be 9.664 seconds in a simulated Arm7 environment with Gem5. Priyanka Chauhan, Kolin Paul |
HealthCom | 2 |
| 2024 | MMTS: Multi-Modal Time Series Based Decision Support System for Ventilator Associated PneumoniaabstractVentilator-associated pneumonia (VAP) stands out as the predominant nosocomial pneumonia among critically ill patients and poses a significant threat to morbidity and mortality in intensive care units (ICUs). The timely identification of individuals susceptible to VAP allows for early intervention, thereby enhancing patient outcomes. We have proposed a multi-modal time series model to facilitate the early prediction of VAP. This model predicts a patient’s VAP condition early, by examining the previous day’s chest X-rays, Clinical and Micro Biological Analysis(CMBA), and the patient’s medical history, enabling early detection before the actual diagnosis. Our approach involves a two-stage architecture. In the first stage, we train two modality-specific models independently for chest X-rays and CMBA respectively. These encoders are then frozen, and in the second stage, a sequence model is trained on the embeddings of previous days’ chest X-rays and CMBA, integrated in the temporal domain, extracted from the encoders from stage 1, to capture patterns and predict the next day’s VAP condition. Our model achieves an AUROC of 89.2 percent. Notably, our model benefits from an increased number of previous day data points for Early VAP diagnosis. Nikki Tiwari, Kolin Paul, Shrawan Kumar Raut, Animesh Ray, Surabhi Vyas, Naveet Wig |
IJCNN | 2 |
| 2024 | NetraDeep: An Integrated Deep Learning and Image Processing System for Precise Detection of Hard ExudatesabstractHard exudate (HE) is a common manifestation of various eye diseases, such as diabetic retinopathy (DR), and a prominent cause of vision loss and blindness. Researchers aim to visualize and quantify these exudates using deep learning (DL) and image processing (IP) models from retinal images. However, the requirement for a large number of labelled image datasets for DL models to work on diverse and poor-quality images makes this task challenging. To address this challenge, we introduce NetraDeep, a system that integrates data-driven DL and rule-based IP techniques for exudate segmentation. Our system uses IP models to detect and extract some features and assists DL models in detecting more advanced features and vice versa. The IP models are rule-based and use predefined rules to process images, while the DL models are data-driven and learn from the input data. NetraDeep provides visual and quantitative assessments while mitigating noise and other confounding factors such as artifacts and noise. The training of DL models of this system requires only a limited number of labelled fundus images from publicly available datasets. It provides accurate pixel-wise segmentation results on the public and private image datasets collected from local eye hospitals. Through extensive evaluation, our system achieved remarkable performance, with a dice coefficient of \(0.84\) for the public dataset and a rating of \(9.78\) and \(9.43\) out of 10, as corroborated by two medical experts with experience of more than 20 and 5 years, respectively, for the private image dataset. Vatsal Agrawal, Rohan Chawla, Kolin Paul |
ACM Trans. Comput. Heal. | 5 |
| 2023 | Fundus Imaging-Based Healthcare: Present and FutureabstractA fundus image is a two-dimensional pictorial representation of the membrane at the rear of the eye that consists of blood vessels, the optical disc, optical cup, macula, and fovea. Ophthalmologists use it during eye examinations to screen, diagnose, and monitor the progress of retinal diseases or conditions such as diabetes, age-marked degeneration (AMD), glaucoma, retinopathy of prematurity (ROP), and many more ocular ailments. Developments in ocular optical systems, image acquisition, processing, and management techniques over the past few years have contributed to the use of fundus images to monitor eye conditions and other related health complications. This review summarizes the various state-of-the-art technologies related to the fundus imaging device, analysis techniques, and their potential applications for ocular diseases such as diabetic retinopathy, glaucoma, AMD, cataracts, and ROP. We also present potential opportunities for fundus imaging–based affordable, noninvasive devices for scanning, monitoring, and predicting ocular health conditions and providing other physiological information, for example, heart rate (HR), blood components, pulse rate, heart rate variability (HRV), retinal blood perfusion, and more. In addition, we present different types of technological, economical, and sociological factors that impact the growth of the fundus imaging–based technologies for health monitoring. Kolin Paul |
ACM Trans. Comput. Heal. | 2 |
| 2023 | Deep Learning-assisted Retinopathy of Prematurity (ROP) ScreeningabstractRetinopathy of prematurity (ROP) is a leading cause of blindness in premature infants worldwide, particularly in developing countries. In this research, we propose a Deep Convolutional Neural Network (DCNN) and image processing-based approach for the automatic detection of retinal features, including the optical disc (OD) and retinal blood vessels (BV), as well as disease classification using a rule-based method for ROP patients. Our DCNN model uses YOLO-v5 for OD detection and either Pix2Pix or a U-Net for BV segmentation. We trained our DCNN models on publicly available fundus image datasets of size 1,117 and 288 for OD detection and BV segmentation, respectively. We evaluated our approach on a dataset of 439 preterm neonatal retinal images, testing for ROP Zone and 6 BV masks. Our proposed system achieved excellent results, with the OD detection module achieving an overall accuracy of 98.94% (when IoU 0.5) and the BV segmentation module achieving an accuracy of 96.69% and a Dice coefficient between 0.60 and 0.64. Moreover, our system accurately diagnosed ROP in Zone-1 with 88.23% accuracy. Our approach offers a promising solution for accurate ROP screening and diagnosis, particularly in low-resource settings, where it has the potential to improve healthcare outcomes. Het Patel, Kolin Paul, Shorya Azad |
ACM Trans. Comput. Heal. | 3 |
| 2023 | Deep-learning-based data-manipulation attack resilient supervisory backup protection of transmission lines
Astha Chawla, Prakhar Agrawal, Bijaya K. Panigrahi, Kolin Paul |
Neural Comput. Appl. | 4 |
| 2022 | SACC: Split and Combine Approach to Reduce the Off-chip Memory Accesses of LSTM AcceleratorsabstractLong Short-Term Memory (LSTM) networks are widely used in speech recognition and natural language processing. Recently, a large number of LSTM accelerators have been proposed for the efficient processing of LSTM networks. The high energy consumption of these accelerators limits their usage in energy-constrained systems. LSTM accelerators repeatedly access large weight matrices from off-chip memory, significantly contributing to energy consumption. Reducing off-chip memory access is the key to improving the energy efficiency of these accelerators. We propose a data reuse approach that splits and combines the LSTM cell computations in a way that reduces the off-chip memory accesses of LSTM hidden state matrices by 50%. In addition, the data reuse efficiency of our approach is independent of on-chip memory size, making it more suitable for small on-chip memory LSTM accelerators. Experimental results show that our approach reduces off-chip memory access by 28% and 32%, and energy consumption by 13% and 16%, respectively, compared to conventional approaches for character level Language Modelling and Speech Recognition LSTM models. Saurabh Tewari, Kolin Paul |
DATE | 3 |
| 2021 | Scheduling Persistent and Fully Cooperative InstructionsabstractParallel, distributed two-level control system has been adopted in streaming application accelerators that implement atomic vector operations. Each instruction of such architecture deals with one aspect (arithmetic, interconnect, storage, etc.) of an atomic vector operation. Such instructions are persistent and fully cooperative. Their lifetimes vary because of the vector size and the degree of parallelism. More complex constraints are also required to express the cooperation among these instructions. The conventional instruction behavior models are no longer suitable for such instructions. Therefore, we develop a novel instruction behavior model to address the scheduling aspect of the instruction set required by such architecture. Based on the behavior model, we formally define the scheduling problem and formulate it as a constraint satisfaction optimization problem (CSOP). However, the naive CSOP formulation quickly becomes unscalable. Thus a heuristic enhanced scheduling algorithm is introduced to make the CSOP approach scalable. The enhanced algorithm’s scalability is validated by a large set of experiments varying in problem size. Yu Yang 0020, Ahmed Hemani, Kolin Paul |
DSD | 3 |
| 2021 | Scheduling Persistent and Fully Cooperative InstructionsabstractThe distributed two-level control (D2LC) system has been adopted in streaming application accelerators that implement atomic vector operations. Each instruction of such architecture deals with one aspect (arithmetic, interconnect, storage, etc.) of an atomic vector operation. The D2LC architecture is different from traditional computer architecture such as MMX or VLIW. The D2LC architecture consists of many cells interconnected via a NoC. In each cell, there is a two-level controller. The level-1 controller sends instructions to configure their level-2 controllers. Each level-2 controller, once configured, works as an independent finite state machine (FSM). It manages a datapath for specific functionality, including computation, interconnection, as well as storage. We can say that the level-1 controllers implement threads while the level-2 controllers implement micro-threads. For generality, these microthreads, even though distributed in different cells, can be grouped by the interconnection units and orchestrated to implement a larger functionality. The compiler of D2LC architecture needs to schedule these micro-threads correctly so that the larger functionality is reached. Yu Yang 0020, Ahmed Hemani, Kolin Paul |
FCCM | 3 |
| 2021 | Global Monitor using SpatioTemporally Correlated Local MonitorsabstractAn exponential increase in the IIoT network leads to complex interdependencies between the network devices. These network devices are designed to perform a fixed set of tasks and log their activities as system logs. These logs act as an excellent source of information to understand a system state. The device networks are prone to sophisticated Multi-host Multistep (MhMs) attacks, which may not be detected using machine learning-based isolated system security solutions and need a large amount of attack data for training. This led to the development of Central Monitoring Systems (CMSs) that need centralized system log collection, hence suffer from latency, network bandwidth and data loss due to network congestion. It leads to the requirement for a global monitoring system to detect ongoing MhMs attacks in real-time with low false positives and low network overhead. In this direction, we propose GLoM: a global monitor using spatio-temporally correlated local monitors to detect ongoing MhMs attacks. It leverages deep learning-based algorithms to detect anomalies with high accuracy and attack graphs to map various anomalous behavior to detect MhMs attacks. GLoM is a two-stage hybrid model, where the workload is divided between Local Monitors (LM) and Global Monitor (GM). LMs use LSTM to detect the abnormal activities of a system leveraging syslogs and forward anomalous logs to the GM. At the same time, GM discovers possible vulnerabilities on the devices followed by generating Possible Attack Graphs (PAG) by mapping the prerequisites and post-conditions required to exploit a vulnerability. GM is responsible for further analysis of the anomalous logs to find whether the logs in the current window resemble to vulnerability (CVE) exploit logs using a rule-based attack pattern repository. We track all the successful CVE exploits observed using anomalous logs followed by the generation of Evidence List (EL)) for each system. The similarity index between attack-paths and EL identifies the most probable attack scenario an adversary may be following. Using LMs, network communication overhead decreased by 88% on the publicly available dataset OpenStack (Loghub). LSTM based anomaly detection shows 99% accuracy in detecting the anomalous logs with an average anomalous log prediction overhead of 0.6 msec. We achieved 98% and 97% accuracy to generate the pre-requisites and post-conditions of a vulnerability. We evaluate GLoM's efficiency to detect MhMs attacks using a case study evaluation on a local testbed. Geeta Yadav, Kolin Paul |
NCA | 2 |
| 2020 | A Security Verification Template to Assess Cache Architecture VulnerabilitiesabstractIn the recent years, cache based side-channel attacks have become a serious threat for computers. To face this issue, researches have been looking at verifying the security policies. However, these approaches are limited to manual security verification and they typically work for a small subset of the attacks. Hence, an effective verification environment to automatically verify the cache security for all side-channel attacks is still missing. To address this shortcoming, we propose a security verification methodology that formally verifies cache designs against cache side-channel vulnerabilities. Results show that this verification template is a straightforward, automated method in verifying cache invulnerability. Tara Ghasempouri, Jaan Raik, Kolin Paul, Cezar Reinbrecht, Said Hamdioui, Mottaqiallah Taouil |
DDECS | 3 |
| 2020 | Early RTL Analysis for SCA Vulnerability in Fuzzy Extractors of Memory-Based PUF Enabled DevicesabstractPhysical Unclonable Functions (PUFs) are gaining attention in the cryptography community because of the ability to efficiently harness the intrinsic variability in the manufacturing process. However, this means that they are noisy devices and require error correction mechanisms, e.g., by employing Fuzzy Extractors (FEs). Recent works demonstrated that applying FEs for error correction may enable new opportunities to break the PUFs if no countermeasures are taken. In this paper, we address an attack model on FEs hardware implementations and provide a solution for early identification of the timing Side-Channel Attack (SCA) vulnerabilities which can be exploited by physical fault injection. The significance of this work stems from the fact that FEs are an essential building block in the implementations of PUF-enabled devices. The information leaked through the timing side-channel during the error correction process can reveal the FE input data and thereby can endanger revealing secrets. Therefore, it is very important to identify the potential leakages early in the process during RTL design. Experimental results based on RTL analysis of several Bose-Chaudhuri-Hocquenghem (BCH) and Reed-Solomon decoders for PUF-enabled devices with FEs demonstrate the feasibility of the proposed methodology. Xinhui Lai, Maksim Jenihhin, Georgios N. Selimis, Sven Goossens, Roel Maes, Kolin Paul |
VLSI-SOC | 6 |
| 2019 | GRanDE: Graphical Representation and Design Space Exploration of Embedded SystemsabstractTasks executing computer vision and machine learning algorithms are becoming popular on embedded platforms. A key characteristic of such tasks is the presence of modes providing different levels of application performance in terms of metrics like accuracy. The system designer has the flexibility to select an appropriate mode for executing such tasks. Secondly, the designer also has the traditional flexibility of choosing suitable components to build the execution platform. Thirdly, the system performance might vary with various external factors (known as context), and during the initial stages of system design, the designer might have the flexibility to support only a subset of the possible contexts. This three-fold flexibility in the hands of the designer has not been explored simultaneously in prior works and raises the complexity of designing embedded systems many-fold. In this paper, we address the design of such systems through a novel framework named GRanDE (Graphical Representation and Design Space Exploration). GRanDE consists of a comprehensive graphical representation to capture the three aspects of the design space discussed earlier. Further, we transform this representation into Constraint Logic Programming (CLP) constructs, which could be used to interactively explore and prune the design space. We demonstrate the applicability of the proposed framework on an embedded system named MAVI having ~1.3 million design points. The generated CLP program could prune up to 99.74% of the design space of MAVI. Rajesh Kedia, M. Balakrishnan, Kolin Paul |
DSD | 3 |
| 2019 | PatchRank: Ordering updates for SCADA systemsabstractSecuring SCADA is a challenging task for the research community as well as the industry. SCADA networks form the basis of industrial productivity. Industry 4.0 is likely to see more expansive use of SCADA & IIoT for enhanced productivity. These complex systems consist of numerous vulnerable subsystems. It is challenging for the timely application of patches to all the vulnerabilities, due to resource constraints and the high cost of the patch process. Usually, the more severe (attack probable) weaknesses are patched first to secure the system. Often organizations ignore the vulnerabilities in the “critical” node in favor of securing a vulnerability in an isolated subsystem. Therefore, the sequence in which patches are applied needs to be prioritized. State of the art indicates that patch prioritization is primarily an art rather than any significant methodology being followed.This paper proposes PatchRank - a patch prioritization method for the SCADA systems based on Viable System Model, Common Vulnerability Scoring System, and Game theory. PatchRank provides a ranking of vulnerable nodes/subsystems as well as a ranking of subsystem vulnerabilities, thereby allowing well-formed strategies for patch management. This paper also proposes a “Usable Secure State” to define a security assurance level. A comparative analysis of PatchRank with other benchmark algorithms, i.e., SecureRank, CVSS, and density based prioritization shows that PatchRank converges to a usable secure state faster. Geeta Yadav, Kolin Paul |
ETFA | 2 |
| 2019 | Assessment of SCADA System VulnerabilitiesabstractSCADA system is an essential component for automated control and monitoring in many of the Critical Infrastructures (CI). Cyber-attacks like Stuxnet, Aurora, Maroochy on SCADA systems give us clear insight about the damage a determined adversary can cause to any country's security, economy, and health-care systems. An in-depth analysis of these attacks can help in developing techniques to detect and prevent attacks. In this paper, we focus on the assessment of SCADA vulnerabilities from the widely used National Vulnerability Database (NVD) until May 2019. We analyzed the vulnerabilities based on severity, frequency, availability, integrity and confidentiality impact, and Common Weaknesses. The number of reported vulnerabilities are increasing yearly. Approximately 89% of the attacks are the network exploits severely impacting availability of these systems. About 19% of the weaknesses are due to buffer errors due to the use of insecure and legacy operating systems. We focus on finding the answer to four key questions that are required for developing new technologies for securing SCADA systems. We believe this is the first study of its kind which looks at correlating SCADA attacks with publicly available vulnerabilities. Our analysis can provide security researchers with useful insights into SCADA critical vulnerabilities and vulnerable components, which need attention. We also propose a domain-specific vulnerability scoring system for SCADA systems considering the interdependency of the various components. Geeta Yadav, Kolin Paul |
ETFA | 2 |
| 2019 | PASCAL: Timing SCA Resistant Design and Verification FlowabstractA large number of crypto accelerators are being deployed with the widespread adoption of IoT. It is vitally important that these accelerators and other security hardware IPs are provably secure. Security is an extra functional requirement and hence many security verification tools are not mature. We propose an approach/flow - PASCAL - that works on RTL designs and discovers potential Timing Side Channel Attack (SCA) vulnerabilities in them. Based on information flow analysis, this is able to identify Timing Disparate Security Paths that could lead to information leakage. This flow also (automatically) eliminates the information leakage caused by the timing channel. The insertion of a lightweight Compensator Block as balancing or compliance FSM removes the timing channel with minimum modifications to the design with no impact on the clock cycle time or combinational delay of the critical path in the circuit. Xinhui Lai, Maksim Jenihhin, Jaan Raik, Kolin Paul |
IOLTS | 4 |
| 2019 | Equivalence Checking and Compaction of n-input Majority Terms Using Implicants of Majority
Rajeswari Devadoss, Kolin Paul, M. Balakrishnan |
J. Electron. Test. | 2 |
| 2019 | DADS: Decentralized Attestation for Device SwarmsabstractWe present a novel scheme called Decentralized Attestation for Device Swarms (DADS), which is, to the best of our knowledge, the first to accomplish decentralized attestation in device swarms. Device swarms are smart, mobile, and interconnected devices that operate in large numbers and are likely to be part of emerging applications in Cyber-Physical Systems (CPS) and Industrial Internet of Things (IIoTs). Swarm devices process and exchange safety, privacy, and mission-critical information. Thus, it is important to have a good code verification technique that scales to device swarms and establishes trust among collaborating devices. DADS has several advantages over current state-of-the-art swarm attestation techniques: It is decentralized, has no single point of failure, and can handle changing topologies after nodes are compromised. DADS assures system resilience to node compromise/failure while guaranteeing only devices that execute genuine code remain part of the group. We conduct performance measurements of communication, computation, memory, and energy using the TrustLite embedded systems architecture in OMNeT++ simulation environment. We show that the proposed approach can significantly reduce communication cost and is very efficient in terms of computation, memory, and energy requirements. We also analyze security and show that DADS is very effective and robust against various attacks. Samuel Wedaj, Kolin Paul, Vinay J. Ribeiro |
ACM Trans. Priv. Secur. | 2 |
| 2017 | MOCHA: Morphable Locality and Compression Aware Architecture for Convolutional Neural NetworksabstractToday, machine learning based on neural networks has become mainstream, in many application domains. A small subset of machine learning algorithms, called Convolutional Neural Networks (CNN), are considered as state-ofthe- art for many applications (e.g. video/audio classification). The main challenge in implementing the CNNs, in embedded systems, is their large computation, memory, and bandwidth requirements. To meet these demands, dedicated hardware accelerators have been proposed. Since memory is the major cost in CNNs, recent accelerators focus on reducing the memory accesses. In particular, they exploit data locality using either tiling, layer merging or intra/inter feature map parallelism to reduce the memory footprint. However, they lack the flexibility to interleave or cascade these optimizations. Moreover, most of the existing accelerators do not exploit compression that can simultaneously reduce memory requirements, increase the throughput, and enhance the energy efficiency. To tackle these limitations, we present a flexible accelerator called MOCHA. MOCHA has three features that differentiate it from the state-of-the-art: (i) the ability to compress input/ kernels, (ii) the flexibility to interleave various optimizations, and (iii) intelligence to automatically interleave and cascade the optimizations, depending on the dimension of a specific CNN layer and available resources. Post layout Synthesis results reveal that MOCHA provides up to 63% higher energy efficiency, up to 42% higher throughput, and up to 30% less storage, compared to the next best accelerator, at the cost of 26-35% additional area. Syed M. A. H. Jafri, Ahmed Hemani, Kolin Paul, Naeem Abbas |
IPDPS | 3 |
| 2016 | Dynamic core allocation for energy efficient video decoding in homogeneous and heterogeneous multicore architectures
Rajesh Kumar Pal, Ierum Shanaya, Kolin Paul, Sanjiva Prasad |
Future Gener. Comput. Syst. | 3 |
| 2016 | Polymorphic Configuration Architecture for CGRAsabstractIn the era of platforms hosting multiple applications with arbitrary reconfiguration requirements, static configuration architectures are neither optimal nor desirable. The static reconfiguration architectures either incur excessive overheads or cannot support advanced features (like time-sharing and runtime parallelism). As a solution to this problem, we present a polymorphic configuration architecture (PCA) that provides each application with a configuration infrastructure tailored to its needs. Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Ahmed Hemani, Kolin Paul, Juha Plosila, Peeter Ellervee, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Architecture and Implementation of Dynamic Parallelism, Voltage and Frequency Scaling (PVFS) on CGRAsabstractIn the era of platforms hosting multiple applications with arbitrary performance requirements, providing a worst-case platform-wide voltage/frequency operating point is neither optimal nor desirable. As a solution to this problem, designs commonly employ dynamic voltage and frequency scaling (DVFS). DVFS promises significant energy and power reductions by providing each application with the operating point (and hence the performance) tailored to its needs. To further enhance the optimization potential, recent works interleave dynamic parallelism with conventional DVFS. The induced parallelism results in performance gains that allow an application to lower its operating point even further (thereby saving energy and power consumption). However, the existing works employ costly dedicated hardware (for synchronization) and rely solely on greedy algorithms to make parallelism decisions. To efficiently integrate parallelism with DVFS, compared to state-of-the-art, we exploit the reconfiguration (to reduce DVFS synchronization overheads) and enhance the intelligence of the greedy algorithm (to make optimal parallelism decisions). Specifically, our solution relies on dynamically reconfigurable isolation cells and an autonomous parallelism, voltage, and frequency selection algorithm. The dynamically reconfigurable isolation cells reduce the area overheads of DVFS circuitry by configuring the existing resources to provide synchronization. The autonomous parallelism, voltage, and frequency selection algorithm ensures high power efficiency by combining parallelism with DVFS. It selects that parallelism, voltage, and frequency trio which consumes minimum power to meet the deadlines on available resources. Synthesis and simulation results using various applications/algorithms (WLAN, MPEG4, FFT, FIR, matrix multiplication) show that our solution promises significant reduction in area and power consumption (23% and 51% ) compared to state-of-the-art. Syed M. A. H. Jafri, Ozan Ozbag, Nasim Farahini, Kolin Paul, Ahmed Hemani, Juha Plosila, Hannu Tenhunen |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2014 | Morphable Compression Architecture for Efficient Configuration in CGRAsabstractToday, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads (up to 50% area of the overall platform). As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties (i.e. compression ratio and decoding time), and is therefore best suited for a particular class of applications (and situation). However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. The proposed architecture allows each application to enjoy a separate compression/decompression hierarchy (consisting of various types and implementations of hardware/software decoders) tailored to its needs. Thereby, our solution offers minimal memory while meeting the required configuration deadlines. Simulation results, using different applications (FFT, Matrix multiplication, and WLAN), reveal that the choice of compression hierarchy has a significant impact on compression ratio (from configware replication to 52%) and configuration cycles (from 33 nsec to 1.5 secs) for the tested applications. Synthesis results reveal that introducing adaptivity incurs negligible additional overheads (1%) compared to the overall platform area. Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen |
DSD | 5 |
| 2014 | High Level Design Approach to Accelerate De Novo Genome Assembly Using FPGAsabstractMany scientific applications take a very long time to execute on general purpose processors. Speedups can be obtained by using specialized hardware in conjunction with the processors. FPGA based accelerators are known to be effective for reducing the execution time of many scientific applications. Since FPGAs are configurable, they can be customized to implement a variety of processing elements as accelerators. The process of mapping algorithm to architecture is complex, as the design space is large. System simulation is usually employed to carry out the exploration, in spite of the fact that simulation models take significantly large amount of time to execute. High level design space exploration helps in taking the required decisions to arrive at an optimal design. In this paper we describe design space exploration carried out for accelerating de novo genome assembly using FPGAs. Three models at various levels of abstraction were used. We discuss how the simulation time of these models influence the choice of design parameters at different levels of abstraction. We illustrate this process by using the high level models to evaluate Hard Embedded Blocks (HEBs) in FPGAs for accelerating the de novo genome assembly application. B. Sharat Chandra Varma 0001, Kolin Paul, M. Balakrishnan |
DSD | 2 |
| 2014 | Customizable Compression Architecture for Efficient Configuration in CGRAsabstractToday, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads. As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties, and is therefore best suited for a particular class of applications. However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen |
FCCM | 5 |
| 2014 | Mapping Tasks to a Dynamically Reconfigurable Coarse Grained ArrayabstractCoarse-Grained Reconfigurable Architectures (CGRAs) have become popular in recent times as the increased transistor densities have enabled greater integration of increasingly complex “compute cores”. These devices pack massive compute power and can be effectively used to build efficient solutions for applications which have a significant degree of parallelism. In many cases, these CGRAs are also partially reconfigurable. Clearly to make effective use of these highly “parallel compute platforms”, a good mapping flow is required to map the parallelism that is present in a target application. Mansureh S. Moghaddam, Kolin Paul, M. Balakrishnan |
FCCM | 2 |
| 2014 | TransPar: Transformation based dynamic Parallelism for low power CGRAsabstractCoarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern applications (e.g. 4G, CDMA, etc.). Recently proposed CGRAs offer runtime parallelism to reduce energy consumption (by lowering voltage/frequency). To implement the runtime parallelism, CGRAs commonly store multiple compile-time generated implementations of an application (with different degree of parallelism) and select the optimal version at runtime. However, the compile-time binding incurs excessive configuration memory overheads and/or is unable to parallelize an application even when sufficient resources are available. As a solution to this problem, we propose Transformation based dynamic Parallelism (TransPar). TransPar stores only a single implementation and applies a series for transformations to generate the bitstream for the parallel version. In addition, it also allows to displace and/or rotate an application to parallelize in resource constrained scenarios. By storing only a single implementation, TransPar offers significant reductions in configuration memory requirements (up to 73% for the tested applications), compared to state of the art compaction techniques. Simulation and synthesis results, using real applications, reveal that the additional flexibility allows up to 33% energy reduction compared to static memory based parallelism techniques. Gate level analysis reveals that TransPar incurs negligible silicon (0.2% of the platform) and timing (6 additional cycles per application) penalty. Syed M. A. H. Jafri, Guilermo Serrano, Masoud Daneshtalab, Naeem Abbas, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen |
FPL | 6 |
| 2014 | An event-triggered based robust control of robot manipulatorabstractThis paper proposes a framework to design an event-triggered based robust control law for nonlinear uncertain robot manipulator. Load variations and unmodeled system dynamics of manipulator are the primary sources of both system and input uncertainties. A static event-triggering rule is employed to realize the proposed robust control law. Derivation of static event-triggering rule with a positive inter-event time and corresponding stability criteria for uncertain manipulator dynamics are the key contribution of this paper. Validation of proposed control technique is carried out numerically on a two-link SCARA type robot manipulator. Simulation results show that measurement error norm is always bounded by the state dependent threshold and also ensures that asymptotic convergence of manipulator states in the presence of both system and input uncertainty. Niladri Sekhar Tripathy, Indra Narayan Kar, Kolin Paul |
ICARCV | 3 |
| 2014 | ReKonf: Dynamically reconfigurable multiCore architecture
Rajesh Kumar Pal, Kolin Paul, Sanjiva Prasad |
J. Parallel Distributed Comput. | 2 |
| 2013 | Distributed Runtime Computation of Constraints for Multiple Inner LoopsabstractThis paper presents hardware solution for runtime computation of loop constraints and synchronizing delays for multiple inner loops in parallel distributed implementation of digital signal processing sub-systems. Methods to map and generate the runtime computation code for loop constraints and synchronizing delays are also presented. Compared to the traditional methods, the proposed solution achieves 55% average code compaction and 32.7% average performance improvement. The solution has modest hardware cost that increases linearly with the dimension of the architecture and has no performance penalty. Results from multiple realistic examples are presented, analyzed and compared to the traditional methods. Nasim Farahini, Ahmed Hemani, Kolin Paul |
DSD | 3 |
| 2013 | Energy-Aware Fault-Tolerant CGRAs Addressing Application with Different Reliability NeedsabstractIn this paper, we propose a polymorphic fault tolerant architecture that can be tailored to efficiently support the reliability needs of multiple applications at run-time. Today, coarse-grained reconfigurable architectures (CGRAs) host multiple applications with potentially different reliability needs. Providing platform-wide worst-case (maximum) protection to all the applications is neither optimal nor desirable. To reduce the fault-tolerance overhead, adaptive fault-tolerance strategies have been proposed. The proposed techniques access the reliability requirements of each application and adjust the fault-tolerance intensity (and hence overhead), accordingly. However, existing flexible reliability schemes only allow to shift between different levels of modular redundancy (duplication, triplication, etc.) and deal with only a single class of faults (e.g. soft errors). To complement these strategies, we propose energy-aware fault-tolerance that, in addition to modular redundancy, can also provide low cost, sub-modular (e.g. residue mod 3) redundancy, to cater both permanent and temporary faults. Our solution relies on an agent based control layer and a configurable fault-tolerance data path. The control layer identifies the application class and configures the data path to provide the needed reliability. Simulation results using a few selected algorithms (FFT, matrix multiplication, and FIR filter) showed that the proposed method provides flexible protection with energy overhead ranging from 3.125% to 107% for different reliability levels. Synthesis results have confirmed that the proposed architecture significantly reduces the area overhead for self-checking (59.1%) and fault tolerant (7.1%) versions, compared to the state of the art adaptive reliability techniques. Syed M. A. H. Jafri, Stanislaw J. Piestrak, Kolin Paul, Ahmed Hemani, Juha Plosila, Hannu Tenhunen |
DSD | 3 |
| 2013 | FAssem: FPGA Based Acceleration of De Novo Genome AssemblyabstractNext generation sequencing technologies produce large amounts of data at very low cost. They produce short reads of DNA fragments. These fragments have many overlaps, lots of repeats and may also include sequencing errors. The assembly process involves merging these sequences to form the original sequences. In recent years many software programs have been developed for this purpose. All of them take significant amount of time to execute. Velvet is a commonly used de novo assembly program. We propose a method to reduce the overall time for assembly by using pre-processing of the short read data on FPGAs and processing its output using Velvet. We show significant speed-ups with slight or no compromise on the quality of the assembled output. B. Sharat Chandra Varma 0001, Kolin Paul, M. Balakrishnan, Dominique Lavenier |
FCCM | 2 |
| 2013 | GAGM: Genome assembly on GPU using mate pairsabstractGenome fragment assembly has long been a time and computation intensive problem in the field of bioinformatics. Many parallel assemblers have been proposed to accelerate the process but there hasn't been any effective approach proposed for GPUs. Also with the increasing power of GPUs, applications from various research fields are being parallelized to take advantage of the massive number of “cores” available in GPUs. In this paper we present the design and development of a GPU based assembler (GAGM) for sequence assembly using Nvidia's GPUs with the CUDA programming model. Our assembler utilizes the mate pair reads produced by the current NGS technologies to build paired de Bruijn graph. Every paired read is broken into paired k-mers and l-mers. Every paired k-mer represents a vertex and paired l-mers are mapped as edges. Contigs are formed by grouping the regions of graph which can be unambiguously connected. We present parallel algorithms for k - mer extraction, paired de Bruijn graph construction and grouping of edges. We have benchmarked GAGM on four bacterial genomes. Our results show that the design on GPU is effective in terms of time as well as the quality of assembly produced. Ashutosh Jain, Anshuj Garg, Kolin Paul |
HiPC | 3 |
| 2013 | High performance 3D-FFT implementationabstract3D FFT is a very data and compute intensive kernel encountered in many applications. We report a high performance design and implementation of 3D-FFT on a CGRA which supports partial reconfiguration. The hardware software multi clock design uses dynamic reconfiguration to reduce the required communication bandwidth to achieve a sustained throughput of 40 GOPS on a wordsize of 48 bits. Performance metrics including overheads and speed over software for implementations of up to 256 point 3D-FFT have been presented in the paper. U. Nidhi, Kolin Paul, Ahmed Hemani |
ISCAS | 2 |
| 2012 | Energy-Aware Fault-Tolerant Network-on-Chips for Addressing Multiple Traffic ClassesabstractThis paper presents an energy efficient architecture to provide on-demand fault tolerance to multiple traffic classes, running simultaneously on single network on chip (NoC) platform. Today, NoCs host multiple traffic classes with potentially different reliability needs. Providing platform-wide worst-case (maximum) protection to all the classes is neither optimal nor desirable. To reduce the overheads incurred by fault tolerance, various adaptive strategies have been proposed. The proposed techniques rely on individual packet fields and operating conditions to adjust the intensity and hence the overhead of fault tolerance. Presence of multiple traffic classes undermines the effectiveness of these methods. To complement the existing adaptive strategies, we propose on-demand fault tolerance, capable of providing required reliability, while significantly reducing the energy overhead. Our solution relies on a hierarchical agent based control layer and a reconfigurable fault tolerance data path. The control layer identifies the traffic class and directs the packet to the path providing the needed reliability. Simulation results using representative applications (matrix multiplication, FFT, wavefront, and HiperLAN) showed up to 95% decrease in energy consumption compared to traditional worst case methods. Synthesis results have confirmed a negligible additional overhead, for providing on-demand protection (up to 5.3% area), compared to the overall fault tolerance circuitry. Syed M. A. H. Jafri, Liang Guang, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen |
DSD | 4 |
| 2012 | reMORPH: A Runtime Reconfigurable ArchitectureabstractProgrammable hardware built on a regular architecture can partially alleviate the problem of increased defect densities associated with transistor scaling by dynamically wiring around the defects [1]. The fine granularity of FPGAs is however unsuitable for effectively exploiting runtime reconfiguration because of the high overheads involved. A coarse grain reconfigurable array with malleable communication links - reMORPH - is proposed in this paper. The compute tile uses DSP48E and BRAM embedded blocks in a Xilinx FPGA and has a very low footprint of about 200 slice LUTs. The semi-systolic near neighbour communication interconnect can be dynamically reconfigured for each “epoch” of computation. The “epoch” or phases of the application are obtained via profiling or static data flow analysis. Some of the links between the compute tiles are changed during the reconfiguration phase which drastically reduces the context switch overhead enabling high performance/area applications to be built on this fabric. Kolin Paul, Chinmaya Dash, Mansureh S. Moghaddam |
DSD | 1 |
| 2012 | ReKonf: A Reconfigurable Adaptive ManyCore ArchitectureabstractThe "one-architecture-fits-all" design philosophy is inadequate for catering to the diverse characteristics of applications running on manyCore architectures. After evaluating various configurations of manyCore architectures for a variety of applications, we designed ReKonf, a reconfigurable adaptive manyCore architecture. ReKonf dynamically configures its reconfigurable components to morph into significantly different configurations on the same architecture by continuously monitoring vital parameters of applications. ReKonf adapts the architecture by tracking core utilization, live cache utilization and cache sharing between threads, at runtime without losing execution state. Our evaluation of various applications on a cycle accurate simulator shows that the (reconfigurable) architecture suitable for such applications can be classified into three main variants: Chip Multiprocessor mode; Symmetric Multiprocessor mode; and Clustered mode, with up to 256 processing cores. Our results show that improvements from 32% to 72% over a baseline configuration can be observed by choosing the right configuration for an application. We also propose architecture components that should be reconfigurable in future manyCore architectures. Rajesh Kumar Pal, Kolin Paul, Sanjiva Prasad |
ISPA | 2 |
| 2011 | Architecture and tools for programmable QCAabstractQuantum-dot Cellular Automata (QCA) is a nano-scale compute fabric being explored by the VLSI research community as the difficulties in shrinking CMOS transistors mount. The paradigm promises high device densities and power-efficiency, and has unique properties that make it an interesting candidate for programmable devices. In this work, we propose a specialized architecture for programmable devices using QCA, present design rules for circuit design using the architecture and introduce a simulation engine tuned to efficiently simulate QCA circuits designed for this architecture. Rajeswari Devadoss, Kolin Paul, M. Balakrishnan |
FPT | 2 |
| 2011 | Compact generic intermediate representation (CGIR) to enable late binding in coarse grained reconfigurable architecturesabstractIn the era of platforms hosting multiple applications, where inter-application communication and concurrency patterns are arbitrary, static compile time decision making is neither optimal nor desirable. As a part of solving this problem, we present a novel method for compactly representing multiple configuration bitstreams of a single application, with varying parallelisms, as a unique, compact, and customizable representation, called CGIR. The representation thus stored is unraveled at runtime to configure the device with optimal (e.g. in terms of energy) implementation. Our goal was to provide optimal decision making capability to the runtime resource manager (RTM) without compromising the runtime behavior or the memory requirements of the system. The presence of multiple binaries enhance optimality by providing the RTM with multiple implementations to choose from. The CGIR ensures minimal increase in memory requirements with the addition of each binary. The low-cost unraveling of CGIR guarantees the runtime behavior. We have chosen the dynamically reconfigurable resource array (DRRA) as a vehicle to study the feasibility of our approach. Simulation results using 16 point decimation in time fast Fourier transform (FFT) has showed massive (up to 18% for 2 versions, 33% for 3 versions) memory savings compared to state of the art. Formal evaluation shows that the savings increase with the increase in the number of implementations stored. Syed M. A. H. Jafri, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen |
FPT | 3 |
| 2011 | p-QCA: A Tiled Programmable Fabric Architecture Using Molecular Quantum-Dot Cellular AutomataabstractQuantum-dot cellular automata is an interesting computation fabric with many never-seen-before properties. However, no programmable fabric scheme has utilized all these properties effectively. We propose an architecture for a programmable device using molecular QCA which exploits all the specialities of the fabric. The architecture taps the flexibility provided by the clocking system of molecular QCA to build a simple tile-based programmable device with the 3-input Majority gate as the fundamental logic element. Observing how a QCA structure can behave as either an interconnect or a logic gate depending on clocking, the proposed architecture merges routing and logic elements, thus drastically changing how programmable fabrics have been designed. Rajeswari Devadoss, Kolin Paul, M. Balakrishnan |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2010 | A high-level synthesis flow for custom instruction set extensions for application-specific processorsabstractCustom instruction set extensions (ISEs) are added to an extensible base processor to provide application-specific functionality at a low cost. As only one ISE executes at a time, resources can be shared. This paper presents a new high-level synthesis flow targeting ISEs. We emphasize a new technique for resource allocation, binding, and port assignment during synthesis. Our method is derived from prior work on datapath merging, and increases area reduction by accounting for the cost of multiplexors that must be inserted into the resulting datapath to achieve multi-operational functionality. Nagaraju Pothineni, Philip Brisk, Paolo Ienne, Kolin Paul |
ASP-DAC | 5 |
| 2010 | A tiled programmable fabric using QCAabstractQuantum-dot Cellular Automata is an interesting computation fabric with many never-seen-before properties. However, no programmable fabric scheme has utilized all these properties effectively. We propose an architecture for a programmable device using QCA which exploits all the specialities of the fabric. The architecture taps the flexibility provided by the clocking system of QCA to build a simple tile based programmable device with the 3-input Majority gate as the fundamental logic element. Observing how a QCA structure can behave as either an interconnect or a logic gate depending on clocking, the proposed architecture merges routing and logic elements, thus drastically changing how programmable fabrics have been designed. Rajeswari Devadoss, Kolin Paul, M. Balakrishnan |
FPT | 2 |
| 2007 | Silicon Compaction/Defragmentation for Partial Runtime ReconfigurationabstractThe effective use of Run Time Reconfiguration (RTR) in modern FPGAs opens up new avenues to design area and power efficient high performance architectures. However the current design flow for exploiting RTR in designs, leads to the problem of silicon Defragmentation. We propose a silicon compaction/ defragmentation technique which works on already placed and routed modules to generate partial bitstreams (programming files) for the device. We have outlined a method which generates these partial bitstreams very fast taking into account the size and position of the "free" silicon when the device is in operation. The other advantage of this method is that the changes in the basic FPGA fabric needed to implement this defragmentation strategy are (almost) trivial. Kolin Paul, Joël Porquet-Lupine, Josep Llosa |
DSD | 1 |
| 2002 | Theory of Extended Linear MachinesabstractThis paper extends the theory of autonomous linear machines (LMs). The theory of the extension field has provided the foundation for the design of such machines referred to as Extended Linear Machines (ELM). An analytical framework has been reported to completely characterize the vector subspace generated by an ELM and also different variations of LMs having cyclic, as well as noncyclic vector subspaces. This formulation has resulted in a single algorithm that characterizes each of the vector subspaces in terms of cyclic and noncyclic subspaces. An ELM significantly reduces the computation time for characterizing the model and study of the behavior of the physical system compared to conventional binary linear machines. Kolin Paul, Dipanwita Roy Chowdhury, Parimal Pal Chaudhuri |
IEEE Trans. Computers | 1 |
| 1999 | Cellular Automata Based Transform Coding for Image Compression
Kolin Paul, Dipanwita Roy Chowdhury, Parimal Pal Chaudhuri |
HiPC | 1 |
| 1998 | Theory and Application of Multiple Attractor Cellular Automata for Fault DiagnosisabstractThis paper reports the use of a class of cellular automata for the testing and diagnosis of faults in analog circuits. The use of the scheme is explained with reference to the testing of OTA based circuits. Kolin Paul, Prasanta Kumar Nandi, B. N. Roy, M. Deb Purkayastha, Santanu Chattopadhyay, Parimal Pal Chaudhuri |
Asian Test Symposium | 1 |