VLDB 2026 Research / reviewers in the wild / expert
Rabi N. Mahapatra
dblp:01/3441
· DBLP profile ↗
92ranked-venue papers
6as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 4 first-author · 1 since 2021Computer networks · 11Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
20 papers |
Distributed systems · 27% Electronic design automation · 16% Reconfigurable computing and FPGAs · 16% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 78% Kernel, tree and ensemble methods · 22% | |
| Computer networks
7 papers |
Internet of things and sensor networks · 32% Routing and switching · 29% Internet architecture and protocols · 20% | |
| Theoretical computer science
2 papers |
Algorithms and data structures · 58% Mathematical optimization · 42% | |
| Network and information security
2 papers |
Cryptographic protocols and secure computation · 89% Network security · 11% |
Topics — the 30 heaviest of 68, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.8 | 2 | 2020 | Distributed Training of Support Vector Machine on a Multiple-FPGA System · IEEE Trans. Computers 2020 Fast and Communication-Efficient Algorithm for Distributed Support Vector Machine Training · IEEE Trans. Parallel Distributed Syst. 2019 |
Distributed systems › distributed machine learning
distributed SVM training |
0.8 | 2 | 2020 | Distributed Training of Support Vector Machine on a Multiple-FPGA System · IEEE Trans. Computers 2020 FPGA-based Distributed Edge Training of SVM · FPGA 2019 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.8 | 2 | 2020 | Distributed Training of Support Vector Machine on a Multiple-FPGA System · IEEE Trans. Computers 2020 FPGA-based Distributed Edge Training of SVM · FPGA 2019 |
Distributed systems
distributed optimization |
0.5 | 1 | 2021 | Householder Sketch for Accurate and Accelerated Least-Mean-Squares Solvers · ICML 2021 |
Mathematical optimization › stochastic optimization › stochastic approximation
least mean squares |
0.5 | 1 | 2021 | Householder Sketch for Accurate and Accelerated Least-Mean-Squares Solvers · ICML 2021 |
Algorithms and data structures › data summarization
sketching and coresets |
0.5 | 1 | 2021 | Householder Sketch for Accurate and Accelerated Least-Mean-Squares Solvers · ICML 2021 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | Distributed Training of Support Vector Machine on a Multiple-FPGA System · IEEE Trans. Computers 2020 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.4 | 1 | 2019 | Fast and Communication-Efficient Algorithm for Distributed Support Vector Machine Training · IEEE Trans. Parallel Distributed Syst. 2019 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.4 | 1 | 2019 | Fast and Communication-Efficient Algorithm for Distributed Support Vector Machine Training · IEEE Trans. Parallel Distributed Syst. 2019 |
Interconnection networks and networks-on-chip
optical network-on-chip |
0.3 | 1 | 2018 | BiGNoC: Accelerating Big Data Computing with Application-Specific Photonic Network-on-Chip Architectures · IEEE Trans. Parallel Distributed Syst. 2018 |
Embedded and real-time systems
real-time scheduling |
0.3 | 3 | 2010 | Reliability aware power management for dual-processor real-time embedded systems · DAC 2010 A Dynamic Slack Management Technique for Real-Time Distributed Embedded Systems · IEEE Trans. Computers 2008 Feedback-controlled reliability-aware power management for real-time embedded systems · DAC 2008 |
Internet of things and sensor networks › wireless sensor network
mobile sink |
0.3 | 2 | 2012 | The Three-Tier Security Scheme in Wireless Sensor Networks with Mobile Sinks · IEEE Trans. Parallel Distributed Syst. 2012 Key Predistribution Schemes for Establishing Pairwise Keys with a Mobile Sink in Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2011 |
Internet of things and sensor networks
wireless sensor network |
0.3 | 2 | 2012 | The Three-Tier Security Scheme in Wireless Sensor Networks with Mobile Sinks · IEEE Trans. Parallel Distributed Syst. 2012 Key Predistribution Schemes for Establishing Pairwise Keys with a Mobile Sink in Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2011 |
Cryptographic protocols and secure computation › key management › key distribution
key predistribution |
0.3 | 2 | 2012 | The Three-Tier Security Scheme in Wireless Sensor Networks with Mobile Sinks · IEEE Trans. Parallel Distributed Syst. 2012 Key Predistribution Schemes for Establishing Pairwise Keys with a Mobile Sink in Sensor Networks · IEEE Trans. Parallel Distributed Syst. 2011 |
Internet architecture and protocols › packet processing
packet classification |
0.2 | 2 | 2011 | A Power and Throughput-Efficient Packet Classifier with n Bloom Filters · IEEE Trans. Computers 2011 A Memory-Efficient Hashing by Multi-Predicate Bloom Filters for Packet Classification · INFOCOM 2008 |
Energy-efficient computing › energy management
reliability-aware power management |
0.2 | 2 | 2010 | Reliability aware power management for dual-processor real-time embedded systems · DAC 2010 Feedback-controlled reliability-aware power management for real-time embedded systems · DAC 2008 |
Routing and switching › IP lookup
hash-based lookup |
0.2 | 2 | 2009 | A Hash-based Scalable IP lookup using Bloom and Fingerprint Filters · ICNP 2009 A Memory-Efficient Hashing by Multi-Predicate Bloom Filters for Packet Classification · INFOCOM 2008 |
Routing and switching › router architecture
high-speed router |
0.1 | 2 | 2011 | A Power and Throughput-Efficient Packet Classifier with n Bloom Filters · IEEE Trans. Computers 2011 A Memory-Efficient Hashing by Multi-Predicate Bloom Filters for Packet Classification · INFOCOM 2008 |
Routing and switching
IP lookup |
0.1 | 2 | 2009 | A Hash-based Scalable IP lookup using Bloom and Fingerprint Filters · ICNP 2009 EaseCAM: An Energy and Storage Efficient TCAM-Based Router Architecture for IP Lookup · IEEE Trans. Computers 2005 |
Distributed systems
fault tolerance |
0.1 | 3 | 2010 | A Dynamic Slack Management Technique for Real-Time Distributed Embedded Systems · IEEE Trans. Computers 2008 Reliability aware power management for dual-processor real-time embedded systems · DAC 2010 Feedback-controlled reliability-aware power management for real-time embedded systems · DAC 2008 |
Internet architecture and protocols
packet processing |
0.1 | 1 | 2011 | A Power and Throughput-Efficient Packet Classifier with n Bloom Filters · IEEE Trans. Computers 2011 |
Electronic design automation
physical design |
0.1 | 2 | 2006 | Antenna Avoidance in Layer Assignment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 Reducing clock skew variability via crosslinks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Machine learning › Efficient and distributed learning › model compression
low-rank approximation |
0.1 | 1 | 2019 | Fast and Communication-Efficient Algorithm for Distributed Support Vector Machine Training · IEEE Trans. Parallel Distributed Syst. 2019 |
Edge and fog computing
edge analytics |
0.1 | 1 | 2019 | FPGA-based Distributed Edge Training of SVM · FPGA 2019 |
Edge and fog computing › edge machine learning
on-device learning |
0.1 | 1 | 2019 | FPGA-based Distributed Edge Training of SVM · FPGA 2019 |
Electronic design automation › physical design
clock network synthesis |
0.1 | 2 | 2006 | Reducing clock skew variability via crosslinks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 Reducing clock skew variability via cross links · DAC 2004 |
Storage systems
data analytics |
0.1 | 1 | 2018 | BiGNoC: Accelerating Big Data Computing with Application-Specific Photonic Network-on-Chip Architectures · IEEE Trans. Parallel Distributed Syst. 2018 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.1 | 1 | 2009 | Hierarchical reconfigurable computing arrays for efficient CGRA-based embedded systems · DAC 2009 |
Algorithms and data structures › probabilistic data structures
bloom filter |
0.1 | 1 | 2009 | A Hash-based Scalable IP lookup using Bloom and Fingerprint Filters · ICNP 2009 |
Algorithms and data structures › data structure design › search structures
hashing |
0.1 | 1 | 2009 | A Hash-based Scalable IP lookup using Bloom and Fingerprint Filters · ICNP 2009 |
Methods — techniques the papers use, named apart from their topics
sketching · 1.0householder transformation · 1.0pipelined training IP core · 0.9parallelization · 0.9pipelined IP core · 0.8distributed training · 0.8QRSVM · 0.8pipelining · 0.4low-rank approximation · 0.4distributed algorithm · 0.4QR decomposition · 0.4photonic waveguide multicasting · 0.3three-tier key pool framework · 0.3pairwise key establishment · 0.3probabilistic key predistribution · 0.2polynomial pool-based key predistribution · 0.2dynamic voltage and frequency scaling · 0.2q-composite scheme · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | L2C: Learn to Clean Time Series Data
Mayuresh Hooli, Rabi N. Mahapatra |
IoTBDS | 2 |
| 2025 | User Locale and Intent Prediction Using Sensors from SmartphoneabstractSmartphone data is becoming increasingly important today as many people own and frequently use smartphones. The smartphone-based data can provide better insight into the user's behavior by providing knowledge related to their daily activities. We can use it in behavior or intent prediction, location prediction, and recommendation of places and activities. We developed an Android application to gather data from the wireless sensors and GPS of the smartphone. Deep learning technology on this complex and diverse behavioral data can help extract useful information. Our research assessed the accuracy and efficiency of four deep learning architectures based on convolutional neural networks and their variants and LSTM. Our study presents a novel deep learning model that integrates residual networks and LSTM. The results of our experiments demonstrate that this model surpasses the traditional approach of combining convolutional networks with LSTM, achieving higher accuracy rates of 90.058% and 90.261% for predicting user locale and user intent, respectively. Overall, this study demonstrates the potential of smartphone-based wireless sensor data and location services for predicting user activity and location. It highlights the importance of utilizing deep learning techniques for effectively processing and analyzing this type of data, which is becoming increasingly prevalent today, and provides a better understanding of how to take advantage of the vast amount of smartphone data to make predictions about human behavior. This can be used in various application areas. Akash Sahoo, Sanjay Naik, Rabi N. Mahapatra |
MDM | 3 |
| 2021 | Householder Sketch for Accurate and Accelerated Least-Mean-Squares SolversabstractLeast-Mean-Squares (\textsc{LMS}) solvers comprise a class of fundamental optimization problems such as linear regression, and regularized regressions such as Ridge, LASSO, and Elastic-Net. Data summarization techniques for big data generate summaries called coresets and sketches to speed up model learning under streaming and distributed settings. For example, \citep{nips2019} design a fast and accurate Caratheodory set on input data to boost the performance of existing \textsc{LMS} solvers. In retrospect, we explore classical Householder transformation as a candidate for sketching and accurately solving LMS problems. We find it to be a simpler, memory-efficient, and faster alternative that always existed to the above strong baseline. We also present a scalable algorithm based on the construction of distributed Householder sketches to solve \textsc{LMS} problem across multiple worker nodes. We perform thorough empirical analysis with large synthetic and real datasets to evaluate the performance of Householder sketch and compare with \citep{nips2019}. Our results show Householder sketch speeds up existing \textsc{LMS} solvers in the scikit-learn library up to $100$x-$400$x. Also, it is $10$x-$100$x faster than the above baseline with similar numerical stability. The distributed algorithm demonstrates linear scalability with a near-negligible communication overhead. Jyotikrishna Dass, Rabi N. Mahapatra |
ICML | 2 |
| 2021 | BPLight-CNN: A Photonics-Based Backpropagation Accelerator for Deep LearningabstractTraining deep learning networks involves continuous weight updates across the various layers of the deep network while using a backpropagation (BP) algorithm. This results in expensive computation overheads during training. Consequently, most deep learning accelerators today employ pretrained weights and focus only on improving the design of the inference phase. The recent trend is to build a complete deep learning accelerator by incorporating the training module. Such efforts require an ultra-fast chip architecture for executing the BP algorithm. In this article, we propose a novel photonics-based backpropagation accelerator for high-performance deep learning training. We present the design for a convolutional neural network (CNN), BPLight-CNN , which incorporates the silicon photonics-based backpropagation accelerator. BPLight-CNN is a first-of-its-kind photonic and memristor-based CNN architecture for end-to-end training and prediction. We evaluate BPLight-CNN using a photonic CAD framework (IPKISS) on deep learning benchmark models, including LeNet and VGG-Net. The proposed design achieves (i) at least 34× speedup, 34× improvement in computational efficiency, and 38.5× energy savings during training; and (ii) 29× speedup, 31× improvement in computational efficiency, and 38.7× improvement in energy savings during inference compared with the state-of-the-art designs. All of these comparisons are done at a 16-bit resolution, and BPLight-CNN achieves these improvements at a cost of approximately 6% lower accuracy compared with the state-of-the-art. Dharanidhar Dang, Sai Vineel Reddy Chittamuru, Sudeep Pasricha, Rabi N. Mahapatra, Debashis Sahoo |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | BPhoton-CNN: An Ultrafast Photonic Backpropagation Accelerator for Deep LearningabstractTraining deep learning networks involves continuous weight updates across its many using a backpropagation algorithm (BP). This results in expensive computation and energy overhead during training. Consequently, most deep learning accelerators today employ pre-trained weights and focus only on improving the design of the inference phase. The recent trend is to develop a complete deep learning accelerator by incorporating the training module. Such efforts require an ultra-fast chip architecture for executing the BP algorithm. In this paper, we introduce a novel photonics-based backpropagation accelerator for high performance deep learning training. We present the design for a convolutional neural network, BPhoton-CNN, which incorporates the silicon photonics-based backpropagation accelerator. BPhoton-CNN is a first-of-its-kind photonic and memristor-based CNN architecture for end-to-end training and prediction. We evaluate BPhoton-CNN using a commercial CAD framework (IPKISS) on deep learning benchmark models including LeNet and VGG-Net. The proposed design achieves at least 35× acceleration in training time, 31× improvement in computational efficiency, and 45× energy savings compared to the state-of-the-art designs, without any loss of accuracy. Dharanidhar Dang, Aurosmita Khansama, Rabi N. Mahapatra, Debashis Sahoo |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | On-chip Parallel Photonic Reservoir Computing using Multiple Delay LinesabstractSilicon-Photonics architectures have enabled high speed hardware implementations of Reservoir Computing (RC). With a delayed feedback reservoir (DFR) model, only one non-linear node can be used to perform RC. However, the delay is often provided by using off-chip fiber optics which is not only space inconvenient but it also becomes architectural bottleneck and hinders to scalability. In this paper, we propose a completely on-chip photonic RC architecture for high performance computing, employing multiple electronically tunable delay lines and micro-ring resonator (MRR) switch for multi-tasking. Proposed architecture provides 84% less error compared to the state-of-the-art standalone architecture in [8] for executing NARMA task. For multi-tasking, the proposed architecture shows 80% better performance than [8]. The architecture outperforms all other proposed architectures as well. The on-chip area and power overhead of proposed architecture due to delay lines and MRR switch are 0.0184mm2and 26mW respectively. Syed Ali Hasnain, Rabi N. Mahapatra |
SBAC-PAD | 2 |
| 2020 | Distributed Training of Support Vector Machine on a Multiple-FPGA SystemabstractSupport Vector Machine (SVM) is a supervised machine learning model for classification tasks. Training SVM on a large number of data samples is challenging due to the high computational cost and memory requirement. Hence, model training is supported on a high-performance server which typically runs a sequential training algorithm on centralized data. However, as we move towards massive workloads, it will be impossible to store all the data in a centralized manner and expect such sequential training algorithms to scale on traditional processors. Moreover, with the growing demands of real-time machine learning for edge analytics, it is imperative to devise an efficient training framework with relatively cheaper computations and limited memory. Therefore, we propose and implement a first-of-its-kind system of multiple FPGAs as a distributed computing framework comprising up to eight FPGA units on Amazon F1 instances with negligible communication overhead to fully parallelize, accelerate, and scale the SVM training on decentralized data. Each FPGA unit has a pipelined SVM training IP logic core operating at 125 MHz with a power dissipation of 39 Watts for accelerating its allocated computations in the overall training process. We evaluate and compare the performance of the proposed system on five real SVM benchmarks. Jyotikrishna Dass, Yashwardhan Narawane, Rabi N. Mahapatra, Vivek Sarin |
IEEE Trans. Computers | 3 |
| 2020 | Adaptive Group-Based Zero Knowledge Proof-Authentication Protocol in Vehicular Ad Hoc NetworksabstractVehicular ad hoc networks (VANETs) are a particular subclass of mobile ad hoc networks that raise a number of security challenges, notably from the way users authenticate the network. Authentication technologies based on existing security policies and access control rules in such networks assume full trust on roadside unit (RSU) and authentication servers. The disclosure of authentication parameters enables user's traceability over the network. VANETs' trusted entities (e.g., RSU) can utilize such information to track a user traveling behavior, violating user privacy and anonymity. In this paper, we proposed a novel, light-weight, adaptive group-based zero knowledge proof-authentication protocol (AGZKP-AP) for VANETs. The proposed authentication protocol is capable of offering various levels of users' privacy settings based on the type of services available on such networks. Our scheme is based on the zero-knowledge-proof crypto approach with the support of tradeoff options. Users have the option to make critical decisions on the level of privacy and the amount of resources usage they prefer such as short system response time versus the number of private information disclosures. Furthermore, AGZKP-AP is incorporated with a distributed privilege control and revoking mechanism that render user's private information to law enforcement in case of a traffic violation. Amar A. Rasheed, Rabi N. Mahapatra, Felix G. Hamza-Lup |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | FPGA-based Distributed Edge Training of SVMabstractSupport Vector Machine (SVM) is a widely used supervised machine learning algorithm for classification. Training SVM is challenging due to high computational cost and memory requirements. More often such training is handled at back end servers leading to significant communication and energy overheads. This approach is unsuitable for edge analytics which is a growing trend with various IoT applications. Enabling efficient training on the edge requires a distributed computing approach that has negligible communication overhead and an energy-efficient hardware design to execute it. In this paper, we present a scalable FPGA-based design for distributed SVM training amenable for edge-based learning. Specifically, we implement a pipelined QRSVM IP logic on Xilinx Virtex UltraScale+ VU9P FPGA. Each synthesized IP core operates at 125 MHz with a power dissipation of 39 Watts. We evaluate the training time, parallel speedup, scalability, and energy efficiency of the proposed design on five SVM benchmarks on a multiple FPGA system comprising up to eight FPGA units. When compared with software implementation on the traditional embedded system edge processors like ARM Cortex-A15, the proposed FPGA implementation is around 3x to 24x faster and 2x to 8x more energy efficient on the above benchmarks. Jyotikrishna Dass, Yashwardhan Narawane, Rabi N. Mahapatra, Vivek Sarin |
FPGA | 3 |
| 2019 | Fast and Communication-Efficient Algorithm for Distributed Support Vector Machine TrainingabstractSupport Vector Machines (SVM) are widely used as supervised learning models to solve the classification problem in machine learning. Training SVMs for large datasets is an extremely challenging task due to excessive storage and computational requirements. To tackle so-called big data problems, one needs to design scalable distributed algorithms to parallelize the model training and to develop efficient implementations of these algorithms. In this paper, we propose a distributed algorithm for SVM training that is scalable and communication-efficient. The algorithm uses a compact representation of the kernel matrix, which is based on the QR decomposition of low-rank approximations, to reduce both computation and storage requirements for the training stage. This is accompanied by considerable reduction in communication required for a distributed implementation of the algorithm. Experiments on benchmark data sets with up to five million samples demonstrate negligible communication overhead and scalability on up to 64 cores. Execution times are vast improvements over other widely used packages. Furthermore, the proposed algorithm has linear time complexity with respect to the number of samples making it ideal for SVM training on decentralized environments such as smart embedded systems and edge-based internet of things, IoT. Jyotikrishna Dass, Vivek Sarin, Rabi N. Mahapatra |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Privacy-Preserving ECG based Active Authentication (PPEA2) for IoT DevicesabstractIoT devices have become essential in our day-to-day life starting from health monitoring to industrial control systems. While the benefits of IoT are undeniable, IoT ecosystem comes with its own set of system vulnerabilities that include malicious actors manipulating the flow of information to and from the IoT devices, which can lead to the capture of sensitive data and loss of data privacy. In this paper, we propose a Privacy-Preserving ECG based Active Authentication (PPEA2) scheme that is deployable on power-limited wearable systems (e.g. fitness tracker systems, health monitoring systems for solider in the battlefield, and large-scale health monitoring infrastructure for rapid response systems). The proposed scheme is capable of supporting active authentication of users by utilizing live stream of electrocardiogram (ECG) signal to derive unique authentication parameters. In addition to providing active authentication, we incorporated a privacy-preserving feature into the design of our system. The scheme preserves the privacy of the users ECG data features by employing a light-weight secure computation approach based on secure weighted hamming distance computation from oblivious transfer to compute a joint set between two participating entities without revealing the authentication parameters to either of them. We demonstrate the feasibility of the system, its performance and resilience against various threats in a semi-honest model. Ghanshyam Bhutra, Amar A. Rasheed, Rabi N. Mahapatra |
IPCCC | 3 |
| 2018 | Hybrid algorithm for accumulated error suppression in open-loop Doppler receiverabstractA hybrid algorithm is proposed to decrease the integration phase measurement error accumulation in open‐loop Doppler receiver which is designed for space orbit determination and positioning applications when cooperated closed‐loop system is unavailable. Firstly, adaptive‐neuro‐fuzzy inference system (ANFIS) module α is implemented for the interrupted closed‐loop measurement data prediction. Then to improve the prediction accuracy of the ANFIS module α , an adaptive Kalman filter is employed for data fusion with predicted data from ANFIS module α and open‐loop measurement data for complementation and corrections. Meanwhile, ANFIS module β is embedded in the Kalman filter for adaptive error compensation to optimise the filter performance. Finally, the time costly ANFIS computations are accelerated by reconfigurable software–hardware co‐design module implemented in the receiver system to improve the computing capability and efficiency of system. Experiment results are analysed to demonstrate the effectiveness of proposed hybrid algorithm. It suppresses error accumulation in open‐loop receiver phase measurement by 85.89% compared to the directly integration results when closed‐loop system is off‐line and has 61.02% enhancement compared to the algorithm without adaptive error compensations. Thus, the whole combinatory system accuracy is improved. Jifei Tang, Rabi N. Mahapatra, Lanhua Xia |
IET Commun. | 2 |
| 2018 | BiGNoC: Accelerating Big Data Computing with Application-Specific Photonic Network-on-Chip ArchitecturesabstractIn the era of big data, high performance data analytics applications are frequently executed on large-scale cluster architectures to accomplish massive data-parallel computations. Often, these applications involve iterative machine learning algorithms to extract information and make predictions from large data sets. Multicast data dissemination is one of the major performance bottlenecks for such data analytics applications in cluster computing, as terabytes of data need to be distributed frequently from a single data source to hundreds of computing nodes. To overcome this bottleneck for big data applications, we proposeBiGNoC, a manycore chip platform with a novel application-specific photonic network-on-chip (PNoC) fabric.BiGNoCis designed for big data computing and exploits multicasting in photonic waveguides. For high performance data analytics applications,BiGNoCimproves throughput by up to${{9.9}}\times$while reducing latency by up to 88 percent and energy-per-bit by up to 98 percent over two state-of-the-art PNoC architectures as well as a broadcast-optimized electrical mesh NoC architecture, and a traditional electrical mesh NoC architecture. Sai Vineel Reddy Chittamuru, Dharanidhar Dang, Sudeep Pasricha, Rabi N. Mahapatra |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Islands of heaters: A novel thermal management framework for photonic NoCsabstractSilicon photonics has become a promising candidate for future networks-on-chip (NoCs) as it can enable high bandwidth density and lower latency with traversal of data at the speed of light. But the operation of photonic NoCs (PNoCs) is very sensitive to temperature variations that frequently occur on a chip. These variations can create significant reliability issues for PNoCs. For example, microring resonators (MRRs) which are the building blocks of PNoCs, may resonate at another wavelength instead of their designated wavelength due to thermal variations, which can lead to bandwidth wastage and data corruption in PNoCs. This paper proposes a novel run-time framework to overcome temperature-induced issues in PNoCs. The framework consists of (i) a PID controlled heater mechanism to nullify the thermal gradient across PNoCs, (ii) a device-level thermal island framework to distribute MRRs across regions of temperatures; and (iii) a system-level proactive thread migration technique to avoid on-chip thermal threshold violations and to reduce MRR tuning/trimming power by migrating threads between cores. Our experimental results with 64-core Corona and Flexishare PNoCs indicate that the proposed approach reliably satisfies on-chip thermal thresholds and maintains high network bandwidth while reducing total power by up to 64.1%. Dharanidhar Dang, Sai Vineel Reddy Chittamuru, Rabi N. Mahapatra, Sudeep Pasricha |
ASP-DAC | 3 |
| 2017 | ConvLight: A Convolutional Accelerator with Memristor Integrated Photonic ComputingabstractNeuromorphic computing is a promising candidate to accelerate big data processing. Recently, several attempts have been made to design neuromorphic accelerators for popular machine learning algorithms, such as reservoir computing, deep learning, spiking neurons etc. Deep learning accelerator which involves convolutional neural networks (CNNs) have received widespread attention for their accuracy and efficiency. This paper proposes ConvLight, a novel deep learning accelerator based on memristor integrated photonic computing framework. While the use of on-chip photonic circuits for analog computing is well known, no prior work has demonstrated a full-fledged accelerator based on photonic components. In particular, this paper makes the following novel contributions: (i) A multilayer deep learning architecture design is proposed using compute efficient memristors and photonic components for the first time. (ii) A pipelined design for each CNN layer is presented for maximizing throughput and enabling parallelism across the layers. (iii) Simulation of ConvLight architecture with standard photonic tools for demonstrating the execution of DNN and CNN workloads yielding 25X, 60X, and 40X improvements in computational efficiency, throughput, and energy efficiency (respectively) compared to state-of-the-art design. Dharanidhar Dang, Jyotikrishna Dass, Rabi N. Mahapatra |
HiPC | 3 |
| 2017 | Distributed QR Decomposition Framework for Training Support Vector MachinesabstractSupport Vector Machines (SVM) belong to a class of supervised machine learning algorithms with applications in classification and regression analysis. SVM training is modeled as a convex optimization problem that is computationally tedious and has large memory requirements. Specifically, it is a quadratic programming problem which scales rapidly with the training set size rather than the dimensionality of the feature space. In this work, we first present a novel QR decomposition framework (QRSVM) to efficiently model and solve a large scale SVM problem by capitalizing on low-rank representations of the full kernel matrix rather than solving the problem as a sequence of smaller sub-problems. The low-rank structure of the kernel matrix is leveraged to transform the dense matrix into one with a sparse and separable structure. The modified SVM problem requires significantly lesser memory and computation. Our approach scales linearly with the training set size which makes it applicable to large datasets. This motivates towards our another contribution; exploring a distributed QRSVM framework to solve large-scale SVM classification problems in parallel across a cluster of computing nodes. We also derive an optimal step size for fast convergence of the dual ascent method which is used to solve the quadratic programming problem. Jyotikrishna Dass, V. N. S. Prithvi Sakuru, Vivek Sarin, Rabi N. Mahapatra |
ICDCS | 4 |
| 2016 | A Relaxed Synchronization Approach for Solving Parallel Quadratic Programming Problems with Guaranteed ConvergenceabstractIn this paper we present a novel numerical algorithm for efficiently solving large-scale quadratic programming problems in massively parallel computing systems. The main challenge in maximizing processor utilization is to reduce idling due to synchronization across processors. Typically, synchronization is necessary after every iteration, which prevents many numerical algorithms from scaling with number of processors. We relax this requirement by synchronizing at a lower rate, which is referred to as lazy synchronization. We show analytically and experimentally that lazy synchronization is numerically stable and converges to the same result as the conventional tightly synchronized implementation. Furthermore, the convergence speed of the proposed algorithm is faster with lazy synchronization. The numerical stability, convergence rate and the optimal rate for synchronization are analytically shown. The proposed algorithm is implemented in a 40-node distributed system in the Amazon Elastic Computing infrastructure. We show a 160 times speedup in solution time for a large-scale quadratic programming problem using a synthetic dataset. The experiments demonstrate that the use of relaxed synchronization technique reduces communication overhead in the distributed systems by 99.65% in comparison to the tightly synchronization implementation. Kooktae Lee, Raktim Bhattacharya, Jyotikrishna Dass, V. N. S. Prithvi Sakuru, Rabi N. Mahapatra |
IPDPS | 5 |
| 2015 | A Multilayered Design Approach for Efficient Hybrid 3D Photonics Network-on-chipabstractIn Chip Multiprocessors, traditional metallic interconnects will soon reach their bandwidth and energy dissipation limits. Photonic NoC (PNoC) is a promising alternative to renew higher performance in the advent of rising number of cores on chip. Efficient PNoC architectures are needed to reduce laser related energy consumption and maintain high performance. In this work we propose a novel sandwich layered approach to design a 3D PNoC architecture that is able to reduce no of hops, cross over points, and no of laser sources using multiplexing techniques. The 3D hybrid PNoC uses high performance 5X5 photonic routers incorporating mode division multiplexing (MDM) along with wavelength division multiplexing (WDM) and time division multiplexing (TDM). Experimental results demonstrates an increase in aggregated bandwidth up to 4x while reducing average energy consumption per router by 83\% as compared to the recently reported results. Dharanidhar Dang, Biplab Patra, Rabi N. Mahapatra |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | PID controlled thermal management in photonic network-on-chipabstractThe communication bandwidth and power consumption of network-on-chip (NoC) are going to meet their limits soon because of traditional metallic interconnects. Photonic NoC (PNoC) is emerging as a promising alternative to address these bottlenecks. However, PNoCs are highly susceptible to thermal fluctuations which is highly common in a manycore chip. This paper first introduces a low power, low cost mesh-based PNoC architecture and provides a quantitative analysis of it's power consumption over varying on-chip temperature. The paper then proposes a proportional-integral-derivative (PID) heater mechanism that minimizes the effect of thermal variation on PNoC's performance and power. Experimental results for a 8 × 8 PNoC shows that the proposed technique offers the maximum network-bandwidth considering the thermal effects. Compared to the recently reported results, the proposed design consumes 40% less power and has a temperature variation as low as 1 °C. Dharanidhar Dang, Rabi N. Mahapatra, Eun Jung Kim 0001 |
ICCD | 2 |
| 2013 | A reconfigurable computing architecture for semantic information filteringabstractThe increasing amount of information accessible to a user digitally makes information retrieval & filtering difficult, time consuming and ineffective. New meaning representation techniques proposed in literature help to improve accuracy but increase problem size exponentially. In this paper, we present a novel reconfigurable computing architecture that addresses this issue, outperforms contemporary many-core processors such as Intel's Single Chip Cloud computer and Nvidia's GPU's by ~20x for semantic information filtering. We validate our design using industry standard System-on-chip virtual prototyping and synthesis tools. Such a high performance reconfigurable architecture can form a template for a wide range of content-based and collaborative filtering engines used for big-data analytics. Aalap Tripathy, Ka Chon Ieong, Atish Patra, Rabi N. Mahapatra |
IEEE BigData | 4 |
| 2013 | Exploring topologies for source-synchronous ring-based network-on-chipabstractThe mesh interconnection network has been preferred by the Network-on-Chip (NoC) community due to its simple implementation, high bandwidth and overall scalability. Most existing mesh-based NoC designs operate the mesh at the same or lower clock speed as the processing elements (PEs). Recently, a new source synchronous ring-based NoC architecture has been proposed, which runs significantly faster than the PEs and offers a significantly higher bandwidth and lower communication latency. The authors implement the NoC topology as a mesh of rings, which occupies the same area as that of a mesh. In this work, we evaluate two alternate source synchronous ring-based NoC topologies called the ring of stars (ROS) and the spine with rings (SWR), which occupy a much lower area, and are able to provide better performance in terms of communication latency compared to a state of the art mesh. In our proposed topologies, the clock and the data NoC are routed in parallel, yielding a fast, synchronous, robust design. Our design allows the PEs to extract a low jitter clock from the high speed ring clock by division. The area and performance of these ring-based NoC topologies is quantified. Experimental results on synthetic traffic show that the new ring-based NoC designs can provide significantly lower latency (upto 4.6×) compared to a state of the art mesh. The proposed floorplan-friendly topologies use fewer buffers (upto 50% less) and lower wire length (upto 64.3% lower) compared to the mesh. Depending on the performance and the area desired, a NoC designer can select among the topologies presented. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
DATE | 3 |
| 2013 | A source-synchronous Htree-based network-on-chipabstractMost existing Network-on-Chip (NoC) designs operate at the same or lower clock speed as the processing elements (PEs). Recently, a new source-synchronous ring-based NoC architecture has been proposed, which runs significantly faster than the PEs and offers a significantly higher bandwidth and lower communication latency. However, the ring-based design assumes a separate clock distribution scheme for the NoC and the PEs, and uses a standard mesh topology for the NoC. In this work, we present a source synchronous ring-based NoC, laid out in an H-tree topology, with each data link being routed parallel to a clock ring. The clock is generated and distributed by multiple standing wave oscillator (SWO) rings, which are also laid out in an H-tree topology. Our design allows the PEs to extract a low jitter clock directly from the high speed ring-based SWO clock by division. Moreover, since the PEs are synchronous with the ring clock, they do not need synchronizers while communicating with the NoC. We also show that by recursively duplicating links in the H-tree based source synchronous NoC (Hnoc), we can obtain new hybrid NoC structures. In the limit, this recursive duplication causes the H-tree based NoC to morph into the meshbased source synchronous NoC (Mnoc). The performance of each such intermediate hybrid NoC structure is quantified in terms of area, link utilization and contention free latency. We also enhance the performance of the hybrid NoCs by widening congested links, and quantify the tradeoffs. Experimental results show that the hybrid NoC designs can provide significantly lower latency (upto 5× lower) and are able to sustain a higher injection rate (upto 6.8× higher) compared to a state of the art mesh. Moreover, these hybrid NoC designs use fewer buffers (upto 19.4% less) and lower wire length (upto 19.7% lower) compared to a mesh. Based on the performance and the area tradeoffs, an NoC designer can select any hybrid NoC structure among the presented. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | Distributed Collaborative Filtering on a Single Chip Cloud ComputerabstractMany-cores on chip have now become a reality. They necessitate the revisit of several layers of a cloud infrastructure. For this to happen, parallel programming runtimes need to be designed for many-cores on chip as the target architecture. In this paper, we show that Map Reduce programming paradigm can be adapted to run on Intel's experimental single chip cloud computer (SCC) with 48-cores on chip. We demonstrate this using a Collaborative Filtering (CF) recommender system as an application. CF is widely used in e-commerce deployments to predict user's preference towards an unknown item from their past ratings. We address scalability with data partitioning, combining and sorting algorithms, maximize data locality to minimize communication cost within the SCC cores. We demonstrate ~2x speedup, ~94% lower power consumption for benchmark workloads as compared to a distributed cluster multi-processor nodes in use today. Aalap Tripathy, Atish Patra, Suneil Mohan, Rabi N. Mahapatra |
IC2E | 4 |
| 2013 | A low-jitter phase-locked resonant clock generation and distribution schemeabstractClock distribution networks have traditionally been optimized to minimize end-to-end delay of the distribution network. However, since most digital ICs have an on-chip PLL, a more relevant design goal is to minimize cycle-to-cycle jitter. In this paper, we present a novel low-jitter phase-locked clock generation and distribution methodology which uses resonant standing wave oscillators (SWOs). In contrast to traveling wave oscillator rings (TWOs or “rotary” clocks), our SWO achieves the same phase at every point in the ring, making it amenable to a synchronous design methodology. The standing wave oscillator is controlled by coarse as well as fine tuning. Coarse tuning is achieved by varying the ring inductance, while fine tuning is accomplished by varying the ring capacitance. Clock distribution is done by routing the resonant ring chip-wide in a “comb” like manner. Experimental results demonstrate that the cycle-to-cycle jitter and skew of our approach is dramatically lower than existing schemes, while the power consumption is significantly lower as well. These benefits occur due to the resonant nature of our SWO-based clock generation and distribution approach. Ayan Mandal, Kalyana C. Bollapalli, Nikhil Jayakumar, Sunil P. Khatri, Rabi N. Mahapatra |
ICCD | 5 |
| 2012 | A fast, source-synchronous ring-based network-on-chip designabstractMost network-on-chip (NoC) architectures are based on a mesh-based interconnection structure. In this paper, we present a new NoC architecture, which relies on source synchronous data transfer over a ring. The source synchronous ring data is clocked by a resonant clock, which operates significantly faster than individual processors that are served by the ring. This allows us to significantly improve the cross section bandwidth and the latency of the NoC. We have validated the design using a 22 nm predictive process. Compared to the state-of-the-art mesh based NoC, our scheme achieves a 4.5× better bandwidth, 7.4× better contention free latency with 11% lower area and 35% lower power. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
DATE | 3 |
| 2012 | Architectural simulations of a fast, source-synchronous ring-based Network-on-Chip designabstractRecently, a new source-synchronous ring-based NoC architecture has been proposed, which runs significantly faster than the PEs and offers a high bandwidth and low contention free latency. Architectural simulations show that the original ring-based NoC design suffers from deadlock. In this paper, we explore the architectural aspects of the fast ring-based NoC after redesigning the routers used in the previous authors' work to avoid deadlock. Architectural results obtained on synthetic traffic demonstrate that the modified ring-based NoC has up to 3.5× lower latency and up to 2.9× higher maximum sustained injection rate compared with a state of the art mesh-based NoC. Ayan Mandal, Sunil P. Khatri, Rabi N. Mahapatra |
ICCD | 3 |
| 2012 | The Three-Tier Security Scheme in Wireless Sensor Networks with Mobile SinksabstractMobile sinks (MSs) are vital in many wireless sensor network (WSN) applications for efficient data accumulation, localized sensor reprogramming, and for distinguishing and revoking compromised sensors. However, in sensor networks that make use of the existing key predistribution schemes for pairwise key establishment and authentication between sensor nodes and mobile sinks, the employment of mobile sinks for data collection elevates a new security challenge: in the basic probabilistic and q-composite key predistribution schemes, an attacker can easily obtain a large number of keys by capturing a small fraction of nodes, and hence, can gain control of the network by deploying a replicated mobile sink preloaded with some compromised keys. This article describes a three-tier general framework that permits the use of any pairwise key predistribution scheme as its basic component. The new framework requires two separate key pools, one for the mobile sink to access the network, and one for pairwise key establishment between the sensors. To further reduce the damages caused by stationary access node replication attacks, we have strengthened the authentication mechanism between the sensor and the stationary access node in the proposed framework. Through detailed analysis, we show that our security framework has a higher network resilience to a mobile sink replication attack as compared to the polynomial pool-based scheme. Amar A. Rasheed, Rabi N. Mahapatra |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | A Power and Throughput-Efficient Packet Classifier with n Bloom FiltersabstractPacket processing is a critical operation in a high-speed router, and in order for this router to achieve memory efficient and fast O(1) lookup operations, Bloom filters (BFs) have been widely used as a packet classifier to reduce expensive hash table accesses. However, it has been identified that a parallel packet classifier (PPC), using all n parallel BFs for a lookup, is neither power nor throughput efficient for high-speed routers. In this paper, we propose a multitiered packet classifier (MPC), both to save power and to improve throughput, with the same memory size as that of a PPC. While a PPC with n BFs consumes Θ(n) BF access complexity for a lookup, our MPC is designed to have the complexity which is probabilistically significantly less than Θ(n). Furthermore, by preprocessing a group of lookups in one cycle in an MPC, we assign each lookup to its associated BF at best effort, and consequently, obtain a higher throughput. With the same reason, as in preprocessing, our MPC design reduces a significant amount of power by preventing accesses to noninvolved BFs during a lookup. In simulation for flow identification with NLANR traces, we observed that the MPC throughput is increased by at most 100 percent, compared to a PPC. Additionally, our MPC shows 4.2 times power efficiency over an equivalent PPC, in terms of power saving. Heeyeol Yu, Rabi N. Mahapatra |
IEEE Trans. Computers | 2 |
| 2011 | Key Predistribution Schemes for Establishing Pairwise Keys with a Mobile Sink in Sensor NetworksabstractSecurity services such as authentication and pairwise key establishment are critical to sensor networks. They enable sensor nodes to communicate securely with each other using cryptographic techniques. In this paper, we propose two key predistribution schemes that enable a mobile sink to establish a secure data-communication link, on the fly, with any sensor nodes. The proposed schemes are based on the polynomial pool-based key predistribution scheme, the probabilistic generation key predistribution scheme, and the Q-composite scheme. The security analysis in this paper indicates that these two proposed predistribution schemes assure, with high probability and low communication overhead, that any sensor node can establish a pairwise key with the mobile sink. Comparing the two proposed key predistribution schemes with the Q-composite scheme, the probabilistic key predistribution scheme, and the polynomial pool-based scheme, our analytical results clearly show that our schemes perform better in terms of network resilience to node capture than existing schemes if used in wireless sensor networks with mobile sinks. Amar A. Rasheed, Rabi N. Mahapatra |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | Reliability aware power management for dual-processor real-time embedded systemsabstractPrimary-Backup (PB) model has been a widely used model for reliability in dual-processor real-time systems. In recent literature, there have been a few works focussing on minimizing energy consumption of periodic task sets executing on such systems. One of the major drawbacks of these works is that they ignore the effects of frequency-scaling on fault arrival rates. In this paper, we present a modified Primary-Backup model for dual-processor systems that aims to maintain the reliability when employing power management techniques to minimize the overall energy consumption. Furthermore, the proposed approach exploits the uncertainties in the execution time of real-time tasks to better predict the available slack for energy management. The proposed modified PB-based Reliability-Aware Power Management (RAPM) approach was tested with synthetic task sets on both homogeneous and heterogeneous dual-processor systems. Simulation results show that it can achieve up to 67% savings in expected energy consumption for low utilization task sets and up to 32% savings for high utilization task sets without any loss in reliability in heterogeneous dual-processor systems. Ranjani Sridharan, Rabi N. Mahapatra |
DAC | 2 |
| 2010 | Low power nanoscale buffer management for network on chip routersabstractNetwork-on-Chip (NoC) is an on-chip communication solution in the future system-on-a-chip (SoC) necessitating high performance operation with low power dissipation. We present a novel dynamic power management technique for low power NoC router buffers using nano CMOS SRAMS. A feedback controller was designed for block level power management and a power aware adaptive controller was designed for low power flit storage encoding to reduce energy consumptions in the router buffers. Experiments with the proposed scheme showed up to 20% reduction in energy consumption while improving throughput by up to 21%. Suman Kalyan Mandal, Ron Denton, Saraju P. Mohanty, Rabi N. Mahapatra |
ACM Great Lakes Symposium on VLSI | 4 |
| 2010 | A parallel architecture for meaning comparisonabstractIn this paper we present a fine grained parallel architecture that performs meaning comparison using vector cosine similarity (dot product). Meaning comparison assigns a similarity value to two objects (e.g. text documents) based on how similar their meanings (represented as two vectors) are to each other. The novelty of our design is the fine grained parallelism which is not exploited in available hardware based dot product processor designs and can not be achieved in traditional server class processors like the Intel Xeon. We compare the performance of our design against that of available hardware based dot product processors as well a server class processor using optimum software code performing the same computation. We show that our hardware design can achieve a speedup of 62,000 times compared to an available hardware design and a speedup of 8866 times with 33% (1.5 times) less power consumption, compared to software code running on Intel Xeon processor for 1024 basis vectors. Our design can significantly reduce the amount of servers required for similarity comparison in a distributed search engine. Thus it can enable reduction in energy consumption, investment, operational costs and floor area in search engine data centers. This design can also be deployed for other applications which require fast dot product computation. Suneil Mohan, Amitava Biswas, Aalap Tripathy, Jagannath Panigrahy, Rabi N. Mahapatra |
IPDPS | 5 |
| 2010 | Dynamic Context Compression for Low-Power Coarse-Grained Reconfigurable ArchitectureabstractMost of the coarse-grained reconfigurable architectures (CGRAs) are composed of reconfigurable ALU arrays and configuration cache (or context memory) to achieve high performance and flexibility. Specially, configuration cache is the main component in CGRA that provides distinct feature for dynamic reconfiguration in every cycle. However, frequent memory-read operations for dynamic reconfiguration cause much power consumption. Thus, reducing power in configuration cache has become critical for CGRA to be more competitive and reliable for its use in embedded systems. In this paper, we propose dynamically compressible context architecture for power saving in configuration cache. This power-efficient design of context architecture works without degrading the performance and flexibility of CGRA. Experimental results show that the proposed approach saves up to 39.72% power in configuration cache with negligible area overhead (2.16%). Yoonjin Kim, Rabi N. Mahapatra |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Design Space Exploration for Efficient Resource Utilization in Coarse-Grained Reconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures (CGRAs) aim to achieve both goals of high performance and flexibility. In addition, power consumption is significant for the reconfigurable architecture to be used as a competitive processing core in embedded systems. However, the existing reconfigurable architectures require too much area and power. In this paper, we propose a new design space exploration flow, optimizing CGRA to reduce area and power with enhancing performance for digital signal processing (DSP) application domain. It reduces the array size through efficient arrangement of array components and customization of their interconnection, exploiting input patterns belonging to the DSP application domain. Such a design flow is based on pipelining and sharing of area/delay-critical resources in the processing element array. Experimental results show that for DSP applications, the proposed approach reduces area by up to 36.75%, average execution time by 36.78%, and average power by 31.85% when compared with the existing CGRA architecture. Yoonjin Kim, Rabi N. Mahapatra, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Hierarchical reconfigurable computing arrays for efficient CGRA-based embedded systemsabstractCoarse-grained reconfigurable architecture (CGRA) based embedded system aims at achieving high system performance with sufficient flexibility to map variety of applications. However, significant area and power consumption in the arrays prohibits its competitive advantage to be used as a processing core. In this work, we propose hierarchical reconfigurable computing array architecture to reduce power/area and enhance performance in configurable embedded system. The CGRA-based embedded systems that consist of hierarchical configurable computing arrays with varying size and communication speed were examined for multimedia and other applications. Experimental results show that the proposed approach reduces on-chip area by 22%, execution time by up to 72% and reduces power consumption by up to 55% when compared with the conventional CGRA-based architectures. Yoonjin Kim, Rabi N. Mahapatra |
DAC | 2 |
| 2009 | Dynamic context management for low power coarse-grained reconfigurable architectureabstractCoarse-grained reconfigurable architectures (CGRA) require many processing elements (PEs) and a configuration memory unit (configuration cache) for reconfiguration of its PE array. Al-though this structure is meant for high performance and flexibility, it consumes significant power. Specially, power consumption by configuration cache is explicit overhead compared to other types of IP cores. Reducing power in configuration cache is very crucial for CGRA to be more competitive and reliable processing core in embedded systems. In this paper, we propose a dynamic context management strategy for power saving in configuration cache. This power-efficient approach works without degrading the per-formance and flexibility of CGRA. Experimental results show that the proposed approach saves 38.24%/38.15% of the power in write/read-operation of configuration cache with negligible area overhead compared to the previous design. Yoonjin Kim, Rabi N. Mahapatra |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Search Co-Ordination by Semantic Routed NetworkabstractSpecialized search engines have potential to provide superior search precision, relevance and recall. A specialized P2P overlay network called semantic routed network, which forwards search queries based on their meanings, can unify a large number of specialized search engines under a single Internet wide search service. Providing fast and successful routing of search queries is a challenging task due to conflicting tradeoff requirements. We present P2P topology generation and semantic routing table compaction techniques to realize fast and successful query routing. Our simulation shows that a SRN coordinating 800 search engines can achieve a competitive 100% query routing success within 6 messaging delays, and have query delivery response of 1.89 messaging delays. Amitava Biswas, Suneil Mohan, Rabi N. Mahapatra |
ICCCN | 3 |
| 2009 | A distributed concurrent on-line test scheduling protocol for many-core NoC-based systemsabstractConcurrent on-line testing (COLT) of many-core systems-on-chip (SoC) has been recently proposed by researchers in response to the growing threat of electronic wear-out to system operational lifetimes and to the increasing reliability and availability demands of safety-critical applications. Previous research in concurrent on-line testing has focused on centralized approaches to manage core testing while the system is available to execute normal user applications. However, as technology scaling allows dozens and hundreds of processing cores to be placed on a single chip, these centralized approaches are not scalable solutions. In this paper, a distributed concurrent on-line test scheduling protocol is proposed and evaluated against previously developed solutions. Our experiments show that a distributed COLT scheduler can test a moderately-sized SoC with a speedup of 3.85 over centralized approaches while consuming 84% less energy, and performance benefits improve as the number of cores per chip increases. This research also presents a core test ordering algorithm - Code-Division Core Test Scheduling - that provides an additional 40% reduction in system test latency compared to other schedulers. Jason D. Lee, Rabi N. Mahapatra, Praveen Bhojwani |
ICCD | 2 |
| 2009 | A Hash-based Scalable IP lookup using Bloom and Fingerprint FiltersabstractSeveral challenges in the IP lookup architecture must be addressed for a high-speed forwarding in a large scale routing table: power, memory, and lookup complexity. Hash-based architectures have lookup schemes that are recognized for being both power and memory efficient due to their O(1) lookup, in contrast to other contemporary architectures. In this paper, we propose a novel hash architecture to address these issues by using pipelined Bloom and fingerprint filters for a binary searching in keys. The proposed hash scheme encodes keys' indexes to an on-chip fingerprint table, approximately returns a few indexes in a key query without pointer overhead, and makes a perfect match in an off-chip key table. Due to a memory banking system in pipeline stages, we can achieve O(1) pipelined throughput complexity of insertion, deletion, and query operations. For the IP lookup, a Lulea bitmap with our hash scheme supports a prefix lookup without inflating the numbers of prefixes and next-hops, so that our scalable hash-based scheme can achieve the worst case O(1) IP lookup. The simulation with large scale routing tables shows that our IP lookup scheme offers 4.5 and 50.1 times memory and power efficiencies than other contemporary hash and TCAM schemes, respectively. Heeyeol Yu, Rabi N. Mahapatra, Laxmi N. Bhuyan |
ICNP | 2 |
| 2009 | Mobile sink using multiple channels to defend against wormhole attacks in wireless sensor networksabstractSecurity is a necessity for many sensor-network applications. A particularly harmful attack against sensor networks is known as the wormhole attack, where an adversary tunnels the messages received in one part of the network over a low-latency link and replays them in a different part of the same network. This article presents the threat posed by wormhole attacks to wireless sensor networks with mobile sinks. A novel technique that involves leveraging channel diversity for defense against the wormhole attack has been proposed. Through quantitative analyses, it is shown that even when 50% of a sensor node's neighbors are malicious devices, the provision of one extra available channel for communication with the mobile sink reduces the probability of a wormhole attack to almost zero. Amar A. Rasheed, Rabi N. Mahapatra |
IPCCC | 2 |
| 2009 | A key pre-distribution scheme for heterogeneous sensor networksabstractKey pre-distribution techniques developed recently to establish pairwise keys between nodes with no or limited mobility. Existing schemes make use of only one key pool to establish secure links between stationary and mobile nodes, allowing an attacker to easily gain control of the network by randomly compromising a small fraction of stationary nodes. A method of preventing this type of security breach is the use of separate key pools for mobile and stationary nodes, in which small fractions of stationary nodes are randomly pre-selected to help the mobile nodes establish links with stationary nodes. Analysis shows that with 10% of stationary nodes carry a key from the mobile key pool. To recover any key from the mobile key pool and gain control of the network, an attacker would have to capture 20.8 times more stationary nodes than if a single key pool is used for both mobile and stationary nodes. Amar A. Rasheed, Rabi N. Mahapatra |
IWCMC | 2 |
| 2009 | Low Power Reconfiguration Technique for Coarse-Grained Reconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures (CGRAs) require many processing elements (PEs) and a configuration memory unit (configuration cache) for reconfiguration of its PE array. Although this structure is meant for high performance and flexibility, it consumes significant power. Specially, power consumption by configuration cache is explicit overhead compared to other types of intellectual property (IP) cores. Reducing power is very crucial for CGRA to be more competitive and reliable processing core in embedded systems. In this paper, we propose a reusable context pipelining (RCP) architecture to reduce power-overhead caused by reconfiguration. It shows that the power reduction can be achieved by using the characteristics of loop pipelining, which is a multiple instruction stream, multiple data stream (MIMD)-style execution model. RCP efficiently reduces power consumption in configuration cache without performance degradation. Experimental results show that the proposed approach saves much power even with reduced configuration cache size. Power reduction ratio in the configuration cache and the entire architecture are up to 86.33% and 37.19%, respectively, compared to the base architecture. Yoonjin Kim, Rabi N. Mahapatra, Ilhyun Park, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | IntellBatt: towards smarter battery designabstractBattery lifetime and safety are primary concerns in the design of battery operated systems. Lifetime management is typically supervised by the system via battery-aware task scheduling, while safety is managed on the battery side via features deployed into smart batteries. This research proposes IntellBatt; an intelligent battery cell array based novel design of a multi-cell battery that offloads battery lifetime management onto the battery. By deploying a battery cell array management unit, IntellBatt exploits various battery related characteristics such as charge recovery effect, to enhance battery lifetime and ensure safe operation. This is achieved by using real-time cell status information to selects cells to deliver the required load current, without the involvement of a complex task scheduler on the host system. The proposed design was evaluated via simulation using accurate cell models and real experimental traces from a portable DVD player. The use of a multi-cell design enhanced battery lifetime by 22% in terms of battery discharge time. Besides a standalone deployment, IntellBatt can also be combined with existing battery-aware task scheduling approaches to further enhance battery lifetime. Suman Kalyan Mandal, Praveen Bhojwani, Saraju P. Mohanty, Rabi N. Mahapatra |
DAC | 4 |
| 2008 | Feedback-controlled reliability-aware power management for real-time embedded systemsabstractIn recent literature it has been reported that Dynamic Power Management (DPM) may lead to decreased reliability in real-time embedded systems. The ever-shrinking device sizes contribute further to this problem. In this paper, we present a reliability aware power management algorithm that aims at reducing energy consumption while preserving the overall system reliability. The idea behind the proposed scheme is to utilize the dynamic slack to scale down processes while ensuring that the overall system reliability does not reduce drastically. The proposed algorithm employs a proportional feedback controller to keep track of the overall miss ratio of a system of tasks and provide additional level of fault-tolerance based on demand. It was tested with both real-world and synthetic task sets and simulation results have been presented. Both fixed and dynamic priority scheduling policies have been considered for demonstration of results. Ranjani Sridharan, Nikhil Gupta 0004, Rabi N. Mahapatra |
DAC | 3 |
| 2008 | A New Array Fabric for Coarse-Grained Reconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures (CGRA) employ square or rectangular arrays composed of many computational resources for high performance. Though these array fabrics are mostly suitable for embedded systems including multimedia applications, they occupy large area and consume much power. Therefore, reducing area and power of CGRA is necessary for the reconfigurable architectures to be used as a competitive IP core in embedded systems. In this paper, we propose a new array fabric for designing CGRA to reduce area and power consumption without any performance degradation. This cost-effective approach is able to reduce the array size through efficient arrangement of array components and their inter-connections. Experimental results show that for multimedia applications, the proposed array fabric reduces area up to 40.32% and saves power by up to 28.35 % when compared with the existing CGRA architecture. Yoonjin Kim, Rabi N. Mahapatra |
DSD | 2 |
| 2008 | A Throughput-Efficient Packet Classifier with n Bloom filtersabstractPacket classification is a critical data path in a highspeed router. Due to memory efficiency and fast lookup, Bloom filters (BFs) have been widely used for packet classification in a high-speed router. However, in a parallel packet classifier (PPC) of n parallel BFs, using all n BFs for a lookup is not throughput efficient in a high speed router. In this paper, we propose a multi-tiered packet classifier (MPC) for high throughput with the same memory size as a PPC. While a PPC of n BFs needs Theta(n) BF access complexity for a lookup, our MPC is geared to have the complexity which is probabilistically far less than Theta(n). Furthermore, by preprocessing a group of lookups in one cycle, each lookup is assigned to its associated BF at best effort, so that a higher throughput in an MPC is obtained. In simulation for flow identification with NLANR traces, we observed that, at most, 2.0 times more throughput was recorded than a PPC . Heeyeol Yu, Rabi N. Mahapatra |
GLOBECOM | 2 |
| 2008 | Optimization of Semantic Routing TableabstractIn a semantic routed network (SRN), messages are routed based on the meaning of the message key. This means network nodes can be addressed by the meaning of the data content. Providing fast and successful semantic routing for any key is a challenging task due to conflicting performance demands. We present semantic routing table optimizing techniques to realize small world topology for fast and successful message routing. Our simulation shows that a SRN of 1000 nodes can achieve a competitive 57% routing success within 6 messaging delays, and have message delivery response of 3.3 messaging delays. Amitava Biswas, Suneil Mohan, Rabi N. Mahapatra |
ICCCN | 3 |
| 2008 | In-field NoC-based SoC testing with distributed test vector storageabstractThe operational lifetimes of SoC and microprocessors face growing threats from technology scaling and increasing device temperature and power density. In-field (or on-line) testing of NoC-based SoC is an important technique in ensuring system integrity throughout this potentially shorter lifetime. Whether in-field testing is conducted concurrently with normal applications or executed in isolation, application intrusion must be minimized in order to maintain system availability. Specialized infrastructure IP have been proposed to manage on-line testing by scheduling tests and delivering test vectors to the various cores within the SoC from a centralized location. However, as the number of cores integrated into a single chip continues to increase, issuing test vectors from a centralized location is not a scalable solution. These increased distances that test vectors must travel have become a major concern for on-line testing because of its direct impact on application intrusion in terms of energy consumption, network load, and latency. In this paper, we apply a distributed storage technique to bound and minimize this distance, thereby minimizing network load, energy consumption, and test delivery latency across the entire network. Our experiments show that test delivery latency and energy consumption is reduced by approximately 90% for moderately sized NoC. Jason D. Lee, Rabi N. Mahapatra |
ICCD | 2 |
| 2008 | A Memory-Efficient Hashing by Multi-Predicate Bloom Filters for Packet ClassificationabstractHash tables (HTs) are poorly designed for multiple off-chip memory accesses during packet classification and critically affect throughput in high-speed routers. Therefore, an HT with fast on-chip memory and high-capacity off-chip memory for predictable lookup-throughput is desirable. Both a legacy HT (LHT) and a recently proposed fast HT (FHT) have the disadvantage of memory overhead due to pointers and duplicate items in linked lists. Also, memory usage for an FHT did not consider the bits in counters for fair comparison with an LHT. In this paper, we propose a novel hash architecture called a Multi-predicate Bloom-filtered HT (MBHT) using parallel Bloom filters and generating off-chip memory addresses in the base- 2xnumber system, xisin{1,2,hellip}, which removes the overhead of pointers. Using a larger base of number system, an MBHT reduces on-chip memory size by a factor of log2b2/ log2b1where b1and b2are bases of number system (b2>b1). Compared to an FHT, the MBHT is approximately x(log2n + 4)/(2 log2n) times more efficient for on-chip memory, where n is the number of keys. This results in a significant reduction in the number of off- chip memory accesses. A simulation with a dataset of packets from NLANR shows the on-chip memory reductions by 1.7 and 2 times over an LHT and an FHT are made. Besides, an MBHT of base-16 needs less off-chip memory accesses by 2117 in total URL queries of NLANR, compared to an FHT. Heeyeol Yu, Rabi N. Mahapatra |
INFOCOM | 2 |
| 2008 | An Efficient Key Distribution Scheme for Establishing Pairwise Keys with a Mobile Sink in Distributed Sensor NetworksabstractSecurity services such as authentication and pair-wise key establishment are critical in sensor networks. They enable sensor nodes to communicate securely with each other using cryptographic techniques. In this paper, we propose a novel key predistribution scheme that enables a mobile sink to establish a secure data communication link with any sensor nodes on the fly. The proposed scheme is based on the polynomial pool-based key pre-distribution scheme and the scheme in [7]. The security analysis in this paper indicates that for a given node density of d sensors within the communication range of the mobile sink and with certain probabilities q and p, our scheme assures, with high probability, that any sensor node can establish a pair-wise key with the mobile sink. It remains perfectly secure up to the capture of a certain fraction of sensor nodes. Amar A. Rasheed, Rabi N. Mahapatra |
IPCCC | 2 |
| 2008 | Reusable context pipelining for low power coarse-grained reconfigurable architectureabstractCoarse-grained reconfigurable architectures (CGRA) require many processing elements and a configuration memory unit (configuration cache) for reconfiguration of the ALU array elements. This structure consumes significant amount of power. Power reduction during reconfiguration is necessary for the reconfigurable architecture to be used as a competitive IP core in embedded systems. In this paper, we propose a power-conscious reusable context pipelining architecture for CGRA that efficiently reduces power consumption in configuration cache without performance degradation. Experimental results show that the proposed approach saves up to 57.97% of the total power consumed in the configuration cache with reduced configuration cache size compared to the previous approach. Yoonjin Kim, Rabi N. Mahapatra |
IPDPS | 2 |
| 2008 | A space- and time-efficient hash table hierarchically indexed by Bloom filtersabstractHash tables (HTs) are poorly designed for multiple memory accesses during IP lookup and this design flow critically affects their throughput in high-speed routers. Thus, a high capacity HT with a predictable lookup throughput is desirable. A recently proposed fast HT (FHT) [20] has drawbacks like low on-chip memory utilization for a high-speed router and substantial memory overheads due to off-chip duplicate keys and pointers. Similarly, a Bloomier filter-based HT (BFHT) [13], generating an index to a key table, suffers from setup failures and static membership testing for keys. In this paper, we propose a novel hash architecture which addresses these issues by using pipelined Bloom filters. The proposed scheme, a hierarchically indexed HT (HIHT), generates indexes to a key table for the given key, so that the on-chip memory size is reduced and the overhead of pointers in a linked list is removed. Secondly, an HIHT demonstrates approximately 5.1 and 2.3 times improvement in on- chip space efficiency with at most one off-chip memory access, compared to an FHT and a BFHT, respectively. In addition to our analyses on access time and memory space, our simulation for IP lookup with 6 BGP tables shows that an HIHT exhibits 4.5 and 2.0 times on-chip memory efficiencies for 160 Gbps router than an FHT and a BFHT, respectively. Heeyeol Yu, Rabi N. Mahapatra |
IPDPS | 2 |
| 2008 | PowerAntz: distributed power sharing strategy for network on chipabstractAdvent of Network-on-Chip (NoC) based complex system designs made on-chip power management a challenging issue. Power management schemes have been proposed to tackle the problem. But they fail to provide optimal sharing when the power budget distribution varies significantly among on-chip components that are placed further apart on a chip. This paper presents PowerAntz, a distributed power management strategy for NoC based systems. This adaptive and distributed approach to power sharing across various components of a large chip is shown to be a scalable solution. Our experiments have demonstrated PowerAntz to be up to 30% more effective in distributing power budget compared to existing strategies. Further, it also achieves up to 21.25% improvement in power utilization while keeping overhead as low as zero in best case. Suman Kalyan Mandal, Rabi N. Mahapatra |
ISLPED | 2 |
| 2008 | C++ Dynamic Cast in Autonomous Space SystemsabstractThe dynamic cast operation allows flexibility in the design and use of data management facilities in object- oriented programs. Dynamic cast has an important role in the implementation of the data management services (DMS) of the mission data system project (MDS), the jet propulsion laboratory's experimental work for providing a state-based and goal-oriented unified architecture for testing and development of mission software. DMS is responsible for the storage and transport of control and scientific data in a remote autonomous spacecraft. Like similar operators in other languages, the C++ dynamic cast operator does not provide the timing guarantees needed for hard real-time embedded systems. In a recent study, Gibbs and Stroustrup (G&S) devised a dynamic cast implementation strategy that guarantees fast constant-time performance. This paper presents the definition and application of a co-simulation framework to formally verify and evaluate the G&S fast dynamic casting scheme and its applicability in the mission data system DMS application. We describe the systematic process of model-based simulation and analysis that has lead to performance improvement of the G&S algorithm's heuristics by about a factor of 2. Damian Dechev, Rabi N. Mahapatra, Bjarne Stroustrup, David A. Wagner 0002 |
ISORC | 2 |
| 2008 | Secure Data Collection Scheme in Wireless Sensor Network with Mobile SinkabstractWireless sensor networks that use a mobile sink to collect sensor data along a predetermined path raise a new security challenge: without verifying the source of the data request message, the network will become vulnerable to attacks. We propose an efficient security scheme, which divides the sinkpsilas data collection path into grids, sensors in each grid, uses secret keying in-formation and collision-resistant hash functions to authenticate the source of beacons. Through probabilistic analysis and definitive simulation, the proposed scheme shows with 60% of the grids under wormhole attacks, the probability that a node reply to a malicious beacon is 0.1. Amar A. Rasheed, Rabi N. Mahapatra |
NCA | 2 |
| 2008 | A Dynamic Slack Management Technique for Real-Time Distributed Embedded SystemsabstractThis work presents a novel slack management technique, the Service Rate Proportionate(SRP) Slack Distribution, for real-time distributed embedded systems to reduce energy consumption. The proposed SRP based Slack Distribution Technique has been considered with EDF and Rate Based scheduling schemes that are most commonly used with embedded systems. A fault tolerance mechanism has also been incorporated into the proposed technique inorder to utilize the available dynamic slack to maintain checkpoints and provide for rollbacks on faults. Results show that in comparion to contemporary techniques, the proposed SRP Slack Distribution Technique provides for about 29% more performance/overhead improvement benefits when validated with real world and random benchmarks. Subrata Acharya, Rabi N. Mahapatra |
IEEE Trans. Computers | 2 |
| 2008 | Robust Concurrent Online Testing of Network-on-Chip-Based SoCsabstractLifetime concerns for complex systems-on-a-chip (SoC) designs due to decreasing levels in reliability motivate the development of solutions to ensure reliable operation. A precursor to any proposed recovery scheme would require the identification of failures in the system. Non-concurrent in-field testing is an impractical solution due to prohibitive costs in terms of test power and test time.This novel research proposes the use of concurrent online testing (COLT) to circumvent these issues. A test infrastructure-intellectual property (TI-IP) is deployed within network-on-chip (NoC)-based SoC designs to provide online test support while managing intrusion of test into executing applications within the system. This research describes the architecture and operation of a TI-IP capable of COLT. To address scalability of this solution, we show how these would operate when more than one is deployed in an SoC. In the absence of benchmarks for the analysis of COLT, two baseline and eight TI-IP configuration variations within SoC test configurations were developed using application and test benchmarks from the research domain. The power profiles from theNoCSimsimulation environment are reported here demonstrating how different configurations of TI-IPs would operate. A robust TI-IP protocol is also specified and possible hazards and their mitigations are identified. Praveen Bhojwani, Rabi N. Mahapatra |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Data Handling Limits of On-Chip InterconnectsabstractWith shrinking feature size and growing integration density in the deep sub-micrometer (DSM) technologies, the global buses are fast becoming the ldquoweakest-linksrdquo in VLSI design. They have large delays and are error-prone. Especially, in system-on-chip (SoC) designs, where parallel interconnects run over large distances, they pose difficult research and design problems. This paper presents a two-fold approach for evaluating the signal and data carrying capacity of on-chip interconnects. In the first approach, the wire is modeled as a linear time invariant (LTI) system and a frequency response is studied. The second approach addresses delay and reliability in interconnects from an information theoretic perspective. Simulation results for an 8-bit-wide bus in 0.1-mum technology are presented for both approaches. The results closely match to a similar optimal bus clock frequency that will result in the maximumdatatransferrate. Moreover, this optimal frequency is higher than that achieved by present day designs which accommodate the worst case delays. The first approach achieves this higher transmission rate using ideal signal shapes, instead of square pulses, while the second approach uses coding techniques to eliminate high delay cases to generate a higher transmission rate. It is seen that the signal delay distribution has a long tail, meaning that most signals arrive at the output much faster than the worst case delay. Using communication theory, these ldquogoodrdquo signals arriving early can be used to predict/correct the ldquofewrdquo signals that arrive late. Rohit Singhal, Gwan S. Choi, Rabi N. Mahapatra |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | A Robust Protocol for Concurrent On-Line Test (COLT) of NoC-based Systems-on-a-ChipabstractConcurrent on-line testing (COLT) of complex systems-on-a-chip (SoC) designs under lowering noise margins and degrading lifetimes of on-chip components, provides the ideal solution for the monitoring of system health while managing intrusion into executing applications. Deploying Test Infrastructure-IPs (TI-IPs) into designs has demonstrated the feasibility of using COLT in SoCs. Identifying potential hazards and ensuring correct operation of COLT is critical to providing reliable health monitoring. With the emergence of networks-on-a-chip (NoC) as communication infrastructures not only suitable for application related on-chip communication, but also test access mechanisms to on-chip cores, the experimental setup in this research, deploys TI-IP in a NoC environment and demonstrates TI-IP operation, its communication protocol specification and other related costs. Praveen Bhojwani, Rabi N. Mahapatra |
DAC | 2 |
| 2007 | Dynamically compressible context architecture for low power coarse-grained reconfigurable arrayabstractMost of the coarse-grained reconfigurable array architectures (CGRAs) are composed of reconfigurable ALU arrays and configuration cache (or context memory) to achieve high performance and flexibility. Specially, configuration cache is the main component in CGRA that provides distinct feature for dynamic reconfiguration in every cycle. However, frequent memory-read operations for dynamic reconfiguration cause much power consumption. Thus, reducing power in configuration cache has become critical for CGRA to be more competitive and reliable for its use in embedded systems. In this paper, we propose dynamically compressible context architecture for power saving in configuration cache. This power-efficient design of context architecture works without degrading the performance and flexibility of CGRA. Experimental results show that the proposed approach saves up to 39.72% power in configuration cache with negligible area overhead. Yoonjin Kim, Rabi N. Mahapatra |
ICCD | 2 |
| 2007 | SAPP: scalable and adaptable peak power management in nocsabstractTo address peak power concerns in networks-on-chip (NoCs), dynamic peak power management schemes that handle varying power requirements are essential. Previous schemes use deterministic peak power budget management techniques that do not scale or adapt efficiently to changing power budget requirements. Using a non-deterministic and independent approach, this research proposes SAPP, a Scalable and Adaptable Peak Power management technique for NoCs. Evaluation of SAPP on uniform and non-uniform varying injection loads demonstrates flit latency and effective throughput improvements averaging 47% and 36%, respectively. Efficient power budget utilization makes SAPP an ideal technique for peak power management in NoCs under varying traffic patterns. Praveen Bhojwani, Jason D. Lee, Rabi N. Mahapatra |
ISLPED | 3 |
| 2007 | An Efficient Approach to On-Chip Logic MinimizationabstractBoolean logic minimization is being applied increasingly to a new variety of applications that demand very fast and frequent minimization services. These applications typically have access to very limited computing and memory resources, rendering the traditional logic minimizers ineffective. We present a new approximate logic minimization algorithm based on ternary trie. We compare its performance with Espresso-II and ROCM logic minimizers for routing table compaction and demonstrate that it is 100 to 1000 times faster and can execute with a data memory as little as 16 KB. We also found that the proposed approach can support up to 25000 incremental updates per second. We also compare its performance for compaction of the routing access control list and demonstrate that the proposed approach is highly suitable for minimizing large access control lists containing several thousand entries. Therefore, the algorithm is ideal for on-chip logic minimization. Seraj Ahmad, Rabi N. Mahapatra |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Information theoretic approach to address delay and reliability in long on-chip interconnectsabstractWith shrinking feature size and growing integration density in the Deep Sub-Micron technologies, the global buses are fast becoming the "weakest-links" in VLSI design. They have large delays and are error-prone. Especially, in system-on-chip (SoC) designs, where parallel interconnects run over large distances, the effects of crosstalk are detrimental to the overall system performance due to the large delays and un-reliability involved. This paper presents an information theoretic approach to address delay and reliability in long interconnects. A framework to calculate the capacity of a physical wire is laid out herein. The results for 8-bit wide buses of varying lengths in 0.1μm technology are also presented. The wires are modeled based on their calculated parasitic (R, L, C) values and the coupling (C, L) parameters. Using this model, results are obtained for the data transfer capacity of long interconnects. It is seen that for wide buses, the signal delay distribution has a long tail, meaning that most signals arrive at the output much faster than the worst case delay. Using communication-theory, these "good" signals arriving early can be used to predict/correct the "few" signals arriving late. Further, results show that for every bus configuration, there exists an optimal frequency of transmission that will result in the maximum data transfer rate. Also, this optimal frequency is higher than the pessimistic worst case delay based clock design. Rohit Singhal, Gwan S. Choi, Rabi N. Mahapatra |
ICCAD | 3 |
| 2006 | Reducing clock skew variability via crosslinksabstractIncreasingly significant variational effects present a great challenge for delivering desired clock skew reliably. Nontree clock network has been recognized as a promising approach to overcome the variation problem. Existing nontree clock routing methods are restricted to a few simple or regular structures, and often consume excessive amounts of wirelength. This paper suggests to construct a low-cost nontree clock network by inserting crosslinks in a given clock tree. The effects of the link insertion on clock skew variability are analyzed. Based on the analysis, this paper proposes two link insertion schemes that can quickly convert a clock tree to a nontree with significantly lower skew variability and very limited wirelength increase. In these schemes, the complicated nontree delay computation is circumvented. Further, they can be applied to the recently popular nonzero skew routing easily. The effectiveness of the proposed techniques has been validated through SPICE-based Monte Carlo simulations. Anand Rajaram, Jiang Hu 0001, Rabi N. Mahapatra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Analytical bound for unwanted clock skew due to wire width variationabstractUnder modern very large-scale integrated technology, process variations greatly affect circuit performance, especially clock skew, which is very timing sensitive. Unwanted skew due to process variation forms a bottleneck, preventing further improvement on clock frequency. Impact from intrachip interconnect variation is becoming remarkable and is difficult to be modeled efficiently due to its distributive nature. Through wire shaping analysis, the authors establish an analytical bound for the unwanted skew due to wire width variation, which is a nonnegligible factor among interconnect variations. Experimental results on benchmark circuits show that this bound is safer, tighter, and computationally faster than similar existing approach. Anand Rajaram, Jiang Hu 0001, Rabi N. Mahapatra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | Antenna Avoidance in Layer AssignmentabstractThe sustained progress of very-large-scale-integration (VLSI) technology has dramatically increased the likelihood of the antenna problem in the manufacturing process and calls for corresponding considerations in the routing stage. In this paper, the authors propose a technique that can handle the antenna problem during the layer-assignment (LA) stage, which is an important step between global routing and detailed routing. The antenna-avoidance problem is modeled as a tree-partitioning problem with a linear-time-optimal-algorithm solution. This algorithm is customized to guide antenna avoidance in the LA stage. A linear-time optimal jumper-insertion algorithm is also derived. Experimental results on benchmark circuits show that the proposed techniques can lead to an average of 76% antenna-violation reduction and 99% via-violation reduction. Di Wu 0017, Jiang Hu 0001, Rabi N. Mahapatra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Timing driven track routing considering coupling capacitanceabstractAs VLSI technology enters the ultra-deep submicron era, wire coupling capacitance starts to dominate self capacitance and can no longer be neglected in timing driven routing. In this paper, a coupling aware timing driven track routing heuristic is proposed. Given a global routing solution and timing constraint for each net, major trunks of wire segments are assigned to routing tracks such that the minimum timing slack among all nets is maximized. Delay penalties from both coupling capacitance and wire detour are considered in a unified graph model. The core problem is formulated and solved as a Sequential Ordering Problem (SOP). Routing blockages are handled in a post processing procedure. The experimental results on benchmark circuits show that the effect of coupling capacitance on timing is significant and the proposed heuristic results in greater improvement on coupling aware timing compared with other approaches. Di Wu 0017, Jiang Hu 0001, Min Zhao 0001, Rabi N. Mahapatra |
ASP-DAC | 4 |
| 2005 | TCAM enabled on-chip logic minimizationabstractThis paper presents an efficient hardware architecture of an on-chip logic minimization coprocessor. The proposed architecture employs TCAM cells to provide fastest and memory efficient implementation suitable for emerging on-chip minimization applications. The paper presents a detailed design of the on-chip minimizer and shows that it requires very little hardware resources to achieve acceptable quality of minimization. An incremental insertion and bulk deletion is achieved in 0.25 μs and 3.8 ms respectively and a compaction of 100000 entries in 25 ms using just 300 TCAM entries. Seraj Ahmad, Rabi N. Mahapatra |
DAC | 2 |
| 2005 | Lifetime Modeling of a Sensor NetworkabstractWe provide a mathematical analysis for the lifetime of a sensor network, when data-generation at individual sensor nodes is a random process. We show that the mathematical results for expected lifetime and its probability distribution closely validate the simulations results, both in linear and planar networks. Vivek Rai, Rabi N. Mahapatra |
DATE | 2 |
| 2005 | Integrated scheduling and buffer management input queued switches under extreme traffic scheme conditionsabstractThis paper addresses scheduling and memory management in input queued switches having finite buffer space to improve the performance in terms of throughput and average delay. Most of the prior works on scheduling related to input queued switches assume infinite buffer space. In practice, buffer space being a finite resource, special memory management scheme becomes essential. We introduce a buffer management scheme called iSMM (integrated scheduling and memory management) that can be employed jointly with any deterministic iterative scheduling algorithm. We applied iSMM over iSLIP, a popular scheduling algorithm, and examined its effect under extreme traffic conditions. Simulation results indicate iSMM to perform better than the raw iSLIP and maximum weighted matching (MWM) scheduling algorithms both in terms of throughput and delay. Rabi N. Mahapatra |
ICC | 2 |
| 2005 | DiCER: distributed and cost-effective redundancy for variation toleranceabstractIncreasingly prominent variational effects impose imminent threat to the progress of VLSI technology. This work explores redundancy, which is a well-known fault tolerance technique, for variation tolerance. It is observed that delay variability can be reduced by making redundant paths distributed or less correlated. Based on this observation, a gate splitting methodology is proposed for achieving distributed redundancy. We show how to avoid short circuit and estimate delay in dual-driver nets which are caused by gate splitting. A spin-off gate placement heuristic is developed to minimize redundancy cost. Monte Carlo simulation results on benchmark circuits show that our method can improve timing yield from 59% to 72% with only 03% increase on cell area and 2.2% increase on wirelength on average. Di Wu 0017, Ganesh Venkataraman, Jiang Hu 0001, Quiyang Li, Rabi N. Mahapatra |
ICCAD | 5 |
| 2005 | X-Routing using Two Manhattan Route InstancesabstractIn deep sub-micron (DSM) technologies, wire delays comprise a dominant fraction of the total delay of a design. As a consequence, routing techniques which reduce the total wire length of a design are highly relevant to such technologies. One such approach which holds promise is that of non-Manhattan routing (or X routing). In this paper, we describe a technique to perform non-Manhattan routing by combining the results of two related Manhattan routing instances. The first is a regular, unrotated routing instance. The second routing instance is derived from the first by rotating the coordinate system by 45/spl deg/. Both instances are routed on the same pair of metal layers. By selectively combining the results of the two instances, we obtain a final routing result that contains non-Manhattan wire segments. Our approach utilizes a powerful Floyd-Warshall based engine to combine the results of the two instances. We demonstrate that our router produces highly efficient results, reducing the total wire length by an average of about 20% (31%) over the unrotated (rotated) results, with a via-count decrease of between 4% (43%). Seraj Ahmad, Nikhil Jayakumar, Vijay Balasubramanian, Edward Hursey, Sunil P. Khatri, Rabi N. Mahapatra |
ICCD | 6 |
| 2005 | Coupling aware timing optimization and antenna avoidance in layer assignmentabstractThe sustained progress of VLSI technology has altered the landscape of routing which is a major physical design stage. For timing driven routings, traditional approaches which consider only wire self capacitance become inadequate since the wire delay is affected more by coupling capacitance in ultra-deep submicron designs. Furthermore, the technology scaling dramatically increases the likelihood of the antenna problem in manufacturing and requests corresponding considerations in the routing stage. In this paper, we propose techniques that can be applied to handle the coupling aware timing and the antenna problem simultaneously during layer assignment which is an important step between global routing and detailed routing. An improved probabilistic coupling capacitance model is suggested for coupling aware timing optimization without performing track assignment. The antenna avoidance problem is modeled as a tree partitioning problem with a linear time optimal algorithm solution. This algorithm is customized to guide antenna avoidance in layer assignment. A linear time optimal jumper insertion algorithm is also derived. Experimental results on benchmark circuits show that the proposed techniques can lead to an average of 270ps timing slack improvement validated by track assignment, 76% antenna violation reduction and 99% via violation reduction. Di Wu 0017, Jiang Hu 0001, Rabi N. Mahapatra |
ISPD | 3 |
| 2005 | An integrated scheduling and buffer management scheme for input queued switches with finite buffer space
Rabi N. Mahapatra |
Comput. Commun. | 2 |
| 2005 | EaseCAM: An Energy and Storage Efficient TCAM-Based Router Architecture for IP LookupabstractTernary content addressable memories (TCAMs) have been emerging as a popular device in designing routers for packet forwarding and classifications. Despite their premise on high-throughput, large TCAM arrays are prohibitive due to their excessive power consumption and lack of scalable design schemes. We present a TCAM-based router architecture that is energy and storage efficient. We introduce prefix aggregation and expansion techniques to compact the effective TCAM size in a router. Pipelined and paging schemes are employed in the architecture to activate a limited number of entries in the TCAM array during an IP lookup. The new architecture provides low power, fast incremental updating, and fast table look-up. Heuristic algorithms for page filling, fast prefix update, and memory management are also provided. Results have been illustrated with two large routers (bbnplanet and attcanada) to demonstrate the effectiveness of our approach. V. C. Ravikumar, Rabi N. Mahapatra, Laxmi N. Bhuyan |
IEEE Trans. Computers | 2 |
| 2005 | An Energy-Efficient Slack Distribution Technique for Multimode Distributed Real-Time Embedded SystemsabstractIn multimode distributed systems, active task sets are assigned to their distributed components for realizing one or more functions. Many of these systems encounter runtime task variations at the input and across the system while processing their tasks in real time. Very few efforts have been made to address energy efficient scheduling in these types of distributed systems. In this paper, we propose an analytical model for energy efficient scheduling in distributed real-time embedded systems to handle time-varying task inputs. A new slack distribution scheme is introduced and adopted during the schedule of the task sets in the system. The slack distribution is made according to the service demand at the nodes which affects the energy consumption in the system. The active component at a node periodically determines the service rate and applies voltage scaling according to the dynamic traffic condition observed at various network nodes. The proposed approach uses a comprehensive traffic description function at nodes and provides adequate information about the worst-case traffic behavior anywhere in the distributed network, thereby enhancing the system power management capabilities. We evaluate the proposed technique using several benchmarks employing an event driven simulator and demonstrate its performance for multimode applications. Experimental results indicate significant energy savings in various examples and case studies. Rabi N. Mahapatra, Wei Zhao 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Layer assignment for crosstalk risk minimization
Di Wu 0017, Jiang Hu 0001, Rabi N. Mahapatra, Min Zhao 0001 |
ASP-DAC | 3 |
| 2004 | Energy characterization of filesystems for diskless embedded systemsabstractThe need for low power, small form-factor, secondary storage devices in embedded systems has led to the widespread use of flash memory. Energy consumption due to processor and flash for such devices is critical to embedded system design. In this paper, we have proposed a quantitative account of energy consumption in both processor and flash due to overhead of filesystem related system calls. A macromodel for such energy consumption is derived using linear regression analysis. The results describing filesystem energy consumption have been obtained from Linux Kernel running Journaling Flash Filesystem 2 (JFFS2) and Extended 3 (Ext3) filesystems on StrongARM processor with flash as secondary storage device. Armed with such a macromodel, a designer can choose to partition filesystem, estimate the application energy consumption (processor and flash) due to filesystem during the early stage of system design. Siddharth Choudhuri, Rabi N. Mahapatra |
DAC | 2 |
| 2004 | Reducing clock skew variability via cross linksabstractIncreasingly significant variational effects present a great challenge for delivering desired clock skew reliably. Non-tree clock network has been recognized as a promising approach to overcome the variation problem. Existing non-tree clock routing methods are restricted to a few simple or regular structures, and often consume excessive amount of wire-length. In this paper, we suggest to construct a low cost non-tree clock network by inserting cross links in a given clock tree. The effects of the link insertion on clock skew variability are analyzed. Based on the analysis, we propose two link insertion schemes that can quickly convert a clock tree to a non-tree with significantly lower skew variability and very limited wirelength increase. In these schemes, the complicated non-tree delay computation is circumvented. Further, they can be applied to the recently popular non-zero skew routing easily. Experimental results on benchmark circuits show that this approach can achieve significant skew variability reduction with less than 2 increase of wirelength. Anand Rajaram, Jiang Hu 0001, Rabi N. Mahapatra |
DAC | 3 |
| 2004 | M-trie: an efficient approach to on-chip logic minimizationabstractBoolean logic minimization is being increasingly applied to new applications which demands very fast and frequent minimization services. These applications typically offer very limited computing and memory resources rendering the traditional logic minimizers ineffective. We present a new approximate logic minimization algorithm based on ternary trie. We compare its performance with Espresso-II and ROCM logic minimizers for routing table compaction and demonstrate that it is 100 to 1000 times faster and can run with a data memory as little as 16KB. It is also found that proposed approach can support up to 25000 incremental updates per seconds positioning itself as an ideal on-chip logic minimization algorithm. Seraj Ahmad, Rabi N. Mahapatra |
ICCAD | 2 |
| 2003 | Analytical Bound for Unwanted Clock Skew due to Wire Width Variation
Anand Rajaram, Rabi N. Mahapatra, Jiang Hu 0001 |
ICCAD | 4 |
| 2000 | Hierarchical Simulation of a Multiprocessor ArchitectureabstractWhen proposing new architectural enhancements, it is also important to account for the hardware complexity. To achieve this goal, we propose to model the new design in a hardware description language (HDL), synthesize the HDL code, and infer a realistic clock cycle which will be used in subsequent simulations. For accurate results, we develop a two-level hierarchical simulation technique, where an execution driven simulator (RSIM) and an HDL simulator (Verilog-XL) are coupled together to evaluate an entire system. We detail the simulation process and show its impact on the design of an interconnect switch architecture for CC-NUMA multiprocessors. Marius Pirvu, Laxmi N. Bhuyan, Rabi N. Mahapatra |
ICCD | 3 |
| 2000 | Mapping of Neural Network Models onto Systolic Arrays
Sudipta Mahapatra, Rabi N. Mahapatra |
J. Parallel Distributed Comput. | 2 |
| 1999 | Mapping of neural network models onto massively parallel hierarchical computer systems
Sudipta Mahapatra, Rabi N. Mahapatra, Biswanath N. Chatterji |
J. Syst. Archit. | 2 |
| 1997 | A Parallel Formulation of Back-Propagation Learning on Distributed Memory Multiprocessors
Sudipta Mahapatra, Rabi N. Mahapatra, Biswanath N. Chatterji |
Parallel Comput. | 2 |
| 1997 | Modelling Hadamard Haar transform algorithm for omega connected multiprocessors
Binoy Kumar Das, Rabi N. Mahapatra, Biswanath N. Chatterji |
Signal Process. | 2 |
| 1996 | Mapping of Neural Network Models Onto Two-Dimensional Processor Arrays
Rabi N. Mahapatra, Sudipta Mahapatra |
Parallel Comput. | 1 |
| 1994 | Parallel and Distributed Processing Research in Some Asian Countries
Richard P. Brent, Yong Kim Chong, Guo-Jie Li, Paul B. S. Lin, Rabi N. Mahapatra, Myong-Soon, Makoto Takizawa 0001 |
ICPADS | 5 |
| 1994 | Implementation of Fast Hartley TransformabstractThe use of multiple bus as interconnection network for multiprocessors has shown attractive features as compared to the existing ones. The addition of cache memory makes the architecture still a high performance one. In this paper we consider the implementation of Hou's FHT on multiple bus cache coherent multiprocessors. The analytical formulas are developed and performances are analysed in terms of speedup using these formulas. We also study the limitations of the inter processor communication overhead and propose a modification to the signal flow graph in order to minimise the multiprocessor execution time and hence to improve the speedup performance of the system. Rabi N. Mahapatra, Jharna Majumdar |
ICPADS | 1 |
| 1993 | Modelling a 2-D inverse fast cosine transform algorithm on a multistage network
Rabi N. Mahapatra, Sudipta Mahapatra |
Signal Process. | 1 |
| 1991 | Modelling a Fast Parallel Thinning Algorithm for Shared Memory SIMD Computers
Rabi N. Mahapatra, Harish Pareek |
Inf. Process. Lett. | 1 |
| 1990 | Performance of Parallel FFT Algorithm on Multiprocessors
Rabi N. Mahapatra, V. Ashok Kumar, Binoy Kumar Das, Biswanath N. Chatterji |
ICPP (3) | 1 |