EDBT 2026 Demo / reviewers in the wild / expert
Nagarajan Kandasamy
dblp:63/6537
· DBLP profile ↗
54ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-4224-6362ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 4 first-author · 14 since 2021Security and privacy · 9 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Computer networks · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HypoxSpike: Ternary Spiking Neural Network for Opioid Overdose DetectionabstractOpioid overdose is a growing global health crisis that claims more than 120,000 lives annually, of which more than half use opioids alone, without access to bystander intervention. Fatal overdose events are marked by motionlessness, respiratory depression, and hypoxemia, yet current wearable systems often rely on a single biomarker, limiting detection speed and accuracy. We present HypoxSpike, a novel ternary spiking neural network designed for real-time, multi-biomarker overdose detection for low-power neuromorphic hardware, optimized for integration into shoulder-based wearables. HypoxSpike combines motion, respiration, and oxygen saturation signals, while accounting for skin tone and body physiology, thus addressing known racial bias in pulse oximetry. Our research leverages an open-source shoulder-worn dataset from 19 patients experiencing sleep apnea, exploiting the shared physiological mechanisms underlying apnea and opioid overdose. This allows a direct comparison of our model with existing overdose detection approaches. HypoxSpike classifies three stages of hypoxemia with an average accuracy of 94%, outperforming state-of-the-art shoulder-based hypoxemia estimation while reducing false positive alert rates by 23.5%. By minimizing false positives, HypoxSpike supports accurate and power-efficient overdose detection, improving trust and usability for high-risk populations often overlooked by conventional systems. Anush Niranjan Lingamoorthy, Abhishek Kumar Mishra 0002, Olumuyiwa Oni, Jacob S. Brenner, Nagarajan Kandasamy, Amanda Watson |
AAAI | 5 |
| 2026 | A Reinforcement Learning Framework for Good Die in Bad Neighborhood AnalysisabstractGood-Die-in-Bad-Neighborhood (GDBN) analysis is a critical challenge in semiconductor manufacturing, where overly aggressive rejection reduces yield, while lenient acceptance in-creases test escapes and outgoing defective parts per million (DPPM). This asymmetric trade-off creates a multi-objective optimization problem spanning defect coverage, yield preservation, and return-material-authorization cost, often beyond the reach of conventional gradient-based methods. In this work, we employ reinforcement learning to develop an attention-based Deep Q-Network (DQN) framework tailored for GDBN-driven decision making. The DQN agent learns an optimal die-level screening policy from local wafer patches along with numerical test parametric data, optimizing actions that maximize cumulative long-term reward. By incorporating an attention mechanism, our model captures neighborhood-aware spatial dependencies across dies, enabling context-sensitive decision-making that balances yield and quality. We evaluated our method on the publicly available WM-811K wafer dataset, demonstrating substantial improvements in DPPM reduction and yield–cost tradeoffs compared to existing approaches. The results demonstrate that reinforcement learning provides a scalable and effective solution for adaptive defect screening in high-volume semiconductor test environments. Mohammad Ershad Shaik, Abhishek Kumar Mishra 0002, Nagarajan Kandasamy, Nur A. Touba |
DATE | 3 |
| 2025 | Hierarchical Model-Based Approach for Concurrent Testing of Neuromorphic ArchitectureabstractNeuromorphic architectures that implement spiking neural networks provide a biologically inspired and energy-efficient approach to processing information. These systems use spike trains, where the timing and frequency of spikes drive computation, offering unique advantages in dynamic and event-driven tasks. This paper develops a concurrent testing methodology for neuromorphic architectures, emphasizing Error Detection and Isolation (EDI) through a hierarchical model-based redundancy framework. Our approach uses a software-based monitoring system that compares the discrepancies between the observed and predicted behavior of hardware-mapped neurons at both the system and the neuron levels. We identify key statistical properties of spike trains that are critical for error detection and develop computationally efficient machine learning models to forecast these properties. By combining real-time observations with predictions of neuron behavior, our EDI methodology ensures robust fault detection and isolation. Experimental evaluations using an open source neuromorphic processor design executing benchmark datasets, MNIST, FashionMNIST, and SVHN, demonstrate the effectiveness. We observe high fault coverage with reduced computational overhead, making the EDI scheme suitable for real-time use in neuromorphic systems. Abhishek Kumar Mishra 0002, Anup Das 0001, Nagarajan Kandasamy |
DSN | 4 |
| 2025 | GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
Yilmaz Ege Gonul, Ceyhun Efe Kayan, Ilknur Mustafazade, Nagarajan Kandasamy, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | Clustering and Allocation of Spiking Neural Networks on Crossbar-Based Neuromorphic ArchitectureabstractNeuromorphic hardware, designed to mimic the neural structure of the human brain, offers an energy-efficient platform for implementing machine-learning models in the form of Spiking Neural Networks (SNNs). Achieving efficient SNN execution on this hardware requires careful consideration of various objectives, such as optimizing utilization of individual neuromorphic cores and minimizing inter-core communication. Unlike previous approaches that overlooked the architecture of the neuromorphic core when clustering the SNN into smaller networks, our approach uses architecture-aware algorithms to ensure that the resulting clusters can be effectively mapped to the core. We base our approach on a crossbar architecture for each neuromorphic core. We start with a basic architecture where neurons can only be mapped to the columns of the crossbar. Our technique partitions the SNN into clusters of neurons and synapses, ensuring that each cluster fits within the crossbar's confines, and when multiple clusters are allocated to a single crossbar, we maximize resource utilization by efficiently reusing crossbar resources. We then expand this technique to accommodate an enhanced architecture that allows neurons to be mapped not only to the crossbar's columns but also to its rows, with the aim of further optimizing utilization. To evaluate the performance of these techniques, assuming a multi-core neuromorphic architecture, we assess factors such as the number of crossbars used and the average crossbar utilization. Our evaluation includes both synthetically generated SNNs and spiking versions of well-known machine-learning models: LeNet, AlexNet, DenseNet, and ResNet. We also investigate how the structure of the SNN impacts solution quality and discuss approaches to improve it. Ilknur Mustafazade, Nagarajan Kandasamy, Anup Das 0001 |
CF | 2 |
| 2024 | Drug Overdose Vital-Signs Evaluator Using Machine Learning
Anush Niranjan Lingamoorthy, Abhishek Kumar Mishra 0002, David Gordon, Jacob S. Brenner, Nagarajan Kandasamy, Amanda Watson |
IJCAI | 6 |
| 2024 | Data Driven Learning of Aperiodic Nonlinear Dynamic Systems Using Spike Based Reservoirs-in-ReservoirabstractMimicking an aperiodic nonlinear dynamic system is challenging as it is difficult to represent it using closed-form equations. A feedback-driven spike-based recurrent spiking neural network is a powerful computational model that can mimic such dynamical systems. We propose reservoirs-in-reservoir (R-i-R), a novel architecture to mimic the frequent pattern changes in space and time of an aperiodic nonlinear dynamic system. Here, a large reservoir is built by connecting multiple small reservoirs to a common output. These small reservoirs are individually specialized to mimic a portion of the input dynamic. The internal recurrent connections of each reservoir and its readout are trained using a recursive least squares (RLS)-based full first-order and reduced control error (full-FORCE) algorithm. To make the entire R-i-R architecture adaptable to the change in periodicity of an input, we implement a new cost function that incorporates a unique forgetting factor to control the fading and wind-up of the covariance matrix of each reservoir during training. We evaluate R-i-R using seven aperiodic nonlinear dynamic systems. We show that R-i-R with rate encoding reduces the error rate by an average 59% with 1.8X reduction in network size compared to state-of-the-art. To improve energy efficiency, we implement a time-to-first-spike encoding and show an average reduction of 39. 5% in the number of spikes. Ankita Paul, Nagarajan Kandasamy, Kapil R. Dandekar, Anup Das 0001 |
IJCNN | 2 |
| 2024 | Wafer2Spike: Spiking Neural Network for Wafer Map Pattern ClassificationabstractIn integrated circuit design, the analysis of wafer map patterns is critical to improve yield and detect manufacturing issues. We develop Wafer2Spike, an architecture for wafer map pattern classification using a spiking neural network (SNN), and demonstrate that a well-trained SNN achieves superior performance compared to deep neural network-based solutions. Wafer2Spike achieves an average classification accuracy of 98% on the WM-811k wafer benchmark dataset. It is also superior to existing approaches for classifying defect patterns that are underrepresented in the original dataset. Wafer2Spike achieves this improved precision with great computational efficiency. Abhishek Kumar Mishra 0002, Anush Niranjan Lingamoorthy, Anup Das 0001, Nagarajan Kandasamy |
ITC | 5 |
| 2024 | Model-Based Approach Towards Correctness Checking of Neuromorphic Computing SystemsabstractNeuromorphic hardware that emulates the neural structure of the human brain can implement machine learning models in an extremely energy-efficient manner. It is especially suitable for executing spiking neural networks (SNNs) which comprise spiking neurons interconnected via synapses. The underlying computation is based on spike trains in which the location and frequency of spikes that occur within the network guide the execution. This paper develops a fault detection and isolation (FDI) methodology to monitor the correctness of a neuromorphic program’s execution using model-based redundancy in which a software-based monitor compares discrepancies between the behavior of neurons mapped to hardware and that predicted by a corresponding mathematical model. We identify properties of spike trains generated by neurons that can be used for fault detection and build machine learning models to forecast these properties. Predictions from these models, which describe the nominal behavior of neurons, when combined with real-time observations, form the basis for FDI. Experiments using CARLSim, a high-fidelity SNN simulator, show that the proposed approach achieves high fault coverage using models that can operate with low computational overhead in real time. Abhishek Kumar Mishra 0002, Anup Das 0001, Nagarajan Kandasamy |
PRDC | 3 |
| 2024 | Efficient Built-In Self-Test Strategy for Neuromorphic Hardware Based On Alarm PlacementabstractNeuromorphic hardware that mimics the neural structure of the brain can efficiently run spiking neural networks (SNNs) to perform various machine learning tasks. Built-in self-test capability (BIST) is critical to ensure the reliability and functionality of these systems when deployed in mission-critical applications such as autonomous vehicles. We introduce an online BIST strategy for neuromorphic hardware that aims to maximize fault coverage while reducing the testing time needed to detect and isolate faulty components. Our key contribution is the adaptation of the alarm placement problem for this purpose. The topology of a typical SNN allows for multiple fault propagation paths through the network, and we take advantage of this property to test (or place alarms on) only the minimum number of neurons necessary to isolate any faulty neuron in the SNN. The alarms themselves detect erroneous behavior by measuring the dissimilarity between the observed and expected spike trains generated by a neuron based on the applied test pattern. The efficacy of the BIST approach is evaluated using multiple SNN models in single- and multiple-fault scenarios. We also evaluate different spike dissimilarity metrics in terms of their fault-detection effectiveness. Ilknur Mustafazade, Anup Das 0001, Nagarajan Kandasamy |
PRDC | 3 |
| 2024 | WaferCap: Open Classification of Wafer Map Patterns using Deep Capsule NetworkabstractIn integrated circuit design, analysis of wafer map patterns is critical to enhance yield and detect manufacturing issues. With the emergence of novel wafer map patterns, there is increasing need for robust artificial intelligence models that can both accurately classify seen patterns and while also detecting ones not seen during training, a capability known as open world classification. We develop a novel solution to this problem: WaferCap, a Deep Capsule Network designed for wafer map pattern classification and equipped with a rejection mechanism. When evaluated using the WM-811k dataset, WaferCap significantly surpasses existing methods, achieving 99% accuracy for fully seen patterns while demonstrating robust performance in open-world settings by effectively detecting unseen wafer map patterns. Abhishek Kumar Mishra 0002, Mohammad Ershad Shaik, Anush Niranjan Lingamoorthy, Anup Das 0001, Nagarajan Kandasamy, Nur A. Touba |
VTS | 6 |
| 2023 | Hardware-Software Co-Design for On-Chip Learning in AI SystemsabstractSpike-based convolutional neural networks (CNNs) are empowered with on-chip learning in their convolution layers, enabling the layer to learn to detect features by combining those extracted in the previous layer. We propose ECHELON, a generalized design template for a tile-based neuromorphic hardware with on-chip learning capabilities. Each tile in ECHELON consists of a neural processing units (NPU) to implement convolution and dense layers of a CNN model, an on-chip learning unit (OLU) to facilitate spike-timing dependent plasticity (STDP) in the convolution layer, and a special function unit (SFU) to implement other CNN functions such as pooling, concatenation, and residual computation. These tile resources are interconnected using a shared bus, which is segmented and configured via the software to facilitate parallel communication inside the tile. Tiles are themselves interconnected using a classical Network-on-Chip (NoC) interconnect. We propose a system software to map CNN models to ECHELON, maximizing the performance. We integrate the hardware design and software optimization within a co-design loop to obtain the hardware and software architectures for a target CNN, satisfying both performance and resource constraints. In this preliminary work, we show the implementation of a tile on a FPGA and some early evaluations. Using 8 STDP-enabled CNN models, we show the potential of our co-design methodology to optimize hardware resources. M. Lakshmi Varshika, Abhishek Kumar Mishra 0002, Nagarajan Kandasamy, Anup Das 0001 |
ASP-DAC | 3 |
| 2023 | Online Performance Monitoring of Neuromorphic Computing SystemsabstractNeuromorphic computation is based on spike trains in which the location and frequency of spikes occurring within the network guide the execution. This paper develops a frame-work to monitor the correctness of a neuromorphic program’s execution using model-based redundancy in which a software-based monitor compares discrepancies between the behavior of neurons mapped to hardware and that predicted by a corresponding mathematical model in real time. Our approach reduces the hardware overhead needed to support the monitoring infrastructure and minimizes intrusion on the executing application. Fault-injection experiments utilizing CARLSim, a high-fidelity SNN simulator, show that the framework achieves high fault coverage using parsimonious models which can operate with low computational overhead in real time. Abhishek Kumar Mishra 0002, Anup Das 0001, Nagarajan Kandasamy |
ETS | 3 |
| 2022 | Design-Technology Co-Optimization for NVM-Based Neuromorphic Processing ElementsabstractAn emerging use case of machine learning (ML) is to train a model on a high-performance system and deploy the trained model on energy-constrained embedded systems. Neuromorphic hardware platforms, which operate on principles of the biological brain, can significantly lower the energy overhead of an ML inference task, making these platforms an attractive solution for embedded ML systems. We present a design-technology tradeoff analysis to implement such inference tasks on the processing elements (PEs) of a non-volatile memory (NVM)-based neuromorphic hardware. Through detailed circuit-level simulations at scaled process technology nodes, we show the negative impact of technology scaling on the information-processing latency, which impacts the quality of service of an embedded ML system. At a finer granularity, the latency inside a PE depends on (1) the delay introduced by parasitic components on its current paths, and (2) the varying delay to sense different resistance states of its NVM cells. Based on these two observations, we make the following three contributions. First, on the technology front, we propose an optimization scheme where the NVM resistance state that takes the longest time to sense is set on current paths having the least delay, and vice versa, reducing the average PE latency, which improves the quality of service. Second, on the architecture front, we introduce isolation transistors within each PE to partition it into regions that can be individually power-gated, reducing both latency and energy. Finally, on the system-software front, we propose a mechanism to leverage the proposed technological and architectural enhancements when implementing an ML inference task on neuromorphic PEs of the hardware. Evaluations with a recent neuromorphic hardware architecture show that our proposed design-technology co-optimization approach improves both performance and energy efficiency of ML inference tasks without incurring high cost-per-bit. Shihao Song, Adarsha Balaji, Anup Das 0001, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic HardwareabstractSpiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches. Shihao Song, Harry Chong, Adarsha Balaji, Anup Das 0001, James A. Shackleford, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2022 | Endurance-Aware Mapping of Spiking Neural Networks to Neuromorphic HardwareabstractNeuromorphic computing systems are embracing memristors to implement high density and low power synaptic storage as crossbar arrays in hardware. These systems are energy efficient in executing Spiking Neural Networks (SNNs). We observe that long bitlines and wordlines in a memristive crossbar are a major source of parasitic voltage drops, which create current asymmetry. Through circuit simulations, we show the significant endurance variation that results from this asymmetry. Therefore, if the critical memristors (ones with lower endurance) are overutilized, they may lead to a reduction of the crossbar's lifetime. We propose eSpine, a novel technique to improve lifetime by incorporating the endurance variation within each crossbar in mapping machine learning workloads, ensuring that synapses with higher activation are always implemented on memristors with higher endurance, and vice versa. eSpine works in two steps. First, it uses the Kernighan-Lin Graph Partitioning algorithm to partition a workload into clusters of neurons and synapses, where each cluster can fit in a crossbar. Second, it uses an instance of Particle Swarm Optimization (PSO) to map clusters to tiles, where the placement of synapses of a cluster to memristors of a crossbar is performed by analyzing their activation within the workload. We evaluate eSpine for a state-of-the-art neuromorphic hardware model with phase-change memory (PCM)-based memristors. Using 10 SNN workloads, we demonstrate a significant improvement in the effective lifetime. Twisha Titirsha, Shihao Song, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Aging-Aware Request Scheduling for Non-Volatile Main MemoryabstractModern computing systems are embracing non-volatile memory (NVM) to implement high-capacity and low-cost main memory. Elevated operating voltages of NVM accelerate the aging of CMOS transistors in the peripheral circuitry of each memory bank. Aggressive device scaling increases power density and temperature, which further accelerates aging, challenging the reliable operation of NVM-based main memory. We propose HEBE, an architectural technique to mitigate the circuit aging-related problems of NVM-based main memory. HEBE is built on three contributions. First, we propose a new analytical model that can dynamically track the aging in the peripheral circuitry of each memory bank based on the bank's utilization. Second, we develop an intelligent memory request scheduler that exploits this aging model at run time to de-stress the peripheral circuitry of a memory bank only when its aging exceeds a critical threshold. Third, we introduce an isolation transistor to decouple parts of a peripheral circuit operating at different voltages, allowing the decoupled logic blocks to undergo long-latency de-stress operations independently and off the critical path of memory read and write accesses, improving performance. We evaluate HEBE with workloads from the SPEC CPU2017 Benchmark suite. Our results show that HEBE significantly improves both performance and lifetime of NVM-based main memory. Shihao Song, Anup Das 0001, Onur Mutlu, Nagarajan Kandasamy |
ASP-DAC | 4 |
| 2021 | A Design Flow for Mapping Spiking Neural Networks to Many-Core Neuromorphic HardwareabstractThe design of many-core neuromorphic hardware is becoming increasingly complex as these systems are now expected to execute large machine-learning models. A predictable design flow is needed to guarantee real-time performance such as latency and throughput without significantly increasing the buffer requirement of computing cores. Synchronous Data Flow Graphs (SDFGs) have been previously used for predictable mapping of streaming applications to multiprocessor systems. We propose an SDFG-based design flow to map spiking neural networks (SNNs) to many-core neuromorphic hardware with the objective of exploring the tradeoff between throughput and buffer-size requirements. The proposed design flow integrates an iterative partitioning approach based on Kernighan-Lin graph partitioning heuristic to create SNN clusters such that each cluster can be mapped to a core of the hardware. The partitioning approach minimizes inter-cluster spike communication, which improves latency on the shared interconnect of the hardware. Next, the design flow maps clusters to cores using Particle Swarm Optimization (PSO), an evolutionary algorithm, while exploring the design space of throughput and buffer size. Pareto-optimal mappings are retained from the design flow, allowing system designers to select a Pareto mapping that satisfies throughput and buffer-size requirements of the design. We evaluated the developed design flow using five large-scale convolutional neural network (CNN) models. Results demonstrate 63% higher maximum throughput and 10% lower buffer-size requirement compared to state-of-the-art dataflow-based mapping solutions. Shihao Song, M. Lakshmi Varshika, Anup Das 0001, Nagarajan Kandasamy |
ICCAD | 4 |
| 2021 | Dynamic Reliability Management in Neuromorphic ComputingabstractNeuromorphic computing systems execute machine learning tasks designed with spiking neural networks. These systems are embracing non-volatile memory to implement high-density and low-energy synaptic storage. Elevated voltages and currents needed to operate non-volatile memories cause aging of CMOS-based transistors in each neuron and synapse circuit in the hardware, drifting the transistor’s parameters from their nominal values. If these circuits are used continuously for too long, the parameter drifts cannot be reversed, resulting in permanent degradation of circuit performance over time, eventually leading to hardware faults. Aggressive device scaling increases power density and temperature, which further accelerates the aging, challenging the reliable operation of neuromorphic systems. Existing reliability-oriented techniques periodically de-stress all neuron and synapse circuits in the hardware at fixed intervals, assuming worst-case operating conditions, without actually tracking their aging at run-time. To de-stress these circuits, normal operation must be interrupted, which introduces latency in spike generation and propagation, impacting the inter-spike interval and hence, performance (e.g., accuracy). We observe that in contrast to long-term aging, which permanently damages the hardware, short-term aging in scaled CMOS transistors is mostly due to bias temperature instability. The latter is heavily workload-dependent and, more importantly, partially reversible. We propose a new architectural technique to mitigate the aging-related reliability problems in neuromorphic systems by designing an intelligent run-time manager (NCRTM), which dynamically de-stresses neuron and synapse circuits in response to the short-term aging in their CMOS transistors during the execution of machine learning workloads, with the objective of meeting a reliability target. NCRTM de-stresses these circuits only when it is absolutely necessary to do so, otherwise reducing the performance impact by scheduling de-stress operations off the critical path. We evaluate NCRTM with state-of-the-art machine learning workloads on a neuromorphic hardware. Our results demonstrate that NCRTM significantly improves the reliability of neuromorphic hardware, with marginal impact on performance. Shihao Song, Jui Hanamshet, Adarsha Balaji, Anup Das 0001, Jeffrey L. Krichmar, Nikil Dutt, Nagarajan Kandasamy, Francky Catthoor |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2020 | Modeling SAT-Attack Search ComplexityabstractIn this paper, a metric based on mathematical modeling is proposed to evaluate the strength in security of a logic-locked circuit against a satisfiability (SAT) based attack. Current approaches estimate the SAT resilience experimentally based on time-to-solve or the number of calls to a SAT-solver. However, the estimate is often based on one sample or a small sample size. Due to the possible variation in the search path length of the SAT-attack, a measure of resilience based on statistical characterization is proposed. A probabilistic model of a SAT-attack search process is developed to properly capture the variation in the path length and report the SAT resilience as an expectation of the computational complexity. An estimator of the expected complexity, assuming an equally likely branching probability, is proposed. The model and the estimator allow for 1) the derivation of a closed-form estimate of the expected security, and 2) characterization of the key search space without experimental bias toward SAT-attack implementation or circuit topology. As a case study, an analysis of the security gain per inserted key gate is performed on a full adder circuit. The study reveals a monotonically increasing resilience and provides insights on the most efficient key gate placement strategy that maximizes the achievable security. Saran Phatharodom, Nagarajan Kandasamy, Ioannis Savidis |
ISCAS | 2 |
| 2020 | Exploiting inter- and intra-memory asymmetries for data mapping in hybrid tiered-memoriesabstractModern computing systems are embracing hybrid memory comprising of DRAM and non-volatile memory (NVM) to combine the best properties of both memory technologies, achieving low latency, high reliability, and high density. A prominent characteristic of DRAM-NVM hybrid memory is that it has NVM access latency much higher than DRAM access latency. We call this inter-memory asymmetry. We observe that parasitic components on a long bitline are a major source of high latency in both DRAM and NVM, and a significant factor contributing to high-voltage operations in NVM, which impact their reliability. We propose an architectural change, where each long bitline in DRAM and NVM is split into two segments by an isolation transistor. One segment can be accessed with lower latency and operating voltage than the other. By introducing tiers, we enable non-uniform accesses within each memory type (which we call intra-memory asymmetry), leading to performance and reliability trade-offs in DRAM-NVM hybrid memory. Shihao Song, Anup Das 0001, Nagarajan Kandasamy |
ISMM | 3 |
| 2020 | Improving phase change memory performance with data content aware accessabstractPhase change memory (PCM) is a scalable non-volatile memory technology that has low access latency (like DRAM) and high capacity (like Flash). Writing to PCM incurs significantly higher latency and energy penalties compared to reading its content. A prominent characteristic of PCM’s write operation is that its latency and energy are sensitive to the data to be written as well as the content that is overwritten. We observe that overwriting unknown memory content can incur significantly higher latency and energy compared to overwriting known all-zeros or all-ones content. This is because all-zeros or all-ones content is overwritten by programming the PCM cells only in one direction, i.e., using either SET or RESET operations, not both. Shihao Song, Anup Das 0001, Onur Mutlu, Nagarajan Kandasamy |
ISMM | 4 |
| 2020 | Compiling Spiking Neural Networks to Neuromorphic HardwareabstractMachine learning applications that are implemented with spike-based computation model, e.g., Spiking Neural Network (SNN), have a great potential to lower the energy consumption when executed on a neuromorphic hardware. How- ever, compiling and mapping an SNN to the hardware is challenging, especially when compute and storage resources of the hardware (viz. crossbars) need to be shared among the neurons and synapses of the SNN. We propose an approach to analyze and compile SNNs on resource-constrained neuromorphic hardware, providing guarantees on key performance metrics such as execution time and throughput. Our approach makes the following three key contributions. First, we propose a greedy technique to partition an SNN into clusters of neurons and synapses such that each cluster can fit on to the resources of a crossbar. Second, we exploit the rich semantics and expressiveness of Synchronous Dataflow Graphs (SDFGs) to represent a clustered SNN and analyze its performance using Max-Plus Algebra, considering the available compute and storage capacities, buffer sizes, and communication bandwidth. Third, we propose a self-timed execution-based fast technique to compile and admit SNN-based applications to a neuromorphic hardware at run-time, adapting dynamically to the available resources on the hard- ware. We evaluate our approach with standard SNN-based applications and demonstrate a significant performance improvement compared to current practices. Shihao Song, Adarsha Balaji, Anup Das 0001, Nagarajan Kandasamy, James A. Shackleford |
LCTES | 4 |
| 2019 | Data Reduction, Compression, and Recovery for Online Performance MonitoringabstractThe volume of data needed for effective monitoring of datacenters poses significant challenges in its collection, transmission, analysis, and storage. Considering a setting wherein data collected locally at a server is sent to a monitoring station for analysis, this paper develops computationally efficient methods for systematic reduction of this data during the transfer and its subsequent recovery at the monitoring station. Specifically, we develop a low-cost method of obtaining a sparse representation of the data collected at each individual server while preserving a specified fidelity with respect to the original signal. The sparsified representation obtained from the data-collection step is amenable to further compression prior to transmission to the monitoring station. Upon receipt of the compressed-data stream at the monitoring station, a method of sparse-signal recovery is utilized to reconstruct the original full-length signal for further analysis. The techniques are validated using workload traces collected from one of Google's production clusters. Experiments show that the achieved data reduction, which is a function of the specified fidelity, is significant: to reconstruct the signal with a fidelity between 90%-95%, the sample size that must be be transferred to the monitoring station is under 10% of the original. We also verify that the recovered signal tracks the target minimum fidelity requirements specified by the operator with high precision. Salvador DeCelles, Matthew C. Stamm, Nagarajan Kandasamy |
CLOUD | 3 |
| 2019 | An Efficient Strategy for Online Performance Monitoring of Datacenters via Adaptive SamplingabstractPerformance monitoring of datacenters provides vital information for dynamic resource provisioning, anomaly detection, and capacity planning decisions. Online monitoring, however, incurs a variety of costs: the very act of monitoring a system interferes with its performance, consuming network bandwidth and disk space. With the goal of reducing these costs, this paper develops and validates a strategy based on adaptive-rate compressive sampling. It exploits the fact that the signals of interest often can be sparsified under an appropriate representation basis and that the sampling rate can be tuned as a function of sparsity. We use the Trade6 application as our experimental platform and measure the signals of interest-in our case, signals pertaining to memory and disk I/O activity-using adaptive sampling. We then evaluate whether the reconstructed signals can be used for trend detection to track the gradual deterioration of system performance associated with software aging. Our experiments show that the signals recovered by our methods can be used to detect, with high confidence, the existence of trends within the original signal. We also evaluate the reconstructed signals for threshold-violation detection wherein the magnitude of the signal exceeds a preset value. Our experiments show that performance bottlenecks and anomalies that manifest themselves in portions of the signal where its magnitude exceeds a threshold value can also be detected using the reconstructed signals. Most importantly, detection of these anomalies is achieved using a substantially reduced sample size-a reduction of more than 70 percent when compared to the standard fixed-rate sampling method. Tingshan Huang, Nagarajan Kandasamy, Harish Sethu, Matthew C. Stamm |
IEEE Trans. Cloud Comput. | 2 |
| 2019 | Enabling and Exploiting Partition-Level Parallelism (PALP) in Phase Change MemoriesabstractPhase-change memory (PCM) devices have multiple banks to serve memory requests in parallel . Unfortunately, if two requests go to the same bank , they have to be served one after another , leading to lower system performance . We observe that a modern PCM bank is implemented as a collection of partitions that operate mostly independently while sharing a few global peripheral structures, which include the sense amplifiers (to read) and the write drivers (to write). Based on this observation, we propose PALP , a new mechanism that enables partition-level parallelism within each PCM bank, and exploits such parallelism by using the memory controller’s access scheduling decisions. PALP consists of three new contributions. First , we introduce new PCM commands to enable parallelism in a bank’s partitions in order to resolve the read-write bank conflicts, with no changes needed to PCM logic or its interface. Second , we propose simple circuit modifications that introduce a new operating mode for the write drivers, in addition to their default mode of serving write requests. When configured in this new mode, the write drivers can resolve the read-read bank conflicts, working jointly with the sense amplifiers. Finally , we propose a new access scheduling mechanism in PCM that improves performance by prioritizing those requests that exploit partition-level parallelism over other requests, including the long outstanding ones. While doing so, the memory controller also guarantees starvation-freedom and the PCM’s running-average-power-limit (RAPL). We evaluate PALP with workloads from the MiBench and SPEC CPU2017 Benchmark suites. Our results show that PALP reduces average PCM access latency by 23%, and improves average system performance by 28% compared to the state-of-the-art approaches. Shihao Song, Anup Das 0001, Onur Mutlu, Nagarajan Kandasamy |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | Reinforcement learning system to mitigate small-cell interference through directionalityabstractBeam-steering techniques using directional antennas are expected to play an important role in wireless network capacity expansion through ubiquitous small-cell deployment. However, integrating directional antennas into the existing wireless PHY and MAC stack of small cells has been challenging due to the added protocol overhead and lack of a robust antenna beam selection technique that can adapt well to environmental changes. This paper presents the design, implementation, and evaluation of LinkPursuit, a novel learning protocol for distributed antenna state selection in directional small-cell networks. LinkPursuit relies on reconfigurable antennas and a synchronous TimeDivision Multiple Access (TDMA) MAC to achieve simultaneous directional transmission and reception. Further, the system employs a practical antenna selection protocol based on the well known adaptive pursuit algorithm from the reinforcement learning literature. We implement a realtime prototype of LinkPursuit on the WARP platform and conduct extensive experiments to evaluate its performance. The empirical results show that appropriate use of directionality in LinkPursuit can result in higher network sum rates than omnidirectional transmission under various degrees of cross-link interference. Anton Paatelma, Danh H. Nguyen, Harri Saarnisaari, Nagarajan Kandasamy, Kapil R. Dandekar |
PIMRC | 4 |
| 2016 | Detecting Incipient Faults in Software Systems: A Compressed Sampling-Based ApproachabstractThe volume of data to be collected and processed for effective real-time monitoring of large-scale computing systems and networks poses significant Big Data challenges, and a scalable solution requires a systematic approach to dimensionality reduction during the data collection, transmission, and analysis phases. Compressive sampling can reduce the dimensionality of the data collected at the source prior to transmission to the monitoring station. Exploiting the fact that the compressed samples preserve in approximate form, the correlation information between data points in the original full-length signal, we develop a low-cost anomaly detection technique based on principal component analysis (PCA) aimed at incipient faults such as software aging—the key idea being PCA is performed directly on the compressed samples without having to reconstruct the original signal. Using case studies involving long-running enterprise benchmark applications, Trade6 and RuBBoS, with injected memory leaks, we show that the performance of the PCA-based detector when using just the compressed data is almost equivalent to the case in which the raw data is completely available, but achieved using significantly fewer samples with a compression rate exceeding 75%. Salvador DeCelles, Tingshan Huang, Matthew C. Stamm, Nagarajan Kandasamy |
CLOUD | 4 |
| 2016 | WiART - visualize and interact with wireless networks using augmented reality: demoabstractWith the increasing programmability and fast-paced dynamics of modern wireless systems, it has become more difficult to gain timely insights into wireless network operations. In this demonstration1 we present WiART, an augmented reality framework to help visualize and interact with wireless network activities in real time. WiART collects real-time radio and network statistics from participating network devices and depicts them on users' mobile devices in an intuitive way, leveraging the virtual information overlay of augmented reality. Specifically in our current implementation, WiART takes inputs from a cognitive radio link controlling beam-steerable reconfigurable antennas and annotates on a live mobile screen the active pre-measured radiation patterns. In the reverse flow, WiART enables users to select desired antenna radiation patterns directly in the mobile app and observe their effects on link performance in real time. These capabilities add an unprecedented level of instant visualization and interaction with wireless activities and provides valuable insights into the dynamics of a reconfigurable antenna-based cognitive radio network. Danh H. Nguyen, James Chacko, Logan Henderson, Anton Paatelma, Harri Saarnisaari, Nagarajan Kandasamy, Kapil R. Dandekar |
MobiCom | 6 |
| 2016 | SASO 2014: Selected, Revised, and Extended Best PapersabstractThe international conference IEEE SASO (Self-Adapting and Self-Organizing Systems) is the main forum for studying and discussing the foundations of a principled approach to engineering systems, networks, and services based on self-adaptation and self-organization. Over the past decade, it has consolidated as the primary scientific conference for sharing ideas on algorithms, technologies, tools, and applications across a wide range of scientific fields. In 2014, the conference was hosted by Imperial College in London, United Kingdom; its scientific program comprised full papers, short papers, poster presentations, demo sessions, workshops, and tutorials. This special issue of ACM TAAS champions some of the most solid research results of SASO 2014, presenting selected, revised, and extended best articles. Mirko Viroli, Ada Diaconescu, Nagarajan Kandasamy |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2016 | A New Approach to Dimensionality Reduction for Anomaly Detection in Data TrafficabstractThe monitoring and management of high-volume feature-rich traffic in large networks offers significant challenges in storage, transmission, and computational costs. The predominant approach to reducing these costs is based on performing a linear mapping of the data to a low-dimensional subspace such that a certain large percentage of the variance in the data is preserved in the low-dimensional representation. This variance-based subspace approach to dimensionality reduction forces a fixed choice of the number of dimensions, is not responsive to real-time shifts in observed traffic patterns, and is vulnerable to normal traffic spoofing. Based on theoretical insights proved in this paper, we propose a new distance-based approach to dimensionality reduction motivated by the fact that the real-time structural differences between the covariance matrices of the observed and the normal traffic is more relevant to anomaly detection than the structure of the training data alone. Our approach, called the distance-based subspace method, allows a different number of reduced dimensions in different time windows and arrives at only the number of dimensions necessary for effective anomaly detection. We present centralized and distributed versions of our algorithm and, using simulation on real traffic traces, demonstrate the qualitative and quantitative advantages of the distance-based subspace approach. Tingshan Huang, Harish Sethu, Nagarajan Kandasamy |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2015 | A fast algorithm for detecting anomalous changes in network trafficabstractAnomalies in communication network traffic caused by malware or denial-of-service attacks manifest themselves in structural changes in the covariance matrix of traffic features. Real-time detection of anomalies in high-dimensional data demands a very efficient algorithm to identify these changes in a compact low-dimensional representation. This paper presents an efficient algorithm for the rapid detection of structural differences between two covariance matrices, as measured by the maximum possible angle between the subspaces specified by subsets of the two sets of principal components of the matrices. We show that our algorithm achieves a significantly lower computational complexity compared to a naive approach. Finally, we apply our results to real traffic traces from Internet backbone links and show that our approach offers a substantial reduction in the computational overhead of anomaly detection. Tingshan Huang, Harish Sethu, Nagarajan Kandasamy |
CNSM | 3 |
| 2015 | Rapid Prototyping of Wireless Physical Layer Modules Using Flexible Software/Hardware Design FlowabstractThis paper describes a step by step approach in designing wireless physical layer modules starting from a software implementation in MATLAB to a hardware implementation using Xilinx SysGen and ModelSim. The described design flow promotes baseband physical layer research by providing high flexibility and speed to the process of module creation verification and deployment. The novelty introduced into our system lies within the flexible components created using this design flow, which enables on-the-fly modification of multiple parameters to suit various wireless protocols. James Chacko, Cem Sahin, Doug Pfeil, Nagarajan Kandasamy, Kapil R. Dandekar |
FPGA | 4 |
| 2015 | FPGA Implementation of Trained Coarse Carrier Frequency Offset Estimation and Correction for OFDM Signals (Abstract Only)abstractThis paper develops an FPGA implementation of a trained coarse Carrier Frequency Offset estimation and correction scheme using MATLAB System Generator. The designed system is capable of supporting variable FFT sizes for Orthogonal Frequency Division Multiplexing signals and different pilot symbol structures making it compatible with a large number of wireless communication standards, unlike other work that is protocol specific. This design stands out from its more common implementations as it requires only one pilot symbol to be considered for synchronization by using a data-aided modified correlation scheme, allowing for an increase in throughput. The Bit Error Rate of the corrected signal received over an Additive White Gaussian Noise channel is compared to the case without correction. This scheme demonstrated increased performance throughput since only a single pilot symbol was used. Marko Jacovic, James Chacko, Doug Pfeil, Nagarajan Kandasamy, Kapil R. Dandekar |
FPGA | 4 |
| 2015 | Anomaly detection in computer systems using compressed measurementsabstractOnline performance monitoring of computer systems incurs a variety of costs: the very act of monitoring a system interferes with its performance and if the information is transmitted to a monitoring station for analysis and logging, this consumes network bandwidth and disk space. Compressive sampling-based schemes can help reduce these costs on the local machine by acquiring data directly from the system in a compressed form, and in a computationally efficient way. This paper focuses on reducing the computational cost associated with recovering the original signal from the transmitted sample set at the monitoring station for anomaly detection. Towards this end, we show that the compressed samples preserve, in an approximate form, properties such as mean, variance, as well as correlation between data points in the original full-length signal. We then use this result to detect changes in the original signal that could be indicative of an underlying anomaly such as abrupt changes in magnitude and gradual trends without the need to recover the full-length data. We illustrate the usefulness of our approach via case studies involving IBM's Trade Performance Benchmark using signals from the disk and memory subsystems. Experiments indicate that abrupt changes can be detected using a compressed sample size of 25% with a hit rate of 95% for a fixed false alarm rate of 5%; trends can be detected within a confidence interval of 95% using a sample size of only 6%. Tingshan Huang, Nagarajan Kandasamy, Harish Sethu |
ISSRE | 2 |
| 2015 | Leveraging an Agile RF Transceiver for Rapid Prototyping of Small-Cell SystemsabstractThis paper describes a new software-defined radio (SDR) platform targeted for rapid prototyping of small-cell systems. The SDR hardware combines the signal processing power of Xilinx ML605 Virtex-6 FPGA board with the Nutaq Radio420X frequency-agile transceiver and reconfigurable antennas to form a highly versatile platform for spectrum sensing, spectrum access, and cooperative communications. We evaluate the platform with two example applications: an offline OFDM physical processing flow based on WARPLab, and a real-time online automatic gain control mechanism. The results show that our SDR platform can reliably handle both offline and online processing demands with the added benefit of frequency agility offered by a state-of-the-art radio transceiver. Danh H. Nguyen, Mikko Rauhanummi, Harri Saarnisaari, Nagarajan Kandasamy, Kapil R. Dandekar |
VTC Fall | 4 |
| 2014 | A modular multi-location anonymized traffic monitoring tool for a WiFi networkabstractNetwork traffic anomaly detection is now considered a surer approach to early detection of malware than signature-based approaches and is best accomplished with traffic data collected from multiple locations. Existing open-source tools are primarily signature-based, or do not facilitate integration of traffic data from multiple locations for real-time analysis, or are insufficiently modular for incorporation of newly proposed approaches to anomaly detection. In this paper, we describe DataMap, a new modular open-source tool for the collection and real-time analysis of sampled, anonymized, and filtered traffic data from multiple WiFi locations in a network and an example of its use in anomaly detection. Justin Hummel, Vatsal Shah, Riju Singh, Bradford D. Boyle, Tingshan Huang, Nagarajan Kandasamy, Harish Sethu, Steven Weber 0001 |
CODASPY | 7 |
| 2013 | Datacenters as Controllable Load Resources in the Electricity MarketabstractDatacenters, being major consumers of power, can play an important role in the efficient operation of electrical grids. This paper develops an optimization framework to allow datacenters to operate as controllable load resources within the demand dispatch regime, a demand response (DR) program in which incentives are designed to induce lower electricity use not just during times of high prices but also when the reliability of the local grid is jeopardized or when the electricity supply and demand are unbalanced. Assuming the availability of geographically distributed and virtualized datacenters situated in multiple regional electrical markets, the basic idea is to migrate the workload in the form of virtual machines (VMs) between these centers to maximize the expected payoff. The proposed framework addresses issues specific to the demand dispatch of datacenters such as timeliness of VM migrations and the impact of geographic distance on migration times. It also explicitly incorporates risks that may cause the load curtailment operation to be ultimately unsuccessful and result in monetary losses to datacenter operators; specifically, variability in network bandwidth that can cause uncertainty in VM migration times as well as the uncertain payoff when participating in DR markets. A set of case studies involving datacenters participating in an economic DR program is used to validate the framework. Nagarajan Kandasamy, Chika O. Nwankpa, David R. Kaeli |
ICDCS | 2 |
| 2012 | Analytic Regularization of Uniform Cubic B-spline Deformation Fields
James A. Shackleford, Ana M. Lourenço, Nadya Shusharina, Nagarajan Kandasamy, Gregory C. Sharp |
MICCAI (2) | 5 |
| 2012 | Evaluating compressive sampling strategies for performance monitoring of data centersabstractPerformance monitoring of data centers provides vital information for dynamic resource provisioning, fault diagnosis, and capacity planning decisions. However, the very act of monitoring a system interferes with its performance, and if the information is transmitted to a monitoring station for analysis and logging, this consumes network bandwidth and disk space. This paper proposes a low-cost monitoring solution using compressive sampling - a technique that allows certain classes of signals to be recovered from the original measurements using far fewer samples than traditional approaches - and evaluates its ability to measure typical signals generated in a data-center setting using a testbed comprising the Trade6 enterprise application. The results open up the possibility of using low-cost compressive sampling techniques to detect performance bottlenecks and anomalies that manifest themselves as abrupt changes exceeding operator-defined threshold values in the underlying signals. Tingshan Huang, Nagarajan Kandasamy, Harish Sethu |
NOMS | 2 |
| 2012 | Entropy-Based Detection of Incipient Faults in Software SystemsabstractThis paper develops and validates a methodology to detect small, incipient faults in software systems. Incipient faults such as memory leaks slowly deteriorate the software's performance over time and if left undetected, the end result is usually a complete system failure. The proposed method combines tools from information theory and statistics: entropy and principal component analysis (PCA). The entropy calculation summarizes the information content associated with the collected low-level metrics and reduces the computational burden incurred by the subsequent PCA step which detects underlying patterns and correlations present in the multivariate data, as well as distortions in the correlations indicative of an incipient fault. We use the technique to detect memory bloat within the Trade6 enterprise application under dynamic workload patterns, showing that small leaks can be detected quickly and with a low false alarm rate. Our method is also robust to the periodic/seasonal patterns affecting the metrics used to detect the fault. Salvador DeCelles, Nagarajan Kandasamy |
PRDC | 2 |
| 2011 | SDC testbed: Software defined communications testbed for wireless radio and optical networkingabstractThis paper describes the development of a new Software Defined Communications (SDC) testbed architecture. SDC aims to generalize the area of software defined radio to include propagation media not exclusively limited to radio frequencies (optical, ultrasonic, etc.). This SDC platform leverages existing and custom hardware in combination with reference software applications in order to provide a complete research and development platform. This platform can be used to implement current and future standards that make use of highly demanding communications techniques, including ultrawideband (UWB) radio and free-space optical communications. This paper describes the commercial and custom hardware that is being integrated into the platform, including the baseband hardware and the modular transceiver frontends. Furthermore, the paper describes the software development currently in progress with this platform, including the integration of available open source designs into the platform, and the development of custom IP for scalable OFDM PHY implementations in radio and optical communications. We seek to create a complete research platform for the commercial and academic wireless communities, capable of delivering the highest possible performance and flexibility while providing the necessary development tools and reference designs in order to minimize system learning curve and development cost. Boris Shishkin, Doug Pfeil, Danh H. Nguyen, Kevin Wanuga, James Chacko, Nagarajan Kandasamy, Timothy P. Kurzweg, Kapil R. Dandekar |
WiOpt | 7 |
| 2011 | Combined Power and Performance Management of Virtualized Computing Environments Serving Session-Based WorkloadsabstractThis paper develops an online resource provisioning framework for combined power and performance management in a virtualized computing environment serving session-based workloads. We pose this management problem as one of sequential optimization under uncertainty and solve it using limited lookahead control (LLC), a form of model-predictive control. The approach accounts for the switching costs incurred when provisioning virtual machines and explicitly encodes the risk of provisioning resources in an uncertain and dynamic operating environment. We experimentally validate the control framework on a server cluster supporting three online services. When managed using LLC, our cluster setup saves, on average, 41% in power-consumption costs over a twenty-four hour period when compared to a system operating without dynamic control. Finally, we use trace-based simulations to analyze LLC performance on server clusters larger than our testbed and show how concepts from approximation theory can be used to further reduce the computational burden of controlling large systems. Dara Kusic, Nagarajan Kandasamy, Guofei Jiang |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2009 | On the application of predictive control techniques for adaptive performance management of computing systemsabstractThis paper addresses adaptive performance management of real-time computing systems. We consider a generic model-based predictive control approach that can be applied to a variety of computing applications in which the system performance must be tuned using a finite set of control inputs. The paper focuses on several key aspects affecting the application of this control technique to practical systems. In particular, we present techniques to enhance the speed of the control algorithm for real-time systems. Next we study the feasibility of the predictive control policy for a given system model and performance specification under uncertain operating conditions. The paper then introduces several measures to characterize the performance of the controller, and presents a generic tool for system modeling and automatic control synthesis. Finally, we present a case study involving a real-time computing system to demonstrate the applicability of the predictive control framework. Sherif Abdelwahed, Jia Bai, Rong Su 0001, Nagarajan Kandasamy |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2008 | Approximation Modeling for the Online Performance Management of Distributed Computing SystemsabstractA promising method of automating management tasks in computing systems is to formulate them as control or optimization problems in terms of performance metrics. For an online optimization scheme to be of practical value in a distributed setting, however, it must successfully tackle the curses of dimensionality and modeling. This paper develops a hierarchical control framework to solve performance management problems in distributed computing systems operating in a data center. Concepts from approximation theory are used to reduce the computational burden of controlling such large-scale systems. The relevant approximations are made in the construction of the dynamical models to predict system behavior and in the solution of the associated control equations. Using a dynamic resource-provisioning problem as a case study, we show that a computing system managed by the proposed control framework with approximation models realizes profit gains that are, in the best case, within 1% of a controller using an explicit model of the system. Dara Kusic, Nagarajan Kandasamy, Guofei Jiang |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Sensor Deployment for Failure Diagnosis in Networked Aerial Robots: A Satisfiability-Based Approach
Fadi A. Aloul, Nagarajan Kandasamy |
SAT | 2 |
| 2006 | A Dependable System Architecture for Safety-Critical Respiratory-Gated Radiation TherapyabstractThis experience report describes the design and implementation of safety-critical software and hardware for respiratory gating of a medical linear accelerator. Respiratory gating refers to a radiotherapy technique for treating cancer in the lung, liver, and abdomen, where tumors move while a patient breathes. A computer software program tracks the position of the tumor within the human body using X-ray fluoroscopy. When the tumor is in the correct position, the linear accelerator is triggered, delivering a beam of radiation toward the target. As part of the gating system, a comprehensive strategy for safety has been developed. This paper describes these safety features, focusing on the online monitoring techniques used to confirm the proper operation of the fluoroscopic imaging panels and the pattern recognition algorithms used for tumor identification Gregory C. Sharp, Nagarajan Kandasamy |
DSN | 2 |
| 2006 | A Hierarchical Optimization Framework for Autonomic Performance Management of Distributed Computing SystemsabstractThis paper develops a scalable online optimization framework for the autonomic performance management of distributed computing systems operating in a dynamic environment to satisfy desired quality-ofservice objectives. To efficiently solve the performance management problems of interest in a distributed setting, we develop a hierarchical structure where a highlevel limited-lookahead controller manages interactions between lower-level controllers using forecast operating and environment parameters. We develop the overall control structure, and as a case study, show how to efficiently manage the power consumed by a computer cluster. Using workload traces from the Soccer World Cup 98 web site, we show via simulations that the proposed method is scalable, has low run-time overhead, and adapts quickly to time-varying workload patterns. Nagarajan Kandasamy, Sherif Abdelwahed, Mohit Khandekar |
ICDCS | 1 |
| 2006 | Sensor Selection and Placement for Failure Diagnosis in Networked Aerial RobotsabstractUnmanned aerial vehicles (UAVs) represent an important class of networked robotic applications that must be both highly dependable and autonomous. This paper addresses sensor selection and placement problems for distributed failure diagnosis in such networks where multiple vehicles must agree on the fault status of another UAV. An integer linear programming (ILP) approach is proposed to solve these problems. The ILP models of interest are developed and solved using two different solvers. Experimental results indicate that the proposed models are tractable for medium-sized topologies Nagarajan Kandasamy, Fadi A. Aloul, Tak-John Koo |
ICRA | 1 |
| 2005 | Time-Constrained Failure Diagnosis in Distributed Embedded Systems: Application to Actuator DiagnosisabstractAdvanced automotive control applications such as steer-by-wire are typically implemented as distributed systems comprising many embedded processors, sensors, and actuators interacting via a communication bus. They have severe cost constraints, but demand a high level of safety and performance. Motivated by the need for timely diagnosis of faulty actuators in such systems, we present a method to achieve distributed failure diagnosis under deadline and resource constraints. Actuators are diagnosed in distributed fashion by processors to provide a global view of their fault status. The integration of software-based tests for actuator diagnosis within the overall control application is studied. These tests are implemented using analytical redundancy and execute concurrently with the control tasks. The test scheduling problem is then formulated and solved to guarantee actuator diagnosis within designer-specified deadlines while meeting control performance goals. As a secondary objective, the scheduling algorithm also reduces the number of processors required for diagnosis. We demonstrate the practicality of the proposed diagnosis approach by applying it to a steer-by-wire example to identify failed actuators in timely fashion. Nagarajan Kandasamy, John P. Hayes, Brian T. Murray |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Online Control for Self-Management in Computing SystemsabstractDependable computer systems hosting critical commerce, transportation, and military applications, among others, must satisfy stringent quality-of-service (QoS) requirements. However, as these systems become increasingly complex, maintaining the desired QoS by manually tuning the numerous performance-related parameters are very difficult. This paper develops a generic online control framework to design self-managing computer systems. The proposed approach explores a limited region of the system state-space at each time step and decides the best control action accordingly. We present two case studies to demonstrate the practicality of the proposed control framework. Sherif Abdelwahed, Nagarajan Kandasamy, Sandeep Neema |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2003 | Dependable Communication Synthesis for Distributed Embedded Systems
Nagarajan Kandasamy, John P. Hayes, Brian T. Murray |
SAFECOMP | 1 |
| 2002 | Time-Constrained Failure Diagnosis in Distributed Embedded SystemsabstractAdvanced automotive control applications such as steer and brake-by-wire are typically implemented as distributed systems comprising many embedded processors, sensors, and actuators interconnected via a communication bus. They have severe cost constraints but demand a high level of safety and performance. Motivated by the need for timely diagnosis of faulty actuators in such systems, we present a general method to implement failure diagnosis under deadline and resource constraints. Actuators are diagnosed in distributed fashion by processors to provide a global view of their fault status. The diagnostic tests are implemented in software using analytical redundancy and execute concurrently with the control tasks. The proposed method solves the test scheduling problem using a static list-based approach which guarantees actuator diagnosis within designer-specified deadlines while meeting control performance goals. As a secondary objective, it also minimizes the number of required processors. We present simulation results evaluating the effectiveness of the proposed method under various design constraints. Nagarajan Kandasamy, John P. Hayes, Brian T. Murray |
DSN | 1 |
| 1999 | Tolerating Transient Faults in Statically Scheduled Safety-Critical Embedded SystemsabstractStatic off-line scheduling ensures predictability of worst-case behavior and high resource utilization for safety-critical applications but lacks the flexibility needed to deal with run-time fault-tolerance. We present a temporal redundancy-based recovery technique that tolerates transient task failures in statically scheduled distributed embedded systems where tasks have timing, resource, and precedence constraints. Task failures are handled using precomputed contingency schedules that introduce adaptive fault tolerance into table-driven dispatchers. Failures are masked using the spare capacity on the affected processor and the recovery scheme requires no hardware overhead. Our approach combines the benefits of static scheduling with the run-time flexibility needed for fault tolerance in low-cost embedded systems. We present a method to obtain contingency schedules and prove its correctness. We also evaluate the effectiveness of the proposed method through simulation. Nagarajan Kandasamy, John P. Hayes, Brian T. Murray |
SRDS | 1 |