C. Mani Krishna 0001

dblp:k/CManiKrishna · also C. M. Krishna 0001 · DBLP profile ↗
← Back
62ranked-venue papers
14as first author
1since 2021 · last 2025
0000-0002-6019-9125ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 45 · 10 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 3 first-authorSecurity and privacy · 5Software engineering, systems software and programming languages · 5 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorComputer networks · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
23 papers
Energy-efficient computing · 24% Distributed systems · 20% Embedded and real-time systems · 11%
Computer networks
3 papers
Optical networks · 46% Internet architecture and protocols · 22% Wireless networking · 21%

Topics — the 30 heaviest of 63, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
power management
0.242011
Utilization-Based Resource Partitioning for Power-Performance Efficiency in SMT Processors · IEEE Trans. Parallel Distributed Syst. 2011
Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems · IEEE Trans. Computers 2003
Cool-cache for hot multimedia · MICRO 2001
Cloud and datacenter computing › resource allocation › resource allocation policy
resource partitioning
0.112011
Utilization-Based Resource Partitioning for Power-Performance Efficiency in SMT Processors · IEEE Trans. Parallel Distributed Syst. 2011
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.112011
Utilization-Based Resource Partitioning for Power-Performance Efficiency in SMT Processors · IEEE Trans. Parallel Distributed Syst. 2011
Embedded and real-time systems
real-time scheduling
0.162004
Addendum to Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems · IEEE Trans. Computers 2004
Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems · IEEE Trans. Computers 2003
Efficient On-Line Processor Scheduling for a Class of IRIS (Increasing Reward with Increasing Service.) Real-Time Tasks · SIGMETRICS 1993
Distributed systems › fault tolerance
failure detection
0.112007
Software-Based Failure Detection and Recovery in Programmable Network Interfaces · IEEE Trans. Parallel Distributed Syst. 2007
Distributed systems › fault tolerance
failure recovery
0.112007
Software-Based Failure Detection and Recovery in Programmable Network Interfaces · IEEE Trans. Parallel Distributed Syst. 2007
Distributed systems › fault tolerance › failure recovery
state restoration
0.112007
Software-Based Failure Detection and Recovery in Programmable Network Interfaces · IEEE Trans. Parallel Distributed Syst. 2007
Electronic design automation
hardware verification and test
0.051996
On the Effect of Defect Clustering on Test Transparency and IC Test Optimization · IEEE Trans. Computers 1996
On optimizing VLSI testing for product quality using die-yield prediction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Optimal Scheduling of Signature Analysis for VLSI Testing · IEEE Trans. Computers 1991
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.012003
Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems · IEEE Trans. Computers 2003
Energy-efficient computing › energy-aware scheduling
energy-aware real-time scheduling
0.012003
Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems · IEEE Trans. Computers 2003
Memory systems
cache design
0.012002
The Minimax Cache: An Energy-Efficient Framework for Media Processors · HPCA 2002
Distributed systems › peer-to-peer systems
network construction
0.012002
Filtering Random Graphs to Synthesize Interconnection Networks with Multiple Objectives · IEEE Trans. Parallel Distributed Syst. 2002
Interconnection networks and networks-on-chip
network topology
0.012002
Filtering Random Graphs to Synthesize Interconnection Networks with Multiple Objectives · IEEE Trans. Parallel Distributed Syst. 2002
Memory systems › cache design
partitioned cache
0.012002
The Minimax Cache: An Energy-Efficient Framework for Media Processors · HPCA 2002
Energy-efficient computing › power management › memory power management
cache energy reduction
0.012001
Cool-cache for hot multimedia · MICRO 2001
Memory systems
cache management
0.012001
Cool-cache for hot multimedia · MICRO 2001
Memory systems › cache › CPU cache
data cache
0.012001
Cool-cache for hot multimedia · MICRO 2001
Electronic design automation › hardware verification and test
VLSI testing
0.041993
On optimizing VLSI testing for product quality using die-yield prediction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Optimal Scheduling of Signature Analysis for VLSI Testing · IEEE Trans. Computers 1991
Optimal Design and Sequential Analysis of VLSI Testing Strategy · IEEE Trans. Computers 1988
Optical networks › optical switching
optical packet switching
0.012000
An in-band signaling protocol for optical packet switching networks · IEEE J. Sel. Areas Commun. 2000
Wireless networking › medium access control › conflict-free multiple access
reservation protocol
0.012000
An in-band signaling protocol for optical packet switching networks · IEEE J. Sel. Areas Commun. 2000
Internet architecture and protocols
signaling protocol
0.012000
An in-band signaling protocol for optical packet switching networks · IEEE J. Sel. Areas Commun. 2000
Optical networks
wavelength-division multiplexing
0.012000
An in-band signaling protocol for optical packet switching networks · IEEE J. Sel. Areas Commun. 2000
Distributed systems
fault tolerance
0.042002
Filtering Random Graphs to Synthesize Interconnection Networks with Multiple Objectives · IEEE Trans. Parallel Distributed Syst. 2002
On the Graceful Degradation of Phase-Locked Clocks · RTSS 1988
On Scheduling Tasks with a Quick Recovery from Failure · IEEE Trans. Computers 1986
Optical networks › wavelength-division multiplexing
wavelength-division multiplexing local area networks
0.011996
A Distributed Adaptive Protocol Providing Real-Time Service on WDM-Based Lans · INFOCOM 1996
Embedded and real-time systems › real-time scheduling › soft real-time scheduling
reward-based scheduling
0.021993
Efficient On-Line Processor Scheduling for a Class of IRIS (Increasing Reward with Increasing Service.) Real-Time Tasks · SIGMETRICS 1993
Optimal Resource Control in Periodic Real-Time Environments · RTSS 1988
Energy-efficient computing
low-power design
0.012004
Addendum to Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems · IEEE Trans. Computers 2004
Interconnection networks and networks-on-chip
network reliability
0.012002
Filtering Random Graphs to Synthesize Interconnection Networks with Multiple Objectives · IEEE Trans. Parallel Distributed Syst. 2002
Energy-efficient computing › memory energy efficiency
power-aware memory system
0.012002
The Minimax Cache: An Energy-Efficient Framework for Media Processors · HPCA 2002
Electronic design automation › hardware verification and test
adaptive test
0.011993
On optimizing VLSI testing for product quality using die-yield prediction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Storage systems › file systems
distributed file system
0.011993
Performance Analysis of Distributed File Systems with Non-Volatile Caches · HPDC 1993

Methods — techniques the papers use, named apart from their topics

simulation · 0.1adaptive resource partitioning · 0.1self-testing · 0.1random graph generation · 0.1filtration process · 0.1voltage scaling · 0.0clock scaling · 0.0cache simulation · 0.0rational interpolation · 0.0sampling probe algorithm · 0.0passive listening · 0.0demand-adaptive scheduling · 0.0defect distribution analysis · 0.0distributed algorithm · 0.0
YearPublicationVenuePosition
2025 Privacy and Traffic Efficiency of CAVs Under Dynamic Conditions in ITS
Nirupama Ravi, C. Mani Krishna 0001, Israel Koren
IEEE Internet Things J.2
2020 Thermal Aware Task Scheduling for Enhanced Cyber-Physical Systems Sustainability
abstract
Cyber-Physical Systems (CPS) are increasingly used in a variety of transportation, healthcare, electricity grid, and other applications. Thermal stress is often a major concern for processors embedded in such systems. High operating temperatures can dramatically shorten processor life. This, in turn, can require provisioning of significant amounts of additional computational hardware to withstand more frequent failures, with obvious implications for sustainability. This paper describes a novel approach to reduce thermally-induced damage in CPS processors by targeting Dynamic Voltage and Frequency Scaling (DVFS) to high-activity task phases. That is, by preferentially slowing down high-activity task phases, significant additional savings in energy and thermal stress can be attained for a given amount of computational slowdown; this approach is shown to be superior to conventional methods that use DVFS without regard to activity levels. Also, our proposed task reassignment across cores is driven by estimates of current core reliability, which is superior to the usual approach of simply using either current temperature or temperature history. Our approach leads to a significant reliability improvement (around 20 percent) over baseline DVFS techniques.
Shikang Xu, Israel Koren, C. Mani Krishna 0001
IEEE Trans. Sustain. Comput.3
2017 AdaFT: A Framework for Adaptive Fault Tolerance for Cyber-Physical Systems
abstract
Cyber-physical systems (CPS) frequently have to use massive redundancy to meet application requirements for high reliability. While such redundancy is required, it can be activated adaptively, based on the current state of the controlled plant. Most of the time, the plant is in a state that allows for a lower level of fault tolerance. Avoiding the continuous deployment of massive fault tolerance will greatly reduce the workload of the CPS, and lower the operating temperature of the cyber sub-system, thus increasing its reliability. In this article, we extend our prior research by demonstrating a software simulation framework Adaptive Fault Tolerance (AdaFT) that can automatically generate the sub-spaces within which our adaptive fault tolerance can be applied. We also show the theoretical benefits of AdaFT and its actual implementation in several real-world CPSs.
Israel Koren, C. Mani Krishna 0001
ACM Trans. Embed. Comput. Syst.3
2015 Ameliorating Thermally Accelerated Aging With State-Based Application of Fault-Tolerance in Cyber-Physical Computers
abstract
Cyber-physical systems have become more prevalent in recent years. Computer control is increasingly being considered for use in applications which are operationally resource-constrained and are very cost-sensitive, while requiring very high reliability. The primary approach to building ultra-reliable systems is to deploy massive redundancy. However, the constraints just mentioned make it very difficult to use the massive everywhere-redundancy approaches used in such traditional (relatively cost-insensitive) applications as aerospace. In this paper, we address such problems by adaptively adjusting fault-tolerance levels according to the current state of the controlled plant. Such an approach imposes far less thermal stress on the computer, thereby enhancing reliability, and thus requiring a smaller number of line-replaceable units to maintain required levels of reliability over any given period of operation.
C. Mani Krishna 0001
IEEE Trans. Reliab.1
2012 Cost Functions for Scheduling Tasks in Cyber-physical Systems
Abhinna Jain, C. Mani Krishna 0001, Israel Koren, Zahava Koren
ICINCO (1)2
2012 Adaptive Data Update Management in Sensor Networks
C. Mani Krishna 0001
ICINCO (1)1
2012 A Study of the Impact of Computational Delays in Missile Interception Systems
Israel Koren, C. Mani Krishna 0001
ICINCO (1)3
2011 Utilization-Based Resource Partitioning for Power-Performance Efficiency in SMT Processors
abstract
Simultaneous multithreading (SMT) increases processor throughput by allowing parallel execution of several threads. However, fully sharing processor resources may cause resource monopolization by a single thread or other misallocations, resulting in overall performance degradation. Static resource partitioning techniques have been suggested, but are not as effective as dynamic ones since program behavior does change over the course of its execution. In this paper, we propose an Adaptive Resource Partitioning Algorithm (ARPA) that dynamically assigns resources to threads according to changes in thread behavior. ARPA analyzes the resource usage efficiency of each thread in a given time period and assigns more resources to threads which can use them more efficiently. Its purpose is to improve the efficiency of resource utilization, thereby improving overall instruction throughput. Our simulation results on a set of 42 multiprogramming workloads show that ARPA outperforms the traditional fetch policy ICOUNT by 55.8 percent with regard to overall instruction throughput and achieves a 33.8 percent improvement over Static Partitioning. It also outperforms the current best dynamic resource allocation technique, Hill-climbing, by 5.7 percent. Considering fairness accorded to each thread, ARPA attains 43.6, 18.5, and 9.2 percent improvements over ICOUNT, Static Partitioning, and Hill-climbing, respectively, using a common fairness metric. We also explore the energy efficiency of dynamically controlling the number of powered-on reorder buffer entries for ARPA. Compared with ARPA, our energy-aware resource partitioning algorithm achieves 10.6 percent energy savings, while the performance loss is negligible.
Huaping Wang, Israel Koren, C. Mani Krishna 0001
IEEE Trans. Parallel Distributed Syst.3
2008 An adaptive resource partitioning algorithm for SMT processors
abstract
Simultaneous Multithreading (SMT) increases processor throughput by allowing the parallel execution of several threads. However, fully sharing processor resources may cause resource monopolization by a single thread or other misallocations, resulting in overall performance degradation. Static resource partitioning techniques have been suggested, but are not as effective as dynamically controlling the resource usage of each thread since program behavior does change during its execution.
Huaping Wang, Israel Koren, C. Mani Krishna 0001
PACT3
2007 Software-Based Failure Detection and Recovery in Programmable Network Interfaces
abstract
Emerging network technologies have complex network interfaces that have renewed concerns about network reliability. In this paper, we present an effective low-overhead fault tolerance technique to recover from network interface failures. Failure detection is based on a software watchdog timer that detects network processor hangs and a self-testing scheme that detects interface failures other than processor hangs. The proposed self-testing scheme achieves failure detection by periodically directing the control flow to go through only active software modules in order to detect errors that affect instructions in the local memory of the network interface. Our failure recovery is achieved by restoring the state of the network interface using a small backup copy containing just the right amount of information required for complete recovery. The paper shows how this technique can be made to minimize the performance impact to the host system and be completely transparent to the user.
Yizheng Zhou, Vijay Lakamraju, Israel Koren, C. Mani Krishna 0001
IEEE Trans. Parallel Distributed Syst.4
2006 Testing and Validation of the CASA DCAS System
abstract
We present a system emulator that provides a versatile environment for testing and validating a distributed system for sensing of the lower atmosphere during its development lifecycle. The emulator, along with its comprehensive RAPIDS toolbox, provide the necessary infrastructure to experiment with different parameters of the system to ensure that the system specifications are being met, while also providing the scope of experimenting with newer configurations and futuristic designs. Extensive monitoring of the system possible through RAPIDS, allows to expose bugs and performance bottlenecks. The remote experimentation feature of the system allows several users to drive/monitor the system without requiring physical access to the emulator.
A. Bekkerman, Vijay Lakamraju, Israel Koren, C. Mani Krishna 0001
IGARSS4
2006 Compiler-based adaptive fetch throttling for energy-efficiency
abstract
Front-end instruction delivery accounts for a significant fraction of energy consumption in dynamically scheduled superscalar processors. Different front-end throttling techniques have been introduced to reduce the chip-wide energy consumption caused by redundant fetching. Hardware-based techniques, such as flow-based throttling, could reduce the energy consumption considerably, but with a high performance loss. On the other hand, compiler-based IPC-estimation-driven software fetch throttling (CFT) techniques result in relatively low performance degradation, which is desirable for high-performance processors. However, their energy savings are limited by the fact that they typically use a predefined fixed low IPC-threshold to control throttling. In this paper, we propose a compiler-based adaptive fetch throttling (CAFT) technique that allows changing the throttling threshold dynamically at runtime. Instead of using a fixed threshold, our technique uses the decode/issue difference (DID) to assist the fetch throttling decision based on the statically estimated IPC. Changing the threshold dynamically makes it possible to throttle at a higher estimated IPC, thus increasing the throttling opportunities and resulting in larger energy savings. We demonstrate that CAFT could increase the energy savings significantly compared to CFT, while preserving its benefit of low performance loss. Our simulation results show that the proposed technique doubles the energy-delay product (EDP) savings compared to the fixed threshold throttling and achieves a 6.1% average EDP saving.
Huaping Wang, Yao Guo 0001, Israel Koren, C. Mani Krishna 0001
ISPASS4
2005 Energy aware kernel for hard real-time systems
abstract
Embedded systems often have severe power and energy constraints. Dynamic voltage scaling (DVS) is a mechanism by which energy consumption may be reduced. In this paper, we implement a dynamic voltage scaling based scheduler in the eCos operating system running on Voltage Scalable Intel Xscale Board and show how energy usage can be reduced while still meeting hard real-time deadlines.
A. Goel, C. Mani Krishna 0001, Israel Koren
CASES2
2004 Energy Characterization of Hardware-Based Data Prefetching
abstract
This paper evaluates several hardware-based data prefetching techniques from an energy perspective, and explores their energy/performance tradeoffs. We present detailed simulation results and make performance and energy comparisons between different configurations. Power characterization is provided based on HSpice circuit-level simulation of state-of-the-art low-power cache designs implemented in deep-submicron process technology. This is combined with architecture-level simulation of switching activities in the memory system. The results show that while aggressive prefetching techniques often help to improve performance, they increase energy consumption in most of the cases. In designs implemented in deep-submicron 100-nm BPTM process technology, cache leakage becomes one of the dominant factors of the energy consumption. We have, however, found that if leakage is optimized with recently-proposed circuit-level techniques, most of the energy degradation is due to prefetch-hardware related costs and unnecessary L1 data cache lookups related to prefetches that hit in the L1 cache. This overhead on the memory system can be as much as 20%.
Yao Guo 0001, Saurabh Chheda, Israel Koren, C. Mani Krishna 0001, Csaba Andras Moritz
ICCD4
2004 Application-Level Fault Tolerance in the Orbital Thermal Imaging Spectrometer
abstract
Systems that operate in extremely volatile environments, such as orbiting satellites, must be designed with a strong emphasis on fault tolerance. Rather than rely solely on the system hardware, it may be beneficial to entrust some of the fault handling to software at the application level, which can utilize semantic information and software communication channels to achieve fault tolerance with considerably less power and performance overhead. We show the implementation and evaluation of such a software-level approach, application-level fault tolerance and detection (ALFTD) into the orbital thermal imaging spectrometer (OTIS).
E. Ciocca, Israel Koren, Zahava Koren, C. Mani Krishna 0001, Daniel S. Katz
PRDC4
2004 Addendum to Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems
C. Mani Krishna 0001, Yann-Hang Lee
IEEE Trans. Computers1
2003 Low Overhead Fault Tolerant Networking in Myrinet
abstract
193-202
Vijay Lakamraju, Israel Koren, C. Mani Krishna 0001
DSN3
2003 Pre-Processing Input Data to Augment Fault Tolerance in Space Applications
abstract
491-500
Zahava Koren, Israel Koren, C. Mani Krishna 0001
DSN4
2003 A Voltage Scheduling Heuristic for Real-Time Task Graphs
abstract
741-750
Diganta Roychowdhury, Israel Koren, C. Mani Krishna 0001, Yann-Hang Lee
DSN3
2003 Scheduling Techniques for Reducing Leakage Power in Hard Real-Time Systems
abstract
Modern embedded systems are often severely resource-constrained. In current research, the reduction of dynamic power has been the focus. However, with increased chip speed and density in submicron scale, the static (leakage) power consumption has become an increasingly significant fraction of the total. Indeed, a five-fold increase in leakage power per technology generation has been observed. At this pace, leakage power could soon equal dynamic power. In this paper, we investigate scheduling policies to reduce leakage power in real-time systems. We show that with simple scheduling techniques, overall leakage energy can be reduced by an order of magnitude.
Yann-Hang Lee, Krishna P. Reddy, C. Mani Krishna 0001
ECRTS3
2003 Constrained Energy Allocation for Mixed Hard and Soft Real-Time Tasks
Yoonmee Doh, Daeyoung Kim 0001, Yann-Hang Lee, C. Mani Krishna 0001
RTCSA4
2003 Scanning the issue - special issue on real-time systems
abstract
Provides an overview of the technical articles and features presented in this issue.
C. Mani Krishna 0001, Yann-Hang Lee
Proc. IEEE1
2003 Voltage-Clock Scaling for Low Energy Consumption in Fixed-Priority Real-Time Systems
Yann-Hang Lee, C. Mani Krishna 0001
Real Time Syst.2
2003 Voltage-Clock-Scaling Adaptive Scheduling Techniques for Low Power in Hard Real-Time Systems
abstract
Many embedded systems operate under severe power and energy constraints. Voltage clock scaling is one mechanism by which energy consumption may be reduced: it is based on the fact that power consumption is a quadratic function of the voltage, while the speed is a linear function. We show how voltage scaling can be scheduled to reduce energy usage while still meeting real-time deadlines.
C. Mani Krishna 0001, Yann-Hang Lee
IEEE Trans. Computers1
2003 Cool-Cache: A compiler-enabled energy efficient data caching framework for embedded/multimedia processors
abstract
The unique characteristics of multimedia/embedded applications dictate media-sensitive architectural and compiler approaches to reduce the power consumption of the data cache. Our goal is exploring energy savings for embedded/multimedia workloads without sacrificing performance. Here, we present two complementary media-sensitive energy-saving techniques that leverage static information. While our first technique is applicable to existing architectures, in our second technique we adopt a more radical approach and propose a new tagless caching architecture by reevaluating the architecture--compiler interface.Our experiments show that substantial energy savings are possible in the data cache. Across a wide range of cache and architectural configurations, we obtain up to 77% energy savings, while the performance varies from 14% improvement to 4% degradation depending on the application.
Osman S. Unsal, Raksit Ashok, Israel Koren, C. Mani Krishna 0001, Csaba Andras Moritz
ACM Trans. Embed. Comput. Syst.4
2002 The Minimax Cache: An Energy-Efficient Framework for Media Processors
abstract
This work is based on our philosophy of providing interlayer system-level power awareness in computing systems. Here, we couple this approach with our vision of multi-partitioned memory systems, where memory accesses are separated based on their static predictability and memory footprint and managed with various compiler controlled techniques. We show that media applications are mapped more efficiently when scalar memory accesses are redirected to a mini-cache. Our results indicate that a partitioned 8K cache with the scalars being mapped to a 512 byte mini-cache can be more efficient than a 16K monolithic cache from both performance and energy point of view for most applications. In extensive experiments, we report 30% to 60% energy-delay product savings over a range of system configurations and different cache sizes.
Osman S. Unsal, Israel Koren, C. Mani Krishna 0001, Csaba Andras Moritz
HPCA3
2002 Towards energy-aware software-based fault tolerance in real-time systems
abstract
Many real-time systems employed in defense, space, and consumer applications have power constraints and high reliability requirements. In this paper, we focus on the relationship between fault tolerance techniques and energy consumption. In particular, we establish the energy efficiency of Application Level Fault Tolerance (ALFT) over other software-based fault tolerance methods. We then develop sensible energy-aware heuristics for ALFT schemes. The heuristics yield up to 40% energy savings.
Osman S. Unsal, Israel Koren, C. Mani Krishna 0001
ISLPED3
2002 Filtering Random Graphs to Synthesize Interconnection Networks with Multiple Objectives
abstract
Synthesizing networks that satisfy multiple requirements, such as high reliability, low diameter, good embeddability, etc., is a difficult problem to which there has been no completely satisfactory solution. We present a simple, yet very effective, approach to this problem. The crux of our approach is a filtration process that takes as input a large set of randomly generated graphs and filters out those that do not meet the specified requirements. Our experimental results show that this approach is both practical and powerful. The use of random regular networks as the raw material for the filtration process was motivated by their surprisingly good performance with regard to almost all properties that characterize a good interconnection network. We provide results related to the generation of networks that have low diameter, high fault tolerance, and good embeddability. Through this, we show that the generated networks are serious competitors to several traditional well-known networks. We also explore how random networks can be used in a packaging hierarchy and comment on the scope of application of these networks.
Vijay Lakamraju, Israel Koren, C. Mani Krishna 0001
IEEE Trans. Parallel Distributed Syst.3
2001 EDF scheduling using two-mode voltage-clock-scaling for hard real-time systems
abstract
Scaling down power supply voltage yields a quadratic reduction in dynamic power dissipation and also requires a reduction in clock frequency. In order to meet task deadlines in hard real-time systems, the delay penalty in voltage scaling needs to be carefully considered to achieve low power consumption. In this paper, we focus on dynamic reclaiming of early released resources in Earliest Deadline First (EDF) scheduling using voltage scaling. In addition to a static voltage assignment, we propose a new dynamic-mode assignment, which has a flexible voltage mode setting at run-time enabling much larger energy savings. Using simulation results and exploiting the interplay between power supply voltage, frequency, and circuit delay in CMOS technology, we find the optimal two-level voltage settings that minimize energy consumption.
Yann-Hang Lee, Yoonmee Doh, C. Mani Krishna 0001
CASES3
2001 Cool-cache for hot multimedia
abstract
We claim that the unique characteristics of multimedia applications dictate media-sensitive architectural and compiler approaches to reduce the power consumption of the data cache. Our motivation is exploring energy savings for real-time multimedia workloads without sacrificing performance. In this paper, we present two complementary media-sensitive energy-saving techniques that leverage static information. While our first technique is applicable to existing architectures, in our second technique we adopt a more radical approach and propose a new caching architecture by re-evaluating the architecture-compiler interface. Our experiments show that substantial energy savings are possible in the data cache. Across a wide range of cache and architectural configurations we obtain up to 77% energy savings, while the performance varies from 14% improvement to 4% degradation depending on the application.
Osman S. Unsal, Raksit Ashok, Israel Koren, C. Mani Krishna 0001, Csaba Andras Moritz
MICRO4
2001 Rational Interpolation Examples in Performance Analysis
abstract
The rational interpolation approach has been applied to performance analysis of computer systems previously. In this paper, we demonstrate the effectiveness of the rational interpolation technique in the analysis of randomized algorithms and the fault probability calculation for some real-time systems.
Wei-Bo Gong, C. Mani Krishna 0001
IEEE Trans. Computers3
2000 Synthesis of Interconnection Networks: A Novel Approach
abstract
The interconnection network is a crucial element in parallel and distributed systems. Synthesizing networks that satisfy a set of desired properties, such as high reliability, low diameter and good scalability is a difficult problem to which there has been no completely satisfactory solution. In this paper, we present a new approach to network synthesis. We start by generating a large number of random regular networks. These networks are then passed through filters, which filter out networks that do not satisfy specified network design requirements. By applying multiple filters in tandem, it is possible to synthesize networks which satisfy a multitude of properties. The filtered output thus constitutes a short-list of "good" networks that the designer can choose from. The use of random regular networks was motivated by their surprisingly good performance with regard to almost all properties that characterize a good interconnection network. Experimental results have shown that this approach is practical and powerful. In this paper we focus on the generation of networks which have low diameter, good scalability and high fault tolerance. These generated networks are shown to compare favorably with several well-known networks.
Vijay Lakamraju, Zahava Koren, C. Mani Krishna 0001
DSN3
2000 An in-band signaling protocol for optical packet switching networks
abstract
Advances in optical networking reveal that all optical networks offering a multigigabit rate per wavelength will soon become economical as the underlying backbone in wide-area networks, in which the optical switch plays a central role. One of the central issues is the design of efficient signaling protocols which can support diversified traffic types, in particular the bursty IP traffic. This paper introduces a novel signaling protocol called the sampling probe algorithm (or SPA) to be used in a class of optical packet switching systems based on wavelength division multiplexing (WDM). The proposed scheme takes a drastically different approach from all existing signaling protocols. The salient features are 1) the pretransmission coordination is using an in-band signaling protocol, and thus does not require separate control channel(s) for transmission coordination; 2) the protocol is based on a reservation (connection) scheme which is capable of supporting multimedia traffic; 3) a gated service is adopted in which each successful reservation allows multiple packets (train of packets) to be transmitted, which can significantly reduce the per packet overhead; 4) the scheduling algorithm is adaptive by allowing flexible assignment of bandwidth on-demand; 5) the channel status gathering is done in a distributed fashion, and uses a passive listening mechanism, which itself does not interfere with packet transmissions. The results demonstrate that the proposed in-band signaling protocol can achieve high throughput and stability under heavy traffic condition.
Bo Li 0001, Aura Ganz, C. Mani Krishna 0001
IEEE J. Sel. Areas Commun.3
2000 Application-Level Fault Tolerance as a Complement to System-Level Fault Tolerance
Joshua Haines, Vijay Lakamraju, Israel Koren, C. Mani Krishna 0001
J. Supercomput.4
1997 Screening for Known Good Die (KGD) Based on Defect Clustering: An Experimental Study
abstract
Die screening based on the locality of defects has long been informally practised in the industry whereby dice from wafers, or parts of the wafer, that display high defect levels are discarded. More recently this approach has been refined such that test results for neighbouring dice on the wafer are also considered in evaluating test results for a particular die. It has been shown in principle, using negative binomial statistics for defect distributions on wafers, that such an approach can much better optimize test costs and screen for low defect levels in bare dice and packaged chips. In this paper we present, for the first time, experimental test data to demonstrate the effectiveness of this new approach. Our results are based on extensive testing of 4784 dice on 23 wafers from an IBM process. We show that bare die screening based on defect clustering considerations can significantly reduce defect levels in dice that pass wafer probe tests. This approach also has the potential to screen out burn-in failures. Thus it offers new low cost strategies for delivering high quality "known- good" die (KGD) for MCM applications.
Adit D. Singh, Phil Nigh, C. Mani Krishna 0001
ITC3
1996 A Distributed Adaptive Protocol Providing Real-Time Service on WDM-Based Lans
abstract
We consider the problem of supporting multiple classes of real-time and non-real-time traffic on a WDM-based local-area network. We present a demand-adaptive algorithm which schedules transmission according to the quality-of-service requirements of the various traffic classes. The algorithm requires one control channel, on which a single token circulates. This token is used both as a means of communicating status information between the nodes and of controlling access to each of the multiple channels. Numerical examples are provided to illustrate the performance of the algorithm.
Anlu Yan, Aura Ganz, C. Mani Krishna 0001
INFOCOM3
1996 On the Effect of Defect Clustering on Test Transparency and IC Test Optimization
abstract
We recently proposed a wafer-based testing approach which for the first time employs defect clustering information on the wafer to optimize test cost and defect levels in the shipped product. Preliminary analysis of this approach had implicitly assumed that the probability that a test detects a faulty circuit is independent of the number of faults in that circuit. This assumption may be optimistic. In this correspondence, we study the effect of clustering and test transparency on defect distributions in individual dice, and its impact on the fault detection capabilities of a given test set. We show here that significant defect-level improvements in the shipped product can indeed be achieved by exploiting defect clustering in optimization testing.
Adit D. Singh, C. Mani Krishna 0001
IEEE Trans. Computers2
1994 Performance benefits of non-volatile caches in distributed file systems
abstract
Abstract We study the use of non‐volatile memory for caching in distributed file systems. This provides an advantage over traditional distributed file systems in that the load is reduced at the server without making the data vulnerable to failures. We propose the use of a small non‐volatile cache for writes, at the client and the file server, together with a larger volatile read cache to keep the cost of the caches reasonable. We use a synthetic workload developed from analysis of file I/O traces from commercial production systems and use a detailed simulation of the distributed environment. The service times for the resources of the system were derived from measurements performed on a typical workstation. We show that non‐volatile write caches at the clients and the file server reduce the write response time and the load on the file server dramatically, thus improving the scalability of the system. We examine the comparative benefits of two alternative writeback policies for the non‐volatile write cache. We show that a proposed threshold based writeback policy is more effective than a periodic writeback policy under heavy load. We also investigate the effect of varying the write cache size and show that introducing a small non‐volatile cache at the client in conjunction with a moderate sized non‐volatile server write cache improves the write response time by a factor of four at all load levels.
Prabuddha Biswas, Don Towsley, K. K. Ramakrishnan, C. Mani Krishna 0001
Concurr. Pract. Exp.4
1993 Performance Analysis of Distributed File Systems with Non-Volatile Caches
abstract
The authors study the use of non-volatile memory for caching in distributed file systems. This provides an advantage over traditional distributed file systems in that the load is reduced at the server without making the data vulnerable to failures. They show that small non-volatile write caches at the clients and the server are quite effective. They reduce the write response time and the load on the file server dramatically, thus improving the scalability of the system. They show that a proposed threshold based writeback policy is more effective than a periodic writeback policy. They use a synthetic workload developed from analysis of file I/O traces from commercial production systems. The study is based on a detailed simulation of the distributed environment. The service times for the resources of the system were derived from measurements performed on a typical workstation.>
Prabuddha Biswas, K. K. Ramakrishnan, Don Towsley, C. Mani Krishna 0001
HPDC4
1993 Efficient On-Line Processor Scheduling for a Class of IRIS (Increasing Reward with Increasing Service.) Real-Time Tasks
abstract
In this paper we consider the problem of on-line scheduling of real-time tasks which receive a that depends on the amount of service received. In our model, tasks have associated deadlines at which they must depart the system. The task computations are such that the longer they are able to execute before their deadline, the greater the value of their computations, i.e., the tasks have the property that they receive increasing reward with increasing service (IRIS). We focus on the problem of scheduling IRIS tasks in a system in which tasks arrive randomly over time, with the goal of maximizing the average reward accrued per task and per unit time. We describe and evaluate a two-level policy for this system. A top-level algorithm executes each time a task arrives and determines the amount of service to allocate to each task in the absence of future arrivals. A lower-level algorithm, an earliest deadline first (EDF) policy in our case, is responsible for the actual selection of tasks to execute. This two-level policy is evaluated through a combination of analysis and simulation, We observe that it provides nearly optimal performance when the variance in the interarrival times and/or laxities is low and that the performance is more sensitive to changes in the arrival process than the deadline distribution.
Jayanta K. Dey, James F. Kurose, Don Towsley, C. Mani Krishna 0001, Mahesh Girkar
SIGMETRICS4
1993 The effect of defect clustering on test transparency and defect levels
abstract
Proposes a wafer based testing approach which for the first time employs defect clustering information on the wafer to optimize test cost and defect levels in the shipped product. Preliminary analysis of this approach had assumed that the probability that a test detects a faulty circuit is independent of the number of faulty dies in the neighborhood of the circuit under test. Here, the authors relax this assumption by making test transparency a function of the number of faults. In this paper they study the effect of clustering on test transparency and defect levels based on maps of particle distributions on test wafers.>
Adit D. Singh, C. Mani Krishna 0001
VTS2
1993 On optimizing VLSI testing for product quality using die-yield prediction
abstract
An adaptive testing procedure that uses spatial defect clustering information and the available test results for neighboring dies to optimize test costs for VLSI testing is proposed. For the same average test costs, the approach shows the potential for better than a factor-of-two improvement in average defect levels. Perhaps more significantly, it allows the separation of high-quality circuits with defect levels more than order of magnitude better than the average for the production run. The proposal is orthogonal to all other approaches for improving defect levels and can be combined with them.>
Adit D. Singh, C. Mani Krishna 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1992 Analysis of the die test optimization algorithm for negative binomial yield statistics
abstract
Introduces a new adaptive testing algorithm that uses spatial defect clustering information and available test from neighbouring dies to optimize test lengths during wafer-probe testing. When applied to the defect distribution data for 12 sample wafers collected by Saji and Armstrong, the new approach showed potential for providing improvement in overall product quality. In this paper, the authors conduct a more general study to evaluate the proposed new test optimization algorithm based on the widely accepted negative binomial model for defect distributions on a wafer. The objective is to obtain a more accurate measure of the magnitude of the defect-level improvements that can be expected under various yield and defect-clustering conditions.>
C. Mani Krishna 0001, Adit D. Singh
VTS1
1992 Workshop Report: 1991 Workshop on Architectural Aspects of Real-Time Systems, San Antonio, Texas, U. S. A
C. Mani Krishna 0001, Yann-Hang Lee
Real Time Syst.1
1991 On Optimizing Wafer-Probe Testing for Product Quality Using Die-Yield Prediction
abstract
We propose a new adaptive testing procedure that uses spatial defect clustering information to optimize test lengths during wafer-probe testing. For the same average test lengths, our approach shows better than a factor-of-two improvement in average defect levels. It further allows the separation of high-quality dies with defect levels more than an order of magnitude better than the average for the production run. Our proposal is orthogonal to all other approaches for improving defect quality and can be combined with them.
Adit D. Singh, C. Mani Krishna 0001
ITC2
1991 A Random Distributed Algorithm to Embed Trees in Partially Faulty Processor Arrays
Dipak Sitaram, Israel Koren, C. Mani Krishna 0001
J. Parallel Distributed Comput.3
1991 Optimal Scheduling of Signature Analysis for VLSI Testing
abstract
A simple algorithm that shows how to optimally schedule the test-application and the signature-analysis phases of VLSI testing is presented. The testing process is broken into subintervals, the signature is analyzed at the end of each subinterval, and future tests are aborted if the circuit is found to be faulty, thus saving test time. The mathematical proofs associated with the algorithm are given.>
Yann-Hang Lee, C. Mani Krishna 0001
IEEE Trans. Computers2
1990 An Adaptive Algorithm to Ensure Differential Service in aToken-Ring Network
abstract
A distributed and adaptive token-passing algorithm that can be used to maintain the values of a designated performance parameter (e.g. mean waiting time) at the hosts of a token-ring network at a prescribed ratio is presented. The algorithm is simple to implement. The authors demonstrate its application in keeping the mean waiting times at the various hosts in a prespecified ratio, in maintaining a designated throughput differential between different classes of customer, and in integrating voice and two classes of data on a token ring.>
Myungwhan Choi, C. Mani Krishna 0001
IEEE Trans. Computers2
1989 Optimal Dynamic Control of Resources in a Distributed System
abstract
1188-1198
Kang G. Shin, C. Mani Krishna 0001, Yann-Hang Lee
IEEE Trans. Software Eng.2
1988 A Parallel Algorithm for Colouring Graphs
abstract
A simple and effective parallel algorithm to color the nodes of a graph is presented. The algorithm is useful in its own right. It also provides a vehicle to demonstrate the tradeoffs between performance and speed implicit in many parallel algorithms. The algorithm is used when the parallel machine has at least as many processors as there are nodes to color.>
Inderpal S. Bhandari, C. Mani Krishna 0001, Daniel P. Siewiorek
ICDCS2
1988 Optimal Scheduling of Signature Analysis for VLSI Testing
abstract
A simple algorithm is presented which minimizes the mean testing time for VLSI circuits. By breaking up the testing process into subintervals, and analyzing the signature are the end of each subinterval, it is possible to abort future tests if the circuit is found to be faulty, thus saving test time. Subdivision of the test process also reduces the probability of aliasing, thus increasing the effective coverage of the signature analysis process. If the process is sufficiently subdivided, it may be possible to use the test results not only to determine if the circuit is faulty or not, but to diagnose the fault.>
C. Mani Krishna 0001, Yann-Hang Lee
ITC1
1988 On the Graceful Degradation of Phase-Locked Clocks
abstract
The use of phase-locked clocks to limit clock skews to fractions of the clock period while keeping the algorithm overhead very small is investigated. The number of clocks in the system required to ensure that up to m arbitrary failures can be tolerated with all the good clocks still in synchrony has been shown to be N>or=3m+1. It is shown here that, if N>
C. Mani Krishna 0001, Inderpal S. Bhandari
RTSS1
1988 Optimal Resource Control in Periodic Real-Time Environments
abstract
Three factors determine the optimum configuration of a multiprocessor at any epoch: the workload, the reward structure, and the state of the computer system. An algorithm is presented for the optimal (more realistically, quasi-optimal) configuration of such systems used in real-time applications with periodic reward rates and workloads. The algorithm is based on Markov decision theory. It is suggested that a change in the workload or the reward structure should be as powerful a motivation for reconfiguration as component failure. Such changes occur naturally over the course of operation: an example of an online transaction processing system with a workload and reward structure that has a period of a day is given.>
Kang G. Shin, C. Mani Krishna 0001, Yann-Hang Lee
RTSS2
1988 Optimal Design and Sequential Analysis of VLSI Testing Strategy
abstract
A method for determining the optimal testing period and measuring the production yield is discussed. With the increased complexity of VLSI circuits, testing has become more costly and time-consuming. The design of a testing strategy, which is specified by the testing period based on the coverage function of the testing algorithm, involves trading off the cost of testing and the penalty of passing a bad chip as good. The optimal testing period is first derived, assuming the production yield is known. Since the yield may not be known a priori, an optimal sequential testing strategy which estimates the yield based on ongoing testing results, which in turn determines the optimal testing period, is developed next. Finally, the optimal sequential testing strategy for batches in which N chips are tested simultaneously is presented. The results are of use whether the yield stays constant or varies from one manufacturing run to another.>
Philip S. Yu, C. Mani Krishna 0001, Yann-Hang Lee
IEEE Trans. Computers2
1987 VLSI Circuit Testing Using an Adaptive Optimization Model
abstract
The purpose of testing is to determine the correctness of the unit under test in come optimal way. One difficulty in meeting the optimality requirement is that the stochastic properties of the unit are usually unknown a priori. For instance, one might not know exactly the yield of a VLSI production line before one tests the chips made as a result. Given the probability of unit failure and the coverage of a test, the optimal test period is easy to obtain. However, the probability of failure is not usually known a priori. We there- fore develop an optimal sequential testing strategy which estimates the production yield based on ongoing test results, and then use it to determine the optimal test period.
Philip S. Yu, C. Mani Krishna 0001, Yann-Hang Lee
DAC2
1987 Processor Tradeoffs in Distributed Real-Time Systems
abstract
Optimizing the design of real-time distributed systems is important since the systems are frequently critical to life. This optimization is a difficult problem, and heuristics and designer judgment are called for in the process. The chief cause of the difficulty is the large number of parameters under the designer's control which impact performance and life-cycle cost. We study the interplay between the more important parameters in this paper using two objective measures, i. e., the mean cost and the probability of dynamic failure in [6], [10]. Among these are the processor burn-in time and processor replacement policy. A central feature of this work is a look at how the application requirements affect the optimality of the distributed systems; indeed, the application requirements are an integral part of the analysis.
C. Mani Krishna 0001, Kang G. Shin, Inderpal S. Bhandari
IEEE Trans. Computers1
1986 On Scheduling Tasks with a Quick Recovery from Failure
abstract
Multiprocessors used in life-critical real-time systems must recover quickly from failure. Part of this recovery consists of switching to a new task schedule that ensures that hard deadlines for critical tasks continue to be met. We present a dynamic programming algorithm that ensures that backup, or contingency, schedules can be efficiently embedded within the original, "primary" schedule to ensure that hard deadlines continue to be met in the face of up to a given maximum number of processor failures. Several illustrative examples are included.
C. Mani Krishna 0001, Kang G. Shin
IEEE Trans. Computers1
1985 The Processor Number-Power Tradeoff in a Class of Multiprocessors
Kang G. Shin, C. Mani Krishna 0001
ICDCS2
1985 Ensuring Fault Tolerance of Phase-Locked Clocks
abstract
Processors within a real-time multiprocessor system must be synchronized with as little overhead as possible. Although synchronization can be achieved via both software (e.g., interactive convergence and interactive consistency algorithms) and hardware (e.g., multistage synchronizers and phase-locked clocks), phase-locked clocks are most attractive due to their small overheads.
C. Mani Krishna 0001, Kang G. Shin, Ricky W. Butler
IEEE Trans. Computers1
1983 Performance Measures for Multiprocessor Controllers
C. Mani Krishna 0001, Kang G. Shin
Performance1
1983 Queueing analysis of a canonical model of real-time multiprocessors
abstract
Multiprocessors are beginning to be regarded increasingly favorably as candidates for controllers in critical real-time control applications such as aircraft. Their considerable tolerance of component failures together with their great potential for high throughput are contributory factors.
C. Mani Krishna 0001, Kang G. Shin
SIGMETRICS1
1980 A Distributed Microprocessor System for Controlling and Managing Fighter Aircraft
Kang G. Shin, C. Mani Krishna 0001
RTSS2