EDBT 2026 Demo / reviewers in the wild / expert
Krishna M. Kavi
dblp:k/KrishnaMKavi · also Krishna Kavi
· DBLP profile ↗
37ranked-venue papers
13as first author
7since 2021 · last 2025
0000-0003-1581-8166ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 9 first-author · 5 since 2021Software engineering, systems software and programming languages · 10 · 7 first-author · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CADOSys: Cache Aware Design Space Optimization for Spatial ML Accelerators
Ruihao Li 0002, Krishna M. Kavi, Gayatri Mehta, Neeraja J. Yadwadkar, Lizy Kurian John |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Performance Implications of Pipelining the Data Transfer in CPU-GPU Heterogeneous SystemsabstractDriven by the increasing demands of machine learning, heterogeneous systems combining CPUs and GPUs have emerged as the dominant architecture for parallel computing in recent years. To optimize memory management and data transfer between CPUs and GPUs, Nvidia GPUs have introduced unified virtual memory ( UVM ) and pinned memory ( PM ) over the last decade. UVM can avoid explicit memory copies and potentially overlap GPU kernel computations with CPU-GPU data transfer. PM ensures that data with high locality remains in the main memory, preventing it from being paged out. In addition to these two techniques, asynchronous memory copy ( Async Memcpy ) was introduced recently in Nvidia GPUs to improve the CPU-GPU pipeline further. By utilizing Async Memcpy , the data transfer from GPU global memory to shared memory can be overlapped with GPU computations, adding an additional stage to the CPU-GPU data transfer pipeline. A thorough performance analysis of how Async Memcpy affects the current UVM and PM CPU-GPU data transfer scheme is desired. In this article, we provide performance implications of the combined effect of UVM , PM , and Async Memcpy , exploring which applications benefit from which combination of these features. We implement all these features on a suite of 25 workloads, including microbenchmarks and realworld applications. We observe an average performance gain of 24% when utilizing UVM and a 34% gain when employing PM on realworld applications, compared to not applying any data transfer optimization techniques. The performance benefits of Async Memcpy vary across different workloads. For workloads featuring extensive shared memory usage and high compute density (e.g., kmeans and lud ), Async Memcpy delivers around a 20% performance improvement over using UVM or PM alone. In other workloads like knn , we note a 20% performance degradation when using Async Memcpy . Furthermore, we conduct an in-depth investigation of the GPU kernel using performance counters to uncover the root causes of performance differences among various data transfer models. We also perform sensitivity analyses to examine how the number of blocks and threads, as well as the L1-cache/shared memory partitioning, impact performance. We explore future research directions aimed at enhancing the data transfer pipeline by overlapping memory allocation with data transfer and computation across GPU kernels. Ruihao Li 0002, Bagus Hanindhito, Sanjana Yadav, Qinzhe Wu, Krishna M. Kavi, Gayatri Mehta, Neeraja J. Yadwadkar, Lizy Kurian John |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | SecurityCloak: Protection against cache timing and speculative memory access attacks
Fernando Mosquera, Ashen Ekanayake, William Hua, Krishna M. Kavi, Gayatri Mehta, Lizy Kurian John |
J. Syst. Archit. | 4 |
| 2023 | NextGen-Malloc: Giving Memory Allocator Its Own Room in the HouseabstractMemory allocation and management have a significant impact on performance and energy of modern applications. We observe that performance can vary by as much as 72% in some applications based on which memory allocator is used. Many current allocators are multi-threaded to support concurrent allocation requests from different threads. However, such multi-threading comes at the cost of maintaining complex metadata that is tightly coupled and intertwined with user data. When memory management functions and other user programs run on the same core, the metadata used by management functions may pollute the processor caches and other resources. Ruihao Li 0002, Qinzhe Wu, Krishna M. Kavi, Gayatri Mehta, Neeraja J. Yadwadkar, Lizy Kurian John |
HotOS | 3 |
| 2022 | Sparse-T: Hardware Accelerator Thread for Unstructured Sparse Data ProcessingabstractSparse matrix-dense vector (SpMV) multiplication is inherent in most scientific, neural networks and machine learning algorithms. To efficiently exploit sparsity of data in SpMV computations, several compressed data representations have been used. However, compressed data representations of sparse data can result in overheads of locating nonzero values, requiring indirect memory accesses which increases instruction count and memory access delays. We call these translations of compressed representations as metadata processing. We propose a memory-side accelerator for metadata (or indexing) computations and supplying only the required nonzero values to the processor, additionally permitting an overlap of indexing with core computations on nonzero elements. In this contribution, we target our accelerator for low-end micro-controllers with very limited memory and processing capabilities. In this paper we will explore two dedicated ASIC designs of the proposed accelerator that handles the indexed memory accesses for compressed sparse row (CSR) format working alongside a simple RISC-like programmable core. One version of the accelerator supplies only vector values corresponding to nonzero matrix values and the second version supplies both nonzero matrix and matching vector values for SpMV computations. Our experiments show speedups ranging between 1.3 and 2.1 times for SpMV for different levels of sparsity. Our accelerator also results in energy savings ranging between 15.8% and 52.7% over different matrix sizes, when compared to the baseline system with primary RISC-V core performing all computations. We use smaller synthetic matrices with different sparsity levels and larger real-world matrices with higher sparsity (below 1% non-zeros) in our experimental evaluations. Pranathi Vasireddy, Krishna M. Kavi, Gayatri Mehta |
ICCAD | 2 |
| 2022 | Memory-Side Acceleration and Sparse Compression for Quantized Packed ConvolutionsabstractNeural Network compression techniques, such as parameter quantization and weight pruning have made deep neural network (DNN) inference more efficient for low-power devices such as MCUs and edge devices by reducing the memory and computation overhead required with minimal impact on model accuracy. To avoid storing and computing zeros, these techniques necessitate the use of sparse data representations, which introduces execution overhead to locate values required by a computation. Sparse matrix formats like Compressed Sparse Row (CSR) and other more recent designs are computationally inefficient when applied to the convolution algorithm as well as inefficient for storing quantized values. In this paper, we outline an intuitive extension of CSR called Partitioned Sparse Representation (PSR) in conjunction with a convolution algorithm that hides the cost of indexing overhead via a simple memory-side RISC-like core. PSR divides the entire weight array for a convolution layer into partitions that allow for smaller (e.g., 8-bit) indexes to reduce storage overhead. We also rely on a memory-side accelerator called HHT, a programmable, near-memory RISC-like co-processor that enables efficient processing of sparse data (including PSR). We show that HHT together with PSR allows the CPU to maximize the advantage of RISC-V packed instructions on sparse quantized data. We show as much as 10x speedup for sparse CONV with HHT over a baseline of the CPU performing all computations on dense data. HHT performs 2.7x faster on end-to-end image classification inference over the baseline and achieves 70% energy savings over sparse CONV with CPU performing all computations. Alex Weaver, Krishna M. Kavi, Pranathi Vasireddy, Gayatri Mehta |
SBAC-PAD | 2 |
| 2021 | Dynamically Adapting Page Migration Policies Based on Applications' Memory Access BehaviorsabstractThere have been numerous studies on heterogeneous memory systems comprised of faster DRAM (e.g., 3D stacked HBM or HMC) and slower non-volatile memories (e.g., PCM, STT-RAM). However, most of these studies focused on static policies for managing data placement and migration among the different memory devices. These policies are based on the average behavior across a range of applications. Results show that these techniques do not always result in higher performance when compared to systems that do not migrate data across the devices: some applications show performance gains, but other applications show performance losses. It is possible to utilize offline analyses to identify which applications benefit from page migration (migration friendly) and use page migration only with those applications. However, we observed that several applications exhibit both migration friendly and migration unfriendly behaviors during different phases of execution supporting a need for adaptive page migration techniques. We introduce and evaluate techniques that dynamically adapt to the behavior of applications and either reduce or increase migrations, or even halt migrations. Our adaptive techniques show performance gains for both migration friendly (on average of 81% over no migrations) and unfriendly workloads (by an average of 3%): it should be remembered that previous migration techniques resulted in performance losses for unfriendly workloads. Shashank Adavally, Mahzabeen Islam, Krishna M. Kavi |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | On-the-fly Page Migration and Address Reconciliation for Heterogeneous Memory SystemsabstractFor efficient placement of data in flat-address heterogeneous memory systems consisting of fast (e.g., 3D-DRAM) and slow memories (e.g., NVM), we present a hardware-based page migration technique. Unlike epoch-based approaches that migrate heavily accessed (“hot”) pages from slow to fast memories at each epoch interval, we migrate a page immediately when it becomes hot (“on-the-fly”), using hardware in user-transparent manner and with minimal OS intervention. The management of physical addresses due to page relocation becomes cumbersome and requires costly OS intervention. We use a small hardware remap table to keep track of new physical addresses of the migrated pages. This limits address reconciliation to occur only at periodic evictions of old remap entries. Also, we propose a hardware-orchestrated light-weight address reconciliation process. For our studied heterogeneous memory system, on-the-fly page migration with hardware-assisted address reconciliation provides 74% and 24% IPC improvements, on average for a set of SPEC CPU2006 workloads when compared to a baseline without any page migration and a system with on-the-fly page migration using OS-based address reconciliation, respectively. Furthermore, we present an analytical model for classifying applications as page migration friendly (applications that show performance gains from page migration) or unfriendly based on memory access behavior. Mahzabeen Islam, Shashank Adavally, Marko Scrbak, Krishna M. Kavi |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2017 | Exploring the Processing-in-Memory design space
Marko Scrbak, Mahzabeen Islam, Krishna M. Kavi, Mike Ignatowski, Nuwan Jayasena |
J. Syst. Archit. | 3 |
| 2015 | Recycling trash in cacheabstractThe disparity between processing and storage speeds can be bridged in part by reducing the traffic into and out of the slower memory components. Some recent studies reduce such traffic by determining dead data in cache, showing that a significant fraction of writes can be squashed before they make the trip toward slower memory. In this paper, we examine a technique for eliminating traffic in the other direction, specifically the traffic induced by dynamic storage allocation. We consider recycling dead storage in cache to satisfy a program's storage-allocation requests. We first evaluate the potential for recycling under favorable circumstances, where the associated logic can run at full speed with no impact on the cache's normal behavior. We then consider a more practical implementation, in which the associated logic executes independently from the cache's critical path. Here, the cache's performance is unfettered by recycling, but the operations necessary to determine dead storage and recycle such storage execute as time is available. Finally, we present the design and analysis of a hardware implementation that scales well with cache size without sacrificing too much performance. Jonathan A. Shidal, Ari J. Spilo, Paul T. Scheid, Ron Cytron, Krishna M. Kavi |
ISMM | 5 |
| 2015 | Memory organizations for 3D-DRAMs and PCMs in processor memory hierarchy
Krishna M. Kavi, Stefano Pianelli, Giandomenico Pisano, Giuseppe Regina, Mike Ignatowski |
J. Syst. Archit. | 1 |
| 2011 | Evaluation of Techniques to Improve Cache Access UniformitiesabstractWhile higher associativities are common at L-2 or Last-Level cache hierarchies, direct-mapped and low associative caches are still used at L-1 level. Lower associativities result in higher miss rates, but have fast access times on hits. Another issue that inhibits cache performance is the non-uniformity of accesses exhibited by most applications: some sets are underutilized while others receive the majority of accesses. Higher associative caches mitigate access non-uniformities, but do not eliminate them. This implies that increasing the size of caches or associativities may not lead to proportionally improved cache hit rates. Several solutions have been proposed in the literature over the past decade to address the non-uniformity of accesses, and each proposal independently claims improvements. However, because the published results use different benchmarks and different experimental setups, it is not easy to compare them. In this paper we report a side-by-side comparison of these techniques. The conclusion of our work is that, each application may benefit from a different technique and no single scheme works universally well for all applications. Our research is investigating the use of multiple techniques within a processor core and across cores in multicore system to improve the performance of cache memory hierarchies. The study reported in this paper allows us to select best possible solutions for each running application. In this paper, we have included some preliminary results of using multiple solutions simultaneously when running multiple threads. Izuchukwu Nwachukwu, Krishna M. Kavi, Ademola Fawibe, Chris Yan |
ICPP | 2 |
| 2011 | Parabilis: Speeding up Single-Threaded Applications by Extracting Fine-Grained Threads for Multi-core ExecutionabstractThe trend in architectural designs has been towards using simple cores for building multicore chips, instead of a single complex out-of-order (OOO) core, due to the increased complexity and energy requirements of out of order processors. Multicore chips provide better performance when compared with OOO cores while executing parallel applications. However, they are not able to exploit the parallelism inherent in single threaded applications. To this end, this paper presents a compiler optimization methodology coupled with minimal hardware extensions to extract simple fine-grained threads from a single-threaded application, for execution on multiple cores of a chip multiprocessor (CMP). These fine-grained threads are independent and eliminate the need for communication between cores, reducing costly communication latencies. This approach, which we call Parabilis is scalable for up to eight cores, and does not require complex hardware additions to simple multicore systems. Our evaluation shows that Parabilis yields an average speedup of 1.51 on an 8-core CMP architecture. Ademola Fawibe, Oghenekarho Okobiah, Oleg Garitselov, Krishna M. Kavi, Izuchukwu Nwachukwu, Mohana Asha Latha Dubasi, Vinay R. Prabhu |
ISPDC | 4 |
| 2008 | A Non-blocking Multithreaded Architecture with Support for Speculative Threads
Krishna M. Kavi, Wentong Li 0003, Ali R. Hurson |
ICA3PP | 1 |
| 2007 | Feasibility of decoupling memory management from the execution pipeline
Wentong Li 0003, Mehran Rezaei, Krishna M. Kavi, Afrin Naz, Philip H. Sweany |
J. Syst. Archit. | 3 |
| 2006 | Intelligent memory manager: Reducing cache pollution due to memory management functions
Mehran Rezaei, Krishna M. Kavi |
J. Syst. Archit. | 2 |
| 2003 | An Unfolding-Based Loop Optimization Technique
Litong Song, Krishna M. Kavi, Ron Cytron |
SCOPES | 2 |
| 2002 | Visual requirement representation
Deng-Jyi Chen, Wu-Chi Chen, Krishna M. Kavi |
J. Syst. Softw. | 3 |
| 2001 | Scheduled Dataflow: Execution Paradigm, Architecture, and Performance EvaluationabstractIn this paper, the scheduled dataflow (SDF) architecture-a decoupled memory/execution, multithreaded architecture using nonblocking threads-is presented in detail and evaluated against superscalar architecture. Recent focus in the field of new processor architectures is mainly on VLIW (e.g., IA-64), superscalar, and superspeculative designs. This trend allows for better performance, but at the expense of increased hardware complexity and, possibly, higher power expenditures resulting from dynamic instruction scheduling. Our research deviates from this trend by exploring a simpler, yet powerful execution paradigm that is based on dataflow and multithreading. A program is partitioned into nonblocking execution threads. In addition, all memory accesses are decoupled from the thread's execution. Data is preloaded into the thread's context (registers) and all results are poststored after the completion of the thread's execution. While multithreading and decoupling are possible with control-flow architectures, SDF makes it easier to coordinate the memory accesses and execution of a thread, as well as eliminate unnecessary dependencies among instructions. We have compared the execution cycles required for programs on SDF with the execution cycles required by programs on SimpleScalar (a superscalar simulator) by considering the essential aspects of these architectures in order to have a fair comparison. The results show that SDF architecture can outperform the superscalar. SDF performance scales better with the number of functional units and allows for a good exploitation of Thread Level Parallelism (TLP) and available chip area. Krishna M. Kavi, Roberto Giorgi, Joseph Arul |
IEEE Trans. Computers | 1 |
| 2000 | Multimedia File Allocation on VC Networks Using Multipath RoutingabstractThe problem of allocating high-volume multimedia files on a virtual circuit network with the objective of maximizing channel throughput (and minimizing data transmission time) is addressed. The problem is formulated as a multicommodity flow problem. We present both the optimal and suboptimal solutions to the problem using novel approaches. Pao-Yuan Chang, Deng-Jyi Chen, Krishna M. Kavi |
IEEE Trans. Computers | 3 |
| 1998 | Design of cache memories for dataflow architecture
Krishna M. Kavi, Ali R. Hurson |
J. Syst. Archit. | 1 |
| 1998 | Cyclic Staggered Scheme: A Loop Allocation Policy for DOACROSS LoopsabstractWithin the scope of the multithreaded dataflow, the problem of scheduling/allocation of DOACROSS loops has been discussed and it was shown that the so called staggered allocation offers higher performance and resource utilization than other schemes described in the literature. The staggered scheme, however, produces an unbalanced load among processors. The paper introduces an extension to the staggered scheme-cyclic staggered scheme-that produces a more balanced distribution of iterations among processors. The cyclic staggered scheme is simulated and its performance improvement is analyzed. Ali R. Hurson, Krishna M. Kavi, Joford T. Lim |
IEEE Trans. Computers | 2 |
| 1996 | Specification and Analysis of Real-Time Systems Using CSP and Petri NetsabstractFormal methods such as CSP (Communicating Sequential Processes) are widely used for reasoning about concurrency, communication, safety, and liveness issues. Some of these models have been extended to permit reasoning about real-time constraints. Yet, the research in formal specification and verification of complex systems has often ignored the specification of stochastic properties of the system under study. We are developing methods and tools to permit stochastic analyses of CSP-based specifications. Our basic objective is to evaluate candidate design specifications by converting formal systems descriptions into the information needed for analysis. In doing so, we translate a CSP-based specification into a Petri net which is analyzed to predict system behavior in terms of reliability and performability as a function of observable parameters (e.g., topology, fault-tolerance, deadlines, communications, and failure categories). This process can give insight into further refinements of the original specification (i.e., identify potential failure processes and recovery actions). Relating the parameters needed for performability analysis to user level specifications is essential for realizing systems that meet user needs in terms of cost, functionality, and other nonfunctional requirements. An example translation is given (in addition, some general examples of CSP → Petri net translations can be viewed in Appendix A). Based on this translation, we report both the discrete and continuous time Markovian analysis which provides reliability predictions for the candidate specification. The term “CSP-based” is used here to distinguish between the notation of Hoare’s original CSP and our textual representations which are similar to occum. Our CSP-based grammar does not restrict consideration of the properties of CSP (traces, refusal sets, livelock, etc.), but we are not considering those properties. We are only interested that the structural properties are preserved. We define performability as a measure of the system’s ability in meeting deadlines, in the presence of failures and variance in task execution times. Krishna M. Kavi, Frederick T. Sheldon, Sherman Reed |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 1995 | Design of Cache Memories for Multi-Threaded Dataflow ArchitectureabstractCache memories have proven their effectiveness in the von Neumann architecture when localities of reference govern the execution loci of programs. A pure dataflow program, in contrast, contains no locality of reference since the execution sequence is enforced only by the availability of arguments. Instruction locality may be enhanced if, dataflow programs are reordered. Enhancing the locality of data references in the dataflow architecture is a more challenging problem. In this paper we report our approaches to the design of instruction, data (operand) and I-Structure cache memories using the Explicit Token Store (ETS) model of dataflow systems. We will present the performance results obtained using various benchmark programs. Krishna M. Kavi, Ali R. Hurson, Phenil Patadia, Elizabeth Abraham, Ponnarasu Shanmugam |
ISCA | 1 |
| 1994 | Specification of Stochastic Properties with CSPabstractThe research in formal specification and verification of complex systems has often ignored the specification of stochastic properties of the system. We are exploring new methodologies and tools to permit stochastic analysis of CSP-based systems specifications. In doing so, we have investigated the relationship between specification models and stochastic models by translating the specification into another form that is amenable to such analyses (e.g., from CSP to stochastic Petri Nets). This process can give insight for further refinements of the original specification (i.e., identify potential failure processes and recovery actions). It does this by relating the parameters needed for reliability analysis to user level specifications which is essential for realizing systems that meet the users needs in terms of cost, functionality, performance and reliability. Krishna M. Kavi, Frederick T. Sheldon |
ICPADS | 1 |
| 1993 | PARSA: A Parallel Program Scheduling and Assessment EnvironmentabstractEfficient partitioning and scheduling of parallel programs and the distribution of data among pmessing elements are very important issues in parallel and distributed systems. Existing tools fdl short in addressing the issues satisfactorily. On one hand, it is believed to k unreasonable to leave the burden of these complex tasks to the programmers. On the other hand, fully automated schedulers have ken shown to be d little practical significance, or suitable only for restricted cases. In this paper we address the issues and algorithms for efficient partitioning and scheduling d parallel programs, including the distribution of data, in dislributed-memory multiprocessor systems, using the PARSA parallel software development environment. PARSA consists of a set of visual, interactive, compiletime tools that will provide automated program partitions and schedules whenever possible, while permitting the user to exert conml over these operations for a htter performance. The support program assessment tool provides the users the opportunity b fine-tune the program and achieve their performance objectives. Behrooz A. Shirazi, Krishna M. Kavi, Ali R. Hurson, Prasenjit Biswas |
ICPP (2) | 2 |
| 1992 | An Efficient Data Interface for Heterogeneous Distributed EnvironmentsabstractA multi-language-based data interface system for heterogeneous distributed processing is introduced. A prototyped environment based on this system is discussed, and an evaluation of the prototyped system is presented. It is shown that by keeping the syntax of the specification language flexible and close to existing high-level languages, a user can learn the interface language quickly. Semantically, this data interface views structured data as consisting of two parts: the data values themselves and the representation of the structure among the data values. Through this separation, it is possible to have pipelined data type checking and data conversion operations.> David D. H. Lin, Behrooz A. Shirazi, Krishna M. Kavi |
ICDCS | 3 |
| 1992 | A New Cache Coherency and Address Translation Consistency Protocol
Jen-Tien Yen, Behrooz A. Shirazi, Krishna M. Kavi |
ICPP (1) | 3 |
| 1992 | Real-time systems design methodologies: An introduction and a survey
Krishna M. Kavi, Seung-Min Yang |
J. Syst. Softw. | 1 |
| 1991 | A Heterogeneous Distributed Processing Interface Specification Language
David D. H. Lin, Behrooz A. Shirazi, Krishna M. Kavi |
ICPP (2) | 3 |
| 1991 | Specification of concurrent processes using a dataflow model of computation and partially ordered events
Krishna M. Kavi, Akshay K. Deshpande |
J. Syst. Softw. | 1 |
| 1990 | An n-grid model for group authorizationabstractThe n-grid model for group authorization and access control extends the NTree representation of two-dimensional partial orders and incorporates the implicit authorizations of the MCC model. The n-grid is a representation of multi-dimensional partial orders, permitting the inclusion of relationships among user (subject) groups, object groups, and access-right groups. Each (unique) element in an authorization class is represented as a vector, facilitating efficient implementation. The model contains two parts: the access control part for mapping implicit authorizations onto an n-grid, and the propagation part for restricting a user's membership to a single subject group.> Wen-Gong Shieh, Bob P. Weems, Krishna M. Kavi |
ACSAC | 3 |
| 1989 | A review of specification and verification methods for parallel programs including the dataflow approachabstractParallel programs are usually described informally, and these descriptions are implemented on parallel computer systems. When a program does not run correctly, it is often very difficult to determine whether the program description or the implementation is incorrect. This has led to a search for more formal descriptions of parallel programs and to proof systems for the verification of the implementations. Formal methods for the specification and verification of parallel programs are reviewed, and a new method that is based on dataflow graphs is described.> Akshay K. Deshpande, Krishna M. Kavi |
Proc. IEEE | 2 |
| 1987 | Isomorphisms Between Petri Nets and Dataflow GraphsabstractDataflow graphs are a generalized model of computation. Uninterpreted dataflow graphs with nondeterminism resolved via probabilities are shown to be isomorphic to a class of Petri nets known as free choice nets. Petri net analysis methods are readily available in the literature and this result makes those methods accessible to dataflow research. Nevertheless, combinatorial explosion can render Petri net analysis inoperative. Using a previously known technique for decomposing free choice nets into smaller components, it is demonstrated that, in principle, it is possible to determine aspects of the overall behavior from the particular behavior of components. Krishna M. Kavi, Bill P. Buckles, U. Narayan Bhat |
IEEE Trans. Software Eng. | 1 |
| 1986 | A Formal Definition of Data Flow Graph ModelsabstractIn this paper, a new model for parallel computations and parallel computer systems that is based on data flow principles is presented. Uninterpreted data flow graphs can be used to model computer systems including data driven and parallel processors. A data flow graph is defined to be a bipartite graph with actors and links as the two vertex classes. Actors can be considered similar to transitions in Petri nets, and links similar to places. The nondeterministic nature of uninterpreted data flow graphs necessitates the derivation of liveness conditions. Krishna M. Kavi, Bill P. Buckles, U. Narayan Bhat |
IEEE Trans. Computers | 1 |
| 1984 | Message Repository Definitional Facility: An Architectural Model for Interprocess CommunicationabstractAn architectural model of a system research tool, the Message Repository Definitional Facility (MRDF), is presented. MRDF provides for interprocess communication and synchronization by message passing, and permits the user a wide selection of policy options, promising general descriptive power. The project's current levels of system definition, design considerations and issues, and plans for future development are briefly described. The current goal for this research is the implementation of the MRDF in the data flow simulation environment for test and evaluation. Potential application of this concept beyond the research tool stage is speculated as incorporation in a computer architecture to support concurrent processes and a “self-optimizing” operating system. Krishna M. Kavi, Edward W. Banios |
ISCA | 1 |
| 1982 | HLL architectures: Pitfalls and predilectionsabstractAn examination of high-level language architectures reveals that some design considerations are based on inaccurate premises. This paper discusses some HLL architecture design misconceptions. Some features that should be considered in the design of HLL architectures are presented. Based on the HLL architecture design considerations, a methodology for quantifying architecture is proposed. Krishna M. Kavi, Boumediene Belkhouche, Evelyn Bullard, Lois M. L. Delcambre, Stephen M. Nemecek |
ISCA | 1 |