Kemal Ebcioglu

dblp:65/6903 · DBLP profile ↗
← Back
36ranked-venue papers
13as first author
4since 2021 · last 2025
0000-0001-6256-4248ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 10 first-author · 4 since 2021Software engineering, systems software and programming languages · 9 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 1,024-FPGA DES Supercomputer on the AWS Cloud
abstract
ABSTRACT We present a 1,024‐FPGA DES supercomputer accelerator that is automatically compiled from a single‐threaded sequential DES key search application by means of our High‐Level Synthesis compiler. Our 1,024‐FPGA supercomputer is deployed on several Amazon Web Services (AWS) EC2 F1 instance platforms from different AWS regions. Consequently, it can be considered the first multi‐chip application‐specific supercomputer that is scattered to multiple geographically distributed data centers around the world. Furthermore, invoking our 1,024‐FPGA DES supercomputer is functionally identical to invoking the single‐threaded sequential DES application the supercomputer accelerator is compiled from. Our 1,024‐FPGA supercomputer achieves 3.016E+12 keys/sec and it performs 5,286,000 times better than an AWS EC2 m5.8xlarge Xeon x86 machine executing the original sequential application with a performance of 5.706 E+5 keys/sec.
Kemal Ebcioglu, Batuhan Bulut, Atakan Dogan, Gürhan Küçük, Ismail San
Concurr. Comput. Pract. Exp.1
2022 A parallel hardware hypervisor for hardware-accelerated cloud computing
abstract
Summary Hardware‐accelerated cloud computing systems based on FPGA or ASIC chips have proved useful in providing power‐efficient acceleration for a variety of software applications. However, these computing systems rely on operating systems and hypervisors, which not only are implemented with inefficient software, but which also are incapable of handling massively parallel systems, due to the lack of parallelism and scalability in their algorithmic designs. As a result, power, performance, and scalability problems will emerge in an exascale cloud computing environment. As a solution to these problems, the present study proposes a parallel hardware hypervisor system implemented entirely in special‐purpose hardware without resorting to energy‐inefficient general‐purpose processors. Furthermore, the proposed hypervisor system virtualizes application‐specific multi‐chip supercomputers, in order to enable the virtual supercomputers to share available FPGA and semi‐configurable ASIC resources in a cloud system. Single‐chip verification studies based on Verilog simulation have been done to verify the functional correctness of the proposed hardware hypervisor system, which consumes only a fraction of hardware resources. This article will in particular focus on the virtualization of multi‐chip message‐passing based supercomputers using limited reconfigurable hardware resources.
Atakan Dogan, Kemal Ebcioglu
Concurr. Comput. Pract. Exp.2
2022 Cloud Building Block Chip for Creating FPGA and ASIC Clouds
abstract
Hardware-accelerated cloud computing systems based on FPGA chips (FPGA cloud) or ASIC chips (ASIC cloud) have emerged as a new technology trend for power-efficient acceleration of various software applications. However, the operating systems and hypervisors currently used in cloud computing will lead to power, performance, and scalability problems in an exascale cloud computing environment. Consequently, the present study proposes a parallel hardware hypervisor system that is implemented entirely in special-purpose hardware, and that virtualizes application-specific multi-chip supercomputers, to enable virtual supercomputers to share available FPGA and ASIC resources in a cloud system. In addition to the virtualization of multi-chip supercomputers, the system’s other unique features include simultaneous migration of multiple communicating hardware tasks, and on-demand increase or decrease of hardware resources allocated to a virtual supercomputer. Partitioning the flat hardware design of the proposed hypervisor system into multiple partitions and applying the chip unioning technique to its partitions, the present study introduces a cloud building block chip that can be used to create FPGA or ASIC clouds as well. Single-chip and multi-chip verification studies have been done to verify the functional correctness of the hypervisor system, which consumes only a fraction of (10%) hardware resources.
Atakan Dogan, Kemal Ebcioglu
ACM Trans. Reconfigurable Technol. Syst.2
2022 Highly Parallel Multi-FPGA System Compilation from Sequential C/C++ Code in the AWS Cloud
abstract
We present a High Level Synthesis compiler that automatically obtains a multi-chip accelerator system from a single-threaded sequential C/C++ application. Invoking the multi-chip accelerator is functionally identical to invoking the single-threaded sequential code the multi-chip accelerator is compiled from. Therefore, software development for using the multi-chip accelerator hardware is simplified, but the multi-chip accelerator can exhibit extremely high parallelism. We have implemented, tested, and verified our push-button system design model on multiple field-programmable gate arrays (FPGAs) of the Amazon Web Services EC2 F1 instances platform, using, as an example, a sequential-natured DES key search application that does not have any DOALL loops and that tries each candidate key in order and stops as soon as a correct key is found. An 8- FPGA accelerator produced by our compiler achieves 44,600 times better performance than an x86 Xeon CPU executing the sequential single-threaded C program the accelerator was compiled from. New features of our compiler system include: an ability to parallelize outer loops with loop-carried control dependences, an ability to pipeline an outer loop without fully unrolling its inner loops, and fully automated deployment, execution and termination of multi-FPGA application-specific accelerators in the AWS cloud, without requiring any manual steps.
Kemal Ebcioglu, Ismail San
ACM Trans. Reconfigurable Technol. Syst.1
2007 Efficient Register Mapping and Allocation in LaTTe, an Open-Source Java Just-in-Time Compiler
Byung-Sun Yang, Junpyo Lee, SeungIl Lee, Seongbae Park, Yoo C. Chung, Suhyun Kim 0001, Kemal Ebcioglu, Erik R. Altman, Soo-Mook Moon
IEEE Trans. Parallel Distributed Syst.7
2005 X10: an object-oriented approach to non-uniform cluster computing
abstract
It is now well established that the device scaling predicted by Moore's Law is no longer a viable option for increasing the clock frequency of future uniprocessor systems at the rate that had been sustained during the last two decades. As a result, future systems are rapidly moving from uniprocessor to multiprocessor configurations, so as to use parallelism instead of frequency scaling as the foundation for increased compute capacity. The dominant emerging multiprocessor structure for the future is a Non-Uniform Cluster Computing (NUCC) system with nodes that are built out of multi-core SMP chips with non-uniform memory hierarchies, and interconnected in horizontally scalable cluster configurations such as blade servers. Unlike previous generations of hardware evolution, this shift will have a major impact on existing software. Current OO language facilities for concurrent and distributed programming are inadequate for addressing the needs of NUCC systems because they do not support the notions of non-uniform data access within a node, or of tight coupling of distributed nodes.We have designed a modern object-oriented programming language, X10, for high performance, high productivity programming of NUCC systems. A member of the partitioned global address space family of languages, X10 highlights the explicit reification of locality in the form of places}; lightweight activities embodied in async, future, foreach, and ateach constructs; a construct for termination detection (finish); the use of lock-free synchronization (atomic blocks); and the manipulation of cluster-wide global data structures. We present an overview of the X10 programming model and language, experience with our reference implementation, and results from some initial productivity comparisons between the X10 and Java™ languages.
Philippe Charles, Christian Grothoff, Vijay A. Saraswat, Christopher Donawa, Allan Kielstra, Kemal Ebcioglu, Christoph von Praun, Vivek Sarkar
OOPSLA6
2005 Programming by sketching for bit-streaming programs
Armando Solar-Lezama, Rodric M. Rabbah, Rastislav Bodík, Kemal Ebcioglu
PLDI4
2005 Selective sweeping
abstract
Traditional mark and sweep garbage collectors use time proportional to the heap size when sweeping memory, since all objects in the heap, dead or alive, must be traversed. Here we introduce a sweeping algorithm which traverses only the live objects. Since this sweeping algorithm is slower when the heap occupancy is high, we also discuss how to avoid this slowdown by using an adaptive algorithm. Copyright © 2004 John Wiley & Sons, Ltd.
Yoo C. Chung, Soo-Mook Moon, Kemal Ebcioglu, Dan Sahlin
Softw. Pract. Exp.3
2005 Lightweight monitors for the Java virtual machine
abstract
Java supports the monitor construct for language-level synchronization in the context of multi-threading. This paper introduces the lightweight monitor, an efficient user-level monitor implementation. The lightweight monitor is useful for single-threaded Java programs as well as for multi-threaded Java programs with little lock contention. A 32-bit lock is embedded in each object for efficient lock access while other monitor data structures are managed using a hash table. We highly optimized the lock manipulation code, which is translated and inlined by a just-in-time (JIT) compiler. In the most probable cases, only nine SPARC instructions are spent for lock acquisition and five instructions are spent for lock release. Our experimental results indicate that the lightweight monitor is faster than the monitor implementation in the SUN JDK 1.2 RC1 by up to 21 times in the absence of lock contention and by up to seven times in the presence of lock contention. Copyright © 2004 John Wiley & Sons, Ltd.
Byung-Sun Yang, Soo-Mook Moon, Kemal Ebcioglu
Softw. Pract. Exp.3
2004 Topic 8: Parallel Computer Architecture and Instruction-Level Parallelism
Kemal Ebcioglu, Wolfgang Karl, André Seznec, Marco Aldinucci
Euro-Par1
2002 A Register File Architecture and Compilation Scheme for Clustered ILP Processors
Krishnan Kailas, Manoj Franklin, Kemal Ebcioglu
Euro-Par3
2002 Unroll-Based Copy Elimination for Enhanced Pipeline Scheduling
abstract
Enhanced pipeline scheduling (EPS) is a software pipelining technique which can achieve a variable initiation interval (II) for loops with control flow via its code motion pipelining. EPS, however, leaves behind many renaming copy instructions that cannot be coalesced due to interferences. These copies take resources and, more seriously, they may cause a stall if they rename a multilatency instruction whose latency is longer than the II aimed for by EPS. This paper proposes a code transformation technique based on loop unrolling which makes those copies coalescible. Two unique features of the technique are its method of determining the precise unroll amount, based on an idea of extended live ranges, and its insertion of special bookkeeping copies at loop exits. The proposed technique enables EPS to avoid a serious slowdown from latency handling and resource pressure, while keeping its variable II and other advantages. In fact, renaming through copies, followed by unroll-based copy elimination, is EPS's solution to the cross-iteration register overwrite problem in software pipelining. It works for loops with arbitrary control flow that EPS must deal with, as well as for straightline loops. Our empirical study performed on a VLIW testbed with a two-cycle load latency shows that 86 percent of the otherwise uncoalescible copies in innermost loops become coalescible when unrolled 2.2 times on average. In addition, it is demonstrated that the unroll amount obtained is precise and the most efficient. The unrolled version of the VLIW code includes fewer no-op VLIW caused by stalls, improving the performance by a geometric mean of 18 percent on a 16-ALU machine.
Suhyun Kim 0001, Soo-Mook Moon, Jinpyo Park, Kemal Ebcioglu
IEEE Trans. Computers4
2001 CARS: A New Code Generation Framework for Clustered ILP Processors
abstract
Clustered ILP processors are characterized by a large number of non-centralized on-chip resources grouped into clusters. Traditional code generation schemes for these processors consist of multiple phases for cluster assignment, register allocation and instruction scheduling. Most of these approaches need additional re-scheduling phases because they often do not impose finite resource constraints in all phases of code generation. These phase-ordered solutions have several drawbacks, resulting in the generation of poor performance code. Moreover the iterative/back-tracking algorithms used in some of these schemes have large turning times. In this paper we present CARS, a code generation framework for Clustered ILP processors, which combines the cluster assignment, register allocation, and instruction scheduling phases into a single code generation phase, thereby eliminating the problems associated with phase-ordered solutions. The CARS algorithm explicitly takes into account all the resource constraints at each cluster scheduling step to reduce spilling and to avoid iterative re-scheduling steps. We also present a new on-the-fly register allocation scheme developed for CARS. We describe an implementation of the proposed code generation framework and the results of a performance evaluation study using the SPEC95/2000 and MediaBench benchmarks.
Krishnan Kailas, Kemal Ebcioglu, Ashok K. Agrawala
HPCA2
2001 Advances and future challenges in binary translation and optimization
abstract
Binary translation and optimization have achieved a high profile in recent years. Binary translation has several potential attractions. While still in its early stages, could binary translation offer a new way to design processors, i.e. is it a disruptive technology? This paper discusses this question, examines some future possibilities for binary translation, and then gives an overview of selected projects (DAISY, Crusoe, Dynamo and LaTTe). One future possibility for binary translation is the Virtual IT Shop. Binary translation offers a possible solution for better utilization of computational resources as services over the World Wide Web. The Internet is radically changing the software landscape, and is fostering platform independence and interoperability. Along the lines of software convergence, recent advances in binary JIT (just-in-time) optimizations also present the future possibility of a convergence virtual machine (CVM). CVM aims to address research challenges in allowing the same standard operating system and application object code to run on different hardware platforms, through state-of-the-art JIT compilation and virtual device emulation.
Erik R. Altman, Kemal Ebcioglu, Michael Gschwind, Sumedh Sathaye
Proc. IEEE2
2001 Dynamic Binary Translation and Optimization
abstract
We describe a VLIW architecture designed specifically as a target for dynamic compilation of an existing instruction set architecture. This design approach offers the simplicity and high performance of statically scheduled architectures, achieves compatibility with an established architecture, and makes use of dynamic adaptation. Thus, the original architecture is implemented using dynamic compilation, a process we refer to as DAISY (Dynamically Architected Instruction Set from Yorktown). The dynamic compiler exploits runtime profile information to optimize translations so as to extract instruction level parallelism. This paper reports different design trade-offs in the DAISY system and their impact on final system performance. The results show high degrees of instruction parallelism with reasonable translation overhead and memory usage.
Kemal Ebcioglu, Erik R. Altman, Michael Gschwind, Sumedh W. Sathaye
IEEE Trans. Computers1
2000 Instruction-Level Parallelism and Processor Architecture
Kemal Ebcioglu
Euro-Par1
2000 Binary translation and architecture convergence issues for IBM system/390
abstract
We describe the design issues in an implementation of the ESA/390 architecture based on binary translation to a very long instruction word (VLIW) processor. During binary translation, complex ESA/390 instructions are decomposed into instruction “primitives” which are then scheduled onto a wide-issue machine. The aim is to achieve high instruction level parallelism due to the increased scheduling and optimization opportunities which can be exploited by binary translation software, combined with the efficiency of long instruction word architectures. A further aim is to study the feasibility of a common execution platform for different instruction set architectures, such as ESA/390, RS?6000, AS/400 and the Java Virtual Machine, so that multiple systems can be built around a common execution platform.
Michael Gschwind, Kemal Ebcioglu, Erik R. Altman, Sumedh W. Sathaye
ICS2
2000 Unroll-based register coalescing
abstract
Aggressive instruction scheduling leaves behind many renaming copy instructions that cannot be coalesced due to interferences. These copies take resources, and more seriously, they may cause a stall if they are generated for renaming of multi-latency instructions. This paper proposes a code transformation technique based on loop unrolling which makes those copies coalescible. Two unique features of the technique are its method of determining the precise unroll amount based on an idea of extended live range, and its insertion of special bookkeeping copies at loop exits. In fact, the technique provides a more general and simpler solution for the cross-iteration register overwrite problem in software pipelining which works for loops with control flows as well as for straight-line loops. In addition, it is applicable to other optimizations including path length reduction and redundant subscripted reference elimination.
Suhyun Kim 0001, Soo-Mook Moon, Jinpyo Park, Kemal Ebcioglu
ICS4
2000 Reducing Sweep Time for a Nearly Empty Heap
abstract
Mark and sweep garbage collectors are known for using time proportional to the heap size when sweeping memory, since all objects in the heap, regardless of whether they are live or not, must be visited in order to reclaim the memory occupied by dead objects. This paper introduces a sweeping method which traverses only the live objects, so that sweeping can be done in time dependent only on the number of live objects in the heap.
Yoo C. Chung, Soo-Mook Moon, Kemal Ebcioglu, Dan Sahlin
POPL3
1999 Execution-Based Scheduling for VLIW Architectures
Kemal Ebcioglu, Erik R. Altman, Sumedh W. Sathaye, Michael Gschwind
Euro-Par1
1999 Optimizations and Oracle Parallelism with Dynamic Translation
abstract
We describe several optimizations which can be employed in a dynamic binary translation (DBT) system, where low compilation/translation overhead is essential. These optimizations achieve a high degree of ILP, sometimes even surpassing a static compiler employing more sophisticated, and more time-consuming algorithms. We present results in which we employ these optimizations in a dynamic binary translation system capable of computing oracle parallelism.
Kemal Ebcioglu, Erik R. Altman, Sumedh W. Sathaye, Michael Gschwind
MICRO1
1998 An eight-issue tree-VLIW processor for dynamic binary translation
abstract
Presented is an 8-issue tree-VLIW processor designed for efficient support of dynamic binary translation. This processor confronts two primary problems faced by VLIW architectures: binary compatibility and branch performance. Binary compatibility with existing architectures is achieved through dynamic binary translation which translates and schedules PowerPC instructions to take advantage of the available instruction level parallelism. Efficient branch performance is achieved through tree instructions that support multi-way path and branch selection within a single VLIW instruction. The processor architecture is described, along with design details of the branch unit, pipeline, register file and memory hierarchy for a 0.25 micron standard-cell design. Performance simulations show that the simplicity of a VLIW architecture allows a wide-issue processor to operate at high frequencies.
Kemal Ebcioglu, Jason Fritts, Stephen V. Kosonocky, Michael Gschwind, Erik R. Altman, Krishnan Kailas, Terry Bright
ICCD1
1998 The Performance Impact of Exploiting Branch ILP with Tree Representation of ILP Code
abstract
Modern single-CPU microprocessors exploit instruction-level parallelism (ILP) by deriving their performance advantage mainly from parallel execution of ALU and memory instructions within a single clock cycle. This performance advantage obtained by exploiting data ILP is severely offset by sequential execution of conditional branches, especially in branch-intensive non-numerical code. Consequently, branch ILP must also be exploited by executing branches and data instructions in parallel. This requires compilation support for scheduling branches as well as architectural support for executing branches and data instructions in the same cycle. This paper performs a comprehensive empirical study aimed at evaluating the performance impact of exploiting branch ILP using a representation of ILP code called tree representation, which has been proposed by Nicolau [A. Nicolau (1985), Technical Report TR-85-678, Cornell University, Ithaca, NY] and Ebcioğlu to exploit branch ILP in the most generalized form. Our results indicate that exploiting branch ILP can enhance performance substantially (i.e., as much as a geometric mean of speedup 4.5 in the 16-ALU machine, compared to the base speedup 3.0) and that the performance benefit comes not only from the intended parallel execution but from the decrease of useless speculative execution due to earlier scheduling of branches.
Soo-Mook Moon, Kemal Ebcioglu
Comput. J.2
1997 Performance Analysis of Tree VLIW Architecture for Exploiting Branch ILP in Non-Numerical Code
abstract
In order to fully exploit instruction-level parallelism (ILP) in non-numerical code, we must exploit branch ILP as well as data ILP.Exploiting branch ILP requires architectural support for executing branches and data instructions in parallel, and the compiler needs to schedule conditional branches.As a VLIW architecture that exploits branch ILP in the most generalized form, we have proposed the tree VLIW architecture [ 1], which exhibits significant performance advantage when combined with appropriate scheduling techniques [2].This paper analyzes the performance advantage, characterizing the performance impact of the tree VLIW architecture.We also provide pertinent insights into its two architectural features (generalized multi-way branching and conditional execution) and describe implementation details of the tree VLIW machine.Our analysis indicates that the performance benefit of the tree VLIW architecture comes not only from the intended branch ILP but from the improvement of data ILP caused by the decrease of useless speculative execution.
Soo-Mook Moon, Kemal Ebcioglu
International Conference on Supercomputing2
1997 DAISY: Dynamic Compilation for 100% Architectural Compatibility
abstract
Although VLIW architectures offer the advantages of simplicity of design and high issue rates, a major impediment to their use is that they are not compatible with the existing software base. We describe new simple hardware features for a VLIW machine we call DAISY (DynamicallyArchitectedInstructionSet fromYorktown). DAISY is specifically intended to emulate existing architectures, so that all existing software for an old architecture (including operating system kernel code) runs without changes on the VLIW. Each time a new fragment of code is executed for the first time, the code is translated to VLIW primitives, parallelized and saved in a portion of main memory not visible to the old architecture, by a Virtual Machine Monitor (software) residing in read only memory. Subsequent executions of the same fragment do not require a translation (unless cast out). We discuss the architectural requirements for such a VLIW, to deal with issues including self-modifying code, precise exceptions, and aggressive reordering of memory references in the presence of strong MP consistency and memory mapped I/O. We have implemented the dynamic parallelization algorithms for the PowerPC architecture. The initial results show high degrees of instruction level parallelism with reasonable translation overhead and memory usage.
Kemal Ebcioglu, Erik R. Altman
ISCA1
1997 Parallelizing Nonnumerical Code with Selective Scheduling and Software Pipelining
abstract
Instruction-level parallelism (ILP) in nonnumerical code is regarded as scarce and hard to exploit due to its irregularity. In this article, we introduce a new code-scheduling technique for irregular ILP called “selective scheduling” which can be used as a component for superscalar and VLIW compilers. Selective scheduling can compute a wide set of independent operations acrossallexecution paths based on renaming and forward-substitution and can compute available operations across loop iterations if combined with software pipelining. This scheduling approach has better heuristics for determining the usefulness of moving one operation versus moving another and can successfully find useful code motions without resorting to branch profiling. The compile-time overhead of selective scheduling is low due to its incremental computation technique and its controlled code duplication. We parallelized the SPEC integer benchmarks and five AIX utilities without using branch probabilities. The experiments indicate that a fivefold speedup is achievable on realistic resources with a reasonable overhead in compilation time and code expansion and that a solid speedup increase is also obtainable on machines with fewer resources. These results improve previously known characteristics of irregular ILP.
Soo-Mook Moon, Kemal Ebcioglu
ACM Trans. Program. Lang. Syst.2
1994 VLIW Compilation Techniques in a Superscalar Environment
abstract
We describe techniques for converting the intermediate code representation of a given program, as generated by a modern compiler, to another representation which produces the same run-time results, but can run faster on a superscalar machine. The algorithms, based on novel parallelization techniques for Very Long Instruction Word (VLIW) architectures, find and place together independently executable operations that may be far apart in the original code. i.e., they may be separated by many conditional branches or belong to different iterations of a loop. As a result, the functional units in the superscalar are presented with more work that can proceed in parallel, thus achieving higher performance than the approach of using hardware instruction dispatch techniques alone.While general scheduling techniques improve performance by removing idle pipeline cycles, to further improve performance on a superscalar with only a few functional units requires a reduction in the pathlength. We have designed a set of new algorithms for reducing pathlength and removing stalls due to branches, namely speculative load-store motion out of loops, unspeculation, limited combining, basic block expansion, and prolog tailoring. These algorithms were implemented in a prototype version of the IBM RS/6000 xlc compiler and have shown significant improvement in SPEC integer benchmarks on the IBM POWER machines.Also, we describe a new technique to obtain profiling information with low overhead, and some applications of profiling directed feedback, including scheduling heuristics, code reordering and branch reversal.
Kemal Ebcioglu, Randy D. Groves, Ki-Chang Kim, Gabriel M. Silberman, Isaac Ziv
PLDI1
1993 On Performance, Efficiency of VLIW and Superscalar
abstract
Instruction-level parallelism in non-numerical code character as leading to small speeduo (as little) due to its irregularity. Recently, we have developed a new static scheduling algorithm called selective scheduling which can be used as a component of VLIW and superscalar compilers to exploit the irregular parallelism.
Soo-Mook Moon, Kemal Ebcioglu
ICPP (2)2
1993 A study on the number of memory ports in multiple instruction issue machines
abstract
Compiler-controlled speculative execution has been shown to be effective in increasing the available instruction level parallelism (ILP) found in non-numeric programs. An important problem associated with compiler-controlled speculative execution is to accurately report and handle exceptions caused by speculatively executed instructions. Previous solutions to this problem incur either excessive hardware overhead or significant register pressure. The paper introduces a new architectural scheme referred to as write-back suppression. This scheme systematically suppresses register file updates for subsequent speculative instructions after an exception condition is detected for a speculatively executed instruction. The authors show that with a modest amount of hardware, write-back suppression supports accurate reporting and handling of exceptions for compiler-controlled speculative execution with minimal additional register pressure. Experiments based on a prototype compiler implementation and hardware simulation indicate that ensuring accurate handling of exceptions with write-back suppression incurs little run-time performance overhead.>
Soo-Mook Moon, Kemal Ebcioglu
MICRO2
1993 Making Compaction-Based Parallelization Affordable
abstract
Compaction-based parallelization suffers from long compile time and large code size because of its inherent code explosion problem. If software pipelining is performed for loop parallelization along with compaction, as in the authors' compiler, the code explosion problem becomes more serious. The authors propose the software lookahead heuristic for use in software pipelining, which allows inter-basic-block movement of code within a prespecified number of operations, called the software lookahead window, on any path emanating from the currently processed instruction at compile time. Software lookahead enables instruction-level parallelism to be exploited in a much greater code area than a single basic block, but the lookahead region is still limited to a constant depth by means of a user-specifiable window, and thus code explosion is restricted. The proposed scheme has been implemented in the authors' VLIW parallelizing compiler. To study the code explosion problem and instruction-level parallelism for branch-intensive code, they compiled five AIX utilities: sort, fgrep, sed, yacc, and compress. It is demonstrated that the software lookahead heuristic effectively alleviates the code explosion problem while successfully extracting a substantial amount of inter-basic-block parallelism.>
Toshio Nakatani, Kemal Ebcioglu
IEEE Trans. Parallel Distributed Syst.2
1992 An architectural framework for migration from CISC to higher performance platforms
abstract
We describe a novel architectural framework that allows software applications written for a given Complex Instruction Set Computer (CISC) to migrate to a different, higher performance architecture, without a significant investment on the part of the application user or developer. The framework provides a hardware mechanism for seamless switching between two instruction sets, resulting in a machine that enhances application performance while keeping the same program behavior (from a user perspective). High execution speed on migrated applications is achieved through automated translation of the object code of one machine to that of the other, using advanced global optimization and scheduling techniques. Issues affecting application behavior, such as precise exceptions, as well as self-modifying code, are addressed. Relaxation of full compatibility on these issues lead to further possible performance gains, encouraging applications to adopt the newer architecture.
Gabriel M. Silberman, Kemal Ebcioglu
ICS2
1992 An efficient resource-constrained global scheduling technique for superscalar and VLIW processors
Soo-Mook Moon, Kemal Ebcioglu
MICRO2
1991 On Optimal Parallelization of Arbitrary Loops
Uwe Schwiegelshohn, Franco Gasperoni, Kemal Ebcioglu
J. Parallel Distributed Comput.3
1989 A global resource-constrained parallelization technique
abstract
This paper presents a new approach to resource-constrained compiler extraction of fine-grain parallelism, targeted towards VLIW supercomputers, and in particular, the IBM VLIW (Very Large Instruction Word) processor. The algorithms described integrate resource limitations into Percolation Scheduling—a global parallelization technique—to deal with resource constraints, without sacrificing the generality and completeness of Percolation Scheduling in the process. This is in sharp contrast with previous approaches which either applied only to conditional-free code, or drastically limited the parallelization process by imposing relatively local heuristic resource constraints early in the scheduling process.
Kemal Ebcioglu, Alexandru Nicolau
ICS1
1987 An Efficient Logic Programming Language and Its Application to Music
Kemal Ebcioglu
ICLP1
1986 An Expert System for Chorale Harmonization
Kemal Ebcioglu
AAAI1