EDBT 2026 Demo / reviewers in the wild / expert
Emre Ozer 0001
dblp:35/4637-1 · also Emre Özer 0001
· DBLP profile ↗
25ranked-venue papers
9as first author
3since 2021 · last 2026
0000-0001-8285-1551ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lifetime-Aware Design for Item-Level Intelligence at the Extreme EdgeabstractWe present FlexiFlow, a lifetime-aware design framework for item-level intelligence (ILI) where computation is integrated directly into disposable products like food packaging and medical patches. Our framework leverages natively flexible electronics which offer significantly lower costs than silicon but are limited to kHz speeds and several thousands of gates. Our insight is that unlike traditional computing with more uniform deployment patterns, ILI applications exhibit 1000× variation in operational lifetime, fundamentally changing optimal architectural design decisions when considering trillion-item deployment scales. To enable holistic design and optimization, we model the trade-offs between embodied carbon footprint and operational carbon footprint based on application-specific lifetimes. The framework includes: (1) FlexiBench, a workload suite targeting sustainability applications from spoilage detection to health monitoring; (2) FlexiBits, area-optimized RISC-V cores with 1/4/8-bit datapaths achieving 2.65× to 3.50× better energy efficiency per workload execution; and (3) a carbon-aware model that selects optimal architectures based on deployment characteristics. We show that lifetime-aware microarchitectural design can reduce carbon footprint by 1.62×, while algorithmic decisions can reduce carbon footprint by 14.5×. We validate our approach through the first tape-out using a PDK for flexible electronics with fully open-source tools, achieving 30.9\,kHz operation. FlexiFlow enables exploration of computing at the Extreme Edge where conventional design methodologies must be reevaluated to account for new constraints and considerations. FlexiFlow is available at https://github.com/harvard-edge/FlexiFlow. Shvetank Prakash, Andrew Cheng, Olof Kindgren, Ashiq Ahamed, Graham Knight, Jedrzej Kufel, Francisco Rodriguez, Arya Tschand, David Kong 0001, Mariam Elgamal, Jerry Huang, Emma Chen, Gage Hills, Richard Price, Emre Ozer 0001, Vijay Janapa Reddi |
ASPLOS (2) | 15 |
| 2026 | Weightless Neural Networks on Flexible Substrates: A Novel Approach to Wearable Machine LearningabstractIn this article, we present a novel approach that seamlessly integrates machine learning (ML) algorithms into wearable technology through the use of weightless neural networks (WNNs) and flexible integrated circuits (FlexICs). Our methodology employs combinational intelligent networks (COIN) for edge inference on resource-constrained devices, highlighting the advantages of WNNs in terms of power efficiency and minimal hardware requirements. We propose an automated design flow for implementing COIN as FlexICs aimed at developing scalable, cost-effective, and environmentally sustainable wearable monitoring solutions. As a proof-of-concept demonstrator, an arrhythmia detection FlexIC was fabricated using COIN to meet the stringent requirements of medium-complexity wearable applications, offering a promising path toward personalized and accessible healthcare solutions. Igor D. S. Miranda, Velu Pillai, Tejas Musale, Mugdha P. Jadhao, Paulo C. R. Souza Neto, Zachary Susskind, Alan T. L. Bacellar, Mael Lhostis, Priscila M. V. Lima, Diego Leonel Cadette Dutra, Eugene John, Maurício Breternitz, Felipe M. G. França, Emre Ozer 0001, Lizy Kurian John |
IEEE Trans. Very Large Scale Integr. Syst. | 14 |
| 2025 | Flexing RISC-V Instruction Subset Processors to Extreme EdgeabstractThis paper presents an automated approach for designing processors that support a subset of the RISC-V instruction set architecture (ISA) for a new class of applications at Extreme Edge.The electronics used in extreme edge applications must be area and power-efficient, but also provide additional qualities, such as low cost, conformability, comfort and sustainability.Flexible electronics, rather than silicon-based electronics, will be able to meet the above qualities.For this purpose, we propose a methodology for generating RISC-V instruction subset processors (RISSPs) tailored to these applications and implementing them as flexible integrated circuits (FlexICs).The methodology makes verification an integral part of the processor design by treating each instruction in the ISA as a discrete, fully functional, pre-verified hardware block.It automatically builds a custom processor by stitching together the instruction hardware blocks required by an application or a set of applications in a specific domain.We generate RISSPs using the proposed methodology for three extreme edge applications, and embedded applications from the Embench benchmark suite.When synthesized, RISSPs can achieve 8-to-43% reduction in area and 3-to-30% reduction in power compared to a processor supporting the full RISC-V ISA, and are also on average ~40 times more energy efficient than Serv -the world's smallest 32-bit RISC-V processor.When physically implemented as FlexICs, the three extreme edge RISSPs achieve up to 42% area and 21% power savings with respect to the full RISC-V processor. Alireza Raisiardali, Konstantinos Iordanou, Jedrzej Kufel, Kowshik Gudimetla, Kris Myny, Emre Ozer 0001 |
MICRO | 6 |
| 2018 | Error Correlation Prediction in Lockstep Processors for Safety-Critical SystemsabstractThis paper presents a new phenomenon called error correlation prediction for lockstep processors. Lockstep processors run the same copy of a program, and their outputs are compared at every cycle to detect divergence, and have been popular in safety-critical systems. When the lockstep error checker detects an error, it alerts the safety-critical system by putting the lockstep processor in a safe state in order to prevent hazards. This is done by running the online diagnostics to identify the cause of the error because the lockstep processor has no knowledge of whether the error is caused by a transient or permanent fault. The online diagnostics can be avoided if the error is caused by a transient fault, and the lockstep processor can recover from it. If, however, it is caused by a permanent fault, having prior knowledge about error's likely location(s) within the CPU speeds up the diagnostics process. We discover that the error's type and likely location(s) inside CPUs from which the fault may have originated can be predicted by analyzing the output signals of the CPU(s) when the error is detected. We design a simple static predictor exploiting this phenomenon and show that system availability can be increased by 42-64% with an overhead of less than 2% in silicon area and power. Emre Ozer 0001, Balaji Venu, Xabier Iturbe, Shidhartha Das, Spyros Lyberis, John Biggs, Peter Harrod, John Penton |
MICRO | 1 |
| 2018 | The Arm Triple Core Lock-Step (TCLS) ProcessorabstractThe Arm Triple Core Lock-Step (TCLS) architecture is the natural evolution of Arm Cortex-R Dual Core Lock-Step (DCLS) processors to increase dependability, predictability, and availability in safety-critical and ultra-reliable applications. TCLS is simple, scalable, and easy to deploy in applications where Arm DCLS processors are widely used (e.g., automotive), as well as in new sectors where the presence of Arm technology is incipient (e.g., enterprise) or almost non-existent (e.g., space). Specifically in space, COTS Arm processors provide optimal power-to-performance, extensibility, evolvability, software availability, and ease of use, especially in comparison with the decades old rad-hard computing solutions that are still in use. This article discusses the fundamentals of an Arm Cortex-R5 based TCLS processor, providing key functioning and implementation details. The article shows that the TCLS architecture keeps the use of rad-hard technology to a minimum, namely, using rad-hard by design standard cell libraries only to protect the critical parts that account for less than 4% of the entire TCLS solution. Moreover, when exposure to radiation is relatively low, such as in terrestrial applications or even satellites operating in Low Earth Orbits (LEO), the system could be implemented entirely using commercial cell libraries, relying on the radiation mitigation methods implemented on the TCLS to cope with sporadic soft errors in its critical parts. The TCLS solution allows thus to significantly reduce chip manufacturing costs and keep pace with advances in low power consumption and high density integration by leveraging commercial semiconductor processes, while matching the reliability levels and improving availability that can be achieved using extremely expensive rad-hard semiconductor processes. Finally, the article describes a TRL4 proof-of-concept TCLS-based System-on-Chip (SoC) that has been prototyped and tested to power the computer on-board an Airbus Defence and Space telecom satellite. When compared to the currently used processor solution by Airbus, the TCLS-based SoC results in a more than 5× performance increase and cuts power consumption by more than half. Xabier Iturbe, Balaji Venu, Emre Ozer 0001, Jean-Luc Poupat, Gregoire Gimenez, Hans-Ulrich Zurek |
ACM Trans. Comput. Syst. | 3 |
| 2017 | A "high resilience" mode to minimize soft error vulnerabilities in ARM cortex-R CPU pipelines: work-in-progressabstractThis paper proposes a "high resilience" execution mode to increase the robustness of CPU pipelines to soft errors when executing critical software routines. The proposed execution mode reduces the error rate by approximately 11% in an ARM Cortex-R5 CPU, and requires only a few minor modifications to be made in its microarchitecture. These modifications do not impact the characteristic area, power consumption and performance features of the original CPU. Xabier Iturbe, Balaji Venu, John Penton, Emre Ozer 0001 |
CASES | 4 |
| 2016 | Predicting room occupancy with a single passive infrared (PIR) sensor through behavior extractionabstractPassive infrared sensors have widespread use in many applications, including motion detectors for alarms, lighting systems and hand dryers. Combinations of multiple PIR sensors have also been used to count the number of humans passing through doorways. In this paper, we demonstrate the potential of the PIR sensor as a tool for occupancy estimation inside of a monitored environment. Our approach shows how flexible nonparametric machine learning algorithms extract useful information about the occupancy from a single PIR sensor. The approach allows us to understand and make use of the motion patterns generated by people within the monitored environment. The proposed counting system uses information about those patterns to provide an accurate estimate of room occupancy which can be updated every 30 seconds. The system was successfully tested on data from more than 50 real office meetings consisting of at most 14 room occupants. Yordan P. Raykov, Emre Ozer 0001, Ganesh Dasika, Alexis Boukouvalas, Max A. Little |
UbiComp | 2 |
| 2015 | A Highly-Efficient, Adaptive and Fault-Tolerant SoC Implementation of a Fourier Transform Spectrometer Data ProcessingabstractWe present here one of the first research efforts conducted at Jet Propulsion Laboratory (JPL) to implement on a single chip (Xilinx Zynq) all the functionality necessary to control a NASA instrument, namely a Fourier Transform Spectrometer (FTS) that is proposed for deployment on future missions to Jupiter's moon Europa. The system requires custom logic to process the data delivered by the instrument, and software, to perform floating-point operations and to drive the interface with the main spacecraft computer. Three features are central in our SoC FTS implementation: (1) Efficiency, as the system achieves a high data processing throughput at relatively low power consumption, (2) Adaptivity, as the system can be configured from Earth based on the data observed while exploring a priori unknown space environments and (3) Fault-Tolerance, as the system needs to operate in the harsh radiation Jupiter magnetosphere where Europa orbits. Xabier Iturbe, Didier Keymeulen, Patrick Yiu, Dan Berisford, Kevin P. Hand, Robert Carlson, Emre Ozer 0001 |
FCCM | 7 |
| 2015 | An integrated SoC for science data processing in next-generation space flight instruments avionicsabstractWe present here an integrated SoC platform, called APEX-SoC, that is aimed at speeding-up the design of next-generation space flight instruments avionics by providing a convenient infrastructure for hardware and software based science data processing. We use a case-study drawn from the JPL Compositional Infrared Imaging Spectrometer (CIRIS) to illustrate the process of integrating instrument-dependent data acquisition and processing stages in this platform. In order to enable the use of APEX-SoC-based instruments in deep space missions, the platform implements Radiation Hardening By Design (RHBD) techniques and offers support for instantiating multiple processing stages that can be used at runtime to increase reliability or performance, based on the requirements of the mission at each particular stage. Finally, in the specific case of CIRIS, the data processing includes a stage to cope with radiation affecting the instrument photo-detector. Xabier Iturbe, Didier Keymeulen, Emre Ozer 0001, Patrick Yiu, Dan Berisford, Kevin P. Hand, Robert Carlson |
VLSI-SoC | 3 |
| 2013 | Memory array protection: check on read or check on write?abstractThis work introduces Check-on-Write: a memory array error protection approach that enables a trade-off between a memory array's fault-coverage and energy. The presented approach checks for error in a value stored in an array before it is overwritten rather than, as currently done, when it is read (check-on-read). This aims at reducing the number and energy of error code checks. This lazy protection approach can be used for caches in systems that support failure-atomicity to recover from corrupted state due to a fault. The paper proposes and evaluates an adaptive memory protection scheme that is capable of both check-on-read and check-on-write and switches between the two protection modes depending on the energy to be saved and fault coverage requirements. Experimental analysis shows that our technique reduces the average dynamic energy of the L1 instruction cache tag and data arrays by 18.6% and 17.7% respectively. For the L1 data cache, this is 17.2% and 2.9%, and the savings are 13.4% for the L2 tag array. The paper also quantifies the implications of the proposed scheme on fault-coverage by analyzing the mean-time-to-failure as a function of the transient failure rate. Panagiota Nikolaou, Yiannakis Sazeides, Lorena Ndreu, Emre Ozer 0001, Sachin Idgunji |
DATE | 4 |
| 2013 | An analytical framework for estimating TCO and exploring data center design spaceabstractIn this paper, we present EETCO: an estimation and exploration tool that provides qualitative assessment of data center design decisions on Total-Cost-of-Ownership (TCO) and environmental impact. It can capture the implications of many parameters including server performance, power, cost, and Mean-Time-To-Failure (MTTF). The tool includes a model for spare estimation needed to account for server failures and performance variability. The paper describes the tool model and its implementation, and presents experiments that explore tradeoffs offered by different server configurations, performance variability, MTTF, 2D vs 3D-stacked processors, and ambient temperature. These experiments reveal, for the data center configurations used in this study, several opportunities for profit and optimization in the datacenter ecosystem: (i) servers with different computing performance and power consumption merit exploration to minimize TCO and the environmental impact, (ii) performance variability is desirable if it comes with a drastic cost reduction, (iii) shorter processor MTTF is beneficial if it comes with a moderate processor cost reduction, (iv) increasing by few degrees the ambient datacenter temperature reduces the environmental impact with a minor increase in the TCO and (v) a higher cost for a 3D-stacked processor with shorter MTTF and higher power consumption can be preferred, over a conventional 2D processor, if it offers a moderate performance increase. Damien Hardy, Marios Kleanthous, Isidoros Sideris, Ali G. Saidi, Emre Ozer 0001, Yiannakis Sazeides |
ISPASS | 5 |
| 2013 | Implicit-storing and redundant-encoding-of-attribute information in error-correction-codesabstractThis paper proposes implicit-storing to extend the logical capacity of a memory array without increasing its physical capacity by leveraging the array's error-correction-codes to infer the implicitly stored bits. Implicit-storing is related to error-code-tagging, a technique that distinguishes between faults in data and invariant attributes of a location when the attributes are not stored in the memory array but are encoded in the error-correction-codes. Both error-code-tagging and implicit-storing cause a code-strength reduction due to their encoding of additional information in the code meant to only protect data. Yiannakis Sazeides, Emre Ozer 0001, Danny Kershaw, Panagiota Nikolaou, Marios Kleanthous, Jaume Abella 0001 |
MICRO | 2 |
| 2012 | Thermal characterization of cloud workloads on a power-efficient server-on-chipabstractWe propose a power-efficient many-core server-on-chip system with 3D-stacked Wide I/O DRAM targeting cloud workloads in datacenters. The integration of 3D-stacked Wide I/O DRAM on top of a logic die increases available memory bandwidth by using dense and fast Through-Silicon Vias (TSVs) instead of off-chip IOs, enabling faster data transfers at much lower energy per bit. We demonstrate a methodology that includes full-system microarchitectural modeling and rapid virtual physical prototyping with emphasis on the thermal analysis. Our findings show that while executing CPU-centric benchmarks (e.g. SPECInt and Dhrystone), the temperature in the server-on-chip (logic+DRAM) is in the range of 175-200°C at a power consumption of less than 20W, exceeding the reliable operating bounds without any cooling solutions, even with embedded cores. However, with real cloud workloads, the power density in the server-on-chip remains much below the temperatures reached by the CPU-centric workloads as a result of much lower power burnt by memory-intensive cloud workloads. We show that such a server-on-chip system is feasible with a low-cost passive heat sink eliminating the need for a high-cost active heat sink with an attached fan, creating an opportunity for overall cost and energy savings in datacenters. Dragomir Milojevic, Sachin Idgunji, Djordje Jevdjic, Emre Ozer 0001, Pejman Lotfi-Kamran, Andreas Panteli, Andreas Prodromou, Chrysostomos Nicopoulos, Damien Hardy, Babak Falsafi, Yiannakis Sazeides |
ICCD | 4 |
| 2012 | Scale-out processorsabstractScale-out datacenters mandate high per-server throughput to get the maximum benefit from the large TCO investment. Emerging applications (e.g., data serving and web search) that run in these datacenters operate on vast datasets that are not accommodated by on-die caches of existing server chips. Large caches reduce the die area available for cores and lower performance through long access latency when instructions are fetched. Performance on scale-out workloads is maximized through a modestly-sized last-level cache that captures the instruction footprint at the lowest possible access latency. In this work, we introduce a methodology for designing scalable and efficient scale-out server processors. Based on a metric of performance-density, we facilitate the design of optimal multi-core configurations, called pods. Each pod is a complete server that tightly couples a number of cores to a small last-level cache using a fast interconnect. Replicating the pod to fill the die area yields processors which have optimal performance density, leading to maximum per-chip throughput. Moreover, as each pod is a stand-alone server, scale-out processors avoid the expense of global (i.e., interpod) interconnect and coherence. These features synergistically maximize throughput, lower design complexity, and improve technology scalability. In 20nm technology, scaleout chips improve throughput by 5x-6.5x over conventional and by 1.6x-1.9x over emerging tiled organizations. Pejman Lotfi-Kamran, Boris Grot, Michael Ferdman, Stavros Volos, Yusuf Onur Koçberber, Javier Picorel, Almutaz Adileh, Djordje Jevdjic, Sachin Idgunji, Emre Ozer 0001, Babak Falsafi |
ISCA | 10 |
| 2011 | Eliminating energy of same-content-cell-columns of on-chip SRAM arrays
Bushra Ahsan, Lorena Ndreu, Isidoros Sideris, Yiannakis Sazeides, Sachin Idgunji, Emre Ozer 0001 |
ISLPED | 6 |
| 2009 | Way guard: a segmented counting bloom filter approach to reducing energy for set-associative cachesabstractThe design trend of caches in modern processors continues to increase their capacity with higher associativity to cope with large data footprint and take advantage of feature size shrink, which, unfortunately, also leads to higher energy consumption. This paper presents a technique using segmented counting Bloom filters called "Way Guard" to reduce the number of redundant way lookups in large set-associative caches to achieve dynamic energy savings. Our Way Guard mechanism only looks up an average of 25-30% of the cache ways and saved up to 65% of the L2 energy and up to 70% of the L1 cache energy. Mrinmoy Ghosh, Emre Ozer 0001, Simon Ford, Stuart Biles, Hsien-Hsin S. Lee |
ISLPED | 2 |
| 2008 | A stochastic bitwidth estimation technique for compact and low-power custom processorsabstractThere is an increasing trend toward compiling from C to custom hardware for designing embedded systems in which the area and power consumption of application-specific functional units, registers, and memory blocks are heavily dependent on the bit-widths of integer operands used in computations. The actual bit-width required to store the values assigned to an integer variable during the execution of a program will not, in general, match the built-in C data types. Thus, precious area is wasted if the built-in data type sizes are used to declare the size of integer operands. In this paper, we introduce stochastic bit-width estimation that follows a simulation-based probabilistic approach to estimate the bit-widths of integer variables using extreme value theory. The estimation technique is also empirically compared to two compile-time integer bit-width analysis techniques. Our experimental results show that the stochastic bit-width estimation technique dramatically reduces integer bit-widths and, therefore, enables more compact and power-efficient custom hardware designs than the compile-time integer bit-width analysis techniques. Up to 37% reduction in custom hardware area and 30% reduction in logic power consumption using stochastic bit-width estimation can be attained over ten integer applications implemented on an FPGA chip. Emre Ozer 0001, Andy Nisbet, David Gregg |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2007 | Low-cost Techniques for Reducing Branch Context Pollution in a Soft Realtime Embedded Multithreaded ProcessorabstractIn this paper, we propose two low-cost and novel branch history buffer handling schemes aiming at skewing the branch prediction accuracy in favor of a real-time thread for a soft real-time embedded multithreaded processor. The processor core accommodates two running threads, one with the highest priority and the other thread is a background thread, and both threads share the branch predictor. The first scheme uses a 3-bit branch history buffer in which the highest priority thread uses the most significant 2 bits to change the prediction state while the background thread uses only the least significant 2 bits. The second scheme uses the shared 2-bit branch history buffer that implements integer updates for the highest priority thread but fractional updates for the background thread in order to achieve relatively higher prediction accuracy in the highest priority thread. The low cost nature of these two schemes, particularly in the second scheme, makes them attractive with moderate improvement in the performance of the highest priority thread. Emre Ozer 0001, Alastair Reid 0001, Stuart Biles |
SBAC-PAD | 1 |
| 2006 | Reducing energy of virtual cache synonym lookup using bloom filtersabstractVirtual caches are employed as L1 caches of both high performance and embedded processors to meet their short latency requirements. However, they also introduce the synonym problem where the same physical cache line can be present at multiple locations in the cache due to their distinct virtual addresses, leading to potential data consistency issues. To guarantee correctness, common hardware solutions either perform serial lookups for all possible synonym locations in the L1 consuming additional energy or employ a reverse map in the L2 cache that incurs a large area overhead. Such preventive mechanisms are nevertheless indispensable even though synonyms may not always be present during the execution.In this paper, we study the synonym issue using Windows applications workload and propose a technique based on Bloom filters to reduce synonym lookup energy. By tracking the address stream using Bloom filters, we can confidently exclude the addresses that were never observed to eliminate unnecessary synonym lookups, thereby saving energy in the L1 cache. Bloom filters have a very small area overhead making our implementation a feasible and attractive solution for synonym detection. Our results show that synonyms in these applications actually constitutes less than 0.1% of the total cache misses. By applying our technique, the dynamic energy consumed in L1 data cache can be reduced up to 32.5%. When taking leakage energy into account, the savings is up to 27.6%. Dong Hyuk Woo, Mrinmoy Ghosh, Emre Ozer 0001, Stuart Biles, Hsien-Hsin S. Lee |
CASES | 3 |
| 2005 | High-Performance and Low-Cost Dual-Thread VLIW Processor Using Weld Architecture ParadigmabstractThis paper presents a cost-effective and high-performance dual-thread VLIW processor model. The dual-thread VLIW processor model is a low-cost subset of the Weld architecture paradigm. It supports one main thread and one speculative thread running simultaneously in a VLIW processor with a register file and a fetch unit per thread along with memory disambiguation hardware for speculative load and store operations. This paper analyzes the performance impact of the dual-thread VLIW processor, which includes analysis of migrating disambiguation hardware for speculative load operations to the compiler and of the sensitivity of the model to the variation of branch misprediction, second-level cache miss penalties, and register file copy time. Up to 34 percent improvement in performance can be attained using the dual-thread VLIW processor when compared to a single-threaded VLIW processor model. Emre Ozer 0001, Thomas M. Conte |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Stochastic Bit-Width Approximation Using Extreme Value Theory for Customizable Processors
Emre Ozer 0001, Andy Nisbet, David Gregg |
CC | 1 |
| 2004 | Automatic Customization of Embedded Applications for Enhanced Performance and Reduced Power Using Optimizing Compiler Techniques
Emre Ozer 0001, Andy Nisbet, David Gregg |
Euro-Par | 1 |
| 2004 | Fine-Tuning Loop-Level Parallelism for Increasing Performance of DSP Applications on FPGAsabstractThis paper discusses the balance between loop-level parallelism and clock rate for enhancing the performance of DSP applications fully implemented on FPGAs. Loop-level parallelism reduces the total cycles of an application at the cost of increased routing complexity that often results in lower clock rates. We analyze loops that can be fully parallelized and show that it is possible to achieve better performance by controlling the number of parallel iterations of the loops than using fully parallel loops. We have implemented loop parallelism in our compilation framework and fine-tune them to enhance the performance of DSP applications that target Xilinx Virtex-II FPGA chip. Our experimental results show that it is possible to reach a performance equilibrium point where the total number of cycles and the overall clock frequency can be adjusted to maximize the overall performance of an application. Emre Ozer 0001, Andy Nisbet, David Gregg |
FCCM | 1 |
| 2001 | Weld: A Multithreading Technique Towards Latency-Tolerant VLIW Processors
Emre Ozer 0001, Thomas M. Conte |
HiPC | 1 |
| 1998 | Unified Assign and Schedule: A New Approach to Scheduling for Clustered Register File MicroarchitecturesabstractRecently, there has been a trend towards clustered microarchitectures to reduce the cycle time for wide issue microprocessors. In such processors, the register file and functional units are partitioned and grouped into clusters. Instruction scheduling for a clustered machine requires assignment and scheduling of operations to the clusters. In this paper, a new scheduling algorithm named unified-assign-and-schedule (UAS) is proposed for clustered, statically-scheduled architectures. UAS merges the cluster assignment and instruction scheduling phases in a natural and straightforward fashion. We compared the performance of UAS with various heuristics to the well-known Bottom-up Greedy (BUG) algorithm and to an optimal cluster scheduling algorithm, measuring the schedule lengths produced by all of the schedulers. Our results show that UAS gives better performance than the BUG algorithm and is quite close to optimal. Emre Ozer 0001, Sanjeev Banerjia, Thomas M. Conte |
MICRO | 1 |