EDBT 2026 Demo / reviewers in the wild / expert
Feihui Li
dblp:l/FeihuiLi
· DBLP profile ↗
33ranked-venue papers
11as first author
0since 2021 · last 2009
0000-0003-3244-2113ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 8 first-authorSoftware engineering, systems software and programming languages · 12 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Memory systems · 25% Processor architecture and microarchitecture · 20% Interconnection networks and networks-on-chip · 17% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
chip multiprocessor |
0.2 | 4 | 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008 A novel migration-based NUCA design for chip multiprocessors · SC 2008 Locality-conscious workload assignment for array-based computations in MPSOC architectures · DAC 2005 |
Memory systems
cache design |
0.2 | 3 | 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008 A novel migration-based NUCA design for chip multiprocessors · SC 2008 Design and Management of 3D Chip Multiprocessors Using Network-in-Memory · ISCA 2006 |
Storage systems
data migration |
0.2 | 2 | 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008 A novel migration-based NUCA design for chip multiprocessors · SC 2008 |
Memory systems › cache design
non-uniform cache architecture |
0.2 | 2 | 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008 A novel migration-based NUCA design for chip multiprocessors · SC 2008 |
Reconfigurable computing and FPGAs
application mapping |
0.1 | 1 | 2008 | Application mapping for chip multiprocessors · DAC 2008 |
Parallel and multicore computing › locality optimization
data locality optimization |
0.1 | 1 | 2008 | Application mapping for chip multiprocessors · DAC 2008 |
Processor architecture and microarchitecture
multicore design |
0.1 | 1 | 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008 |
Compilers and program optimization › dynamic optimization
profile-guided optimization |
0.1 | 1 | 2007 | Profile-driven energy reduction in network-on-chips · PLDI 2007 |
Energy-efficient computing › energy-efficient communication
network-on-chip power management |
0.1 | 1 | 2007 | Profile-driven energy reduction in network-on-chips · PLDI 2007 |
Compilers and program optimization › compiler optimization
compiler-directed optimization |
0.1 | 1 | 2006 | Compiler-directed channel allocation for saving power in on-chip networks · POPL 2006 |
Compilers and program optimization › compiler optimization
energy-aware compilation |
0.1 | 1 | 2006 | Compiler-directed channel allocation for saving power in on-chip networks · POPL 2006 |
Interconnection networks and networks-on-chip
channel assignment |
0.1 | 1 | 2006 | Compiler-directed channel allocation for saving power in on-chip networks · POPL 2006 |
Interconnection networks and networks-on-chip › network-on-chip design
noc power management |
0.1 | 1 | 2006 | Reducing NoC energy consumption through compiler-directed channel voltage scaling · PLDI 2006 |
Energy-efficient computing
voltage and frequency scaling |
0.1 | 1 | 2006 | Reducing NoC energy consumption through compiler-directed channel voltage scaling · PLDI 2006 |
Compilers and program optimization › memory optimization
data locality optimization |
0.1 | 1 | 2005 | Locality-conscious workload assignment for array-based computations in MPSOC architectures · DAC 2005 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessors · SIGMETRICS 2008 |
Distributed systems
communication link |
0.0 | 1 | 2007 | Profile-driven energy reduction in network-on-chips · PLDI 2007 |
Embedded and real-time systems › embedded software
embedded applications |
0.0 | 1 | 2007 | Profile-driven energy reduction in network-on-chips · PLDI 2007 |
Energy-efficient computing
leakage power reduction |
0.0 | 1 | 2007 | Profile-driven energy reduction in network-on-chips · PLDI 2007 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 2005 | Locality-conscious workload assignment for array-based computations in MPSOC architectures · DAC 2005 |
Methods — techniques the papers use, named apart from their topics
profile-driven compiler analysis · 0.1link shutdown · 0.1trace-driven simulation · 0.1task scheduling · 0.1processor mapping · 0.1packet routing · 0.1facility location modeling · 0.1data mapping · 0.1router architecture · 0.1network simulation · 0.1channel reuse optimization · 0.13d stacking · 0.1loop nest analysis · 0.1data locality optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Compiler-assisted soft error detection under performance and energy constraints in embedded systemsabstractSoft errors induced by terrestrial radiation are becoming a significant concern in architectures designed in newer technologies. If left undetected, these errors can result in catastrophic consequences or costly maintenance problems in different embedded applications. In this article, we focus on utilizing the compiler's help in duplicating instructions for error detection in VLIW datapaths. The instruction duplication mechanism is further supported by a hardware enhancement for efficient result verification, which avoids the need of additional comparison instructions. In the proposed approach, the compiler determines the instruction schedule by balancing the permissible performance degradation and the energy constraint with the required degree of duplication. Our experimental results show that our algorithms allow the designer to perform trade-off analysis between performance, reliability, and energy consumption. Jie S. Hu, Feihui Li, Vijay Degalahal, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | Application mapping for chip multiprocessorsabstractThe problem attacked in this paper is one of automatically mapping an application onto a Network-on-Chip (NoC) based chip multiprocessor (CMP) architecture in a locality-aware fashion. The proposed compiler approach has four major steps: task scheduling, processor mapping, data mapping, and packet routing. In the first step, the application code is parallelized and the resulting parallel threads are assigned to virtual processors. The second step implements a virtual processor-to-physical processor mapping. The goal of this mapping is to ensure that the threads that are expected to communicate frequently with each other are assigned to neighboring processors as much as possible. In the third step, data elements are mapped to memories attached to CMP nodes. The main objective of this mapping is to place a given data item into a node which is close to the nodes that access it. The last step of our approach determines the paths (between memories and processors) for data to travel in an energy efficient manner. In this paper, we describe the compiler algorithms we implemented in detail and present an experimental evaluation of the framework. In our evaluation, we test our entire framework as well as the impact of omitting some of its steps. This experimental analysis clearly shows that the proposed framework reduces energy consumption of our applications significantly (27.41% on average over a pure performance oriented application mapping strategy) as a result of improved locality of data accesses. Guangyu Chen, Feihui Li, Seung Woo Son 0001, Mahmut T. Kandemir |
DAC | 2 |
| 2008 | Ring data location prediction scheme for Non-Uniform Cache ArchitecturesabstractIncreases in cache capacity are accompanied by growing wire delays due to technology scaling. Non-uniform cache architecture (NUCA) is one of proposed solutions to reducing the average access latency in such cache designs. While most of the prior NUCA work focuses on data placement, data replacement, and migration related issues, this paper studies the problem of data search (access) in NUCA. In our architecture we arrange sets of banks with equal access latency into rings. Our last access based (LAB) prediction scheme predicts the ring that is expected to contain the required data and checks the banks in that ring first for the data block sought. We compare our scheme to two alternate approaches: searching all rings in parallel, and searching rings sequentially. We show that our LAB ring prediction scheme reduces L2 energy significantly over the sequential and parallel schemes, while maintaining similar performance. Our LAB scheme reduces energy consumption by 15.9% relative to the sequential lookup scheme, and 53.8% relative to the parallel lookup scheme. Sayaka Akioka, Feihui Li, Konrad Malkowski, Padma Raghavan, Mahmut T. Kandemir, Mary Jane Irwin |
ICCD | 2 |
| 2008 | A novel migration-based NUCA design for chip multiprocessorsabstractChip Multiprocessors (CMPs) and Non-Uniform Cache Architectures (NUCAs) represent two emerging trends in computer architecture. Targeting future CMP based systems with NUCA type L2 caches, this paper proposes a novel data migration algorithm for parallel applications and evaluates it. The goal of this migration scheme is to determine a suitable location for each data block within a large L2 space at any given point during execution. A unique characteristic of the proposed scheme is that it models the problem of optimal data placement in the L2 cache space as a two-dimensional post office placement problem, presents a practical architectural implementation of this model, and gives a detailed evaluation of the proposed implementation. In our experimental evaluation, we also compare our approach to a previously-proposed NUCA management scheme using applications from the specomp suite, oltp, specjbb, and specweb. These experiments show that our migration approach generates about 35% improvement, on average, in average L2 access latency over the previous migration scheme, and these L2 latency savings translate, on average, to 9.5% improvement in IPC (instructions per cycle).We also observed during our experiments that both the careful initial placement of data (which itself triggers migrations within the L2 space) and subsequent migrations (due to inter-processor data sharing) play an important role in achieving our performance improvements. Mahmut T. Kandemir, Feihui Li, Mary Jane Irwin, Seung Woo Son 0001 |
SC | 2 |
| 2008 | Implementation and evaluation of a migration-based NUCA design for chip multiprocessorsabstractChip Multiprocessors (CMPs) and Non-Uniform Cache Architectures (NUCAs) represent two emerging trends in computer architecture. Targeting future CMP based systems with NUCA type L2 caches, this paper proposes a novel data migration algorithm for parallel applications and evaluates it. The goal of this migration scheme is to determine a suitable location for each data block within a large L2 space at any given point during execution. A unique characteristic of the proposed scheme is that it models the problem of optimal data placement in the L2 cache space as a two dimensional post office placement problem, presents a practical architectural implementation of this model, and gives an evaluation of the proposed implementation. Feihui Li, Mahmut T. Kandemir, Mary Jane Irwin |
SIGMETRICS | 1 |
| 2007 | Ring Prediction for Non-Uniform Cache Architectures
Sayaka Akioka, Feihui Li, Mahmut T. Kandemir, Padma Raghavan, Mary Jane Irwin |
PACT | 2 |
| 2007 | Reducing Energy Consumption of On-Chip Networks Through a Hybrid Compiler-Runtime Approach
Guangyu Chen, Feihui Li, Mahmut T. Kandemir |
PACT | 2 |
| 2007 | Compiler-directed application mapping for NoC based chip multiprocessorsabstractThe problem attacked in this paper is one of automatically mapping an application onto a Network-on-Chip (NoC) based chip multi-processor architecture in a locality-aware fashion. The proposed compiler approach has four major steps: task scheduling, processor mapping, data mapping, and packet routing. Our experimental result clearly shows that the proposed framework reduces energy consumption of our applications significantly (27.41% on average over a pure performance oriented application mapping strategy) as a result of improved locality of data accesses. Guangyu Chen, Feihui Li, Mahmut T. Kandemir |
LCTES | 2 |
| 2007 | Profile-driven energy reduction in network-on-chipsabstractReducing energy consumption of a Network-on-Chip (NoC) is a critical design goal, especially for power-constrained embedded systems.In response, prior research has proposed several circuit/architectural level mechanisms to reduce NoC power consumption. This paper considers the problem from a different perspective and demonstrates that compiler analysis can be very helpful for enhancing the effectiveness of a hardware-based link power management mechanism by increasing the duration of communication links' idle periods. The proposed profile-based approach achieves its goal by maximizing the communication link reuse through compiler-directed, static message re-routing. That is, it clusters the required data communications into a small set of communication links at any given time, which increases the idle periods for the remaining communication links in the network. This helps hardware shut down more communication links and their corresponding buffers to reduce leakage power. The current experimental evaluation, with twelve data-intensive embedded applications, shows that the proposed profile-driven compiler approach reduces leakage energy by more than 35% (on average) as compared to a pure hardware-based link power management scheme. Feihui Li, Guangyu Chen, Mahmut T. Kandemir, Ibrahim Kolcu |
PLDI | 1 |
| 2006 | Energy-aware computation duplication for improving reliability in embedded chip multiprocessorsabstractCompilers designed for current embedded systems must be capable of addressing multiple constraints such as low power, high performance, small memory footprint and form factor, and high reliability at the same time. In particular, optimizing for one constraint should be performed carefully, considering its impact on other constraints. Recent trends indicate that transient errors are becoming increasingly important in embedded systems. Focusing on an embedded chip multiprocessor and array-intensive applications, this paper demonstrates how reliability against transient errors can be improved without impacting execution time by utilizing idle processors for duplicating some of the computations of the active processors. It also shows how a balance between power savings and reliability improvement can be struck using a metric called the energy-delay-fallibility product. Our experimental results indicate that the "percentage of duplicated computations" is a useful high-level metric for studying the tradeoffs among performance, power, and reliability Guilin Chen, Mahmut T. Kandemir, Feihui Li |
ASP-DAC | 3 |
| 2006 | Prefetching-aware cache line turnoff for saving leakage energyabstractWhile numerous prior studies focused on performance and energy optimizations for caches, their interactions have received much less attention. This paper studies this interaction and demonstrates how performance and energy optimizations can affect each other. More importantly, we propose three optimization schemes that turn off cache lines in a prefetching-sensitive manner. These schemes treat prefetched cache lines differently from the lines brought to the cache in a normal way (i.e., through a load operation) in turning off the cache lines. Our experiments with applications from the SPEC2000 suite indicate that the proposed approaches save significant leakage energy with very small degradation on performance. Ismail Kadayif, Mahmut T. Kandemir, Feihui Li |
ASP-DAC | 3 |
| 2006 | Maximizing data reuse for minimizing memory space requirements and execution cyclesabstractEmbedded systems in the form of vehicles and mobile devices such as wireless phones, automatic banking machines and new multi-modal devices operate under tight memory and power constraints. Therefore, their performance demands must be balanced very well against their memory space requirements and power consumption. Automatic tools that can optimize for memory space utilization and performance are expected to be increasingly important in the future as increasingly larger portions of embedded designs are being implemented in software. In this paper, we describe a novel optimization framework that can be used in two different ways: (i) deciding a suitable on-chip memory capacity for a given code, and (ii) restructuring the application code to make better use of the available on-chip memory space. While prior proposals have addressed these two questions, the solutions proposed in this paper are very aggressive in extracting and exploiting all data reuse in the application code, restricted only by inherent data dependences Mahmut T. Kandemir, Guangyu Chen, Feihui Li |
ASP-DAC | 3 |
| 2006 | Energy savings through embedded processing on disk systemabstractMany of today's data-intensive applications manipulate disk-resident data sets. As a result, their overall behavior is tightly coupled with their disk performance. Unfortunately, most of these applications quickly become disk bound since disk I/O times, the communication latencies, and energy consumption required to transfer disk data to the host machine can be very large. A promising solution to this problem is to embed computational power into the disk storage system. This paper concentrates on such a smart disk based architecture and proposes an automated approach that partitions a given application code between the host machine and the smart disk. The main goal is to perform data filterings, identified at compile time, on the smart disk, thereby reducing the energy spent in communicating disk data to the host unit for processing. To achieve this, the proposed approach uses integer linear programming to identify the code fragments that perform significant data filtering and assigns such fragments to the smart disk for execution. In addition to the communication energy benefits of the proposed approach, we show in this paper that this approach can also help us better exploit the low-power management capabilities provided by the system. Our experiments with four data-intensive applications indicate significant energy savings. Seung Woo Son 0001, Guangyu Chen, Mahmut T. Kandemir, Feihui Li |
ASP-DAC | 4 |
| 2006 | Reducing dynamic compilation overhead by overlapping compilation and executionabstractAn important problem in executing applications in energy-sensitive embedded environments is to tune their behavior based on dynamic variations in energy constraints. One option for achieving this is dynamic compilation /spl sim/ compiling code fragments on the fly to adapt to changing energy demands. While dynamic compilation can be very beneficial in many embedded environments where multiple criteria need to be satisfied during execution, it can also incur a significant performance overhead since compilation takes place at runtime. The goal in this work is to reduce this performance overhead of dynamic compilation by overlapping it with application execution. Specifically, provided that we have available hardware resources to perform dynamic compilation concurrently with application execution, our approach compiles the next code fragment to be executed while we are executing the current code fragment. The experimental results from our implementation indicate significant savings in execution times. Our experimental results also indicate that the proposed strategy performs consistently well under different parameters. Priya Unnikrishnan, Mahmut T. Kandemir, Feihui Li |
ASP-DAC | 3 |
| 2006 | Activity clustering for leakage management in SPMsabstractThis paper proposes compiler-based leakage optimization strategy for on-chip scratch-pad memories (SPMs). The idea is to keep only a small set of SPM regions active at a given time and pre-activate SPM regions based on the compiler-extracted data access pattern. Our strategy, called activity clustering, increases the length of the idle periods of SPM regions by clustering accesses to a small set of regions at a time. It thus allows an SPM to take better advantage of the underlying leakage optimization mechanism Mahmut T. Kandemir, Guangyu Chen, Feihui Li, Mary Jane Irwin, Ibrahim Kolcu |
DATE | 3 |
| 2006 | Dynamic partitioning of processing and memory resources in embedded MPSoC architecturesabstractCurrent trends indicate that multiprocessor-system-on-chip (MPSoC) architectures are being increasingly used in building complex embedded systems. While circuit/architectural support for MPSoC based systems are making significant strides, programming these devices and providing suitable software support (e.g., compiler and operating systems) seem to be a tougher problem. This is because either programmers or compilers will have to make code explicitly parallel to run on these systems. An additional difficulty occurs when multiple applications use an MPSoC at the same time, because MPSoC resources should be partitioned across these applications carefully. This paper explores a proactive resource partitioning scheme for parallel applications simultaneously exercising the same MPSoC system. The proposed approach has two major components. The first component includes an offline preprocessing of applications which gives us an estimated profile for each application. Each application to be executed on our MPSoC is profiled and annotated with the profile information. The second component of our approach is an online resource partitioning, which partitions both the processing cores (i.e., computation resources) and on-chip memory space (i.e., storage resource) among simultaneously-executing applications. Our experimental evaluation with this partitioner shows that it generates much better results than conventional operating system based resource management. The results also reveal that both memory partitioning and processor partitioning are very important for obtaining the best results Liping Xue, Ozcan Ozturk 0001, Feihui Li, Mahmut T. Kandemir, Ibrahim Kolcu |
DATE | 3 |
| 2006 | Design and Management of 3D Chip Multiprocessors Using Network-in-MemoryabstractLong interconnects are becoming an increasingly important problem from both power and performance perspectives. This motivates designers to adopt on-chip network-based communication infrastructures and three-dimensional (3D) designs where multiple device layers are stacked together. Considering the current trends towards increasing use of chip multiprocessing, it is timely to consider 3D chip multiprocessor design and memory networking issues, especially in the context of data management in large L2 caches. The overall goal of this paper is to study the challenges for L2 design and management in 3D chip multiprocessors. Our first contribution is to propose a router architecture and a topology design that makes use of a network architecture embedded into the L2 cache memory. Our second contribution is to demonstrate, through extensive experiments, that a 3D L2 memory architecture generates much better results than the conventional two-dimensional (2D) designs under different number of layers and vertical (inter-wafer) connections. In particular, our experiments show that a 3D architecture with no dynamic data migration generates better performance than a 2D architecture that employs data migration. This also helps reduce power consumption in L2 due to a reduced number of data movements. Feihui Li, Chrysostomos Nicopoulos, Thomas D. Richardson, Yuan Xie 0001, Narayanan Vijaykrishnan, Mahmut T. Kandemir |
ISCA | 1 |
| 2006 | Compiler-directed thermal management for VLIW functional unitsabstractAs processors, memories, and other components of today's embedded systems are pushed to higher performance in more enclosed spaces, processor thermal management is quickly becoming a limiting design factor. While previous proposals mostly approached this thermal management problem from circuit and architecture angles, software can also play an important role in identifying and eliminating thermal hotspots as it is the main factor that shapes the order and frequency of accesses to different hardware components in the chip. This is particularly true for compiler-scheduled Very Long Instruction Word (VLIW) datapath.In this paper, we focus on a compiler-based approach to make the thermal profile more balanced in the integer functional units of VLIW architectures. For balanced thermal behavior and peak temperature minimization, we propose techniques based on load balancing across the integer functional units with or without rotation of functional unit usage. As leakage power is exponentially dependent on temperature and temperature is dependent on total power (i.e., switching and leakage), in our techniques, we also consider leakage power optimization by IPC tuning (instructions issued per cycle). By taking a code that is already scheduled for maximum performance as input, our scheduling strategies modify this performance-oriented schedule for balanced thermal behavior with negligible performance degradation. We simulate our scheduling strategies using a framework that consists of the Trimaran infrastructure, a power model, and the HotSpot. Our experimental results using several benchmark programs reveal that the peak temperature can be reduced through compiler scheduling. Madhu Mutyam, Feihui Li, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin |
LCTES | 2 |
| 2006 | Reducing NoC energy consumption through compiler-directed channel voltage scalingabstractWhile scalable NoC (Network-on-Chip) based communication architectures have clear advantages over long point-to-point communication channels, their power consumption can be very high. In contrast to most of the existing hardware-based efforts on NoC power optimization, this paper proposes a compiler-directed approach where the compiler decides the appropriate voltage/frequency levels to be used for each communication channel in the NoC. Our approach builds and operates on a novel graph based representation of a parallel program and has been implemented within an optimizing compiler and tested using 12 embedded benchmarks. Our experiments indicate that the proposed approach behaves better - from both performance and power perspectives - than a hardwarebased scheme and the energy savings it achieves are very close to the savings that could be obtained from an optimal, but hypothetical voltage/frequency scaling scheme. Guangyu Chen, Feihui Li, Mahmut T. Kandemir, Mary Jane Irwin |
PLDI | 2 |
| 2006 | Compiler-directed channel allocation for saving power in on-chip networksabstractIncreasing complexity in the communication patterns of embedded applications parallelized over multiple processing units makes it difficult to continue using the traditional bus-based on-chip communication techniques. The main contribution of this paper is to demonstrate the importance of compiler technology in reducing power consumption of applications designed for emerging multi processor, NoC (Network-on-Chip) based embedded systems. Specifically, we propose and evaluate a compiler-directed approach to NoC power management in the context of array-intensive applications, used frequently in embedded image/video processing. The unique characteristic of the compiler-based approach proposed in this paper is that it increases the idle periods of communication channels by reusing the same set of channels for as many communication messages as possible. The unused channels in this case take better advantage of the underlying power saving mechanism employed by the network architecture. However, this channel reuse optimization should be applied with care as it can hurt performance if two or more simultaneous communications are mapped onto the same set of channels. Therefore, the problem addressed in this paper is one of reducing the number of channels used to implement a set of communications without increasing the communication latency significantly. To test the effectiveness of our approach, we implemented it within an optimizing compiler and performed experiments using twelve application codes and a network simulation environment. Our experiments show that the proposed compiler-based approach is very successful in practice and works well under both hardware based and software based channel turn-off schemes. Guangyu Chen, Feihui Li, Mahmut T. Kandemir |
POPL | 2 |
| 2005 | Increasing FPGA resilience against soft errors using task duplicationabstractReconfigurable computing systems are becoming increasingly widespread as they bring the flexibility of programmable systems and approach the performance of ASICs. While the prior research on FPGAs mainly studied issues such as performance, power, and area optimization, reliability related issues have not taken much attention. However, with increasing soft error rates, providing resilience to soft errors in FPGA based embedded platforms is becoming an increasingly important issue. This paper proposes an OS-directed task duplication scheme for increasing reliability by providing resilience against soft errors. The idea is to exploit the unused portions of the FPGA space to schedule duplicates of active tasks. The outputs of the primary and duplicate tasks are compared to check for the existence of soft errors. Guangyu Chen, Feihui Li, Mahmut T. Kandemir, I. Demirkiran |
ASP-DAC | 2 |
| 2005 | Using data replication to reduce communication energy on chip multiprocessorsabstractChip multiprocessors are gaining popularity as they are very suitable for data-intensive embedded and high-end processing. In particular, array-intensive embedded image and video applications can benefit a lot from these architectures due to coarse-grain parallelization they offer. However, if not optimized, interprocessor communication can be a major energy consumer. Focusing on a distributed memory chip multiprocessor architecture and array-intensive embedded applications, this paper proposes a compiler-based communication minimization strategy based on data replication. The proposed scheme replicates shared data items across the memories of the processors in a controlled fashion (i.e., under a memory limit), with the goal of eliminating the otherwise necessary interprocessor communication. Mahmut T. Kandemir, Guangyu Chen, Feihui Li, I. Demirkiran |
ASP-DAC | 3 |
| 2005 | Using loop invariants to fight soft errors in data cachesabstractEver scaling process technology makes embedded systems more vulnerable to soft errors than in the past. One of the generic methods used to fight soft errors is based on duplicating instructions either in the spatial or temporal domain and then comparing the results to see whether they are different. This full duplication based scheme, though effective, is very expensive in terms of performance, power, and memory space. In this paper, we propose an alternate scheme based on loop invariants and present experimental results which show that our approach catches 62% of the errors caught by full duplication, when averaged over all benchmarks tested. In addition, it reduces the execution cycles and memory demand of the full duplication strategy by 80% and 4%, respectively. Sri Hari Krishna Narayanan, Seung Woo Son 0001, Mahmut T. Kandemir, Feihui Li |
ASP-DAC | 4 |
| 2005 | Compiler-directed proactive power management for networksabstractIncreasing use of parallel computation platforms (both off-chip and on-chip) makes communication analysis and optimization an important target. While there have been numerous studies that target network performance of parallel architectures, the efforts that target network power consumption (in terms of both modeling and optimization) are relatively new. One of the common characteristics of most of the prior approaches to network power management is that they are hardware-based and reactive in the sense that they manage power consumption of the network as a response to observed message traffic. Consequently, they can miss important opportunities for saving power and can incur performance penalties due to inaccuracies in predicting future idle and active times of communication links. Motivated by this observation, this paper proposes a compiler-directed proactive approach to network power management for the class of loop-intensive applications running on small-sized networks used exclusively by a single embedded application at a time.As compared to hardware-based approaches, the proposed compiler-directed approach has two potential benefits. First, based on high-level communication analysis, it determines the points at which a given communication link is idle and can be turned off (i.e., powered down) to save power. Therefore, an idle link can be put in the low-power state without waiting for a certain period of time to make sure that the link has really become idle (as in the case of hardware schemes). Second, since the compiler can also determine the point at which a turned-off link will be needed in the future, it can pre-activate it (i.e., before it is actually needed) to eliminate the turn on (reactivation) performance penalty. Our simulations with seven array-intensive applications and an embedded on-chip network clearly show that the proposed compiler-directed approach is better than a hardware-based scheme from both power and performance perspectives. Feihui Li, Guangyu Chen, Mahmut T. Kandemir, Mary Jane Irwin |
CASES | 1 |
| 2005 | A Compiler-Based Approach to Data Security
Feihui Li, Guilin Chen, Mahmut T. Kandemir, Richard R. Brooks |
CC | 1 |
| 2005 | Locality-conscious workload assignment for array-based computations in MPSOC architecturesabstractWhile the past research discussed several advantages of multipro-cessor-system-on-a-chip (MPSOC) architectures from both area uti-lization and design verification perspectives over complex single core based systems, compilation issues for these architectures have relatively received less attention. Programming MPSOCs can be challenging as several potentially conflicting issues such as data locality, parallelism and load balance across processors should be considered simultaneously. Most of the compilation techniques discussed in the literature for parallel architectures (not necessar-ily for MPSOCs) are loop based, i.e., they consider each loop nest in isolation. However, one key problem associated with such loop based techniques is that they fail to capture the interactions be-tween the different loop nests in the application. This paper takes a more global approach to the problem and proposes a compiler-driven data locality optimization strategy in the context of embed-ded MPSOCs. An important characteristic of the proposed ap-proach is that, in deciding the workloads of the processors (i.e., in parallelizing the application) it considers all the loop nests in the application simultaneously. Our experimental evaluation with eight embedded applications shows that the global scheme brings signif-icant power/performance benefits over the conventional loop based scheme. Feihui Li, Mahmut T. Kandemir |
DAC | 1 |
| 2005 | Compiler-Directed Instruction Duplication for Soft Error DetectionabstractWe experiment with compiler-directed instruction duplication to detect soft errors in VLIW datapaths. In the proposed approach, the compiler determines the instruction schedule by balancing the permissible performance degradation with the required degree of duplication. Our experimental results show that our algorithms allow the designer to perform tradeoff analysis between performance and reliability. Jie S. Hu, Feihui Li, Vijay Degalahal, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Mary Jane Irwin |
DATE | 2 |
| 2005 | Studying Storage-Recomputation Tradeoffs in Memory-Constrained Embedded ProcessingabstractFueled by an unprecedented desire for convenience and self-service, consumers are embracing embedded technology solutions that enhance their mobile lifestyles. Consequently, we witness an unprecedented proliferation of embedded/mobile applications. Most of the environments that execute these applications have severe power, performance, and memory space constraints that need to be accounted for. In particular, memory limitations can present serious challenges to embedded software designers. The current solutions to this problem include sophisticated packaging techniques and code optimizations for effective memory utilization. While the first solution is not scalable, the second one is restricted by intrinsic data dependences in the code that prevent code restructuring. In this paper, we explore an alternate approach for reducing memory space requirements of embedded applications. The idea is to re-compute the result of a code block (potentially multiple times) instead of storing it in memory and performing a memory operation whenever needed. The main benefit of this approach is that it reduces memory space requirements, that is, no memory space is reserved for storing the result of the code block in question. Mahmut T. Kandemir, Feihui Li, Guilin Chen, Guangyu Chen, Ozcan Ozturk 0001 |
DATE | 2 |
| 2005 | Exploiting last idle periods of links for network power managementabstractNetwork power optimization is becoming increasingly important as the sizes of the data manipulated by parallel applications and the complexity of inter-processor data communications are continuously increasing. Several hardware-based schemes have been proposed in the past for reducing network power consumption, either by turning off unused communication links or by lowering voltage/frequency in links with low usage. While the prior research shows that these schemes can be effective in certain cases, they share the common drawback of not being able to predict the link active and idle times very accurately. This paper, instead, proposes a compiler-based scheme that determines the last use of communication links at each loop nest and inserts explicit link turn-off calls in the application source. Specifically, for each loop nest, the compiler inserts a turn-off call per communication link. Each turned-off link is reactivated upon the next access to it. We automated this approach within a parallelizing compiler and applied it to eight array-intensive embedded applications. Feihui Li, Guilin Chen, Mahmut T. Kandemir, Mustafa Karaköy |
EMSOFT | 1 |
| 2005 | Compiler-directed voltage scaling on communication links for reducing power consumptionabstractReducing power consumption of communication networks is an important optimization goal in many application domains, ranging from large-scale simulation codes to embedded multi-media applications. Most of the prior efforts on network power optimization are hardware-based schemes. These schemes are predictive by definition as they control communication link status based on the observations made In the past. Since prediction may not be very accurate most of the time, these approaches can result in overheads in terms of both performance and power. This paper proposes a compiler driven approach to communication link voltage management. In this approach, an optimizing compiler analyzes the application code and extracts the data communication pattern among parallel processors. This information along with network topology is used for identifying the link access patterns. These patterns and the inherent data dependence information of the underlying code help the compiler decide the optimum voltages/frequencies to be used for communication links at a given time frame. Our focus in this work is on loop-intensive codes which frequently appear in data intensive video and image processing. We exploit the regularity in data accesses of these codes to abstract out their inter-processor communication patterns, which In turn enable us select the most appropriate voltage/frequency level to employ for each communication link at any time. Feihui Li, Guilin Chen, Mahmut T. Kandemir |
ICCAD | 1 |
| 2005 | Improving scratch-pad memory reliability through compiler-guided data block duplicationabstractRecent trends in embedded computing indicates an increasing use of scratch-pad memories (SPMs) as on-chip store for instructions and data. An important characteristic of these memory components is that they are managed by software, instead of hardware. Ever-scaling process technology and employment of several power-saving techniques in embedded systems (e.g., voltage scaling) make these systems particularly vulnerable to soft errors and other transient errors. Therefore, it is very important in practice to consider the impact of soft errors in SPMs. While it is possible to employ classical memory protection mechanisms such as parity checks and ECC, each of these has its drawbacks. Specifically, a pure parity-based protection cannot correct any errors, and ECCs can be an overkill in the normal operation state when no soft error is experienced. This paper proposes an alternate approach to protect SPMs against soft errors. The proposed approach is based on data block duplication under compiler control. More specifically, an optimizing compiler duplicates data blocks within the SPM and protects each data block by parity if such a duplication does not hurt performance. The goal of this scheme is to provide only parity protection for data blocks (and reduce the overheads at runtime when no error occurs) but correct errors using the duplicate (when an error occurs in the primary copy), provided that the duplicate is not corrupted. Feihui Li, Guilin Chen, Mahmut T. Kandemir, Ibrahim Kolcu |
ICCAD | 1 |
| 2004 | Improving Memory Performance of Embedded Java Applications by Dynamic Layout ModificationsabstractSummary form only given. Unlike desktop systems, embedded systems use different user interface technologies; have significantly smaller form factors; use a wide variety of processors; and have very tight constraints on power/energy consumption, user response time, and physical space. With its platform independence and secure execution environment, Java is fast becoming the language of choice for programming embedded systems. In order to extend its use to array based applications from embedded image and video processing, Java programmers need to employ several optimizations. Unfortunately, due to its precise exception mechanism and bytecode distribution form, it is not generally possible to use classical loop based optimization techniques for array based embedded Java applications. Observing this, this paper proposes a dynamic memory layout optimization strategy for Java applications. The strategy is based on observing the cache behavior dynamically and transforming memory layouts of arrays, when necessary, during the course of execution. This is in contrast to many previously proposed memory layout optimization strategies, which are static in nature (i.e., they are applied at compile time). Our results indicate large performance improvements on a suite of seven array based applications. Feihui Li, Pyush Agrawal, Grace Eberhardt, Eren Manavoglu, Secil Ugurel, Mahmut T. Kandemir |
IPDPS | 1 |
| 2004 | Improving Performance of Java Applications Using a CoprocessorabstractSummary form only given. Java technology is becoming a significant force stimulating the evolution of embedded systems. One of the problems in front of it is its poor performance. An important component of this performance problem is the overhead time spent during dynamic compilation. We demonstrate that a dual-processor heterogeneous architecture that employs both a conventional CPU and a Java coprocessor dedicated to compilation can be very effective in hiding the time spent in dynamic compilation. We propose two different profile-driven algorithms that make use of such a coprocessor based system, and evaluate them using a set of Java benchmarks. Our results clearly indicate that it is possible to improve overall performance by intelligently scheduling dynamic compilations on the coprocessor. Feihui Li, Mahmut T. Kandemir |
IPDPS | 1 |