Yu Du 0002

dblp:27/6228-2 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
0since 2021 · last 2016
0000-0002-2239-4659ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 3 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Memory systems · 80% Storage systems · 6% Interconnection networks and networks-on-chip · 5%

Topics — the 27 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
0.742016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Bit mapping for balanced PCM cell programming · ISCA 2013
Memory systems
cache management
0.742016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
Memory systems › non-volatile memory
phase change memory
0.532013
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Bit mapping for balanced PCM cell programming · ISCA 2013
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
Memory systems › cache management › cache partitioning
last-level cache partitioning
0.422016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
Memory systems › DRAM › DRAM architecture
3D-stacked DRAM
0.322013
Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013
Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access
0.322013
Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013
Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009
Memory systems
memory bandwidth management
0.212016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Memory systems › memory management › virtual memory
address translation
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Memory systems › memory management › virtual memory
huge pages
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Memory systems › memory management
physical memory management
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Memory systems › memory management
virtual memory
0.212015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Storage systems › data compression
delta compression
0.212013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Memory systems › cache
DRAM cache
0.212013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Distributed systems
fault tolerance
0.212013
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Memory systems › hybrid memory
hybrid main memory
0.212013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Storage systems › flash and SSD › flash memory management
wear leveling
0.212013
Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013
Interconnection networks and networks-on-chip
3d network-on-chip
0.112009
A low-radix and low-diameter 3D interconnection network design · HPCA 2009
Interconnection networks and networks-on-chip › network topology
low-diameter topology
0.112009
A low-radix and low-diameter 3D interconnection network design · HPCA 2009
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology
0.112009
A low-radix and low-diameter 3D interconnection network design · HPCA 2009
Processor architecture and microarchitecture
chip multiprocessor
0.122013
Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013
Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009
Parallel and multicore computing › multiprocessor system
multicore resource management
0.112016
Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016
Hardware reliability and fault tolerance
memory fault tolerance
0.112015
Supporting superpages in non-contiguous physical memory · HPCA 2015
Processor architecture and microarchitecture › multi-chip architecture
3d stacking
0.012013
Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013
Memory systems
memory compression
0.012013
Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013
Memory systems › cache management
cache replacement
0.012012
Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012
Processor architecture and microarchitecture › chip multiprocessor
3d chip multiprocessor
0.012009
Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009
Hardware reliability and fault tolerance
process variation
0.012009
Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009

Methods — techniques the papers use, named apart from their topics

theoretical model · 0.2exhaustive search · 0.2approximate scheme · 0.2page table compression · 0.2gap-tolerant sequential mapping · 0.2predictive compression · 0.2line-level mapping · 0.2delta compression · 0.2cache migration · 0.2writeback-aware partitioning · 0.1
YearPublicationVenuePosition
2016 Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems
abstract
In a multicore system, many applications share the last-level cache (LLC) and memory bandwidth. These resources need to be carefully managed in a coordinated way to maximize performance. DRAM is still the technology of choice in most systems. However, as traditional DRAM technology faces energy, reliability, and scalability challenges, nonvolatile memory (NVM) technologies are gaining traction. While DRAM is read/write symmetric (a read operation has comparable latency and energy consumption as a write operation), many NVM technologies (such as Phase-Change Memory, PCM) experience read/write asymmetry: write operations are typically much slower and more power hungry than read operations. Whether the memory’s characteristics are symmetric or asymmetric influences the way shared resources are managed. We propose two symmetry-agnostic schemes to manage a shared LLC through way partitioning and memory through bandwidth allocation. The proposals work well for both symmetric and asymmetric memory. First, an exhaustive search is proposed to find the best combination of a cache way partition and bandwidth allocation. Second, an approximate scheme, derived from a theoretical model, is proposed without the overhead of exhaustive search. Simulation results show that the approximate scheme improves weighted speedup by at least 14% on average (regardless of the memory symmetry) over a state-of-the-art way partitioning and memory bandwidth allocation. Simulation results also show that the approximate scheme achieves comparable weighted speedup as a state-of-the-art multiple resource management scheme, XChange, for symmetric memory, and outperforms it by an average of 10% for asymmetric memory.
Miao Zhou, Yu Du 0002, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
ACM Trans. Archit. Code Optim.2
2015 Supporting superpages in non-contiguous physical memory
abstract
For memory-intensiv e workloads with large memory footprints, superpages are effective to avoid address translation overhead, which can be a critical performance bottleneck. A superpage is a large virtual memory page that is mapped to an equivalently-sized amount of contiguous physical memory pages. Superpage mapping assumes physical memory does not contain retired pages, which is an important technique to improve memory resilience: the OS avoids allocating physical pages that have detected errors. Retired pages create unusable "holes" in the physical memory. We show that even a small percentage of retired pages makes it very difficult to find enough contiguous memory to form superpages. To address this problem, we propose GTSM, or gap-tolerant sequential mapping, that allows superpages to be formed even in the presence of retired physical pages. A new page table format is also proposed to support GTSM. This format has similar storage efficiency as traditional superpaging to hold address translations in the last-level cache. To further compress the page table and improve cache hit rates for address translation in large memory footprint workloads, we also propose an extended format that reduces the page table size by 50%. In comparison to an ideal memory without any retired physical pages, we show that our technique, with retired pages, achieves nearly 96.8% of the performance of traditional 2MB superpaging.
Yu Du 0002, Miao Zhou, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
HPCA1
2014 Errata to "Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor"
abstract
In the above-named articlt that appeared in ibid., vol. 62, no. 11, pp. 2252-2265, 2013, a production error occurred which resulted in the misalignment of Fig. 13, Fig. 14, Fig. 15, Fig. 16, Fig. 17, and Fig. 18 with their captions, starting from Fig. 13 to Fig. 18. As a result, a correct Fig. 13 is missing, and Fig. 18 repeats Fig. 19. We regret that this has happened. The correct figures with their corresponding captions are shown here.
Bo Zhao 0007, Yu Du 0002, Jun Yang 0002, Youtao Zhang
IEEE Trans. Computers2
2013 Writeback-aware bandwidth partitioning for multi-core systems with PCM
abstract
Phase-Change Memory (PCM) has emerged as a promising low-power candidate to replace DRAM in main memory. Hybrid memory architecture comprised of a large PCM and a small DRAM is a popular solution to mitigate undesirable characteristics of PCM writes. Because PCM writes are much slower than reads, writebacks from the last-level cache consume a large portion of memory bandwidth, and thus, impact performance. Effectively utilizing shared resources, such as the last-level cache and the memory bandwidth, is crucial to achieving high performance for multi-core systems. Although existing memory bandwidth allocation schemes improve system performance, no current approach uses writeback information to partition bandwidth for hybrid memory. We use a writeback-aware analytic model to derive the allocation strategy for bandwidth partitioning of phase-change memory. From the derivation of the model, Writeback-aware Bandwidth Partitioning (WBP) is proposed as a new runtime mechanism to partition PCM service cycles among applications. WBP uses a partitioning weight to indicate the importance of writebacks (in addition to LLC misses) to bandwidth allocation. A companion Dynamic Weight Adjustment (DWA) scheme dynamically selects the partitioning weight to maximize system performance. Simulation results show that WBP and DWA improve performance by 24.9% (weighted speedup) over bandwidth partitioning schemes that do not take writebacks into consideration in a 8-core system.
Miao Zhou, Yu Du 0002, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
PACT2
2013 Bit mapping for balanced PCM cell programming
abstract
Write bandwidth is an inherent performance bottleneck for Phase Change Memory (PCM) for two reasons. First, PCM cells have long programming time, and second, only a limited number of PCM cells can be programmed concurrently due to programming current and write circuit constraints,
Yu Du 0002, Miao Zhou, Bruce R. Childers, Daniel Mossé, Rami G. Melhem
ISCA1
2013 Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory
abstract
Limited PCM write bandwidth is a critical obstacle to achieve good performance from hybrid DRAM/PCM memory systems. The write bandwidth is severely restricted in PCM devices, which harms application performance. Indeed, as we show, it is more important to reduce PCM write traffic than to reduce PCM read latency for application performance. To reduce the number of PCM writes, we propose a DRAM cache organization that employs compression. A new delta compression technique for modified data is used to achieve a large compression ratio. Our approach can selectively and predictively apply compression to improve its efficiency and performance. Our approach is designed to facilitate adoption in existing main memory compression frameworks. We describe an instance of how to incorporate delta compression in IBM's MXT memory compression architecture when used for DRAM cache in a hybrid main memory. For fourteen representative memory-intensive workloads, on average, our delta compression technique reduces the number of PCM writes by 54.3%, and improves IPC performance by 24.4%.
Yu Du 0002, Miao Zhou, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
ACM Trans. Archit. Code Optim.1
2013 Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory
abstract
Phase Change Memory (PCM) has recently emerged as a promising memory technology. However, PCM’s limited write endurance restricts its immediate use as a replacement for DRAM. To extend the lifetime of PCM chips, wear-leveling and salvaging techniques have been proposed. Wear-leveling balances write operations across different PCM regions while salvaging extends the duty cycle and provides graceful degradation for a nonnegligible number of failures. Current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a noncontiguous PCM space generated from salvaging complicates wear-leveling and incurs large overhead. In this article, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS integrates wear-leveling and salvaging and copes well with modern OSes. Our experimental results show that LLS achieves 31% longer lifetime than the state-of-the-art. It has negligible hardware cost and performance overhead.
Lei Jiang 0001, Yu Du 0002, Bo Zhao 0007, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
ACM Trans. Archit. Code Optim.2
2013 Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor
abstract
Process variations in integrated circuits have significant impact on their performance, leakage, and stability. This is particularly evident in large, regular, and dense structures such as DRAMs. DRAMs are built using minimized transistors with presumably uniform speed in an organized array structure. Process variation can introduce latency disparity among different memory arrays. With the proliferation of 3D stacking technology, DRAMs become a favorable choice for stacking on top of a multicore processor as a last level cache for large capacity, high bandwidth, and low power. Hence, variations in bank speed create a unique problem of nonuniform cache accesses in 3D space. In this paper, we investigate cache management techniques for tolerating process variation in a 3D DRAM stacked onto a multicore processor. We modeled the process variation in a four-layer DRAM memory, including cell transistor, capacitor trench, and peripheral circuit, to characterize the latency and retention time variations among different banks. As a result, the notion of fast and slow banks from the core's standpoint is no longer associated with their physical distances with the banks. They are determined by the different bank latencies due to process variation. We develop cache migration schemes that utilize fast banks while limiting the cost due to migration. Our experiments show that there is a great performance benefit in exploiting fast memory banks through migration. On average, a variation-aware management can improve the performance of a workload over the baseline (where one of the slowest bank speed is assumed for all banks) by 16.5 percent. We are also only 0.8 percent away in performance from an ideal memory where no process variation is present.
Bo Zhao 0007, Yu Du 0002, Jun Yang 0002, Youtao Zhang
IEEE Trans. Computers2
2012 Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems
abstract
Phase-Change Memory (PCM) has emerged as a promising low-power main memory candidate to replace DRAM. The main problems of PCM are that writes are much slower and more power hungry than reads, write bandwidth is much lower than read bandwidth, and limited write endurance. Adding an extra layer of cache, which is logically the last-level cache (LLC), can mitigate the drawbacks of PCM. However, writebacks from the LLC might (a) overwhelm the limited PCM write bandwidth and stall the application, (b) shorten lifetime, and (c) increase energy consumption. Cache partitioning and replacement schemes are important to achieve high throughput for multi-core systems. However, we noted that no existing partitioning and replacement policy takes into account the writeback information. This paper proposes two writeback-aware schemes to manage the LLC for PCM main memory systems. Writeback-aware Cache Partitioning (WCP) is a runtime mechanism that partitions a shared LLC among multiple applications. Unlike past partitioning schemes, our scheme considers the reduction in cache misses as well as writebacks. Write Queue Balancing (WQB) replacement policy manages the cache partition of each application intelligently so that the writebacks are distributed evenly among PCM write queues. In this way, applications rarely stall due to unbalanced PCM write traffic among write queues. Our evaluation shows that WCP and WQB result in, on average, 21% improvement in throughput, 49% reduction in PCM writes, and 14% reduction in energy over a state-of-the-art cache partitioning scheme.
Miao Zhou, Yu Du 0002, Bruce R. Childers, Rami G. Melhem, Daniel Mossé
ACM Trans. Archit. Code Optim.2
2011 LLS: Cooperative integration of wear-leveling and salvaging for PCM main memory
abstract
Phase change memory (PCM) has emerged as a promising technology for main memory due to many advantages, such as better scalability, non-volatility and fast read access. However, PCM's limited write endurance restricts its immediate use as a replacement for DRAM. Recent studies have revealed that a PCM chip which integrates millions to billions of bit cells has non-negligible variations in write endurance. Wear leveling techniques have been proposed to balance write operations to different PCM regions. To further prolong the lifetime of a PCM device after the failure of weak cell, techniques have been proposed to remap failed lines to spares and to salvage a PCM device that has a large number of failed lines or pages with graceful degradation. However, current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a non-contiguous PCM space generated from salvaging complicates wear leveling and incurs large overhead. In this paper, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS seamlessly integrates wear leveling and salvaging and copes well with modern OSs, including ones that support multiple page sizes. Our experimental results show that LLS achieves 24% longer lifetime than a state-of-the-art technique. It has negligible hardware cost and performance overhead.
Lei Jiang 0001, Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
DSN2
2011 A composite and scalable cache coherence protocol for large scale CMPs
abstract
The number of on-chip cores of modern chip multiprocessors (CMPs) is growing fast with technology scaling. However, it remains a big challenge to efficiently support cache coherence for large scale CMPs. The conventional snoopy and directory coherence protocols cannot be smoothly scaled to many-core or thousand-core processors. Snoopy protocols introduce large power overhead due to enormous amount of cache tag probing triggered by broadcast. Directory protocols introduce performance penalty due to indirection, and large storage overhead due to storing directories. This paper addresses the efficiency problem when supporting cache coherency for large-scale CMPs. By leveraging emerging optical on-chip interconnect (OP-I) technology to provide high bandwidth density, low propagation delay and natural support for multicast/broadcast in a hierarchical network organization, we propose a composite cache coherence (C 3) protocol that benefits from direct cache-to-cache accesses as in snoopy protocol and small amount of cache probing as in directory protocol. Targeting at quickly completing coherence transactions, C 3 organizes accesses in a three-tier hierarchy by combining a mix of designs including local broadcast prediction, filtering, and a coarse-grained directory. Compared to directory-based protocol [18], our evaluations on a thousand-core CMP show that C 3 improves performance by 21%, reduces network latency of coherence messages by 41 % and saves network energy consumption by 5.5 % on average for PARSEC applications.
Yi Xu 0010, Yu Du 0002, Youtao Zhang, Jun Yang 0002
ICS2
2010 Fine-grained QoS scheduling for PCM-based main memory systems
abstract
With wide adoption of chip multiprocessors (CMPs) in modern computers, there is an increasing demand for large capacity main memory systems. The emerging PCM (Phase Change Memory) technology has unique power and scalability advantages and is regarded as a promising candidate among new memory technologies. When scheduling a mix of applications of different priority levels, it is often important to provide tunable QoS (Quality-of-Service) for the applications with high priority. However due to the slow PCM cell access, and the destructive interferences among concurrent applications, existing memory scheduling schemes lack the flexibility to tune QoS in a wide range, in particular to the level close or equal to that of standalone execution. In this paper we propose a novel QoS scheduling scheme that utilizes request preemption and row buffer partition that enable QoS tuning at a fine-granularity. That is, they can tune the request queuing time and the PCM bank service time for the high priority requests. Our experimental results show that the proposed scheme achieves 1.7× ~10× QoS tuning range while introducing negligible area and energy overheads.
Yu Du 0002, Youtao Zhang, Jun Yang 0002
IPDPS2
2009 Frequent value compression in packet-based NoC architectures
abstract
The proliferation of chip multiprocessors (CMPs) has led to the integration of large on-chip caches. For scalability reasons, a large on-chip cache is often divided into smaller banks that are interconnected through packet-based network-on-chip (NoC). With increasing number of cores and cache banks integrated on a single die, the on-chip network introduces significant communication latency and power consumption. In this paper, we propose a novel scheme that exploitsfrequentvaluecompression to optimize the power and performance of NoC. Our experimental results show that the proposed scheme reduces the router power by up to 16.7%, with CPI reduction as much as 23.5% in our setting. Comparing to the recent zero pattern compression scheme, thefrequentvaluescheme saves up to 11.0% more router power and has up to 14.5% more CPI reduction. Hardware design of the FV table and its overhead are also presented.
Bo Zhao 0007, Yu Du 0002, Yi Xu 0010, Youtao Zhang, Jun Yang 0002, Li Zhao 0002
ASP-DAC3
2009 A low-radix and low-diameter 3D interconnection network design
abstract
Interconnection plays an important role in performance and power of CMP designs using deep sub-micron technology. The network-on-chip (NoCs) has been proposed as a scalable and high-bandwidth fabric for interconnect design. The advent of the 3D technology has provided further opportunity to reduce on-chip communication delay. However, the design of the 3D NoC topologies has important distinctions from 2D NoCs or off-chip interconnection networks. First, current 3D stacking technology allows only vertical inter-layer links. Hence, there cannot be direct connections between arbitrary nodes in different layers — the vertical connection topology are essentially fixed. Second, the 3D NoC is highly constrained by the complexity and power of routers and links. Hence, low-radix routers are preferred over high-radix routers for lower power and better heat dissipation. This implies long network latency due to high hop counts in network paths. In this paper, we design a low-diameter 3D network using low-radix routers. Our topology leverages long wires to connect remote intra-layer nodes. We take advantage of the start-of-the-art one-hop vertical communication design and utilize lateral long wires to shorten network paths. Effectively, we implement a small-to-medium sized clique network in different layers of a 3D chip. The resulting topology generates a diameter of 3-hop only network, using routers of the same radix as 3D mesh routers. The proposed network shows up to 29% of network latency reduction, up to 10% throughput improvement, and up to 24% energy reduction, when compared to a 3D mesh network.
Yi Xu 0010, Yu Du 0002, Bo Zhao 0007, Xiuyi Zhou, Youtao Zhang, Jun Yang 0002
HPCA2
2009 Variation-tolerant non-uniform 3D cache management in die stacked multicore processor
abstract
Process variations in integrated circuits have significant impact on their performance, leakage and stability. This is particularly evident in large, regular and dense structures such as DRAMs. DRAMs are built using minimized transistors with presumably uniform speed in an organized array structure. Process variation can introduce latency disparity among different memory arrays. With the proliferation of 3D stacking technology, DRAMs become a favorable choice for stacking on top of a multicore processor as a last level cache for large capacity, high bandwidth, and low power. Hence, variations in bank speed creates a unique problem of non-uniform cache accesses in 3D space.
Bo Zhao 0007, Yu Du 0002, Youtao Zhang, Jun Yang 0002
MICRO2
2008 Adaptive Buffer Management for Efficient Code Dissemination in Multi-Application Wireless Sensor Networks
abstract
Future wireless sensor networks (WSNs) are projected to run multiple applications in the same network infrastructure. While such multi-application WSNs (MA-WSNs) are economically more efficient and adapt better to the changing environments than traditional single-application WSNs, they usually require frequent code redistribution on wireless sensors, making it critical to design energy efficient post-deployment code dissemination protocols in MA-WSNs. Different applications in MA-WSNs often share some common code segments. Therefore when there is a need to disseminate a new application from the sink node, it is possible to disseminate its shared code segments from peer sensors instead of disseminating everything from the sink node. While dissemination protocols have been proposed to handle code of each single type, it is challenging to achieve energy efficiency when the code contains both types and needs simultaneous dissemination. In this paper we utilize an adaptive buffer management approach to achieve efficient code dissemination in MA-WSNs. Our experimental results show that adaptive buffer management can reduce the completion time and the message overhead up to 10% and 20% respectively.
Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002
EUC (1)2
2008 Thermal Management for 3D Processors via Task Scheduling
abstract
A rising horizon in chip fabrication is the 3D integration technology. It stacks two or more dies vertically with a dense, high-speed interface to increase the device density and reduce the delay of interconnects across the dies. However, a major challenge in 3D technology is the increased power density which brings the concern of heat dissipation within the processor. High temperatures trigger voltage and frequency throttlings in hardware which degrade the chip performance. Moreover, high temperatures impair the processorpsilas reliability and reduce its lifetime. To alleviate this problem, we propose in this paper an OS-level scheduling algorithm that performs thermal-aware task scheduling on a 3D chip. Our algorithm leverages the inherent thermal variations within and across different tasks, and schedules them to keep the chip temperature low. We observed that vertically adjacent dies have strong thermal correlations, and the scheduler should consider them jointly. Our proposed algorithm can remove on average 54% of hardware DTMs and result in 7.2% performance improvement over the base case.
Xiuyi Zhou, Yi Xu 0010, Yu Du 0002, Youtao Zhang, Jun Yang 0002
ICPP3