VLDB 2026 Research / reviewers in the wild / expert
Yu Du 0002
dblp:27/6228-2
· DBLP profile ↗
17ranked-venue papers
3as first author
0since 2021 · last 2016
0000-0002-2239-4659ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Memory systems · 80% Storage systems · 6% Interconnection networks and networks-on-chip · 5% |
Topics — the 27 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
0.7 | 4 | 2016 | Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016 Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013 Bit mapping for balanced PCM cell programming · ISCA 2013 |
Memory systems
cache management |
0.7 | 4 | 2016 | Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016 Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013 Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › non-volatile memory
phase change memory |
0.5 | 3 | 2013 | Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013 Bit mapping for balanced PCM cell programming · ISCA 2013 Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › cache management › cache partitioning
last-level cache partitioning |
0.4 | 2 | 2016 | Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016 Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › DRAM › DRAM architecture
3D-stacked DRAM |
0.3 | 2 | 2013 | Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013 Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009 |
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access |
0.3 | 2 | 2013 | Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013 Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009 |
Memory systems
memory bandwidth management |
0.2 | 1 | 2016 | Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016 |
Memory systems › memory management › virtual memory
address translation |
0.2 | 1 | 2015 | Supporting superpages in non-contiguous physical memory · HPCA 2015 |
Memory systems › memory management › virtual memory
huge pages |
0.2 | 1 | 2015 | Supporting superpages in non-contiguous physical memory · HPCA 2015 |
Memory systems › memory management
physical memory management |
0.2 | 1 | 2015 | Supporting superpages in non-contiguous physical memory · HPCA 2015 |
Memory systems › memory management
virtual memory |
0.2 | 1 | 2015 | Supporting superpages in non-contiguous physical memory · HPCA 2015 |
Storage systems › data compression
delta compression |
0.2 | 1 | 2013 | Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013 |
Memory systems › cache
DRAM cache |
0.2 | 1 | 2013 | Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013 |
Distributed systems
fault tolerance |
0.2 | 1 | 2013 | Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013 |
Memory systems › hybrid memory
hybrid main memory |
0.2 | 1 | 2013 | Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013 |
Storage systems › flash and SSD › flash memory management
wear leveling |
0.2 | 1 | 2013 | Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change Memory · ACM Trans. Archit. Code Optim. 2013 |
Interconnection networks and networks-on-chip
3d network-on-chip |
0.1 | 1 | 2009 | A low-radix and low-diameter 3D interconnection network design · HPCA 2009 |
Interconnection networks and networks-on-chip › network topology
low-diameter topology |
0.1 | 1 | 2009 | A low-radix and low-diameter 3D interconnection network design · HPCA 2009 |
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology |
0.1 | 1 | 2009 | A low-radix and low-diameter 3D interconnection network design · HPCA 2009 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2013 | Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013 Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009 |
Parallel and multicore computing › multiprocessor system
multicore resource management |
0.1 | 1 | 2016 | Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore Systems · ACM Trans. Archit. Code Optim. 2016 |
Hardware reliability and fault tolerance
memory fault tolerance |
0.1 | 1 | 2015 | Supporting superpages in non-contiguous physical memory · HPCA 2015 |
Processor architecture and microarchitecture › multi-chip architecture
3d stacking |
0.0 | 1 | 2013 | Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor · IEEE Trans. Computers 2013 |
Memory systems
memory compression |
0.0 | 1 | 2013 | Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memory · ACM Trans. Archit. Code Optim. 2013 |
Memory systems › cache management
cache replacement |
0.0 | 1 | 2012 | Writeback-aware partitioning and replacement for last-level caches in phase change main memory systems · ACM Trans. Archit. Code Optim. 2012 |
Processor architecture and microarchitecture › chip multiprocessor
3d chip multiprocessor |
0.0 | 1 | 2009 | Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009 |
Hardware reliability and fault tolerance
process variation |
0.0 | 1 | 2009 | Variation-tolerant non-uniform 3D cache management in die stacked multicore processor · MICRO 2009 |
Methods — techniques the papers use, named apart from their topics
theoretical model · 0.2exhaustive search · 0.2approximate scheme · 0.2page table compression · 0.2gap-tolerant sequential mapping · 0.2predictive compression · 0.2line-level mapping · 0.2delta compression · 0.2cache migration · 0.2writeback-aware partitioning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Symmetry-Agnostic Coordinated Management of the Memory Hierarchy in Multicore SystemsabstractIn a multicore system, many applications share the last-level cache (LLC) and memory bandwidth. These resources need to be carefully managed in a coordinated way to maximize performance. DRAM is still the technology of choice in most systems. However, as traditional DRAM technology faces energy, reliability, and scalability challenges, nonvolatile memory (NVM) technologies are gaining traction. While DRAM is read/write symmetric (a read operation has comparable latency and energy consumption as a write operation), many NVM technologies (such as Phase-Change Memory, PCM) experience read/write asymmetry: write operations are typically much slower and more power hungry than read operations. Whether the memory’s characteristics are symmetric or asymmetric influences the way shared resources are managed. We propose two symmetry-agnostic schemes to manage a shared LLC through way partitioning and memory through bandwidth allocation. The proposals work well for both symmetric and asymmetric memory. First, an exhaustive search is proposed to find the best combination of a cache way partition and bandwidth allocation. Second, an approximate scheme, derived from a theoretical model, is proposed without the overhead of exhaustive search. Simulation results show that the approximate scheme improves weighted speedup by at least 14% on average (regardless of the memory symmetry) over a state-of-the-art way partitioning and memory bandwidth allocation. Simulation results also show that the approximate scheme achieves comparable weighted speedup as a state-of-the-art multiple resource management scheme, XChange, for symmetric memory, and outperforms it by an average of 10% for asymmetric memory. Miao Zhou, Yu Du 0002, Bruce R. Childers, Daniel Mossé, Rami G. Melhem |
ACM Trans. Archit. Code Optim. | 2 |
| 2015 | Supporting superpages in non-contiguous physical memoryabstractFor memory-intensiv e workloads with large memory footprints, superpages are effective to avoid address translation overhead, which can be a critical performance bottleneck. A superpage is a large virtual memory page that is mapped to an equivalently-sized amount of contiguous physical memory pages. Superpage mapping assumes physical memory does not contain retired pages, which is an important technique to improve memory resilience: the OS avoids allocating physical pages that have detected errors. Retired pages create unusable "holes" in the physical memory. We show that even a small percentage of retired pages makes it very difficult to find enough contiguous memory to form superpages. To address this problem, we propose GTSM, or gap-tolerant sequential mapping, that allows superpages to be formed even in the presence of retired physical pages. A new page table format is also proposed to support GTSM. This format has similar storage efficiency as traditional superpaging to hold address translations in the last-level cache. To further compress the page table and improve cache hit rates for address translation in large memory footprint workloads, we also propose an extended format that reduces the page table size by 50%. In comparison to an ideal memory without any retired physical pages, we show that our technique, with retired pages, achieves nearly 96.8% of the performance of traditional 2MB superpaging. Yu Du 0002, Miao Zhou, Bruce R. Childers, Daniel Mossé, Rami G. Melhem |
HPCA | 1 |
| 2014 | Errata to "Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor"abstractIn the above-named articlt that appeared in ibid., vol. 62, no. 11, pp. 2252-2265, 2013, a production error occurred which resulted in the misalignment of Fig. 13, Fig. 14, Fig. 15, Fig. 16, Fig. 17, and Fig. 18 with their captions, starting from Fig. 13 to Fig. 18. As a result, a correct Fig. 13 is missing, and Fig. 18 repeats Fig. 19. We regret that this has happened. The correct figures with their corresponding captions are shown here. Bo Zhao 0007, Yu Du 0002, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Computers | 2 |
| 2013 | Writeback-aware bandwidth partitioning for multi-core systems with PCMabstractPhase-Change Memory (PCM) has emerged as a promising low-power candidate to replace DRAM in main memory. Hybrid memory architecture comprised of a large PCM and a small DRAM is a popular solution to mitigate undesirable characteristics of PCM writes. Because PCM writes are much slower than reads, writebacks from the last-level cache consume a large portion of memory bandwidth, and thus, impact performance. Effectively utilizing shared resources, such as the last-level cache and the memory bandwidth, is crucial to achieving high performance for multi-core systems. Although existing memory bandwidth allocation schemes improve system performance, no current approach uses writeback information to partition bandwidth for hybrid memory. We use a writeback-aware analytic model to derive the allocation strategy for bandwidth partitioning of phase-change memory. From the derivation of the model, Writeback-aware Bandwidth Partitioning (WBP) is proposed as a new runtime mechanism to partition PCM service cycles among applications. WBP uses a partitioning weight to indicate the importance of writebacks (in addition to LLC misses) to bandwidth allocation. A companion Dynamic Weight Adjustment (DWA) scheme dynamically selects the partitioning weight to maximize system performance. Simulation results show that WBP and DWA improve performance by 24.9% (weighted speedup) over bandwidth partitioning schemes that do not take writebacks into consideration in a 8-core system. Miao Zhou, Yu Du 0002, Bruce R. Childers, Rami G. Melhem, Daniel Mossé |
PACT | 2 |
| 2013 | Bit mapping for balanced PCM cell programmingabstractWrite bandwidth is an inherent performance bottleneck for Phase Change Memory (PCM) for two reasons. First, PCM cells have long programming time, and second, only a limited number of PCM cells can be programmed concurrently due to programming current and write circuit constraints, Yu Du 0002, Miao Zhou, Bruce R. Childers, Daniel Mossé, Rami G. Melhem |
ISCA | 1 |
| 2013 | Delta-compressed caching for overcoming the write bandwidth limitation of hybrid main memoryabstractLimited PCM write bandwidth is a critical obstacle to achieve good performance from hybrid DRAM/PCM memory systems. The write bandwidth is severely restricted in PCM devices, which harms application performance. Indeed, as we show, it is more important to reduce PCM write traffic than to reduce PCM read latency for application performance. To reduce the number of PCM writes, we propose a DRAM cache organization that employs compression. A new delta compression technique for modified data is used to achieve a large compression ratio. Our approach can selectively and predictively apply compression to improve its efficiency and performance. Our approach is designed to facilitate adoption in existing main memory compression frameworks. We describe an instance of how to incorporate delta compression in IBM's MXT memory compression architecture when used for DRAM cache in a hybrid main memory. For fourteen representative memory-intensive workloads, on average, our delta compression technique reduces the number of PCM writes by 54.3%, and improves IPC performance by 24.4%. Yu Du 0002, Miao Zhou, Bruce R. Childers, Rami G. Melhem, Daniel Mossé |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change MemoryabstractPhase Change Memory (PCM) has recently emerged as a promising memory technology. However, PCM’s limited write endurance restricts its immediate use as a replacement for DRAM. To extend the lifetime of PCM chips, wear-leveling and salvaging techniques have been proposed. Wear-leveling balances write operations across different PCM regions while salvaging extends the duty cycle and provides graceful degradation for a nonnegligible number of failures. Current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a noncontiguous PCM space generated from salvaging complicates wear-leveling and incurs large overhead. In this article, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS integrates wear-leveling and salvaging and copes well with modern OSes. Our experimental results show that LLS achieves 31% longer lifetime than the state-of-the-art. It has negligible hardware cost and performance overhead. Lei Jiang 0001, Yu Du 0002, Bo Zhao 0007, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore ProcessorabstractProcess variations in integrated circuits have significant impact on their performance, leakage, and stability. This is particularly evident in large, regular, and dense structures such as DRAMs. DRAMs are built using minimized transistors with presumably uniform speed in an organized array structure. Process variation can introduce latency disparity among different memory arrays. With the proliferation of 3D stacking technology, DRAMs become a favorable choice for stacking on top of a multicore processor as a last level cache for large capacity, high bandwidth, and low power. Hence, variations in bank speed create a unique problem of nonuniform cache accesses in 3D space. In this paper, we investigate cache management techniques for tolerating process variation in a 3D DRAM stacked onto a multicore processor. We modeled the process variation in a four-layer DRAM memory, including cell transistor, capacitor trench, and peripheral circuit, to characterize the latency and retention time variations among different banks. As a result, the notion of fast and slow banks from the core's standpoint is no longer associated with their physical distances with the banks. They are determined by the different bank latencies due to process variation. We develop cache migration schemes that utilize fast banks while limiting the cost due to migration. Our experiments show that there is a great performance benefit in exploiting fast memory banks through migration. On average, a variation-aware management can improve the performance of a workload over the baseline (where one of the slowest bank speed is assumed for all banks) by 16.5 percent. We are also only 0.8 percent away in performance from an ideal memory where no process variation is present. Bo Zhao 0007, Yu Du 0002, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Computers | 2 |
| 2012 | Writeback-aware partitioning and replacement for last-level caches in phase change main memory systemsabstractPhase-Change Memory (PCM) has emerged as a promising low-power main memory candidate to replace DRAM. The main problems of PCM are that writes are much slower and more power hungry than reads, write bandwidth is much lower than read bandwidth, and limited write endurance. Adding an extra layer of cache, which is logically the last-level cache (LLC), can mitigate the drawbacks of PCM. However, writebacks from the LLC might (a) overwhelm the limited PCM write bandwidth and stall the application, (b) shorten lifetime, and (c) increase energy consumption. Cache partitioning and replacement schemes are important to achieve high throughput for multi-core systems. However, we noted that no existing partitioning and replacement policy takes into account the writeback information. This paper proposes two writeback-aware schemes to manage the LLC for PCM main memory systems. Writeback-aware Cache Partitioning (WCP) is a runtime mechanism that partitions a shared LLC among multiple applications. Unlike past partitioning schemes, our scheme considers the reduction in cache misses as well as writebacks. Write Queue Balancing (WQB) replacement policy manages the cache partition of each application intelligently so that the writebacks are distributed evenly among PCM write queues. In this way, applications rarely stall due to unbalanced PCM write traffic among write queues. Our evaluation shows that WCP and WQB result in, on average, 21% improvement in throughput, 49% reduction in PCM writes, and 14% reduction in energy over a state-of-the-art cache partitioning scheme. Miao Zhou, Yu Du 0002, Bruce R. Childers, Rami G. Melhem, Daniel Mossé |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | LLS: Cooperative integration of wear-leveling and salvaging for PCM main memoryabstractPhase change memory (PCM) has emerged as a promising technology for main memory due to many advantages, such as better scalability, non-volatility and fast read access. However, PCM's limited write endurance restricts its immediate use as a replacement for DRAM. Recent studies have revealed that a PCM chip which integrates millions to billions of bit cells has non-negligible variations in write endurance. Wear leveling techniques have been proposed to balance write operations to different PCM regions. To further prolong the lifetime of a PCM device after the failure of weak cell, techniques have been proposed to remap failed lines to spares and to salvage a PCM device that has a large number of failed lines or pages with graceful degradation. However, current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a non-contiguous PCM space generated from salvaging complicates wear leveling and incurs large overhead. In this paper, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS seamlessly integrates wear leveling and salvaging and copes well with modern OSs, including ones that support multiple page sizes. Our experimental results show that LLS achieves 24% longer lifetime than a state-of-the-art technique. It has negligible hardware cost and performance overhead. Lei Jiang 0001, Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
DSN | 2 |
| 2011 | A composite and scalable cache coherence protocol for large scale CMPsabstractThe number of on-chip cores of modern chip multiprocessors (CMPs) is growing fast with technology scaling. However, it remains a big challenge to efficiently support cache coherence for large scale CMPs. The conventional snoopy and directory coherence protocols cannot be smoothly scaled to many-core or thousand-core processors. Snoopy protocols introduce large power overhead due to enormous amount of cache tag probing triggered by broadcast. Directory protocols introduce performance penalty due to indirection, and large storage overhead due to storing directories. This paper addresses the efficiency problem when supporting cache coherency for large-scale CMPs. By leveraging emerging optical on-chip interconnect (OP-I) technology to provide high bandwidth density, low propagation delay and natural support for multicast/broadcast in a hierarchical network organization, we propose a composite cache coherence (C 3) protocol that benefits from direct cache-to-cache accesses as in snoopy protocol and small amount of cache probing as in directory protocol. Targeting at quickly completing coherence transactions, C 3 organizes accesses in a three-tier hierarchy by combining a mix of designs including local broadcast prediction, filtering, and a coarse-grained directory. Compared to directory-based protocol [18], our evaluations on a thousand-core CMP show that C 3 improves performance by 21%, reduces network latency of coherence messages by 41 % and saves network energy consumption by 5.5 % on average for PARSEC applications. Yi Xu 0010, Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
ICS | 2 |
| 2010 | Fine-grained QoS scheduling for PCM-based main memory systemsabstractWith wide adoption of chip multiprocessors (CMPs) in modern computers, there is an increasing demand for large capacity main memory systems. The emerging PCM (Phase Change Memory) technology has unique power and scalability advantages and is regarded as a promising candidate among new memory technologies. When scheduling a mix of applications of different priority levels, it is often important to provide tunable QoS (Quality-of-Service) for the applications with high priority. However due to the slow PCM cell access, and the destructive interferences among concurrent applications, existing memory scheduling schemes lack the flexibility to tune QoS in a wide range, in particular to the level close or equal to that of standalone execution. In this paper we propose a novel QoS scheduling scheme that utilizes request preemption and row buffer partition that enable QoS tuning at a fine-granularity. That is, they can tune the request queuing time and the PCM bank service time for the high priority requests. Our experimental results show that the proposed scheme achieves 1.7× ~10× QoS tuning range while introducing negligible area and energy overheads. Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
IPDPS | 2 |
| 2009 | Frequent value compression in packet-based NoC architecturesabstractThe proliferation of chip multiprocessors (CMPs) has led to the integration of large on-chip caches. For scalability reasons, a large on-chip cache is often divided into smaller banks that are interconnected through packet-based network-on-chip (NoC). With increasing number of cores and cache banks integrated on a single die, the on-chip network introduces significant communication latency and power consumption. In this paper, we propose a novel scheme that exploitsfrequentvaluecompression to optimize the power and performance of NoC. Our experimental results show that the proposed scheme reduces the router power by up to 16.7%, with CPI reduction as much as 23.5% in our setting. Comparing to the recent zero pattern compression scheme, thefrequentvaluescheme saves up to 11.0% more router power and has up to 14.5% more CPI reduction. Hardware design of the FV table and its overhead are also presented. Bo Zhao 0007, Yu Du 0002, Yi Xu 0010, Youtao Zhang, Jun Yang 0002, Li Zhao 0002 |
ASP-DAC | 3 |
| 2009 | A low-radix and low-diameter 3D interconnection network designabstractInterconnection plays an important role in performance and power of CMP designs using deep sub-micron technology. The network-on-chip (NoCs) has been proposed as a scalable and high-bandwidth fabric for interconnect design. The advent of the 3D technology has provided further opportunity to reduce on-chip communication delay. However, the design of the 3D NoC topologies has important distinctions from 2D NoCs or off-chip interconnection networks. First, current 3D stacking technology allows only vertical inter-layer links. Hence, there cannot be direct connections between arbitrary nodes in different layers — the vertical connection topology are essentially fixed. Second, the 3D NoC is highly constrained by the complexity and power of routers and links. Hence, low-radix routers are preferred over high-radix routers for lower power and better heat dissipation. This implies long network latency due to high hop counts in network paths. In this paper, we design a low-diameter 3D network using low-radix routers. Our topology leverages long wires to connect remote intra-layer nodes. We take advantage of the start-of-the-art one-hop vertical communication design and utilize lateral long wires to shorten network paths. Effectively, we implement a small-to-medium sized clique network in different layers of a 3D chip. The resulting topology generates a diameter of 3-hop only network, using routers of the same radix as 3D mesh routers. The proposed network shows up to 29% of network latency reduction, up to 10% throughput improvement, and up to 24% energy reduction, when compared to a 3D mesh network. Yi Xu 0010, Yu Du 0002, Bo Zhao 0007, Xiuyi Zhou, Youtao Zhang, Jun Yang 0002 |
HPCA | 2 |
| 2009 | Variation-tolerant non-uniform 3D cache management in die stacked multicore processorabstractProcess variations in integrated circuits have significant impact on their performance, leakage and stability. This is particularly evident in large, regular and dense structures such as DRAMs. DRAMs are built using minimized transistors with presumably uniform speed in an organized array structure. Process variation can introduce latency disparity among different memory arrays. With the proliferation of 3D stacking technology, DRAMs become a favorable choice for stacking on top of a multicore processor as a last level cache for large capacity, high bandwidth, and low power. Hence, variations in bank speed creates a unique problem of non-uniform cache accesses in 3D space. Bo Zhao 0007, Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
MICRO | 2 |
| 2008 | Adaptive Buffer Management for Efficient Code Dissemination in Multi-Application Wireless Sensor NetworksabstractFuture wireless sensor networks (WSNs) are projected to run multiple applications in the same network infrastructure. While such multi-application WSNs (MA-WSNs) are economically more efficient and adapt better to the changing environments than traditional single-application WSNs, they usually require frequent code redistribution on wireless sensors, making it critical to design energy efficient post-deployment code dissemination protocols in MA-WSNs. Different applications in MA-WSNs often share some common code segments. Therefore when there is a need to disseminate a new application from the sink node, it is possible to disseminate its shared code segments from peer sensors instead of disseminating everything from the sink node. While dissemination protocols have been proposed to handle code of each single type, it is challenging to achieve energy efficiency when the code contains both types and needs simultaneous dissemination. In this paper we utilize an adaptive buffer management approach to achieve efficient code dissemination in MA-WSNs. Our experimental results show that adaptive buffer management can reduce the completion time and the message overhead up to 10% and 20% respectively. Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
EUC (1) | 2 |
| 2008 | Thermal Management for 3D Processors via Task SchedulingabstractA rising horizon in chip fabrication is the 3D integration technology. It stacks two or more dies vertically with a dense, high-speed interface to increase the device density and reduce the delay of interconnects across the dies. However, a major challenge in 3D technology is the increased power density which brings the concern of heat dissipation within the processor. High temperatures trigger voltage and frequency throttlings in hardware which degrade the chip performance. Moreover, high temperatures impair the processorpsilas reliability and reduce its lifetime. To alleviate this problem, we propose in this paper an OS-level scheduling algorithm that performs thermal-aware task scheduling on a 3D chip. Our algorithm leverages the inherent thermal variations within and across different tasks, and schedules them to keep the chip temperature low. We observed that vertically adjacent dies have strong thermal correlations, and the scheduler should consider them jointly. Our proposed algorithm can remove on average 54% of hardware DTMs and result in 7.2% performance improvement over the base case. Xiuyi Zhou, Yi Xu 0010, Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
ICPP | 3 |