EDBT 2026 Demo / reviewers in the wild / expert
Youtao Zhang
dblp:z/YoutaoZhang
· DBLP profile ↗
178ranked-venue papers
15as first author
44since 2021 · last 2026
0000-0001-8425-8743ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 142 · 8 first-author · 35 since 2021Software engineering, systems software and programming languages · 31 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Security and privacy · 5 · 2 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Pointer Nullification for Use-After-Free Prevention
Yubo Du, Youtao Zhang, Jun Yang 0002 |
NDSS | 2 |
| 2025 | Cascade: A Dependency-aware Efficient Training Framework for Temporal Graph Neural NetworkabstractTemporal graph neural networks (TGNN) have gained significant momentum in many real-world dynamic graph tasks. These models use graph changes (i.e., events) as inputs to update nodes' status vectors (i.e., memories), which are then exploited to assist predictions. Despite their improved accuracies, the efficiency of TGNN training is significantly limited due to the inherent temporal relationship between the input events. Although larger training batches can improve parallelism and speed up TGNN training, they lead to infrequent memory updates, which cause outdated information and reduced accuracy. This trade-off forces current methods to use small batches, resulting in high latency and underutilized hardware. To address this, we propose an efficient TGNN training framework, Cascade, to adaptively boost TGNN training parallelism based on nodes' spatial and temporal dependencies. Cascade adopts a topology-aware scheduler that includes as many spatial-independent events in the same batches. Moreover, it leverages node memories' similarities to break temporal dependencies on stabilized nodes, enabling it to pack more temporal-independent events in the same batches. Additionally, Cascade adaptively decides nodes' update frequencies based on runtime feedback. Compared to prior state-of-the-art TGNN training frameworks, our approach can averagely achieve 2.3x (up to 5.1x) speed up without jeopardizing the resulted models' accuracy. Yue Dai 0005, Xulong Tang, Youtao Zhang |
ASPLOS (2) | 3 |
| 2025 | STMC: Small-Tile Multiple-Copy Compilation for Reliable Measurement-Based Quantum ComputingabstractMeasurement-based Quantum Computing (MBQC) achieves universal quantum computing by applying measurements on the photonic architectures. While it has many advantages, such as long qubit decoherence time and strong scalability, the success rate of MBQC execution is constrained by imperfect photon control, measurement, and fusion operations. Both fusion failure and photon loss necessitate the re-execution of the entire quantum circuit, leading to significant overhead in terms of additional execution time and increased consumption of resource state layers. Recent studies mainly focus on mitigating fusion failures and little attention has been paid to photon loss. In this paper, we propose STMC (Small-Tile Multiple-Copy) compilation framework to reduce the re-execution overhead caused by both the fusion failure and photon loss. Specifically, STMC first transforms a quantum circuit into a fusion graph and partitions the fusion graph into subgraphs. Then, STMC generates compact subgraph mappings that are appropriate for the size of a subportion in the resource state layer, referred to as a tile. Finally, STMC employs multiple copies of each subgraph when mapping to tiles, duplicates the execution of tiles in parallel, and finishes the whole circuit execution in order. The experimental results demonstrate that STMC achieves an average execution time speedup of 65.68× for successfully executing the circuit under a 75% fusion success rate, compared to prior work. Additionally, STMC reduces the number of resource state layers by three orders of magnitude and decreases the number of resource states by an average of 36.40×. Rongchao Dong, Zewei Mo, Yingheng Li, Aditya Pawar, Jun Yang 0002, Youtao Zhang, Xulong Tang |
ICCAD | 6 |
| 2025 | MemFreezing: A Novel Adversarial Attack on Temporal Graph Neural Networks under Limited Future KnowledgeabstractTemporal graph neural networks (TGNN) have achieved significant momentum in many real-world dynamic graph tasks.
While most existing TGNN attack methods assume worst-case scenarios where attackers have complete knowledge of the input graph, the assumption may not always hold in real-world situations, where attackers can, at best, access information about existing nodes and edges but not future ones after the attack.
However, studying adversarial attacks under these constraints is crucial, as limited future knowledge can reveal TGNN vulnerabilities overlooked in idealized settings.
Nevertheless, designing effective attacks in such scenarios is challenging: the evolving graph can weaken their impact and make it hard to affect unseen nodes.
To address these challenges, we introduce MemFreezing, a novel adversarial attack framework that delivers long-lasting and spreading disruptions in TGNNs without requiring post-attack knowledge of the graph.
MemFreezing strategically injects fake nodes or edges to push node memories into a stable “frozen state,” reducing their responsiveness to subsequent graph changes and limiting their ability to convey meaningful information.
As the graph evolves, these affected nodes maintain and propagate their frozen state through their neighbors.
Experimental results show that MemFreezing persistently degrades TGNN performance across various tasks, offering a more enduring adversarial strategy under limited future knowledge. Yue Dai 0005, Xulong Tang, Youtao Zhang, Jun Yang 0002 |
ICML | 4 |
| 2025 | Reinforcement Learning-Guided Graph State Generation in Photonic Quantum ComputersabstractThe photonic quantum computer (PQC) is an emerging and promising quantum computing paradigm that has gained momentum in recent years.In PQC, computations are executed by performing measurements on photons in graph states (i.e., a collection of entangled photons).The graph state generation process is fulfilled by applying a sequence of quantum gates to quantum emitters, referred to as the "generation sequence".In a generation sequence, i) the time required to complete the generation sequence, ii) the number of quantum emitters used, and iii) the number of CZ gates performed between emitters greatly affect the fidelity of the generated graph state.In this paper, we propose RLGS (Reinforcement Learningguided Graph State generation), a novel compilation framework to identify optimal generation sequences that optimize the three fidelity metrics.Experimental results show that RLGS achieves an average reduction in generation time of 31.1%,49.6%, and 57.5% for small, medium, and large graph states compared to the baseline.Additionally, the reductions in the number of quantum emitters are 13.9%, 16.7%, and 17.5%, whereas the reductions in the number of CZ gates are 37.7%, 53.4%, and 57.8%, respectively. Yingheng Li, Yue Dai 0005, Aditya Pawar, Rongchao Dong, Jun Yang 0002, Youtao Zhang, Xulong Tang |
ISCA | 6 |
| 2025 | LightML: A Photonic Accelerator for Efficient General Purpose Machine LearningabstractThe rapid integration of AI technologies into everyday life across sectors such as healthcare, autonomous driving, and smart home applications requires extensive computational resources, placing strain on server infrastructure and incurring significant costs.We present LightML, the first system-level photonic crossbar design, optimized for high-performance machine learning applications.This work provides the first complete memory and buffer architecture carefully designed to support the high-speed photonic crossbar, achieving over 80% utilization.LightML also introduces solutions for key ML functions, including large-scale matrix multiplication (MMM), element-wise operations, non-linear functions, and convolutional layers.Delivering 325 TOP/s at only 3 watts, LightML offers significant improvements in speed and power efficiency, making it ideal for both edge devices and dense data center workloads. Sadra Rahimi Kari, Xin Xin 0008, Nathan Youngblood, Youtao Zhang, Jun Yang 0002 |
ISCA | 5 |
| 2025 | Adaptive Kernel Fusion for Improving the GPU Utilization While Ensuring QoSabstractThe prosperity of machine learning applications has promoted the rapid development of GPU architecture. It continues to integrate more CUDA Cores, larger L2 cache and memory bandwidth within SM. Moreover, the GPU integrates Tensor Core dedicated to matrix multiplication. Although studies have shown that task co-location could effectively improve system throughput, existing works only focus on resource scheduling at the SM level and cannot improve resource utilization within the SM. In this paper, we propose Aker, a static kernel fusion and scheduling approach to improve resource utilization inside the SM while ensuring the QoS (Quality-of-Service) of co-located tasks. Aker consists of a static kernel fuser, a duration predictor for fused kernels, an adaptive fused kernel selector, and an enhanced QoS-aware kernel manager. The kernel fuser enables the static and flexible fusion for a kernel pair. The kernel pair could be Tensor Core kernel and CUDA Core kernel, or computing-prefer CUDA Core kernel and memory-prefer CUDA Core kernel. After the kernel fuser provides multiple fused kernel versions for a kernel pair, the duration predictor precisely predicts the duration of the fused kernels and the adaptive fused kernel selector locates the optimal fused kernel version. Finally, the kernel manager invokes the fused kernel or the original kernel based on the QoS headroom of latency-critical tasks to improve the system throughput. Our experimental results show that Aker improves the throughput of best-effort applications compared with state-of-the-art solutions by 50.1% on average, while ensuring the QoS of latency-critical tasks. Han Zhao 0005, Junxiao Deng, Weihao Cui, Quan Chen 0002, Youtao Zhang, Deze Zeng, Minyi Guo |
IEEE Trans. Computers | 5 |
| 2024 | FMCC: Flexible Measurement-based Quantum Computation over Cluster StateabstractMeasurement-based quantum computing (MBQC) is a promising quantum computing paradigm that performs computation through "one-way" measurements on entangled quantum qubits. It is widely used in photonic quantum computing (PQC), where the computation is carried out on photonic cluster states (i.e., a 2-D mesh of entangled photons). In MBQC-based PQC, the cluster state depth (i.e., the length of one-way measurements) plays an important role in the overall execution time and circuit error. In this paper, we propose FMCC, a compilation framework that employs dynamic programming with heuristics to efficiently minimize the cluster state depth. Experimental results on six quantum applications show that FMCC achieves 51.7%, 57.4%, and 56.8% average depth reductions in small, medium, and large qubit counts compared to the state-of-the-art MBQC compilations. Yingheng Li, Aditya Pawar, Zewei Mo, Youtao Zhang, Jun Yang 0002, Xulong Tang |
ASPLOS (4) | 4 |
| 2024 | QRCC: Evaluating Large Quantum Circuits on Small Quantum Computers through Integrated Qubit Reuse and Circuit CuttingabstractQuantum computing has recently emerged as a promising computing paradigm for many application domains. However, the size of quantum circuits that can be run with high fidelity is constrained by the limited quantity and quality of physical qubits. Recently proposed schemes, such as wire cutting and qubit reuse, mitigate the problem but produce sub-optimal results as they address the problem individually. In addition, gate cutting, an alternative circuit-cutting strategy that is suitable for circuits computing expectation values, has not been fully explored in the field. Aditya Pawar, Yingheng Li, Zewei Mo, Yanan Guo 0002, Xulong Tang, Youtao Zhang, Jun Yang 0002 |
ASPLOS (4) | 6 |
| 2024 | FCM: A Fusion-aware Wire Cutting Approach for Measurement-based Quantum ComputingabstractMeasurement-based quantum computing (MBQC) is a promising quantum computing paradigm that carries out computation through one-way measurements on entangled photon qubits. Practical photonic hardware first generates a 2D mesh of resource states with each being a small number of entangled photon qubits and then exploits fusion operations to connect resource states to scale up the computation. Given that the fusion operation is highly error-prone, it is important to reduce the number of fusions for an MBQC circuit. Zewei Mo, Yingheng Li, Aditya Pawar, Xulong Tang, Jun Yang 0002, Youtao Zhang |
DAC | 6 |
| 2024 | RTT-UAF: Reuse Time Tracking for Use-After-Free DetectionabstractMemory safety continues to be a critical challenge in modern computing, with approximately 70% of reported vulnerabilities annually attributed to memory-related issues. Among these issues, Use-After-Free (UAF) vulnerabilities or bugs, where a program accesses memory through a dangling pointer, pose significant threats. Existing UAF detection methods, such as Key-And-Lock (KAL) mechanisms, incur notable performance overhead due to explicitly propagating the key and lock address (metadata). We identify that approximately 67% of KAL’s performance overhead is introduced by this metadata propagation. This paper introduces RTT, which significantly reduces performance overhead by reducing the number of memory accesses related to metadata propagation. RTT achieves an average performance overhead of 170%, substantially lower than existing KAL methods, and incurs only an 8% memory overhead on SPEC CPU 2017. The results of experimental evaluations on real-world UAF bugs further demonstrate that RTT’s UAF bug detection rate is equivalent to other KAL methods. Yubo Du, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
ICS | 3 |
| 2024 | Space-efficient and high-performance inline deduplication for emerging hybrid storage system with Libra+
Renhui Chen, Tianmeng Zhang, Zijing Li, Congming Gao, Youtao Zhang, Qiao Li 0001, Jun Yang 0002, Jiwu Shu |
J. Syst. Archit. | 5 |
| 2023 | Orchestrating Measurement-Based Quantum Computation over Photonic Quantum ProcessorsabstractQuantum computing has rapidly evolved in recent years and has established its supremacy in many application domains. While matter-based qubit platforms such as superconducting qubits have received the most attention so far, there is a rising interest in photonic qubits lately, which show advantages in parallelism, speed, and scalability. Photonic qubits are best served by the paradigm of measurement-based quantum computation (MBQC). To deliver the promise of measurement-based photonic quantum computing (MBPQC), the photon cluster state depth and photon utilization are two of the most important metrics. However, little attention has been paid to optimizing the depth and utilization when mapping quantum circuits to the photon clusters. In this paper, we propose a compiler framework that achieves automatic and dynamic depth and utilization optimizations. Our approach consists of an MBPQC mapping mechanism that maps optimized measurement patterns on a cluster state and a cluster state pruning strategy that removes all possible redundancies without impacting the circuit functions. Experimental results on five quantum benchmark with three different qubit numbers indicate our approach achieves an average of 63.4% cluster depth reduction and 22.8% photon utilization improvements. Yingheng Li, Aditya Pawar, Mohadeseh Azari, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002, Kaushik Parasuram Seshadreesan, Xulong Tang |
DAC | 5 |
| 2023 | EP-ORAM: Efficient NVM-Friendly Path Eviction for Ring ORAM in Hybrid MemoryabstractRecent studies showed that only ORAM (oblivious RAM) can securely protect memory access patterns (i.e., data privacy) on modern computer systems. Ring ORAM is a promising ORAM protocol as it demands O(1) memory accesses for servicing each user memory request. However, Ring ORAM exhibits low memory utilization, i.e., its memory requirement is 4.8× of the protected user space. While adopting NVM (non-volatile memory) can alleviate the memory requirement, a simple implementation tends to introduce large performance degradation, preventing its adoption in practice.In this paper, we propose EP-ORAM, an NVM-friendly Ring ORAM implementation on DRAM/NVM hybrid memory. EP-ORAM is developed based on two key observations: (1) for tree-based Ring ORAM memory organization, saving bottom levels in NVM can dramatically reduce the DRAM memory requirement; (2) the tradeoffs among Ring ORAM operations expose design opportunities without security compromise. We, therefore, propose to save the bottom levels of the ORAM tree in NVM and shorten the path of EvictPath operation, which not only mitigates the number of NVM writes but also speeds up the execution. Our experimental results show that, under the design constraints of similar performance as the baseline that saves two bottom levels in NVM, EP-ORAM helps to save three levels in NVM, achieving 50% DRAM space reduction. In addition, EP-ORAM reduces the NVM writes by 15%. Mehrnoosh Raoufi, Jun Yang 0002, Xulong Tang, Youtao Zhang |
DAC | 4 |
| 2023 | CEGMA: Coordinated Elastic Graph Matching Acceleration for Graph Matching NetworksabstractThe recently proposed Graph Matching Network models (GMNs) effectively improve the inference accuracy of graph similarity analysis tasks. GMNs often take graph pairs as input, embed nodes features, and match nodes between graphs for similarity analysis. While GMNs deliver high inference accuracy, the all-to-all node matching stage in GMNs introduces quadratic computing complexity with excessive memory accesses, resulting in significant computing and memory overhead that cannot be handled by existing approaches. In this paper, we propose the Coordinated Elastic Graph Matching Accelerator (CEGMA), a software and hardware co-design accelerator to address the challenges of GMNs. Specifically, by exploiting duplicate subgraphs in the input graphs, we develop an elastic matching filter to significantly reduce the quadratic computing overhead. By exploring the substantial data reuses oriented from accessing node features, we propose a cross-graph coordinator that fuses cross-graph similarity computing with intra-graph computing to enhance data locality. Experimental results show that, on average, CEGMA achieves 353× and 6.5× speedups in GMN computing compared to state-of-the-art GPU implementation and GNN accelerators, respectively. Yue Dai 0005, Youtao Zhang, Xulong Tang |
HPCA | 2 |
| 2023 | Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote ForwardingabstractMulti-GPU systems have become a popular platform to meet the ever-growing application demands. However, employing multiple GPUs does not guarantee proportional performance improvements. While prior works have extensively studied the optimizations to mitigate the non-uniform memory accesses (NUMA) overheads, the address translation process also plays an important role in shaping the overall execution performance. In this paper, we investigate the address translation process in multi-GPU systems under unified virtual memory (UVM). We specifically focus on the efficiency of page table walk and identify three major latency penalties: i) queuing for available page table walk threads, ii) memory accesses for page walk cache misses, and iii) handling page faults. Based on our observations, we propose Trans-FW, which short circuits the page table walk by leveraging substantial translation sharing and eager remote translation forwarding. Experimental results on 10 representative multi-GPU applications show that our proposed approach improves the overall performance by 53.8% on average. Bingyao Li 0001, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang 0002, Xulong Tang |
HPCA | 4 |
| 2023 | MGC: Multiple-Gray-Code for 3D NAND Flash based High-Density SSDsabstractQLC (4-bit-per-cell) and more-bit-per-cell 3D NAND flash memories are increasingly adopted in large storage systems. While achieving significant cost reduction, these memories face degraded performance and reliability issues. The industry has adopted two-step programming (TSP), rather than one-step programming, to perform fine-granularity program control and choose gray-code encoding, as well as LDPC (Low-Density Parity-Check Code) for error correction. Different flash manufacturers often integrate different gray-codes in their products, which exhibit different performance and reliability characteristics. Unfortunately, a fixed gray-code encoding design lacks the ability to meet the dynamic read and program performance requirements at both application and device levels.In this paper, we propose MGC, a multiple-gray-code encoding strategy, that adaptively chooses the best gray-code to meet the optimization goals at runtime. In particular, MGC first extracts the performance and reliability requirements based on application-level access patterns and detects the reliability degree of SSD. It then determines the appropriate gray-code to encode the data, either from host/user application or due to garbage collection, before writing the pages to the flash memory. MGC is integrated in FTL (flash translation layer) and enhances the flash controller to enable runtime gray-code arbitration. We evaluate the proposed MGC scheme. The results show that MGC achieves better performance and lifetime guarantee compared with state-of-the-arts and introduces little overhead. Yina Lv, Liang Shi 0001, Qiao Li 0001, Congming Gao, Yunpeng Song, Longfei Luo, Youtao Zhang |
HPCA | 7 |
| 2023 | AB-ORAM: Constructing Adjustable Buckets for Space Reduction in Ring ORAMabstractRing ORAM (Oblivious RAM) is a secure primitive that mitigates the large performance degradation of ORAM through reduced online memory bandwidth demand, i.e., the number of memory accesses at servicing a real memory request. Ring ORAM requires 4× or more of the protected data space to enable the optimization and thus presents high capacity pressure on modern memory systems. While recent studies strive to reduce its space consumption through bucket compaction, the large space consumption remains a major design challenge for Ring ORAM.In this paper, we propose AB-ORAM to reduce the space capacity demand in Ring ORAM. AB-ORAM identifies two inefficient use of memory space in Ring ORAM: (i) accessed blocks hold useless data until the next reshuffle operation; and (ii) large buckets provide a diminishing performance benefit for tree levels close to the leaves. AB-ORAM then proposes two schemes to exploit the optimization opportunities, respectively. Specifically, it reclaims accessed blocks early by allocating them to buckets that need a reshuffle; and shrinks the bucket size for tree level close to the leaves for a better space/performance trade-off. We evaluate the proposed AB-ORAM design and compare it to the state-of-the-art. Our results show that AB-ORAM achieves an average of 36% space reduction over the state-of-the-art while introducing very low performance overhead. Mehrnoosh Raoufi, Jun Yang 0002, Xulong Tang, Youtao Zhang |
HPCA | 4 |
| 2023 | FlexGM: An Adaptive Runtime System to Accelerate Graph Matching Networks on GPUsabstractGMNs (Graph Matching Networks) exploit recently developed GNNs (Graph Neural Networks) to analyze the similarity between two graphs. They are increasingly deployed in many application domains due to their improved inference accuracy. A GMN consists of two stages, i.e., node-embedding and node-matching stages. The node-matching stage matches node features from two graphs for similarity, which accounts for over 90% of the total execution time. However, it is challenging to accelerate GMNs on GPUs due to their diverse computing patterns for different graph inputs. For large graphs, the overhead comes mainly from the high computation overhead, which increases quadratically to the size of the graphs; for small graphs, the overhead comes from the low parallelism and resource utilization.In this paper, we propose FlexGM, a flexible runtime, to adaptively accelerate GMNs on GPUs. For large graphs, we exploit the massive computation redundancy in GMNs and develop a low-overhead deduplication module to mitigate the high computation overhead. For small graphs, we develop a unified matching module to optimize GPU hardware resource usage. An adaptive module manager is then developed to judiciously select beneficial optimization strategies. Experimental results show that the FlexGM system achieves 2.5× (up to 7.6 ×) average speedup over existing methods. Yue Dai 0005, Xulong Tang, Youtao Zhang |
ICCD | 3 |
| 2023 | SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing
Sheng Li 0019, Geng Yuan, Yue Dai 0005, Youtao Zhang, Yanzhi Wang 0001, Xulong Tang |
ICLR | 4 |
| 2023 | Understanding and Defending Patched-based Adversarial Attacks for Vision TransformerabstractVision Transformer (ViT) is an attention-based model architecture that has demonstrated superior performance on many computer vision tasks. However, its security properties, in particular, the robustness against adversarial attacks, are yet to be thoroughly studied. Recent works have shown that ViT is vulnerable to attention-based adversarial patch attacks, which cover 1-3% area of the input image using adversarial patches and degrades the model accuracy to 0%. This work provides a generic study targeting the attention-based patch attack. First, we experimentally observe that adversarial patches only activate in a few layers and become lazy during attention updating. According to experiments, we study the theory of how a small adversarial patch perturbates the whole model. Based on understanding adversarial patch attacks, we propose a simple but efficient defense that correctly detects more than 95% of adversarial patches. Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
ICML | 3 |
| 2023 | Uncore Encore: Covert Channels Exploiting Uncore Frequency ScalingabstractModern processors dynamically adjust clock frequencies and voltages to reduce energy consumption. Recent Intel processors separate the uncore frequency from the core frequency, using Uncore Frequency Scaling (UFS) to adapt the uncore frequency to various workloads. While UFS improves power efficiency, it also introduces security vulnerabilities. In this paper, we study the feasibility of covert channels exploiting UFS. First, we conduct a series of experiments to understand the details of UFS, such as the factors that can cause uncore frequency variations. Then, based on the results, we build the first UFS-based covert channel, UF-variation, which works both across-cores and across-processors. Finally, we analyze the robustness of UF-variation under known defense mechanisms against uncore covert channels, and show that UF-variation remains functional even with those defenses in place. Yanan Guo 0002, Dingyuan Cao 0002, Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
MICRO | 4 |
| 2023 | Generating Robust DNN With Resistance to Bit-Flip Based Adversarial Weight AttackabstractRowhammer Attack, a new DRAM-based attack, was developed exploiting weak cells to alter their content. Such attacks can be launched at the user level without requiring access permission to the victim memory cells. Leveraging such attacks, a new bit-flip-based adversarial weights attack (BFA) was developed targeting deep neural network models. When BFA attackers acquire a DNN model, they manipulate the existing DNN adversarial attack into locating vulnerable bits in the target DNN model. By flipping a subset of them using Rowhammer, they can crash that model within 30 trails. In this paper, we propose a lightweight and easy-to-deploy defense mechanism in the bit-level, Randomized Rotated and Nonlinear Encoding (RREC), which generates both robustness and fault-tolerant against BFA. Since flipping the most significant bit (MSB) in quantized data is too dangerous, we introduce randomized Rotation to obfuscate the bit order of model data and efficiently hide truly vulnerable bits with less vulnerable ones. Further, RREC reduces the average bit-flipped distance by more than 3x from the nonlinear encoding. It decreases the bit-flip distance among the majority of bits (including those vulnerable bits). Theoretically, RREC minimized the impact of a single bit BFA to 1/24 compared with baseline. Experimentally, RREC tolerates more than 17x flipped bits versus baseline model and 4.8x and 5.7x more bits compared with the existing BFA defenses (4B QAT and WR) with 0.01x to 0.08x of runtime latency. Moreover, we evaluate RREC against a newly emerged attack, Targeted-BFA, and it improves the defense rate from$5\%$to$95\%$. Yanan Guo 0002, Yueqiang Cheng, Youtao Zhang, Jun Yang 0002 |
IEEE Trans. Computers | 4 |
| 2022 | SRA: a secure ReRAM-based DNN acceleratorabstractDeep Neural Network (DNN) accelerators are increasingly developed to pursue high efficiency in DNN computing. However, the IP protection of the DNNs deployed on such accelerators is an important topic that has been less addressed. Although there are previous works that targeted this problem for CMOS-based designs, there is still no solution for ReRAM-based accelerators which pose new security challenges due to their crossbar structure and non-volatility. ReRAM's non-volatility retains data even after the system is powered off, making the stored DNN model vulnerable to attacks by simply reading out the ReRAM content. Because the crossbar structure can only compute on plaintext data, encrypting the ReRAM content is no longer a feasible solution in this scenario. Youtao Zhang, Jun Yang 0002 |
DAC | 2 |
| 2022 | IR-ORAM: Path Access Type Based Memory Intensity Reduction for Path-ORAMabstractPath ORAM is an effective ORAM (Oblivious RAM) primitive for protecting memory access patterns. Path ORAM converts each off-chip memory request from user program to tens to hundreds of memory accesses. While several schemes have been proposed to mitigate the total number of memory accesses, Path ORAM remains a highly memory intensive primitive that leads to large memory bandwidth occupation and performance degradation.In this paper, we propose IR-ORAM to reduce the memory intensity based on path access types in Path ORAM. Path accesses in Path ORAM, while being kept oblivious to ensure privacy protection, can be categorized to three types: paths for requested data blocks, paths for position map blocks, and dummy paths. We develop a set of techniques to reduce the memory intensity of each type while ensuring the obliviousness at the same time — we reduce the number of data blocks to access for each tree path, reduce the number of path accesses for position maps, and convert many dummy path accesses to early write-backs of dirty data in LLC. Our experimental results show that IR-ORAM achieves on average 42% performance improvement over the state-of-the-art while effectively enforcing the memory access obliviousness and the same level of security protection. Mehrnoosh Raoufi, Youtao Zhang, Jun Yang 0002 |
HPCA | 2 |
| 2022 | Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSabstractThe proliferation of machine learning applications has promoted both CUDA Cores and Tensor Cores’ integration to meet their acceleration demands. While studies have shown that co-locating multiple tasks on the same GPU can effectively improve system throughput and resource utilization, existing schemes focus on scheduling the resources of traditional CUDA Cores and thus lack the ability to exploit the parallelism between Tensor Cores and CUDA Cores.In this paper, we propose Tacker, a static kernel fusion and scheduling approach to improve GPU utilization of both types of cores while ensuring the QoS (Quality-of-Service) of co-located tasks. Tacker consists of a Tensor-CUDA Core kernel fuser, a duration predictor for fused kernels, and a runtime QoS-aware kernel manager. The kernel fuser enables the flexible fusion of kernels that use Tensor Cores and CUDA Cores, respectively. The duration predictor precisely predicts the duration of the fused kernels. Finally, the kernel manager invokes the fused kernel or the original kernel based on the QoS headroom of latency-critical tasks to improve the system throughput. Our experimental results show that Tacker improves the throughput of best-effort applications compared with state-of-the-art solutions by 18.6% on average, while ensuring the QoS of latency-critical tasks. Han Zhao 0005, Weihao Cui, Quan Chen 0002, Youtao Zhang, Yanchao Lu, Chao Li 0009, Jingwen Leng, Minyi Guo |
HPCA | 4 |
| 2022 | Q-GPU: A Recipe of Optimizations for Quantum Circuit Simulation Using GPUsabstractIn recent years, quantum computing has undergone significant developments and has established its supremacy in many application domains. While quantum hardware is accessible to the public through the cloud environment, a robust and efficient quantum circuit simulator is necessary to investigate the constraints and foster quantum computing development, such as quantum algorithm development and quantum device architecture exploration. In this paper, we observe that most of the publicly available quantum circuit simulators (e.g., QISKit from IBM, QDK from Microsoft, and Qsim-Cirq from Google) suffer from slow simulation and poor scalability when the number of qubits increases. To this end, we systematically investigate the deficiencies in quantum circuit simulation (QCS) and propose Q-GPU, a framework that leverages GPUs with comprehensive optimizations to allow efficient and scalable QCS. Specifically, Q-GPU features i) proactive state amplitude transfer, ii) zero state amplitude pruning, iii) delayed qubit involvement, and iv) lossless nonzero state amplitude compression. Experimental results across nine representative quantum circuits indicate that Q-GPU significantly reduces the execution time of the state-of-the-art GPU-based QCS by 71.89% (3.55× speedup). Q-GPU also outperforms the state-of-the-art OpenMP CPU implementation, the Google Qsim-Cirq simulator, and the Microsoft QDK simulator by 1.49×, 2.02×, and 10.82×, respectively. Yilun Zhao 0002, Yanan Guo 0002, Amanda Dumi, Devin M. Mulvey, Shiv Upadhyay, Youtao Zhang, Kenneth D. Jordan, Jun Yang 0002, Xulong Tang |
HPCA | 7 |
| 2022 | Leaky Way: A Conflict-Based Cache Covert Channel Bypassing Set AssociativityabstractModern $\times$86 processors feature many prefetch instructions that developers can use to enhance performance. However, with some prefetch instructions, users can more directly manipulate cache states which may result in powerful cache covert channel and side channel attacks. In this work, we reverse-engineer the detailed cache behavior of PREFETCHNTA on various Intel processors. Based on the results, we first propose a new conflict-based cache covert channel named NTP+NTP. Prior conflict-based channels often require priming the cache set in order to cause cache conflicts. In contrast, in NTP+NTP, the data of the sender and receiver can compete for one specific way in the cache set, achieving cache conflicts without cache set priming for the first time. As a result, NTP+NTP has higher bandwidth than prior conflict-based channels such as Prime+Probe. The channel capacity of NTP+NTP is 302 KB/s. Second, we found that PREFETCHNTA can also be used to boost the performance of existing side channel attacks that utilize cache replacement states, making those attacks much more efficient than before. Yanan Guo 0002, Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
MICRO | 3 |
| 2022 | Adversarial Prefetch: New Cross-Core Cache Side Channel AttacksabstractModern x86 processors have many prefetch instructions that can be used by programmers to boost performance. However, these instructions may also cause security problems. In particular, we found that on Intel processors, there are two security flaws in the implementation of PREFETCHW, an instruction for accelerating future writes. First, this instruction can execute on data with read-only permission. Second, the execution time of this instruction leaks the current coherence state of the target data. Based on these two design issues, we build two cross-core private cache attacks that work with both inclusive and non-inclusive LLCs, named Prefetch+Reload and Prefetch+Prefetch. We demonstrate the significance of our attacks in different scenarios. First, in the covert channel case, Prefetch+Reload and Prefetch+Prefetch achieve 782 KB/s and 822 KB/s channel capacities, when using only one shared cache line between the sender and receiver, the largest-to-date single-line capacities for CPU cache covert channels. Further, in the side channel case, our attacks can monitor the access pattern of the victim on the same processor, with almost zero error rate. We show that they can be used to leak private information of real-world applications such as cryptographic keys. Finally, our attacks can be used in transient execution attacks in order to leak more secrets within the transient window than prior work. From the experimental results, our attacks allow leaking about 2 times as many secret bytes, compared to Flush+Reload, which is widely used in transient execution attacks. Yanan Guo 0002, Andrew Zigerelli, Youtao Zhang, Jun Yang 0002 |
SP | 3 |
| 2022 | An efficient segmented quantization for graph neural networks
Yue Dai 0005, Xulong Tang, Youtao Zhang |
CCF Trans. High Perform. Comput. | 3 |
| 2022 | Reprogramming 3D TLC Flash Memory based Solid State DrivesabstractNAND flash memory-based SSDs have been widely adopted. The scaling of SSD has evolved from plannar (2D) to 3D stacking. For reliability and other reasons, the technology node in 3D NAND SSD is larger than in 2D, but data density can be increased via increasing bit-per-cell. In this work, we develop a novel reprogramming scheme for TLCs in 3D NAND SSD, such that a cell can be programmed and reprogrammed several times before it is erased. Such reprogramming can improve the endurance of a cell and the speed of programming, and increase the amount of bits written in a cell per program/erase cycle, i.e., effective capacity. Our work is the first to perform a real 3D NAND SSD test to validate the feasibility of the reprogram operation. From the collected data, we derive the restrictions of performing reprogramming due to reliability challenges. Furthermore, a reprogrammable SSD (ReSSD) is designed to structure reprogram operations. ReSSD is evaluated in a case study in RAID 5 system (RSS-RAID). Experimental results show that RSS-RAID can improve the endurance by 35.7%, boost write performance by 15.9%, and increase effective capacity by 7.71%, with negligible overhead compared with conventional 3D SSD-based RAID 5 system. Congming Gao, Chun Jason Xue, Youtao Zhang, Liang Shi 0001, Jiwu Shu, Jun Yang 0002 |
ACM Trans. Storage | 4 |
| 2022 | Leveraging Multimodal Semantic Fusion for Gastric Cancer Screening via Hierarchical Attention MechanismabstractGastroscopy is a widely adopted method for locating gastric lesions and performing the early screening and diagnosis of gastric cancer (GC). However, the effectiveness of traditional GC screening methods depends on the medical skills of the gastroscopy specialist. A lack of knowledge and experience may lead to misdiagnosis and mistreatment, especially in small-scale hospitals. Recently, there has been a significant increase in studies on data-driven computer-aided diagnosis techniques. In this article, we propose a novel intelligent decision-making method for GC screening (ID-GCS), a multimodal semantic fusion-based data-driven decision-making system. ID-GCS exploits a hybrid attention mechanism to extract textual semantics from multimodal gastroscopy reports and performs semantic fusion to integrate the semantics of textual gastroscopy reports and images, resulting in improved interpretability of gastroscopy findings. We evaluated ID-GCS using a real gastroscopy report dataset, and experimental results show that compared with state-of-the-art methods, ID-GCS achieves better sensitivity and accuracy in GC screening. Shuai Ding 0001, Shikang Hu, Xiaojian Li 0003, Youtao Zhang, Desheng Dash Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | IVcache: Defending Cache Side Channel Attacks via Invisible AccessesabstractThe sharing of last-level cache (LLC) among different CPU cores makes cache vulnerable to side channel attacks. An attacker can get private information about co-running applications (victims) by monitoring their accesses in LLC. Cache side channel attacks can be mitigated by partitioning cache between the victim and attacker. However, previous partition works either make weak assumptions about the attacker's strength or force their security mechanisms and thus overhead to every user on the system, regardless of their security requirement. Yanan Guo 0002, Andrew Zigerelli, Youtao Zhang, Jun Yang 0002 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2021 | ScaleDNN: Data Movement Aware DNN Training on Multi-GPUabstractTraining Deep Neural Networks (DNNs) models is a time-consuming process that requires immense amount of data and computation. To this end, GPUs are widely adopted to accelerate the training process. However, the delivered training performance rarely scales with the increase in the number of GPUs. The major reason behind this is the large amount of data movement that prevents the system from providing the GPUs with the required data in a timely fashion. In this paper, we propose ScaleDNN, a framework that systematically and comprehensively investigates and optimizes data-parallel training on two types of multi-GPU systems (PCIe-based and NVLink-based). Specifically, ScaleDNN performs: i) CPU-centric input batch splitting, ii) mini-batch data pre-loading, and iii) model parameter compression to effectively a) reduce the data movement between the CPU and multiple GPUs, and b) hide the data movement overheads by overlapping the data transfer with the GPU computation. Our experimental results show that ScaleDNN achieves up to 39.38%, with an average of 17.96% execution time saving over modern data parallelism on PCIe-based multi-GPU system. The corresponding execution time reduction on NVLink-based multi-GPU system is up to 19.20% with an average of 10.26%. Weizheng Xu, Ashutosh Pattnaik, Geng Yuan, Yanzhi Wang 0001, Youtao Zhang, Xulong Tang |
ICCAD | 5 |
| 2021 | ModelShield: A Generic and Portable Framework Extension for Defending Bit-Flip based Adversarial Weight AttacksabstractBit-flip attack (BFA) has become one of the most serious threats to Deep Neural Network (DNN) security. By utilizing Rowhammer to flip the bits of DNN weights stored in memory, the attacker can turn a functional DNN into a random output generator. In this work, we propose ModelShield, a defense mechanism against BFA, based on protecting the integrity of weights using hash verification. ModelShield performs real-time integrity verification on DNN weights. Since this can slow down a DNN inference by up to 7×, we further propose two optimizations for ModelShield. We implement ModelShield as a lightweight software extension that can be easily installed into popular DNN frameworks. We test both the security and performance of ModelShield, and the results show that it can effectively defend BFA with less than 2% performance overhead. Yanan Guo 0002, Yueqiang Cheng, Youtao Zhang, Jun Yang 0002 |
ICCD | 4 |
| 2021 | Flipping Bits to Share Crossbars in ReRAM-Based DNN AcceleratorabstractFuture deep neural networks (DNNs) tend to grow deeper and contain more trainable weights. Although methods such as pruning and quantization are widely adopted to reduce DNN’s model size and computation, they are less applicable in the area of ReRAM-based DNN accelerators. On the one hand, because the cells in crossbars are accessed uniformly, it is difficult to explore fine-grained pruning in ReRAM-based DNN accelerators. On the other hand, aggressive quantization results in poor accuracy coupled with the low precision of ReRAM cells to represent weight values.In this paper, we propose BFlip – a novel model size and computation reduction technique – to share crossbars among multiple bit matrices. BFlip clusters similar bit matrices together, and finds a combination of row and column flips for each bit matrix to minimize its distance to the centroid of the cluster. Therefore, only the centroid bit matrix is stored in the crossbar, which is shared by all other bit matrices in that cluster. We also propose a calibration method to improve the accuracy as well as a ReRAM-based DNN accelerator to fully reap the storage and computation benefits of BFlip. Our experiments show that BFlip effectively reduces model size and computation with negligible accuracy impact. The proposed accelerator achieves 2.45 × speedup and 85% energy reduction over the ISAAC baseline. Youtao Zhang, Jun Yang 0002 |
ICCD | 2 |
| 2021 | ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory based SSDsabstractProcessing-in-memory (PIM) and in-storage-computing (ISC) architectures have been constructed to implement computation inside memory and near storage, respectively. While effectively mitigating the overhead of data movement from memory and storage to the processor, due to the limited bandwidth of existing systems, these architectures still suffer from the large data movement overhead between storage and memory, in particular, if the amount of required data is large. It has become a major constraint for further improving the computation efficiency in PIM and ISC architectures. Congming Gao, Xin Xin 0008, Youyou Lu, Youtao Zhang, Jun Yang 0002, Jiwu Shu |
MICRO | 4 |
| 2021 | AutoBraid: A Framework for Enabling Efficient Surface Code Communication in Quantum ComputingabstractQuantum computers can solve problems that are intractable using the most powerful classical computer. However, qubits are fickle and error prone. It is necessary to actively correct errors in the execution of a quantum circuit. Quantum error correction (QEC) codes are developed to enable fault-tolerant quantum computing. With QEC, one logical circuit is converted into an encoded circuit. Yan-Hao Chen, Yuwei Jin, Chi Zhang 0041, Ari B. Hayes, Youtao Zhang, Eddy Z. Zhang |
MICRO | 6 |
| 2021 | Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignabstractIn recent years, the ever-growing application complexity and input dataset sizes have driven the popularity of multi-GPU systems as a desirable computing platform for many application domains. While employing multiple GPUs intuitively exposes substantial parallelism for the application acceleration, the delivered performance rarely scales with the number of GPUs. One of the major challenges behind is the address translation efficiency. Many prior works focus on CPUs or single GPU execution scenarios while the address translation in multi-GPU systems receives little attention. In this paper, we conduct a comprehensive investigation of the address translation efficiency in both “single-application-multi-GPU” and “multi-application-multi-GPU” execution paradigms. Based on our observations, we propose a new TLB hierarchy design, called least-TLB, tailored for multi-GPU systems and effectively improves the TLB performance with minimal hardware overheads. Experimental results on 9 single-application workloads and 10 multi-application workloads indicate the proposed least-TLB improves the performances, on average, by 23.5% and 16.3%, respectively. Bingyao Li 0001, Jieming Yin, Youtao Zhang, Xulong Tang |
MICRO | 3 |
| 2021 | SAM: Accelerating Strided Memory AccessesabstractStrided memory accesses are an important type of operations for In-Memory Databases (IMDB) applications. Strided memory accesses often demand data at word granularity with fixed strides. Hence, they tend to produce sub-optimal performance on DRAM memory (the de facto standard memory in modern computer systems) that accesses data at cacheline granularity. Recently proposed optimizations either introduce significant reliability degradation or are limited to non-volatile crossbar memory structures. Xin Xin 0008, Yanan Guo 0002, Youtao Zhang, Jun Yang 0002 |
MICRO | 3 |
| 2021 | CacheTree: Reducing Integrity Verification Overhead of Secure Nonvolatile MemoriesabstractEmerging nonvolatile memories (NVMs), while exhibiting great potential to be DRAM alternatives, are vulnerable to security attacks. Secure NVM designs demand data persistence on top of traditional confidentiality and integrity protection. A simple adaption of existing secure memory designs would incur non-negligible overheads, including performance degradation, NVM lifetime reduction, and energy consumption increase. In this article, we propose CacheTree to address the integrity verification overhead for secure NVMs. By constructing extra Merkle trees (MTs) on top of metadata cache, CacheTree helps to authenticate the volatile cache contents, which enables the adoption of write-back policy and prevents frequent NVM writes in persisting metadata. We then adopt CacheTree to address the integrity verification in secure NVM, in particular, the overheads in persisting message authentication codes (for protecting the integrity of user data at memory line level) and persisting the main MT (for protecting the integrity of the whole memory space). Our experimental results show that CacheTree, with less than 0.5% storage overhead, achieves up to 20.1% performance improvement, 44.3% lifetime increase, and 43.7% energy consumption reduction over the state-of-the-art solutions. Zhengguo Chen, Youtao Zhang, Nong Xiao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Automatic Acetowhite Lesion Segmentation via Specular Reflection Removal and Deep Attention NetworkabstractAutomatic acetowhite lesion segmentation in colposcopy images (cervigrams) is essential in assisting gynecologists for the diagnosis of cervical intraepithelial neoplasia grades and cervical cancer. It can also help gynecologists determine the correct lesion areas for further pathological examination. Existing computer-aided diagnosis algorithms show poor segmentation performance because of specular reflections, insufficient training data and the inability to focus on semantically meaningful lesion parts. In this paper, a novel computer-aided diagnosis algorithm is proposed to segment acetowhite lesions in cervigrams automatically. To reduce the interference of specularities on segmentation performance, a specular reflection removal mechanism is presented to detect and inpaint these areas with precision. Moreover, we design a cervigram image classification network to classify pathology results and generate lesion attention maps, which are subsequently leveraged to guide a more accurate lesion segmentation task by the proposed lesion-aware convolutional neural network. We conducted comprehensive experiments to evaluate the proposed approaches on 3045 clinical cervigrams. Our results show that our method outperforms state-of-the-art approaches and achieves better Dice similarity coefficient and Hausdorff Distance values in acetowhite legion segmentation. Zijie Yue, Shuai Ding 0001, Xiaojian Li 0003, Shanlin Yang, Youtao Zhang |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Hierarchical Physician Recommendation via Diversity-enhanced Matrix FactorizationabstractRecent studies have shown that there exhibits significantly imbalanced medical resource allocation across public hospitals. Patients, regardless of their diseases, tend to choose hospitals and physicians with a better reputation, which often overloads major hospitals while leaving others underutilized. Guiding patients to hospitals that can serve their treatment needs both timely and with good quality can make the best use of precious medical resources. Unfortunately, it remains one of the major challenges both for research and in practice. In this article, we propose a novel diversity-enhanced hierarchical physician recommendation approach to address this issue. We adopt matrix factorization to estimate physician competency and exploit implicit similarity relationships to improve the competency estimation of physicians that we are of little information of. We then balance the patient preference and physician diversity using two novel heuristic algorithms. We evaluate our proposed approach and compare it with the state of the art. Experiments show that our approach significantly improves both accuracy and recommendation diversity over existing approaches. Hao Wang 0081, Shuai Ding 0001, Yeqing Li, Xiaojian Li 0003, Youtao Zhang |
ACM Trans. Knowl. Discov. Data | 5 |
| 2021 | Privacy-preserving Time-series Medical Images Analysis Using a Hybrid Deep Learning FrameworkabstractTime-series medical images are an important type of medical data that contain rich temporal and spatial information. As a state-of-the-art, computer-aided diagnosis (CAD) algorithms are usually used on these image sequences to improve analysis accuracy. However, such CAD algorithms are often required to upload medical images to honest-but-curious servers, which introduces severe privacy concerns. To preserve privacy, the existing CAD algorithms support analysis on each encrypted image but not on the whole encrypted image sequences, which leads to the loss of important temporal information among frames. To meet this challenge, a convolutional-LSTM network, named HE-CLSTM, is proposed for analyzing time-series medical images encrypted by a fully homomorphic encryption mechanism. Specifically, several convolutional blocks are constructed to extract discriminative spatial features, and LSTM-based sequence analysis layers (HE-LSTM) are leveraged to encode temporal information from the encrypted image sequences. Moreover, a weighted unit and a sequence voting layer are designed to incorporate both spatial and temporal features with different weights to improve performance while reducing the missed diagnosis rate. The experimental results on two challenging benchmarks (a Cervigram dataset and the BreaKHis public dataset) provide strong evidence that our framework can encode visual representations and sequential dynamics from encrypted medical image sequences; our method achieved AUCs above 0.94 both on the Cervigram and BreaKHis datasets, constituting a significant margin of statistical improvement compared with several competing methods. Zijie Yue, Shuai Ding 0001, Youtao Zhang, Zehong Cao, Muhammad Tanveer 0001, Alireza Jolfaei, James Xi Zheng |
ACM Trans. Internet Techn. | 4 |
| 2020 | Layer RBER Variation Aware Read Performance Optimization for 3D Flash Memoriesabstract3D NAND flash enables the construction of large capacity Solid-State Drives (SSDs) for modern computer systems. While effectively reducing per bit cost, 3D NAND flash exhibits non-negligible process variations and thus RBER (raw bit error rate) difference across layers, which leads to sub-optimal read performance for applications with either small or large I/O requests. In this paper, we propose LRR, Layer RBER variation aware Read optimization schemes, to address the challenge. LRR consists of two schemes - LRR subpage read scheduling (SRS) and LRR fullpage allocation (FPA). SRS groups small read requests from the layers with similar RBERs to reduce the average read latency of subpage sized read requests. FPA distributes the data of a large write to multiple layers, which improves the read latency when reading from layers with large RBERs. Our experimental results show that our proposed scheme LRR reduces 46% read latency on average over the state-of-the-art. Shiqiang Nie, Youtao Zhang, Weiguo Wu, Jun Yang 0002 |
DAC | 2 |
| 2020 | Reducing DRAM Access Latency via Helper RowsabstractThe DRAM technology advancement has seen success in memory density and throughput improvement, but less in access latency reduction. This is mainly due to the intrinsic limitation of capacitance based bit store and access mechanism. The reduction of access latency has been well explored in literature. However, the recently proposed DRAM techniques, such as RowClone and Half-DRAM, offer new opportunities to further optimise the access latency.In this paper, we propose an efficient access strategy to improve the performance of DRAM by optionally discarding the restore. When activating a new row, our technique makes a copy of the row leveraging the RowClone method. Next time when accessing the same row, the cloned row is opened for sensing but it is not restored as the data is preserved in the original row. To improve the efficiency of our proposed strategy, we further exploit three schemes to minimize the copy overhead and increase the reuse of the cloned row. Experimental results show that our proposed strategy can achieve 11% performance improvement on average. Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
DAC | 2 |
| 2020 | SCA: A Secure CNN Accelerator for Both Training and InferenceabstractConvolutional neural networks (CNNs), while being widely deployed to edge devices, face increasingly requirements for IP protection, i.e., the protection of the models and their weights. This becomes particularly challenging for those that demand post-deployment training to enhance inference performance. Existing schemes focus mainly on IP protection at the inference phase, and lack the ability to extend to the training phase. In this paper, we propose SCA, a secure CNN accelerator that exploits stochastic computing to achieve IP protection at both training and inference phases. We propose hybrid stochastic addition and weight remapping to further optimize space utilization and design robustness. Our experimental results show that SCA effectively prevents pirating the CNN IP from the authorized devices. In addition, it achieves 4.8× and 34.2× speedups and 84.3% and 98.5% energy reductions over a non-secure baseline and an inference-only secure baseline, respectively. Youtao Zhang, Jun Yang 0002 |
DAC | 2 |
| 2020 | ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMabstractRecently proposed DRAM based memory-centric architectures have demonstrated their great potentials in addressing the memory wall challenge of modern computing systems. Such architectures exploit charge sharing of multiple rows to enable in-memory bitwise operations. However, existing designs rely heavily on reserved rows to implement computation, which introduces high data movement overhead, large operation latency, large energy consumption, and low operation reliability. In this paper, we propose ELP2IM, an efficient and low power processing in-memory architecture, to address the above issues. ELP2IM utilizes two stable states of sense amplifiers in DRAM subarrays so that it can effectively reduce the number of intra-subarray data movements as well as the number of concurrently opened DRAM rows, which exhibits great performance and energy consumption advantages over existing designs. Our experimental results show that the power efficiency of ELP2IM is more than 2X improvement over the state-of-the-art DRAM based memory-centric designs in real application. Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
HPCA | 2 |
| 2020 | Accelerating 3D Vertical Resistive Memories with Opportunistic Write Latency ReductionabstractThe 3D Vertical Resistive Memory (3D-VRAM) is emerging as a promising non-volatile memory (NVM) technology to construct large capacity main memories in future computing systems. In addition to low energy-efficiency and non-volatility, ReRAM (Resistive Memory) achieves excellent density and scalability by vertically stacking multiple layers of cross-point arrays. However, 3D-VRAM arrays have enormous amount of sneaky paths, which degrade write performance and reliability dramatically. Wen Wen 0003, Youtao Zhang, Jun Yang 0002 |
ICCAD | 2 |
| 2020 | Leveraging partial-refresh for performance and lifetime improvement of 3D NAND flash memory in cyber-physical systems
Jinhua Cui 0001, Youtao Zhang, Liang Shi 0001, Chun Jason Xue, Jun Yang 0002, Laurence T. Yang |
J. Syst. Archit. | 2 |
| 2020 | FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File DesignabstractModern graphics processing units (GPUs) exhibit increasing demands for register files (RFs) with larger capacity and bank sizes, which jeopardize the traditional SRAM-based RF designs due to their large die area and long access latency. Recent hybrid RF designs, e.g., SRAM and spin-transfer torque random access memory (STT-RAM)-based RFs, mitigate the issue by exploiting the density and performance advantages in STT-RAM and SRAM, respectively. However, existing hybrid RF designs adopt coarse integration that has limited write bandwidth between SRAM and STT-RAM, which restricts the adoption of different warp schedulers at runtime. In this article, we propose FRF, a warp-scheduler friendly fine-grained hybrid RF design using SRAM/STT-RAM hybrid cell (HC) structures. By integrating one SRAM cell and N STT-RAM cells as one HC, FRF exploits internal write paths to enlarge the access bandwidth between SRAM and STT-RAM and thus greatly optimizes the area and performance. FRF enables the concurrent context-switching such that different warp schedulers may be adopted at runtime. FRF adopts interleaved register mapping (IRM) and on-demand register remapping to further improve the utilization of SRAM in each HC. Our experimental results show that, on average, FRF achieves 50% performance improvement and 40% energy consumption reduction over the coarse-grained hybrid design when adopting loose round-robin (LRR), and achieves 159% efficiency improvement over pure STT-RAM-based RF. Quan Deng 0003, Youtao Zhang, Shuzheng Zhang, Minxuan Zhang, Jun Yang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Aging Capacitor Supported Cache Management Scheme for Solid-State DrivesabstractSolid-state drives (SSDs) have been widely adopted in embedded systems, data centers, and cloud storage due to its well-identified advantages. Inside SSD, random access memory (RAM) is adopted as the built-in cache for achieving better performance. However, due to the volatility characteristic of RAM, data loss may happen when sudden power interrupts. In order to solve this issue, a capacitor has been equipped inside emerging SSDs as an interim power supplier. But due to the capacitor aging issue, which will result in capacitance decreases over time, there still may exist data loss when power interruption occurs. Once the remaining capacitance drops to the threshold value where all dirty pages in the cache can not be written back to flash memory, data loss happens. To solve the above issue, an efficient cache management scheme for capacitor equipped SSDs is proposed in this article. The basic idea of this scheme is to bound the number of dirty pages in a cache within the capability of the equipped capacitor. The proposed scheme includes three steps: 1) a periodical dirty page budget detection (DPBD) scheme is proposed to acquire the maximal number of dirty pages that can be written back within current capability of equipped capacitor; 2) a smart dirty page synchronizing scheme is proposed during normal run time to bound the number of dirty pages in the cache; and 3) when power supply interrupts, an efficient writing back method is applied to further reduce the capacitance consumption of capacitor. The simulation results show that the proposed scheme achieves encouraging improvement on lifetime and performance while power interruption induced data loss is avoided. Congming Gao, Liang Shi 0001, Qiao Li 0001, Kai Liu 0001, Chun Jason Xue, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2020 | A Dynamic and Proactive GPU Preemption Mechanism Using CheckpointingabstractThe demand for multitasking GPUs increases whenever the GPU may be shared by multiple applications, either spatially or temporally. This requires that GPUs can be preempted and switch context to a new application while already executing one. Unlike CPUs, context switching in GPUs is prohibitively expensive due to the large context states to swap out. There have been a number of efforts on reducing the overhead of preemption, through reducing the context sizes or overlapping context switching with execution. All those techniques are reactive approaches, meaning that context switching occurs when the preemption request arrives. In this paper, we propose a dynamic and proactive mechanism to reduce the latency of preemption. We observe that kernel execution is almost always preceded by known commands in both CUDA and OpenCL implementations. Hence, a preemption can be anticipated before the actual request arrives. We study such lead time and develop a prediction scheme to perform an early state saving. When the actual preemption is invoked, an incremental update relative to the previous saved state is performed, much like the conventional checkpointing mechanism. Our design can also choose to drain or checkpointing dynamically and accurately according to the feature of kernels in the runtime. This design effectively reduces the stall time of the preempting kernel due to context switching by 58.6%. Moreover, through careful handling of the saved state, we can also reduce the overall size of saved state by an average of 23.3%, compared with a full context switching. Chen Li 0015, Andrew Zigerelli, Jun Yang 0002, Youtao Zhang, Sheng Ma, Yang Guo 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Exploiting In-Memory Data Patterns for Performance Improvement on Crossbar Resistive MemoryabstractResistive memory (ReRAM) has emerged as a promising nonvolatile memory technology that may replace a significant portion of DRAM in future computer systems. ReRAM has many advantages, such as high density, low standby power, and good scalability. When adopting crossbar architecture, ReRAM cell can achieve the smallest theoretical size in fabrication, which is ideal for constructing dense memory with large capacity. However, crossbar cell structure suffers from a variety of reliability issues, which come from large voltage drops on long wires. To ensure operation reliability, ReRAM writes conservatively use the worst-case access latency of all cells in ReRAM arrays, which leads to significant performance degradation and dynamic energy waste. In this article, we study the correlation between the ReRAM cell switching latency and the number of cells in low-resistance state (LRS) along bitlines, and propose to dynamically speed up write operations based on bitline data patterns, i.e., the number of LRS cells presented in bitlines. We leverage the intrinsic in-memory processing capability of ReRAM crossbar and propose a low-overhead runtime profiler that effectively tracks the data patterns in different bitlines. To achieve further write latency reduction, we employ data compression and row address dependent memory data layout to reduce the numbers of LRS cells on bitlines. Moreover, we further present two optimization techniques, i.e., selective profiling and fine-grained profiling, to mitigate energy overhead brought by bitline data patterns tracking. The experimental results show that, on average, our design improves system performance by 20.5% and 14.2%, and reduces memory dynamic energy by 20.3% and 12.6%, compared to the baseline and the state-of-the-art crossbar design, respectively. Wen Wen 0003, Youtao Zhang, Jun Yang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Introduction to the Special Issue on Languages, Compilers, Tools, and Theory of Embedded Systems: Part 1abstractintroduction Share on Introduction to the Special Issue on Languages, Compilers, Tools, and Theory of Embedded Systems: Part 1 Authors: Aviral Shrivastava View Profile , Jian-Jia Chen View Profile , Youtao Zhang View Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 19Issue 5September 2020 Article No.: 30pp 1–3https://doi.org/10.1145/3417732Online:26 September 2020Publication History 0citation62DownloadsMetricsTotal Citations0Total Downloads62Last 12 Months22Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Aviral Shrivastava, Jian-Jia Chen, Youtao Zhang |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Introduction to the Special Issue on Languages, Compilers, Tools, and Theory of Embedded Systems: Part 2
Aviral Shrivastava, Jian-Jia Chen, Youtao Zhang |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Automatic CIN Grades Prediction of Sequential Cervigram Image Using LSTM With Multistate CNN FeaturesabstractCervical cancer ranks as the second most common cancer in women worldwide. In clinical practice, colposcopy is an indispensable part of screening for cervical intraepithelial neoplasia (CIN) grades and cervical cancer but exhibits high misdiagnosis rate. Existing computer-assisted algorithms for analyzing cervigram images have neglected that colposcopy is a sequential and multistate process, which is unsuitable for clinical applications. In this work, we construct a cervigram-based recurrent convolutional neural network (C-RCNN) to classify different CIN grades and cervical cancer. Convolutional neural networks are leveraged to extract spatial features. We develop a sequence-encoding module to encode discriminative temporal features and a multistate-aware convolutional layer to integrate features from different states of cervigram images. To train and evaluate the performance of C-RCNN, we leveraged a dataset of 4,753 real cervigrams and obtained 96.13% test accuracy with a specificity and sensitivity of 98.22% and 95.09%, respectively. Areas under each receiver operating characteristic curves are above 0.94, proving that visual representations and sequential dynamics can be jointly and effectively optimized in the training phase. Comparative analysis demonstrated the effectiveness of the proposed C-RCNN against competing methods, showing significant improvement over only focusing on a single frame. This architecture can be extended to other applications in medical image analysis. Zijie Yue, Shuai Ding 0001, Hao Wang 0081, Youtao Zhang, Yanchun Zhang |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | A Novel Trust Model Based Overlapping Community Detection Algorithm for Social NetworksabstractWith the fast advances in Internet technologies, social networks have become a major platform for social interaction, lifestyle demonstration, and message dissemination. Effective community detection in social networks helps to assess public sentiment, identify community leaders, and produce personalized recommendation. While different community detection approaches have been proposed in the literature, the trust model based detection schemes model user interactions as trust transfer, which helps to capture the implicit relation in the network. Unfortunately, trust model based detection schemes face acold startproblem, i.e., they cannot accurately model newly joined users as these users have few interactions for a duration after joining the network. In this paper, we propose TLCDA, a novel trust model based community detection algorithm. By enhancing the traditional trust computation with inter-node relation strength and similarity in social networks, TLCDA detects communities through coarse-grained K-Mediods clustering. Our evaluation on real social networks shows that the communities detected by TLCDA exhibit superior preference cohesion while satisfying the topology cohesion. Shuai Ding 0001, Zijie Yue, Shanlin Yang, Feng Niu, Youtao Zhang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Boosting the Performance of SSDs via Fully Exploiting the Plane Level ParallelismabstractSolid state drives (SSDs) are constructed with multiple level parallel organization, including channels, chips, dies, and planes. Among these parallel levels, plane level parallelism, which is the last level parallelism of SSDs, has the most strict restrictions. Only the same type of operations that access the same address in different planes can be processed in parallel. In order to maximize the access performance, several previous works have been proposed to exploit the plane level parallelism for host accesses and internal operations of SSDs. However, our preliminary studies show that the plane level parallelism is farfrom well utilized and should be further improved. The reason is that the strict restrictions of plane level parallelism are hard to be satisfied. In this article, a from plane to die parallel optimization framework is proposed to exploit the plane level parallelism through smartly satisfying the strict restrictions all the time. In order to achieve the objective, there are at least two challenges. First, due to that host access patterns are always complex, receiving multiple same-type requests to different planes at the same time is uncommon. Second, there are many internal activities, such as garbage collection (GC), which may destroy the restrictions. In order to solve above challenges, two schemes are proposed in the SSD controller: First, a die level write construction scheme is designed to make sure there are always N pages of data written by each write operation. Second, in a further step, a die level GC scheme is proposed to activate GC in the unit of all planes in the same die. Combing the die level write and die level GC, write accesses from both host write operations and GC induced valid page movements can be processed in parallel at all time. To further improve the performance of SSDs, host write operations blocked by GCs are suggested to be processed in parallel with GC induced valid page movements, bringing lesser waiting time cost of host write operations. As a result, the GC cost and average write latency can be significantly reduced. Experiment results show that the proposed framework is able to significantly improve the write performance without read performance impact. Congming Gao, Liang Shi 0001, Kai Liu 0001, Chun Jason Xue, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | Boosting chipkill capability under retention-error induced reliability emergencyabstractThe DRAM based main memory of high embedded systems faces two design challenges: (i) degrading reliability; and (ii) increasing power and energy consumption. While chipkill ECC (error correction code) and multi-rate refresh may be adopted to address them, respectively, a simple integration of the two results in 3x or more SDC (silent data corruption) errors and failing to meet the system reliability guarantee. This is referred to as reliability emergency. Xianwei Zhang 0001, Rujia Wang, Youtao Zhang, Jun Yang 0002 |
ASP-DAC | 3 |
| 2019 | A Framework for Memory Oversubscription Management in Graphics Processing UnitsabstractModern discrete GPUs support unified memory and demand paging. Automatic management of data movement between CPU memory and GPU memory dramatically reduces developer effort. However, when application working sets exceed physical memory capacity, the resulting data movement can cause great performance loss. Chen Li 0015, Rachata Ausavarungnirun, Christopher J. Rossbach, Youtao Zhang, Onur Mutlu, Yang Guo 0003, Jun Yang 0002 |
ASPLOS | 4 |
| 2019 | LAcc: Exploiting Lookup Table-based Fast and Accurate Vector Multiplication in DRAM-based CNN AcceleratorabstractPIM (Processing-in-memory)-based CNN (Convolutional neural network) accelerators leverage the characteristics of basic memory cells to enable simple logic and arithmetic operations so that the bandwidth constraint can be effectively alleviated. However, it remains a major challenge to support multiplication operations efficiently on PIM accelerators, in particular, DRAM-based PIM accelerators. This has prevented PIM-based accelerators from being immediately adopted for accurate CNN inference. Quan Deng 0003, Youtao Zhang, Minxuan Zhang, Jun Yang 0002 |
DAC | 2 |
| 2019 | Leveraging Approximate Data for Robust Flash StorageabstractWith the increasing bit density and adoption of 3D NAND, flash memory suffers from increased errors. To address the issue, flash devices adopt error correction codes (ECC) with strong error correction capability, like low-density parity-check (LDPC) code, to correct errors. The drawback of LDPC is that, to correct data with a high raw bit error rate (RBER), read latency will be amplified. This work proposes to address this issue with the assistance of approximate data. First, studies have been conducted and show there are ample amount of approximate data available in flash storage. Second, a novel data organization is proposed to fortify the reliability of regular data by leaving approximate data unprotected. Finally, a new data allocation strategy and modified garbage collection scheme are presented to complete the design. The experimental results show that the proposed approach can improve read performance by 30% on average comparing to current techniques. Qiao Li 0001, Liang Shi 0001, Jun Yang 0002, Youtao Zhang, Chun Jason Xue |
DAC | 4 |
| 2019 | H-ORAM: A Cacheable ORAM Interface for Efficient I/O AccessesabstractOblivious RAM (ORAM) is an effective security primitive to prevent access pattern leakage. By adding redundant memory accesses, ORAM prevents attackers from revealing the patterns in the access sequences. However, ORAM tends to introduce a huge degradation on the performance. With growing address space to be protected, ORAM has to store the majority of data in the lower level storage, which further degrades the system performance. Rujia Wang, Youtao Zhang, Jun Yang 0002 |
DAC | 3 |
| 2019 | ROC: DRAM-based Processing with Reduced Operation CyclesabstractDRAM based memory-centric computing architectures are promising solutions to tackle the challenges of memory wall. In this paper, we develop a novel design of DRAM-based processing-in-memory (PIM) architecture which achieves lower cycles in every basic operation than prior arts. Our small yet fast in-memory computing units support basic logic operations including NOT, AND, and OR. Using those operations, along with shift and propagation, bitwise operations can be extended to word-wise operations, e.g. increment and comparison, with high efficiency. We also optimize the designs to exploit parallelism and data reuse to further improve the performance of compound operations. Compared with the most powerful state-of-the-art PIM architecture, we can achieve comparable or even better performance while consuming only 6% of its area overhead. Xin Xin 0008, Youtao Zhang, Jun Yang 0002 |
DAC | 2 |
| 2019 | ReNEW: Enhancing Lifetime for ReRAM Crossbar Based Neural Network AcceleratorsabstractWith analog current accumulation feature, resistive memory (ReRAM) crossbars are widely studied to accelerate neural network applications. The ReRAM crossbar based accelerators have many advantages over conventional CMOS-based accelerators, such as high performance and energy efficiency. However, due to the limited cell endurance, these accelerators suffer from short programming cycles when weights that stored in ReRAM cells are frequently updated during the neural network training phase. In this paper, by exploiting the wearing out mechanism of ReRAM cell, we propose a novel comprehensive framework, ReNEW, to enhance the lifetime of the ReRAM crossbar based accelerators, particularly for neural network training. Evaluation results show that, our proposed schemes reduce the total effective writes to ReRAM crossbar based accelerators by up to 500.3×, 50.0×, 2.83× and 1.60× over two MLC ReRAM crossbar baselines, one SLC ReRAM crossbar baseline and an SLC ReRAM crossbar design with optimal timing, respectively. Wen Wen 0003, Youtao Zhang, Jun Yang 0002 |
ICCD | 2 |
| 2019 | RFAcc: a 3D ReRAM associative array based random forest acceleratorabstractRandom forest (RF) is a widely adopted machine learning method for solving classification and regression problems. Training a random forest demands a large number of relational comparison and data movement operations, which take long time when using modern CPUs. Accelerating random forest training using either GPUs or FPGAs achieves only modest speedups. Quan Deng 0003, Youtao Zhang, Jun Yang 0002 |
ICS | 3 |
| 2019 | Constructing Large, Durable and Fast SSD System via Reprogramming 3D TLC Flash MemoryabstractNAND flash memory based SSDs have been widely studied and adopted. The scaling of SSD has evolved from plannar (2D) to 3D stacking. Compared with 2D SSD, 3D SSD stacks more layers into one block, constructing one block with more flash pages. For reliability and other reasons, technology node in 3D NAND SSD is larger than in 2D, but data density can be increased via increasing bit-per-cell. However, representing multiple bits per cell encounters additional challenges such as endurance and access latency. In this work, we develop a novel reprogramming scheme for TLCs in 3D NAND SSD, such that a cell can be programmed and reprogrammed several times before it is erased. Such reprogramming can reduce the frequency of erases which determines the endurance of a cell, improve the speed of programming, and increase the amount of bits written in a cell per program/erase cycle, i.e., effective capacity. Our work is the first to perform real 3D NAND SSD test to validate the feasibility of the reprogram operation. From the collected data, we derive the restrictions of performing reprogramming due to reliability challenges. Further, a reprogrammable SSD (ReSSD) is designed to structure reprogram operations, and when they should be applied. ReSSD is evaluated in a case study in 3D TLC SSD based RAID 5 system (RSS-RAID). Experimental results show that RSS-RAID can improve the endurance by 30.3%, boost write performance by 16.7%, and increase effective capacity by 7.71%, with negligible overhead compared with conventional 3D SSD based RAID 5 system. Congming Gao, Qiao Li 0001, Chun Jason Xue, Youtao Zhang, Liang Shi 0001, Jun Yang 0002 |
MICRO | 5 |
| 2019 | Parallel all the time: Plane Level Parallelism Exploration for High Performance SSDsabstractSolid state drives (SSDs) are constructed with multiple level parallel organization, including channels, chips, dies and planes. Among these parallel levels, plane level parallelism, which is the last level parallelism of SSDs, has the most strict restrictions. Only the same type of operations which access the same address in different planes can be processed in parallel. In order to maximize the access performance, several previous works have been proposed to exploit the plane level parallelism for host accesses and internal operations of SSDs. However, our preliminary studies show that the plane level parallelism is far from well utilized and should be further improved. The reason is that the strict restrictions of plane level parallelism are hard to be satisfied. In this work, a from plane to die parallel optimization framework is proposed to exploit the plane level parallelism through smartly satisfying the strict restrictions all the time. In order to achieve the objective, there are at least two challenges. First, due to that host access patterns are always complex, receiving multiple same-type requests to different planes at the same time is uncommon. Second, there are many internal activities, such as garbage collection (GC), which may destroy the restrictions. In order to solve above challenges, two schemes are proposed in the SSD controller: First, a die level write construction scheme is designed to make sure there are always N pages of data written by each write operation. Second, in a further step, a die level GC scheme is proposed to activate GC in the unit of all planes in the same die. Combing the die level write and die level GC, write accesses from both host write operations and GC induced valid page movements can be processed in parallel at all time. As a result, the GC cost and average write latency can be significantly reduced. Experiment results show that the proposed framework is able to significantly improve the write performance without read performance impact. Congming Gao, Liang Shi 0001, Chun Jason Xue, Cheng Ji 0002, Jun Yang 0002, Youtao Zhang |
MSST | 6 |
| 2019 | DWMAcc: Accelerating Shift-based CNNs with Domain Wall MemoriesabstractPIM (processing-in-memory) based hardware accelerators have shown great potentials in addressing the computation and memory access intensity of modern CNNs (convolutional neural networks). While adopting NVM (non-volatile memory) helps to further mitigate the storage and energy consumption overhead, adopting quantization, e.g., shift-based quantization, helps to tradeoff the computation overhead and the accuracy loss, integrating both NVM and quantization in hardware accelerators leads to sub-optimal acceleration. In this paper, we exploit the natural shift property of DWM (domain wall memory) to devise DWMAcc, a DWM-based accelerator with asymmetrical storage of weight and input data, to speed up the inference phase of shift-based CNNs. DWMAcc supports flexible shift operations to enable fast processing with low performance and area overhead. We then optimize it with zero-sharing , input-reuse , and weight-share schemes. Our experimental results show that, on average, DWMAcc achieves 16.6× performance improvement and 85.6× energy consumption reduction over a state-of-the-art SRAM based design. Zhengguo Chen, Quan Deng 0003, Nong Xiao 0001, Kirk Pruhs, Youtao Zhang |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2018 | DrAcc: a DRAM based accelerator for accurate CNN inferenceabstractModern Convolutional Neural Networks (CNNs) are computation and memory intensive. Thus it is crucial to develop hardware accelerators to achieve high performance as well as power/energy-efficiency on resource limited embedded systems. DRAM-based CNN accelerators exhibit great potentials but face inference accuracy and area overhead challenges. Quan Deng 0003, Lei Jiang 0001, Youtao Zhang, Minxuan Zhang, Jun Yang 0002 |
DAC | 3 |
| 2018 | Wear leveling for crossbar resistive memoryabstractResistive Memory (ReRAM) is an emerging non-volatile memory technology that has many advantages over conventional DRAM. ReRAM crossbar has the smallest 4F2 planar cell size and thus is widely adopted for constructing dense memory with large capacity. However, ReRAM crossbar suffers from large sneaky currents and IR drop. To ensure write reliability, ReRAM write drivers choose larger than ideal write voltages, which over-SET/over-RESET many cells at runtime and lead to severely degraded chip lifetime. Wen Wen 0003, Youtao Zhang, Jun Yang 0002 |
DAC | 2 |
| 2018 | ShadowGC: Cooperative garbage collection with multi-level buffer for performance improvement in NAND flash-based SSDsabstractGarbage collection, an essential background activity in NAND flash based SSDs, often introduces large runtime overhead. Recent studies showed that it is beneficial to separate the flash pages that have dirty copies in the write buffers from those that do not. However, the existing schemes exploring this observation have limitations, which prevent them from maximizing the performance improvement. In this paper, we address the above challenge through ShadowGC, a novel GC design that exploits the pages in both host-side and device-side write buffers and adopts different read and write strategies to minimize the GC overhead. When garbage collecting flash pages that have dirty copies in the device-side write buffer, ShadowGC reads data from the write buffer. When garbage collecting flash pages that have dirty copies in the host-side write buffer, ShadowGC moves them to dedicated blocks and speeds up the movement with fast-write operations. Our experimental results show that, on average, ShadowGC reduces the write amplification by 16.2% and the GC latency by 20.5% over the state-of-the-art. Jinhua Cui 0001, Youtao Zhang, Jianhang Huang, Weiguo Wu, Jun Yang 0002 |
DATE | 2 |
| 2018 | D-ORAM: Path-ORAM Delegation for Low Execution Interference on Cloud Servers with Untrusted MemoryabstractCloud computing has evolved into a promising computing paradigm. However, it remains a challenging task to protect application privacy and, in particular, the memory access patterns, on cloud servers. The Path ORAM protocol achieves high-level privacy protection but requires large memory bandwidth, which introduces severe execution interference. The recently proposed secure memory model greatly reduces the security enhancement overhead but demands the secure integration of cryptographic logic and memory devices, a memory architecture that is yet to prevail in mainstream cloud servers. In this paper, we propose D-ORAM, a novel Path ORAM scheme for achieving high-level privacy protection and low execution interference on cloud servers with untrusted memory. D-ORAM leverages the buffer-on-board (BOB) memory architecture to offload the Path ORAM primitives to a secure engine in the BOB unit, which greatly alleviates the contention for the off-chip memory bus between secure and non-secure applications. D-ORAM upgrades only one secure memory channel and employs Path ORAM tree split to extend the secure application flexibly across multiple channels, in particular, the non-secure channels. D-ORAM optimizes the link utilization to further improve the system performance. Our evaluation shows that D-ORAM effectively protects application privacy on mainstream computing servers with untrusted memory, with an improvement of NS-App performance by 22.5% on average over the Path ORAM baseline. Rujia Wang, Youtao Zhang, Jun Yang 0002 |
HPCA | 2 |
| 2018 | Enabling Intra-Plane Parallel Block Erase in NAND Flash to Alleviate the Impact of Garbage CollectionabstractGarbage collection (GC) in NAND flash can significantly decrease I/O performance in SSDs by copying valid data to other locations, thus blocking incoming I/O requests. To help improve performance, NAND flash utilizes various advanced commands to increase internal parallelism. Currently, these commands only parallelize operations across channels, chips, dies, and planes, neglecting the block level due to risk of disturbances that can compromise valid data by inducing errors. However, due to the triple-well structure of the NAND flash plane architecture, it is possible to erase multiple blocks within a plane, in parallel, without diminishing the integrity of the valid data. The number of page movements due to multiple block erases can be restrained so as to bound the overhead per GC. Moreover, more capacity can be reclaimed per GC which delays future GCs and effectively reduces their frequency. Such an Intra-Plane Parallel Block Erase (IPPBE) in turn diminishes the impact of GC on incoming requests, improving their response times. Experimental results show that IPPBE can reduce the time spent performing GC by up to 50.7% and 33.6% on average, read/write response time by up to 47.0%/45.4% and 16.5%/14.8% on average respectively, page movements by up to 52.2% and 26.6% on average, and blocks erased by up to 14.2% and 3.6% on average. An energy analysis conducted indicates that by reducing the number of page copies and the number of block erases, the energy cost of garbage collection can be reduced up to 44.1% and 19.3% on average. Tyler Garrett, Jun Yang 0002, Youtao Zhang |
ISLPED | 3 |
| 2018 | Time-aware cloud service recommendation using similarity-enhanced collaborative filtering and ARIMA model
Shuai Ding 0001, Yeqing Li, Desheng Dash Wu, Youtao Zhang, Shanlin Yang |
Decis. Support Syst. | 4 |
| 2018 | ApproxFTL: On the Performance and Lifetime Improvement of 3-D NAND Flash-Based SSDsabstract3-D NAND flash is one of the most prospective advances in flash memory industry. While 3-D flash improves cell density and reduces lithography cost through die stacking, it suffers from severe program disturbance, which leads to significant performance and lifetime degradation for 3-D flash-based SSDs. To address the above challenge, we propose ApproxFTL, an approximate-write aware flash translation layer design, that uses approximate-write operations to store error-resilient data of modern applications. By reducing the maximal threshold voltage and tightening the guard bands between multilevel cell states, approximate write operations not only finish early but also exhibit large disturbance reduction, which can be exploited to alleviate disturbance in physical blocks that save both precise and approximate data. ApproxFTL maximizes the disturbance mitigation through approximate-write aware data placement, wear leveling, and garbage collection enhancements. Our experimental results show that ApproxFTL, while preserving high data quality, improves the read and write response time of flash accesses by 41.38% and 45.64% on average, respectively, and extends the lifetime of 3-D flash-based SSDs by 5.75% when comparing to the state-of-the-art. Jinhua Cui 0001, Youtao Zhang, Liang Shi 0001, Chun Jason Xue, Weiguo Wu, Jun Yang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | DLV: Exploiting Device Level Latency Variations for Performance Improvement on Flash Memory Storage SystemsabstractNAND flash has been widely adopted in storage systems due to its better read and write performance and lower power consumption over traditional mechanical hard drives. To meet the increasing performance demand of modern applications, recent studies speed up flash accesses by exploiting access latency variations at the device level. Unfortunately, existing flash access schedulers are still oblivious to such variations, leading to suboptimal I/O performance improvements. In this paper, we propose DLV, a novel flash access scheduler for exploring scheduling opportunities due to device level access latency variations. DLV improves flash access speeds based on process variations and data retention time difference across flash blocks. More importantly, DLV integrates access speed optimization with access scheduling such that the average access response time can be effectively reduced on flash memory storage systems. Our experimental results show that DLV achieves an average of 41.5% performance improvement over the state-of-the-art. Jinhua Cui 0001, Youtao Zhang, Weiguo Wu, Jun Yang 0002, Yinfeng Wang, Jianhang Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | DrMP: Mixed Precision-Aware DRAM for High Performance Approximate and Precise ComputingabstractRecent studies showed that DRAM restore time degrades as technology scales, which imposes large performance and energy overheads. This problem, prolonged restore time (PRT), has been identified by the DRAM industry as one of three major scaling challenges. This paper proposes DrMP, a novel fine-grained precision-aware DRAM restore scheduling approach, to mitigate PRT. The approach exploits process variations (PVs) within and across DRAM rows to save data with mixed precision. The paper describes three variants of the approach: DrMP-A, DrMP-P, and DrMP-U. DrMP-A supports approximate computing by mapping important data bits to fast row segments to reduce restore time for improved performance at a low application error rate. DrMP-P pairs memory rows together to reduce the average restore time for precise computing. DrMP-U combines DrMP-A and DrMP-P to better trade performance, energy consumption, and computation precision. Our experimental results show that, on average, DrMP achieves 20% performance improvement and 15% energy reduction over a precision-oblivious baseline. Further, DrMP achieves an error rate less than 1% at the application level for a suite of benchmarks, including applications that exhibit unacceptable error rates under simple approximation that does not differentiate the importance of different bits. Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
PACT | 2 |
| 2017 | Cooperative Path-ORAM for Effective Memory Bandwidth Sharing in Server SettingsabstractPath ORAM (Oblivious RAM) is a recently proposed ORAM protocol for preventing information leakage from memory access sequences. It receives wide adoption due to its simplicity, practical efficiency and asymptotic efficiency. However, Path ORAM has extremely large memory bandwidth demand, leading to severe memory competition in server settings, e.g., a server may service one application that uses Path ORAM and one or multiple applications that do not. While Path ORAM synchronously and intensively uses all memory channels, the non-secure applications often exhibit low access intensity and large channel level imbalance. Traditional memory scheduling schemes lead to wasted memory bandwidth to the system and large performance degradation to both types of applications. In this paper, we propose CP-ORAM, a Cooperative Path ORAM design, to effectively schedule the memory requests from both types of applications. CP-ORAM consists of three schemes: P-Path,R-Path, and W-Path. P-Path assigns and enforces scheduling priority for effective memory bandwidth sharing. R-Path maximizes bandwidth utilization by proactively scheduling read operations from the next Path ORAM access. W-Path mitigates contention on busy memory channels with write redirection. We evaluate CP-ORAM and compare it to the state-of-the-art. Our results show that CP-ORAM helps to achieve 20% performance improvement on average over the baseline Path ORAM for the secure application in a four-channel server setting. Rujia Wang, Youtao Zhang, Jun Yang 0002 |
HPCA | 2 |
| 2017 | Towards warp-scheduler friendly STT-RAM/SRAM hybrid GPGPU register file designabstractModern Graphics Processing Units (GPUs) widely adopt large SRAM based register file (RF) to enable fast context-switch. A large SRAM RF may consume 20% to 40% GPU power, which has become one of the major design challenges for GPUs. Recent studies mitigate the issue through hybrid RF designs that architect a large STT-RAM (Spin Transfer Torque Magnetic memory) RF and a small SRAM buffer. However, the long STT-RAM write latency throttles the data exchange between STT-RAM and SRAM, which deprecates warp scheduler with frequent context switches, e.g., round robin scheduler. In this paper, we propose HC-RF, a warp-scheduler friendly hybrid RF design using novel SRAM/STT-RAM hybrid cell (HC) structure. HC-RF exploits cell level integration to improve the effective bandwidth between STT-RAM and SRAM. By enabling silent data transfer from SRAM to STT-RAM without blocking RF banks, HC-RF supports concurrent context-switching and decouples its dependency on warp scheduler. Our experimental results show that, on average, HC-RF achieves 50% performance improvement and 44% energy consumption reduction over the coarse-grained hybrid design when adopting LRR(Loose Round Robin) warp scheduler. Quan Deng 0003, Youtao Zhang, Minxuan Zhang, Jun Yang 0002 |
ICCAD | 2 |
| 2017 | Speeding up crossbar resistive memory by exploiting in-memory data patternsabstractResistive Memory (ReRAM) has emerged as a promising non-volatile memory technology that may replace a significant portion of DRAM in future computer systems. ReRAM has many advantages such as high density, low standby power and good scalability. ReRAM, when adopting crossbar architecture, has the smallest 4F2planar cell size, which is ideal for constructing dense memory with large capacity. However, crossbar cell structure suffers from large sneak leakage and IR drop on long wires. To ensure operation reliability, ReRAM writes, in particular, RESET operations, conservatively use the worst-case access latency of all cells in ReRAM arrays, which leads to significant performance degradation and dynamic energy waste. In this paper, we study the correlation between the RESET latency and the number of cells in low resistant state (LRS) along bitlines, and propose to dynamically speed up ReRAM RESET operations for the rows that have small numbers of LRS cells. We leverage the intrinsic in-memory processing capability of ReRAM crossbar and propose a low overhead runtime profiler that effectively tracks the data patterns in different bitlines. To achieve further RESET latency reduction, we employ data compression and row address dependent data layout to reduce LRS cells on bitlines. The experimental results show that, on average, our design improves system performance by 20.5% and 14.2%, and reduces memory dynamic energy by 15.7% and 7.6%, compared to the baseline and the state-of-the-art crossbar design. Wen Wen 0003, Youtao Zhang, Jun Yang 0002 |
ICCAD | 3 |
| 2017 | AEP: An error-bearing neural network accelerator for energy efficiency and model protectionabstractNeural Networks (NNs) have recently gained popularity in a wide range of modern application domains due to its superior inference accuracy. With growing problem size and complexity, modern NNs, e.g., CNNs (Convolutional NNs) and DNNs (Deep NNs), contain a large number of weights, which require tremendous efforts not only to prepare representative training datasets but also to train the network. There is an increasing demand to protect the NN weight matrices, an emerging Intellectual Property (IP) in NN field. Unfortunately, adopting conventional encryption method faces significant performance and energy consumption overheads. In this paper, we propose AEP, a DianNao based NN accelerator design for IP protection. AEP aggressively reduces DRAM timing to generate a device dependent error mask, i.e., a set of erroneous cells while the distribution of these cells are device dependent due to process variations. AEP incorporates the error mask in the NN training process so that the trained weights are device dependent, which effectively defects IP piracy as exporting the weights to other devices cannot produce satisfactory inference accuracy. In addition, AEP speeds up NN inference and achieves significant energy reduction due to the fact that main memory dominates the energy consumption in DianNao accelerator. Our evaluation results show that by injecting 0.1% to 5% memory errors, AEP has negligible inference accuracy loss on the target device while exhibiting unacceptable accuracy degradation on other devices. In addition, AEP achieves an average of 72% performance improvement and 44% energy reduction over the DianNao baseline. Youtao Zhang, Jun Yang 0002 |
ICCAD | 2 |
| 2017 | AEP: An error-bearing neural network accelerator for energy efficiency and model protectionabstractNeural Networks (NNs) have recently gained popularity in a wide range of modern application domains due to its superior inference accuracy. With growing problem size and complexity, modern NNs, e.g., CNNs (Convolutional NNs) and DNNs (Deep NNs), contain a large number of weights, which require tremendous efforts not only to prepare representative training datasets but also to train the network. There is an increasing demand to protect the NN weight matrices, an emerging Intellectual Property (IP) in NN field. Unfortunately, adopting conventional encryption method faces significant performance and energy consumption overheads. In this paper, we propose AEP, a DianNao based NN accelerator design for IP protection. AEP aggressively reduces DRAM timing to generate a device dependent error mask, i.e., a set of erroneous cells while the distribution of these cells are device dependent due to process variations. AEP incorporates the error mask in the NN training process so that the trained weights are device dependent, which effectively defects IP piracy as exporting the weights to other devices cannot produce satisfactory inference accuracy. In addition, AEP speeds up NN inference and achieves significant energy reduction due to the fact that main memory dominates the energy consumption in DianNao accelerator. Our evaluation results show that by injecting 0.1% to 5% memory errors, AEP has negligible inference accuracy loss on the target device while exhibiting unacceptable accuracy degradation on other devices. In addition, AEP achieves an average of 72% performance improvement and 44% energy reduction over the DianNao baseline. Youtao Zhang, Jun Yang 0002 |
ICCAD | 2 |
| 2017 | Read Error Resilient MLC STT-MRAM Based Last Level CacheabstractSTT-MRAM is a promising non-volatile memory technology for building large LLCs (Last Level Caches). Multi-level cell (MLC) STT-MRAM can further enlarge cache capacity with reduced per bit cost. However, due to fast technology scaling, STT-MRAM, in particular, MLC STT-MRAM, suffers from significant read errors, including read disturbance errors and sensing errors, which lead to unreliable accesses that prevent the adoption of MLC STT-MRAM in LLCs. In this paper, we propose R2M, a read error resilient 2T4J (two transistor four MTJ) MLC based LLC design. R2M leverages the recently industry-proposed 2T2J single-level cell (SLC) structure to achieve good tradeoff between reliability and capacity. It consists of two schemes: R2M-S and R2M-C. R2M-S improves read reliability by sensing the resistance difference of two cells, which effectively mitigates both sensing and disturbance errors for the soft bit of MLC. R2M-C further enhances error resiliency by exploiting access locality and data redundancy. We evaluate the proposed R2M design and compare it to the state-of-the-art. Our experimental results show that, on average, R2M achieves 54.8% performance improvement and 42.0% energy consumption reduction over the state-of-the-art MLC design. Wen Wen 0003, Youtao Zhang, Jun Yang 0002 |
ICCD | 2 |
| 2017 | Quality of Service Support for Fine-Grained Sharing on GPUsabstractGPUs have been widely adopted in data centers to provide acceleration services to many applications. Sharing a GPU is increasingly important for better processing throughput and energy efficiency. However, quality of service (QoS) among concurrent applications is minimally supported. Previous efforts are too coarse-grained and not scalable with increasing QoS requirements. We propose QoS mechanisms for a fine-grained form of GPU sharing. Our QoS support can provide control over the progress of kernels on a per cycle basis and the amount of thread-level parallelism of each kernel. Due to accurate resource management, our QoS support has significantly better scalability compared with previous best efforts. Evaluations show that, when the GPU is shared by three kernels, two of which have QoS goals, the proposed techniques achieve QoS goals 43.8% more often than previous techniques and have 20.5% higher throughput. Zhenning Wang, Jun Yang 0002, Rami G. Melhem, Bruce R. Childers, Youtao Zhang, Minyi Guo |
ISCA | 5 |
| 2017 | Multi-objective optimization based ranking prediction for cloud service recommendation
Shuai Ding 0001, Chengyi Xia, Chengjiang Wang, Desheng Dash Wu, Youtao Zhang |
Decis. Support Syst. | 5 |
| 2017 | On the Restore Time Variations of Future DRAM MemoryabstractAs the de facto main memory standard, DRAM (Dynamic Random Access Memory) has achieved dramatic density improvement in the past four decades, along with the advancements in process technology. Recent studies reveal that one of the major challenges in scaling DRAM into the deep sub-micron regime is its significant variations on cell restore time, which affect timing constraints such as write recovery time. Adopting traditional approaches results in either low yield rate or large performance degradation. In this article, we propose schemes to expose the variations to the architectural level. By constructing memory chunks with different access speeds and, in particular, exploiting the performance benefits of fast chunks, a variation-aware memory controller can effectively mitigate the performance loss due to relaxed timing constraints. We then proposed restore-time-aware rank construction and page allocation schemes to make better use of fast chunks. Our experimental results show that, compared to traditional designs such as row sparing and Error Correcting Codes, the proposed schemes help to improve system performance by about 16% and 20%, respectively, for 20nm and 14nm technology nodes on a four-core multiprocessor system. Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | ReadDuo: Constructing Reliable MLC Phase Change Memory through Fast and Robust ReadoutabstractPhase change memory (PCM) has emerged as a promising non-volatile memory technology. Multi-level cell (MLC) PCM, while effectively reducing per bit fabrication cost, suffers from resistance drift based soft errors. It is challenging to construct reliable MLC chips that achieve high performance, high storage density, and low energy consumption simultaneously. In this paper, we propose ReadDuo, a fast and robust readout solution to address resistance drift in MLC PCM. We first integrate fast current sensing and resistance drift resilient voltage sensing, which exposes performance optimization opportunities without sacrificing reliability. We then devise last writes tracking and selective different write schemes to minimize performance and energy consumption overhead in scrubbing. Our experimental results show that ReadDuo achieves 37% improvement on average over existing solutions when considering performance, dynamic energy consumption, and storage density all together. Rujia Wang, Youtao Zhang, Jun Yang 0002 |
DSN | 2 |
| 2016 | Simultaneous Multikernel GPU: Multi-tasking throughput processors via fine-grained sharingabstractStudies show that non-graphics programs can be less optimized for the GPU hardware, leading to significant resource under-utilization. Sharing the GPU among multiple programs can effectively improve utilization, which is particularly attractive to systems where many applications require access to the GPU (e.g., cloud computing). However, current GPUs lack proper architecture features to support sharing. Initial attempts are preliminary: They either provide only static sharing, which requires recompilation or code transformation, or they do not effectively improve GPU resource utilization. We propose Simultaneous Multikernel (SMK), a fine-grain dynamic sharing mechanism, that fully utilizes resources within a streaming multiprocessor by exploiting heterogeneity of different kernels. We propose several resource allocation strategies to improve system throughput while maintaining fairness. Our evaluation shows that for shared workloads with complementary resource occupancy, SMK improves GPU throughput by 52% over non-shared execution and 17% over a state-of-the-art design. Zhenning Wang, Jun Yang 0002, Rami G. Melhem, Bruce R. Childers, Youtao Zhang, Minyi Guo |
HPCA | 5 |
| 2016 | Restore truncation for performance improvement in future DRAM systemsabstractScaling DRAM below 20nm has become a major challenge due to intrinsic limitations in the structure of a bit cell. Future DRAM chips are likely to suffer from significant variations and degraded timings, such as taking much more time to restore cell data after read and write access. In this paper, we propose restore truncation (RT), a low-cost restore strategy to improve performance of DRAM modules that adopt relaxed restore timing. After an access, RT restores a bit cell's voltage only to the level required to persist data to the next scheduled refresh rather than to the default full voltage. Because restore time is shortened, the performance of the cell is improved under process variations. We devise two schemes to balance performance, energy consumption, and hardware overhead. We simulate our proposed RT schemes and compare them with the state of the art. Experimental results show that, on average, RT improves performance by 19.5% and reduces energy consumption by 17%. Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
HPCA | 2 |
| 2015 | SD-PCM: Constructing Reliable Super Dense Phase Change Memory under Write DisturbanceabstractPhase Change Memory (PCM) has better scalability and smaller cell size comparing to DRAM. However, further scaling PCM cell in deep sub-micron regime results in significant thermal based write disturbance (WD). Naively allocating large inter-cell space increases cell size from 4F2 ideal to 12F2. While a recent work mitigates WD along word-lines through disturbance resilient data encoding, it is ineffective for WD along bit-lines, which is more severe due to widely adopted $\mu$Trench structure in constructing PCM cell arrays. Without mitigating WD along bit-lines, a PCM cell still has 8F2, which is 100% larger than the ideal. In this paper, we propose SD-PCM for achieving reliable write operations in super dense PCM. In particular, we focus on mitigating WD along bit-lines such that we can construct super dense PCM chips with 4F2 cell size, i.e., the minimal for diode-switch based PCM. Based on simple verification-n-correction (VnC), we propose LazyCorrection and PreRead to effectively reduce VnC overhead and minimize cascading verification during write. We further propose (n:m)-Alloc for achieving good tradeoff between VnC overhead minimization and memory capacity loss. Our experimental results show that, comparing to a WD-free low density PCM, SD-PCM achieves 80% capacity improvement in cell arrays while incurring around 0-10% performance degradation when using different (n:m) allocators. Rujia Wang, Lei Jiang 0001, Youtao Zhang, Jun Yang 0002 |
ASPLOS | 3 |
| 2015 | Selective restore: an energy efficient read disturbance mitigation scheme for future STT-MRAMabstractSTT-MRAM (Spin-Transfer Torque Magnetic RAM) has recently emerged as one of the most promising memory technologies for constructing large capacity last level cache (LLC) of low power mobile processors. With fast technology scaling, STT-MRAM read operations will become destructive such that post-read restores are inevitable to ensure data reliability. However, frequent restores introduce large energy overheads. In this paper, we propose Selective Restore (SR), an energy efficient scheme to mitigate the restore overheads. Given a L2 cacheline disturbed from a read operation, SR postpones its restore till the cacheline being evicted from the upper level cache L1. Based on the status of the line at the eviction time, SR selectively restores the disturbed cells to achieve energy efficiency. Our experimental results show that SR improves system performance by 5% and reduces dynamic energy consumption by 62%. Rujia Wang, Lei Jiang 0001, Youtao Zhang, Linzhang Wang, Jun Yang 0002 |
DAC | 3 |
| 2015 | Exploit imbalanced cell writes to mitigate write disturbance in dense phase change memoryabstractRecent studies have shown that Phase Change Memory faces significant write disturbance (WD) when scaling in deep submicron regime, i.e., resetting a cell may disturb the values of its adjacent cells if these cells are in amorphous state. A preventive approach to mitigate WD errors is to allocate sufficient inter-cell thermal band. However, this approach greatly reduces chip capacity due to low cell density. A cost effective approach VnC (verify-and-correct), relies on Verification after each write and Correction if errors do happen. Simple VnC improves chip capacity but introduces large performance degradation. Rujia Wang, Lei Jiang 0001, Youtao Zhang, Linzhang Wang, Jun Yang 0002 |
DAC | 3 |
| 2015 | Exploiting DRAM restore time variations in deep sub-micron scaling
Xianwei Zhang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
DATE | 2 |
| 2015 | DLB: Dynamic lane borrowing for improving bandwidth and performance in Hybrid Memory CubeabstractThe Hybrid Memory Cube (HMC) is an innovative DRAM architecture that adopts 3D-stacking to improve bandwidth and save energy. An HMC module adopts separate receive and transmit lanes and thus may achieve the maximal memory bandwidth only if data can be driven at full speed in both directions. However, due to the natural read and write imbalance in modern applications, the effective memory bandwidth utilization is often low, leading to suboptimal system performance. In this paper, we propose DLB (dynamic lane borrowing) that dynamically tracks link utilization and partitions the lanes in one link between receive and transmit directions. DLB allocates more lanes to transmit if servicing read-intensive applications. With more lanes allocated to either direction, DLB reduces the lane contention along that direction and thus the average memory access latency. Our experimental results show that DLB improves the bandwidth utilization by 10.4% on average, reduces the average utilization gap in two directions from 35.6% to 12.8%, and saves execution time by as much as 22.3%. Xianwei Zhang 0001, Youtao Zhang, Jun Yang 0002 |
ICCD | 2 |
| 2015 | TriState-SET: Proactive SET for improved performance of MLC phase change memoriesabstractThe emerging Phase Change Memory (PCM) has many advantages such as good scalability and low leakage. MLC (Multi-Level Cell) PCM further extends the benefits by storing two or more bits per cell and thus reducing the per bit cost. However, adopting MLC PCM in main memory often leads to long write latency, high energy consumption, and degraded performance. In this paper, we propose TriState-SET, a proactive-SET based write strategy for improving MLC PCM write performance. TriState-SET proactively places device cells of a dirty memory line in full SET state. By utilizing only three states of 2bit MLC PCM, TriState-SET involves only fast state transitions when writing such a line at write-back time. Our experimental results show that TriState-SET increases performance by 11% and saves system energy by 6.7% (up to 12.2%), while achieving up to 25% (average 14.1%) energy-delay-product improvement. Xianwei Zhang 0001, Youtao Zhang, Jun Yang 0002 |
ICCD | 2 |
| 2015 | Exploit common source-line to construct energy efficient domain wall memory based cachesabstractDomain wall memory (DWM) is an emerging memory technology that utilizes magnetic domains along a nanowire to achieve high density, short latency and low power. Recent studies showed that it is promising to replace SRAM and STT-MRAM to construct DWM based on-chip caches. However, accessing DWM requires frequent shift operations, which leads to large energy consumption for DWM caches. In this paper, we propose DWM-SSL, an architectural innovation to achieve energy efficiency for multiple-head based DWM caches. DWM-SSL adopts common source line design to re-organize DWM cell arrays such that accessing an N-bit cache line from M-head DWM based cache activates N/M tracks instead of N tracks in the baseline. Our experimental results show that, on average, DWM-SSL reduces around 5.1x track shifts and up to 63% cache energy consumption for a 4-head DWM cache design. Xianwei Zhang 0001, Youtao Zhang, Jun Yang 0002 |
ICCD | 3 |
| 2015 | Wear Relief for High-Density Phase Change Memory Through Cell Morphing Considering Process VariationabstractDue to the scalability and large leakage power, dynamic random-access memory (DRAM) has a lot of challenges in scaling. As an alternative, phase change memory (PCM) has demonstrated promising potential to serve as the main memory in deep submicrometer regime. The broad resistance range of PCM cells enables several cell modes with various densities, pertaining to multiple level cell (MLC), triple state cell (TSC), and single level cell (SLC). High-density mode outperforms low-density ones in terms of capacity and cost-per-bit, but suffers from a weaker cell endurance. Wear leveling strategies are proposed to enhance the memory endurance but encounter more challenges with the aggravating process variation. Due to endurance variations, physical domains are fabricated with irregular tenacity. As a result, balanced write traffic, which is the objective of traditional wear leveling, cannot fully exploit the PCM endurance since the weak parts will be worn out sooner than others. In this paper, considering process variation, we propose a cell morphing based wear leveling scheme. Cell morphing refers to the cell mode transformation between high density (e.g., MLC) and low densities (e.g., TSC and SLC). Instead of redistributing write operations, the proposed wear leveling scheme dynamically transforms weak and frequently written portions into low-density mode for endurance benefits. Multitier cell morphing schemes are proposed to support mode transformation among multiple density levels. The experimental results show 236% endurance improvement for single-tier cell morphing and 209% for two-tier cell morphing with 2% low-density page percentage, when compared with the most related work. Mengying Zhao, Lei Jiang 0001, Liang Shi 0001, Youtao Zhang, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Constructing Large and Fast On-Chip Cache for Mobile Processors with Multilevel Cell STT-MRAM TechnologyabstractModern mobile processors integrating an increasing number of cores into one single chip demand large-capacity, on-chip, last-level caches (LLCs) in order to achieve scalable performance improvements. However, adopting traditional memory technologies such as SRAM and embedded DRAM (eDRAM) leakage and scalability problems. Spin-transfer torque magnetic RAM (STT-MRAM) is a novel nonvolatile memory technology that has emerged as a promising alternative for constructing on-chip caches in high-end mobile processors. STT-MRAM has many advantages, such as short read latency, zero leakage from the memory cell, and better scalability than eDRAM and SRAM. Multilevel cell (MLC) STT-MRAM further enlarges capacity and reduces per-bit cost by storing more bits in one cell. However, MLC STT-MRAM has long write latency which limits the effectiveness of MLC STT-MRAM-based LLCs. In this article, we address this limitation with three novel designs: line pairing (LP), line swapping (LS), and dynamic LP/LS enabler (DLE). LP forms fast cache lines by reorganizing MLC soft bits which are faster to write. LS dynamically stores frequently-written data into these fast cache lines. We then propose a dynamic LP/LS enabler (DLE) to enable LP and LS only if they help to improve the overall cache performance. Our experimental results show that the proposed designs improve system performance by 9--15% and reduce energy consumption by 14--21% for various types of mobile processors. Lei Jiang 0001, Bo Zhao 0007, Jun Yang 0002, Youtao Zhang |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2014 | CD-RAIS: Constrained dynamic striping in redundant array of independent SSDsabstractSolid state drives (SSDs) are increasingly deployed to construct storage arrays (RAIDs) in enterprise environments. The design decisions in RAID are traditionally devised for HDD RAIDs, which cannot fully exploit the characteristics of SSDs. In particular, SSD lacks the ability to update pages in-place. Random writes in traditional parity-based RAIS (SSD RAID) systems that has static striping result in significantly more writes, degraded performance, and shortened SSD lifetime. By dynamically forming full stripes, log-based design originally proposed in HDD RAID can mitigate the write-hole problem caused by random writes. However, it needs a directory to record locations for all data blocks, resulting in large space overhead and consequently sacrificing addressing efficiency. In this paper, we propose CD-RAIS, a compromise between static striping and dynamic striping. CD-RAIS groups requests that are from different SSD drives and places their corresponding unnecessarily consecutive logical blocks in one stripe. It mitigates the write-hole problem, meanwhile remains the same addressing efficiency as static striping. To enable dynamic data striping fit SSDs, CD-RAIS performs lazy data invalidation and consolidates updates to parity blocks. CD-RAIS greatly alleviates the write request increase due to parity block update in RAIS. Our experimental results show that, for random write dominated workloads, CD-RAIS achieves 65% response time improvement and 31% longer SSD lifespan over traditional RAIS schemes. Yimo Du, Youtao Zhang, Nong Xiao 0001, Fang Liu 0002 |
CLUSTER | 2 |
| 2014 | SLC-enabled Wear Leveling for MLC PCM Considering Process VariationabstractPhase change memory is becoming one of the most promising candidates to replace DRAM as main memory in deep silicon regime. Multi-level cell (MLC) PCM outperforms single level cell (SLC) in terms of capacity while suffering from a weaker cell endurance. Wear leveling strategies are proposed to enhance the endurance but encounters more challenges with the aggravating process variation. Due to endurance variations, balanced write traffic cannot fully exploit the PCM endurance since the weak parts will be worn out sooner than others. In this work, considering process variation, we propose an SLC-enabled wear leveling scheme through dynamic and adaptive mode transformation from MLC to SLC. Instead of redistributing write operations, the proposed scheme dynamically transforms weak and write-dense parts into SLC mode for endurance benefits. The experimental results show that the proposed scheme can improve the endurance by 215% with 4% storage overhead while maintaining the capacity advantage of MLC, compared with the most related work. Mengying Zhao, Lei Jiang 0001, Youtao Zhang, Chun Jason Xue |
DAC | 3 |
| 2014 | Mitigating Write Disturbance in Super-Dense Phase Change MemoriesabstractConstructing a highly scalable and dense main memory subsystem with large access bandwidth has become a major challenge for modern computing systems. Traditional memory technologies, like DRAM and NAND Flash, suffer from either poor scalability or limited access bandwidth. Recent studies have identified emerging Phase Change Memory (PCM) as one of the most promising low power main memory technology candidates, because of its short read latency and good scalability. However, PCM still faces serious write disturbance problem below 20nm technology. Write disturbance leads to more cell programming errors, and thus degrades write reliability. Simple solutions, such as allocating large inter-cell space and adopting strong error correction code (ECC), either reduce memory density or incur large performance overhead. In this paper, we propose DIN, a Data encoding based Insulation technique, to mitigate write disturbance in highly dense PCMs. DIN improves memory density by eliminating inter-cell thermal band along a word line. The non-negligible disturbance errors, are then minimized by disturbance-aware data encoding, based on how PCM cells are programmed at device level. Our experimental results show that DIN gains write disturbance resistance in high density PCM chips while achieving comparable performance for a wide range of applications. Lei Jiang 0001, Youtao Zhang, Jun Yang 0002 |
DSN | 2 |
| 2014 | R-Dedup: Content Aware Redundancy Management for SSD-Based RAID SystemsabstractWhile high density SSDs are increasingly adopted in enterprise computing environment, it remains a challenge to meet the high performance and reliability demands of server applications as well as the demands for longer system lifetime and high space utilization in such environment. Existing schemes often address these issues separately. In particular, deduplication schemes improve write performance and SSD lifetime while SSDbased RAID designs improve reliability and read performance. Naively integration of deduplication and RAID results in suboptimal designs. In this paper, we propose R-Dedup, a content-aware redundancy management scheme for SSD-based RAID storage. By combining deduplication with replication, R-Dedup evaluates system performance, reliability, endurance and space utilization, and dynamically manages replicas to achieve better trade off. Our experimental results show that R-Dedup achieves 18% and 20% improvements on read and write performance, respectively, and extends SSD lifetime by 20% with no reliability compromise. Yimo Du, Youtao Zhang, Nong Xiao 0001 |
ICPP | 2 |
| 2014 | A low power and reliable charge pump design for Phase Change MemoriesabstractThe emerging Phase Change Memory (PCM) technology exhibits excellent scalability and density potentials. At the same time, they require high current and high voltages to switch cell states. Their working voltages are provided by CMOS-compatible on-chip charge pumps (CPs). Unfortunately, CPs and particularly those for RESET, have a large parasitic power (a dominant component in total power loss) during operations, which significantly degrades their energy efficiency. In addition, CPs seriously suffer from the Time-Dependent Dielectric Breakdown (TDDB) problem due to their boosted operation voltage. To maintain a reasonable lifetime of CPs, existing solutions actively switch them on per-operation basis, resulting in large performance degradation. In this paper, we address the above issues through two designs - Reset_Sch (RESET scheduling) and CP_Sch (CP scheduling). Reset_Sch schedules when to perform a RESET for different cells upon writing a PCM line. It significantly reduces the power loss, and peak working power of RESET CP. CP_Sch incorporates a fast READ CP design to provide fast charge-up time for reads and minimize performance penalty. Our experimental results show that on average, 70%of power loss for RESET CP can be reduced; and performance loss can be reduced from 16% to 2% while achieving a 16% improvement in reliability. Lei Jiang 0001, Bo Zhao 0007, Jun Yang 0002, Youtao Zhang |
ISCA | 4 |
| 2014 | Combining QoS prediction and customer satisfaction estimation to solve cloud service trustworthiness evaluation problems
Shuai Ding 0001, Shanlin Yang, Youtao Zhang, Changyong Liang, Chenyi Xia |
Knowl. Based Syst. | 3 |
| 2014 | Errata to "Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore Processor"abstractIn the above-named articlt that appeared in ibid., vol. 62, no. 11, pp. 2252-2265, 2013, a production error occurred which resulted in the misalignment of Fig. 13, Fig. 14, Fig. 15, Fig. 16, Fig. 17, and Fig. 18 with their captions, starting from Fig. 13 to Fig. 18. As a result, a correct Fig. 13 is missing, and Fig. 18 repeats Fig. 19. We regret that this has happened. The correct figures with their corresponding captions are shown here. Bo Zhao 0007, Yu Du 0002, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Computers | 4 |
| 2014 | Throughput Enhancement for Phase Change MemoriesabstractPhase Change Memory (PCM) has emerged as a promising candidate for future memories. PCM has high cell density, zero cell leakage, and high stability in deep sub-micron technologies. Although PCM has limited endurance, recent endeavors have shown that its lifetime can be improved by orders of magnitude. However, a major hurdle for PCM is the long write latency and high write power. For this reason, PCM cannot deliver satisfactory memory bandwidth for high-end computing environment such as multi-processing and server systems. In this paper, we develop a non-blocking PCM bank design such that subsequent reads or writes can be carried in parallel with an on-going write. This is effective in removing long blocking time due to serial operations. Moreover, we propose novel memory request scheduling algorithms to exploit intra-bank parallelism brought by our non-blocking hardware. Our non-blocking hardware with scheduling enhancement improves PCM memory throughput by 51% on average. Finally, we propose a fine-grained power budgeting scheme to achieve more throughput improvement under power budgets. Experiments show that our scheduler enhanced with power budgeting scheme can achieve a throughput improvement of 118% on average. Bo Zhao 0007, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Computers | 4 |
| 2013 | Low cost power failure protection for MLC NAND flash storage systems with PRAM/DRAM hybrid bufferabstractIn the latest PRAM/DRAM hybrid MLC NAND flash storage systems (NFSS), DRAM is used to temporarily store file system data for system response time reduction. To ensure data integrity, super-capacitors are deployed to supply the backup power for moving the data from DRAM to NAND flash during power failures. However, the capacitance degradation of super-capacitor severely impairs system robustness. In this work, we proposed a low cost power failure protection scheme to reduce the energy consumption of power failure protection and increase the robustness of the NFSS with PRAM/DRAM hybrid buffer. Our scheme enables the adoption of the more reliable regular capacitor to replace the super capacitor as the backup power. The experimental result shows that our scheme can substantially reduce the capacitance budget of power failure protection circuitry by 75.1% with very marginal performance and energy overheads. Jie Guo 0002, Jun Yang 0002, Youtao Zhang, Yiran Chen 0001 |
DATE | 3 |
| 2013 | The design of sustainable wireless sensor network node using solar energy and phase change memoryabstractSustainability of wireless sensor network (WSN) is crucial to its economy and efficiency. While previous works have focused on solving the energy source limitation through solar energy harvesting, we reveal in this paper that sensor node's lifespan could also be limited by memory wear-out and battery cycle life. We propose a sustainable sensor node design that takes all three limiting factors into consideration. Our design uses Phase Change Memory (PCM) to solve Flash memory's endurance issue. By leveraging PCM's adjustable write width, we propose a low-cost, fine-grained load tuning technique that allows the sensor node to match current MPP of solar panel and reduces the number of discharge/charge cycles on battery. Our modeling and experiments show that our sustainable sensor node design can achieve on average 5.1 years of node lifetime, more than 2× over the baseline. Youtao Zhang, Jun Yang 0002 |
DATE | 2 |
| 2013 | WoM-SET: Low power proactive-SET-based PCM write using WoM codeabstractThe emerging Phase Change Memory (PCM), while having many advantages, suffers from slow write operations. This is mainly due to its asymmetric write characteristic, i.e., for two types of write operations of PCM, SET is much slower than RESET. Recent study has shown that proactively setting dirty memory lines to all ‘1’s can enable RESET-only writes when these lines are written back from the cache, which helps to reduce the effective write latency. Unfortunately, it results in higher write power demand. In this paper, we propose WoM-SET, a low power proactive-SET-based write strategy. By exploiting the WoM (write-once memory) code, we greatly reduce the number of RESETs per write and hence the write power demand. By applying our design only to write-intensive pages, we restrict the extra space requirement in WoM-SET. Our experiments show that WoM-SET achieves 40% RESET bit reduction, 40% write power reduction, and 12% energy-delay-product improvement over the PreSET scheme. Xianwei Zhang 0001, Le Jang, Youtao Zhang, Chuanjun Zhang, Jun Yang 0002 |
ISLPED | 3 |
| 2013 | Compiler directed write-mode selection for high performance low power volatile PCMabstractMicro-Controller Units (MCUs) are widely adopted ubiquitous computing devices. Due to tight cost and energy constraints, MCUs often integrate very limited internal RAM memory on top of Flash storage, which exposes Flash to heavy write traffic and results in short system lifetime. Architecting emerging Phase Change Memory (PCM) is a promising approach for MCUs due to its fast read speed and long write endurance. Qing'an Li, Lei Jiang 0001, Youtao Zhang, Yanxiang He, Chun Jason Xue |
LCTES | 3 |
| 2013 | A speculative arbiter design to enable high-frequency many-VC router in NoCsabstractHigh-performance network-on-chip routers usually prefer a large number of Virtual Channels (VC) for high throughput. However, the growth in VC count results in increased arbitration complexity and reduced router clock frequency. In this paper, we propose a novel high-frequency many-input arbiter design for many-VC routers. It is based on the speculation on short and thus fast arbitrations in case of high VC occupancy. We further enhance it to reduce arbitration latency and promote speculation opportunity. Simulation results show that using the proposed arbiter, a 16-VC router achieves almost the same performance as an ideal design, showing improvements of around 48% on zero-load latency and 100% on network throughput over a naive 16-VC design. Bo Zhao 0007, Youtao Zhang, Jun Yang 0002 |
NOCS | 2 |
| 2013 | Hardware-Assisted Cooperative Integration of Wear-Leveling and Salvaging for Phase Change MemoryabstractPhase Change Memory (PCM) has recently emerged as a promising memory technology. However, PCM’s limited write endurance restricts its immediate use as a replacement for DRAM. To extend the lifetime of PCM chips, wear-leveling and salvaging techniques have been proposed. Wear-leveling balances write operations across different PCM regions while salvaging extends the duty cycle and provides graceful degradation for a nonnegligible number of failures. Current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a noncontiguous PCM space generated from salvaging complicates wear-leveling and incurs large overhead. In this article, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS integrates wear-leveling and salvaging and copes well with modern OSes. Our experimental results show that LLS achieves 31% longer lifetime than the state-of-the-art. It has negligible hardware cost and performance overhead. Lei Jiang 0001, Yu Du 0002, Bo Zhao 0007, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
ACM Trans. Archit. Code Optim. | 4 |
| 2013 | Process Variation-Aware Nonuniform Cache Management in a 3D Die-Stacked Multicore ProcessorabstractProcess variations in integrated circuits have significant impact on their performance, leakage, and stability. This is particularly evident in large, regular, and dense structures such as DRAMs. DRAMs are built using minimized transistors with presumably uniform speed in an organized array structure. Process variation can introduce latency disparity among different memory arrays. With the proliferation of 3D stacking technology, DRAMs become a favorable choice for stacking on top of a multicore processor as a last level cache for large capacity, high bandwidth, and low power. Hence, variations in bank speed create a unique problem of nonuniform cache accesses in 3D space. In this paper, we investigate cache management techniques for tolerating process variation in a 3D DRAM stacked onto a multicore processor. We modeled the process variation in a four-layer DRAM memory, including cell transistor, capacitor trench, and peripheral circuit, to characterize the latency and retention time variations among different banks. As a result, the notion of fast and slow banks from the core's standpoint is no longer associated with their physical distances with the banks. They are determined by the different bank latencies due to process variation. We develop cache migration schemes that utilize fast banks while limiting the cost due to migration. Our experiments show that there is a great performance benefit in exploiting fast memory banks through migration. On average, a variation-aware management can improve the performance of a workload over the baseline (where one of the slowest bank speed is assumed for all banks) by 16.5 percent. We are also only 0.8 percent away in performance from an ideal memory where no process variation is present. Bo Zhao 0007, Yu Du 0002, Jun Yang 0002, Youtao Zhang |
IEEE Trans. Computers | 4 |
| 2013 | Common-source-line array: An area efficient memory architecture for bipolar nonvolatile devicesabstractTraditional array organization of bipolar nonvolatile memories such as STT-MRAM and memristor utilizes two bitlines for cell manipulations. With technology scaling, such bitline pair will soon become the bottleneck for further density improvement. In this article we propose a novel common-source-line array architecture, which uses a shared source-line along the row, leaving only one bitline per column. We elaborate the array design to ensure reliability, and demonstrate its effectiveness on STT-MRAM and memristor memory arrays. Our study results show that with comparable latency and energy, the proposed common-source-line array can save 34% and 33% area for Memristor-RAM and STT-MRAM respectively, compared with corresponding dual-bitline arrays. Bo Zhao 0007, Jun Yang 0002, Youtao Zhang, Yiran Chen 0001, Hai Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2012 | Constructing large and fast multi-level cell STT-MRAM based cache for embedded processorsabstractMLC STT-MRAM (Multi-level Cell Spin-Transfer Torque Magnetic RAM), an emerging non-volatile memory technology, has become a promising candidate to construct L2 caches for high-end embedded processors. However, the long write latency limits the effectiveness of MLC STT-MRAM based L2 caches. In this paper, we address this limitation with two novel designs: Line Pairing (LP) and Line Swapping (LS). LP forms fast cachelines by re-organizing MLC soft bits which are faster to write. LS dynamically stores frequently written data into these fast cachelines. Our experimental results show that LP and LS improve system performance by 15% and reduce energy consumption by 21%. Lei Jiang 0001, Bo Zhao 0007, Youtao Zhang, Jun Yang 0002 |
DAC | 3 |
| 2012 | Architecting a common-source-line array for bipolar non-volatile memory devicesabstractTraditional array organization of bipolar non-volatile memories such as STT-MRAM and memristor utilizes two bitlines for cell manipulations. With technology scaling, such bitline pair will soon become the bottleneck of density improvement. In this paper we propose a novel common-source-line array architecture, which uses a shared source-line along the row, leaving only one bitline per column. We also elaborate our design flow towards a reliable common-source-line array design, and demonstrate its effectiveness on STT-MRAM and memristor memory arrays. Our study results show that with comparable latency and energy, the proposed common-source-line array can save 33% and 21.8% area for Memristor-RAM and STT-MRAM respectively, comparing with corresponding traditional dual-bitline array designs. Bo Zhao 0007, Jun Yang 0002, Youtao Zhang, Yiran Chen 0001, Hai Li 0001 |
DATE | 3 |
| 2012 | Improving write operations in MLC phase change memoryabstractPhase change memory (PCM) recently has emerged as a promising technology to meet the fast growing demand for large capacity memory in modern computer systems. In particular, multi-level cell (MLC) PCM that stores multiple bits in a single cell, offers high density with low per-byte fabrication cost. However, despite many advantages, such as good scalability and low leakage, PCM suffers from exceptionally slow write operations, which makes it challenging to be integrated in the memory hierarchy. In this paper, we propose architectural innovations to improve the access time of MLC PCM. Due to cell process variation, composition fluctuation and the relatively small differences among resistance levels, MLC PCM typically employs an iterative write scheme to achieve precise control, which suffers from large write access latency. To address this issue, we propose write truncation (WT) to reduce the number of write iterations with the assistance of an extra error correction code (ECC). We also propose form switch (FS) to reduce the storage overhead of the ECC. By storing highly compressible lines in SLC form, FS improves read latency as well. Our experimental results show that WT and FS improve the effective write/read latency by 57%/28% respectively, and achieve 26% performance improvement over the state of the art. Lei Jiang 0001, Bo Zhao 0007, Youtao Zhang, Jun Yang 0002, Bruce R. Childers |
HPCA | 3 |
| 2012 | ER: elastic RESET for low power and long endurance MLC based phase change memoryabstractPhase Change Memory (PCM) has recently emerged as a promising nonvolatile memory technology. To effectively increase memory capacity and reduce per bit fabrication cost, multi-level cell (MLC) PCM stores more than one bit per cell by differentiating multiple intermediate resistance levels. However, MLC PCM suffers from significantly shortened endurance due to its large RESET current that initiates the cell state. In this paper, we propose elastic RESET (ER) to construct non-2n-state MLC PCM, e.g., 3-state MLC PCM instead of 4-state one for 2-bit MLC. We then adopt data compression and propose fraction encoding to store compressed data using non-2n-state MLC. By reducing RESET energy, ER significantly reduces write power and prolongs PCM lifetime. On average, we observed 17% RESET power reduction and 32x endurance improvement for 2-bit MLC. Lei Jiang 0001, Youtao Zhang, Jun Yang 0002 |
ISLPED | 2 |
| 2012 | FPB: Fine-grained Power Budgeting to Improve Write Throughput of Multi-level Cell Phase Change MemoryabstractAs a promising nonvolatile memory technology, Phase Change Memory (PCM) has many advantages over traditional DRAM. Multi-level Cell PCM (MLC) has the benefit of increased memory capacity with low fabrication cost. Due to high per-cell write power and long write latency, MLC PCM requires careful power management to ensure write reliability. Unfortunately, existing power management schemes applied to MLC PCM result in low write throughput and large performance degradation. In this paper, we propose Fine-grained write Power Budgeting (FPB) for MLC PCM. We first identify two major problems for MLC write operations: (i) managing write power without consideration of the iterative write process used by MLC is overly pessimistic, (ii) a heavily written (hot) chip may block the memory from accepting further writes due to chip power restrictions, although most chips may be available. To address these problems, we propose two FPB schemes. First, FPB-IPM observes a global power budget and regulates power across write iterations according to the step-down power demand of each iteration. Second, FPB-GCP integrates a global charge pump on a DIMM to boost power for hot PCM chips while staying within the global power budget. Our experimental results show that these techniques achieve significant improvement on write throughput and system performance. Our schemes also interact positively with PCM effective read latency reduction techniques, such as write cancellation, write pausing and write truncation. Lei Jiang 0001, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
MICRO | 2 |
| 2011 | MRAC: A Memristor-based Reconfigurable Framework for Adaptive Cache ReplacementabstractMemristor, a long postulated yet missing circuit element, has recently emerged as a promising device in non-volatile memory technologies. However, beyond its use as memory cell, it is challenging to integrate memristor in modern architectures for general purpose computation. In this paper we propose a non-conventional use of memristor and demonstrate its applicability to enhancing cache replacement policy. We design a memristor-based saturation counter which can track cache access history at low cost. Based on our counter design, we develop a cache replacement framework that is both reconfigurable and adaptive (MRAC). Our evaluation demonstrates MRAC's reconfigurability and adaptivity, which result in better performance and more robust performance improvement. Bo Zhao 0007, Youtao Zhang, Jun Yang 0002, Yiran Chen 0001 |
PACT | 3 |
| 2011 | Proactive recovery for BTI in high-k SRAM cellsabstractRecent studies of BTI behavior in SRAM cells showed that for high-K metal gate stack technology, PBTI induced Vthshift in NMOS is as significant as NBTI induced Vthshift in PMOS. Previous techniques of mitigating NBTI in SRAM focus mainly on PMOS and thus lack the ability to mitigate PBTI of NMOS transistors. In this paper, we propose a novel design to recover 4 internal gates within a SRAM cell simultaneously to mitigate both NBTI and PBTI effects. In the evaluated L2 cache, our technique effectively slows down the cell failure probability increase, and achieves 4.64/2.86× (best/worst case) lifetime improvement over normal design. Youtao Zhang, Jun Yang 0002 |
DATE | 2 |
| 2011 | LLS: Cooperative integration of wear-leveling and salvaging for PCM main memoryabstractPhase change memory (PCM) has emerged as a promising technology for main memory due to many advantages, such as better scalability, non-volatility and fast read access. However, PCM's limited write endurance restricts its immediate use as a replacement for DRAM. Recent studies have revealed that a PCM chip which integrates millions to billions of bit cells has non-negligible variations in write endurance. Wear leveling techniques have been proposed to balance write operations to different PCM regions. To further prolong the lifetime of a PCM device after the failure of weak cell, techniques have been proposed to remap failed lines to spares and to salvage a PCM device that has a large number of failed lines or pages with graceful degradation. However, current wear-leveling and salvaging schemes have not been designed and integrated to work cooperatively to achieve the best PCM device lifetime. In particular, a non-contiguous PCM space generated from salvaging complicates wear leveling and incurs large overhead. In this paper, we propose LLS, a Line-Level mapping and Salvaging design. By allocating a dynamic portion of total space in a PCM device as backup space, and mapping failed lines to backup PCM, LLS constructs a contiguous PCM space and masks lower-level failures from the OS and applications. LLS seamlessly integrates wear leveling and salvaging and copes well with modern OSs, including ones that support multiple page sizes. Our experimental results show that LLS achieves 24% longer lifetime than a state-of-the-art technique. It has negligible hardware cost and performance overhead. Lei Jiang 0001, Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
DSN | 3 |
| 2011 | A composite and scalable cache coherence protocol for large scale CMPsabstractThe number of on-chip cores of modern chip multiprocessors (CMPs) is growing fast with technology scaling. However, it remains a big challenge to efficiently support cache coherence for large scale CMPs. The conventional snoopy and directory coherence protocols cannot be smoothly scaled to many-core or thousand-core processors. Snoopy protocols introduce large power overhead due to enormous amount of cache tag probing triggered by broadcast. Directory protocols introduce performance penalty due to indirection, and large storage overhead due to storing directories. This paper addresses the efficiency problem when supporting cache coherency for large-scale CMPs. By leveraging emerging optical on-chip interconnect (OP-I) technology to provide high bandwidth density, low propagation delay and natural support for multicast/broadcast in a hierarchical network organization, we propose a composite cache coherence (C 3) protocol that benefits from direct cache-to-cache accesses as in snoopy protocol and small amount of cache probing as in directory protocol. Targeting at quickly completing coherence transactions, C 3 organizes accesses in a three-tier hierarchy by combining a mix of designs including local broadcast prediction, filtering, and a coarse-grained directory. Compared to directory-based protocol [18], our evaluations on a thousand-core CMP show that C 3 improves performance by 21%, reduces network latency of coherence messages by 41 % and saves network energy consumption by 5.5 % on average for PARSEC applications. Yi Xu 0010, Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
ICS | 3 |
| 2011 | Enhancing phase change memory lifetime through fine-grained current regulation and voltage upscaling
Lei Jiang 0001, Youtao Zhang, Jun Yang 0002 |
ISLPED | 2 |
| 2011 | Analyzing the impact of useless write-backs on the endurance and energy consumption of PCM main memoryabstractPhase Change Memory (PCM) is an emerging technology that has been recently considered as a cost-effective and energy-efficient alternative to traditional DRAM main memory. Due to the high energy consumption of writes and limited number of write cycles, reducing the number of writes to PCM can result in considerable energy savings and endurance improvement. In this paper, we introduce the concept of useless write-backs, which occur when a dirty cache line that belongs to a dead memory region is evicted from the cache (a dead region is a memory location that is not used again by a program). Since the evicted data is not used again, the write-back can be safely avoided to improve endurance and energy consumption. This paper presents a limit study on the improvement that passing information to the memory system about useless writebacks has on the endurance and energy consumption of systems based on PCM main memory. We developed algorithms to measure the number of useless write-backs to PCM for three different types of memory regions and we present an energy model to determine the maximum energy savings that could potentially be achieved through such a scheme. Our results show that avoiding useless write-backs can save up to 19.8% of energy and improve endurance by up to 26.2%. Santiago Bock, Bruce R. Childers, Rami G. Melhem, Daniel Mossé, Youtao Zhang |
ISPASS | 5 |
| 2011 | A co-commitment based secure data collection scheme for tiered wireless sensor networks
Youtao Zhang, Zhiguang Qin, Taieb Znati |
J. Syst. Archit. | 2 |
| 2010 | Proactive NBTI mitigation for busy functional units in out-of-order microprocessorsabstractDue to fast technology scaling, negative bias temperature instability (NBTI) has become a major reliability concern in designing modern integrated circuits. In this paper, we present a simple and proactive NBTI recovery scheme targeting at critical and busy functional units with storage cells in modern microprocessors. Existing schemes have limitations when recovering these functional units. By exploiting the idle time of busy functional units at per-buffer-entry level, our scheme achieves on average 5.57x MTTF (Mean Time To Failure) improvement at the cost of <1% IPC degradation and <1% area overhead. Youtao Zhang, Jun Yang 0002 |
DATE | 2 |
| 2010 | Simple virtual channel allocation for high throughput and high frequency on-chip routersabstractTechnology scaling has led to the integration of many cores into a single chip. As a result, on-chip interconnection networks start to play a more and more important role in determining the performance and power of the entire chip. Packet-switched network-on-chip (NoC) has provided a scalable solution to the communications for tiled multi-core processors. However the virtual-channel (VC) buffers in the NoC consume significant dynamic and leakage power of the system. To improve the energy efficiency of the router design, it is advantageous to use small buffer sizes while still maintaining throughput of the network. This paper proposes two new virtual channel allocation (VA) mechanisms, termed Fixed VC Assignment with Dynamic VC Allocation (FVADA) and Adjustable VC Assignment with Dynamic VC Allocation (AVADA). The idea is that VCs are assigned based on the designated output port of a packet to reduce the Head-of-Line (HoL) blocking. Also, the number of VCs allocated for each output port can be adjusted dynamically. Unlike previous buffer-pool based designs, we only use a small number of VCs to keep the arbitration latency low. Simulation results show that FVADA and AVADA can improve the network throughput by 41% on average, compared to a baseline design with the same buffer size. AVADA can still outperform the baseline even when our buffer size is halved. Moreover, we are able to achieve comparable or better throughput than a previous dynamic VC allocator while reducing its critical path delay by 60%. Our results prove that the proposed VA mechanisms are suitable for low-power, high-throughput, and high-frequency on-chip network designs. Yi Xu 0010, Bo Zhao 0007, Youtao Zhang, Jun Yang 0002 |
HPCA | 3 |
| 2010 | Fine-grained QoS scheduling for PCM-based main memory systemsabstractWith wide adoption of chip multiprocessors (CMPs) in modern computers, there is an increasing demand for large capacity main memory systems. The emerging PCM (Phase Change Memory) technology has unique power and scalability advantages and is regarded as a promising candidate among new memory technologies. When scheduling a mix of applications of different priority levels, it is often important to provide tunable QoS (Quality-of-Service) for the applications with high priority. However due to the slow PCM cell access, and the destructive interferences among concurrent applications, existing memory scheduling schemes lack the flexibility to tune QoS in a wide range, in particular to the level close or equal to that of standalone execution. In this paper we propose a novel QoS scheduling scheme that utilizes request preemption and row buffer partition that enable QoS tuning at a fine-granularity. That is, they can tune the request queuing time and the PCM bank service time for the high priority requests. Our experimental results show that the proposed scheme achieves 1.7× ~10× QoS tuning range while introducing negligible area and energy overheads. Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
IPDPS | 3 |
| 2010 | An efficient code update scheme for DSP applications in mobile embedded systemsabstractDSP processors usually provide dedicated address generation units (AGUs) to assist address computation. By carefully allocating variables in the memory, DSP compilers take advantage of AGUs and generate efficient code with compact size and improved performance. However, DSP applications running on mobile embedded systems often need to be updated after their initial releases. Studies showed that small changes at the source code level may significantly change the variable layout in the memory and thus the binary code, which causes large energy overheads to mobile embedded systems that patch through wireless or satellite communication, and often pecuniary burden to the users. Youtao Zhang |
LCTES | 2 |
| 2010 | An authentication scheme for locating compromised sensor nodes in WSNs
Youtao Zhang, Jun Yang 0002, Linzhang Wang, Lingling Jin |
J. Netw. Comput. Appl. | 1 |
| 2010 | A low-cost memory remapping scheme for address bus protection
Jun Yang 0002, Lan Gao 0002, Youtao Zhang, Marek Chrobak, Hsien-Hsin S. Lee |
J. Parallel Distributed Comput. | 3 |
| 2010 | Performance-aware thermal management via task schedulingabstractHigh on-chip temperature impairs the processor's reliability and reduces its lifetime. Hardware-level dynamic thermal management (DTM) techniques can effectively constrain the chip temperature, but degrades the performance. We propose an OS-level technique that performs thermal-aware job scheduling to reduce DTMs. The algorithm is based on the observation that hot and cool jobs executed in a different order can make a difference in resulting temperature. Real-system implementation in Linux shows that our scheduler can remove 10.5% to 73.6% of the hardware DTMs in a medium thermal environment. The CPU throughput is improved by up to 7.6% (4.1%, on average) in a severe thermal environment. Xiuyi Zhou, Jun Yang 0002, Marek Chrobak, Youtao Zhang |
ACM Trans. Archit. Code Optim. | 4 |
| 2010 | Thermal-Aware Task Scheduling for 3D Multicore ProcessorsabstractA rising horizon in chip fabrication is the 3D integration technology. It stacks two or more dies vertically with a dense, high-speed interface to increase the device density and reduce the delay of interconnects significantly across the dies. However, a major challenge in 3D technology is the increased power density, which gives rise to the concern of heat dissipation within the processor. High temperatures trigger voltage and frequency throttlings in hardware, which degrade the chip performance. Moreover, high temperatures impair the processor's reliability and reduce its lifetime. To alleviate this problem, we propose in this paper an OS-level scheduling algorithm that performs thermal-aware task scheduling on a 3D chip. Our algorithm leverages the inherent thermal variations within and across different tasks, and schedules them to keep the chip temperature low. We observed that vertically adjacent dies have strong thermal correlations and the scheduler should consider them jointly. Compared with other intuitive algorithms such as a Random and a Round-Robin algorithm, our proposed algorithm brings lower peak temperature and average temperature on-chip. Moreover, it can remove, on average, 46 percent of thermal emergency time and result in 5.11 percent (4.78 percent) performance improvement over the base case on thermally homogeneous (heterogeneous) floorplans. Xiuyi Zhou, Jun Yang 0002, Yi Xu 0010, Youtao Zhang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2010 | The design and evaluation of interleaved authentication for filtering false reports in multipath routing WSNs
Youtao Zhang, Jun Yang 0002, Hai Trong Vu, Yizhi Wu |
Wirel. Networks | 1 |
| 2009 | Frequent value compression in packet-based NoC architecturesabstractThe proliferation of chip multiprocessors (CMPs) has led to the integration of large on-chip caches. For scalability reasons, a large on-chip cache is often divided into smaller banks that are interconnected through packet-based network-on-chip (NoC). With increasing number of cores and cache banks integrated on a single die, the on-chip network introduces significant communication latency and power consumption. In this paper, we propose a novel scheme that exploitsfrequentvaluecompression to optimize the power and performance of NoC. Our experimental results show that the proposed scheme reduces the router power by up to 16.7%, with CPI reduction as much as 23.5% in our setting. Comparing to the recent zero pattern compression scheme, thefrequentvaluescheme saves up to 11.0% more router power and has up to 14.5% more CPI reduction. Hardware design of the FV table and its overhead are also presented. Bo Zhao 0007, Yu Du 0002, Yi Xu 0010, Youtao Zhang, Jun Yang 0002, Li Zhao 0002 |
ASP-DAC | 5 |
| 2009 | MCP: An Energy-Efficient Code Distribution Protocol for Multi-Application WSNs
Youtao Zhang, Bruce R. Childers |
DCOSS | 2 |
| 2009 | A low-radix and low-diameter 3D interconnection network designabstractInterconnection plays an important role in performance and power of CMP designs using deep sub-micron technology. The network-on-chip (NoCs) has been proposed as a scalable and high-bandwidth fabric for interconnect design. The advent of the 3D technology has provided further opportunity to reduce on-chip communication delay. However, the design of the 3D NoC topologies has important distinctions from 2D NoCs or off-chip interconnection networks. First, current 3D stacking technology allows only vertical inter-layer links. Hence, there cannot be direct connections between arbitrary nodes in different layers — the vertical connection topology are essentially fixed. Second, the 3D NoC is highly constrained by the complexity and power of routers and links. Hence, low-radix routers are preferred over high-radix routers for lower power and better heat dissipation. This implies long network latency due to high hop counts in network paths. In this paper, we design a low-diameter 3D network using low-radix routers. Our topology leverages long wires to connect remote intra-layer nodes. We take advantage of the start-of-the-art one-hop vertical communication design and utilize lateral long wires to shorten network paths. Effectively, we implement a small-to-medium sized clique network in different layers of a 3D chip. The resulting topology generates a diameter of 3-hop only network, using routers of the same radix as 3D mesh routers. The proposed network shows up to 29% of network latency reduction, up to 10% throughput improvement, and up to 24% energy reduction, when compared to a 3D mesh network. Yi Xu 0010, Yu Du 0002, Bo Zhao 0007, Xiuyi Zhou, Youtao Zhang, Jun Yang 0002 |
HPCA | 5 |
| 2009 | Energy reduction for STT-RAM using early write terminationabstractThe emerging Spin Torque Transfer memory (STT-RAM) is a promising candidate for future on-chip caches due to STT-RAM's high density, low leakage, long endurance and high access speed. However, one of the major challenges of STT-RAM is its high write current, which is disadvantageous when used as an on-chip cache since the dynamic power generated is too high. Bo Zhao 0007, Jun Yang 0002, Youtao Zhang |
ICCAD | 4 |
| 2009 | A durable and energy efficient main memory using phase change memory technologyabstractUsing nonvolatile memories in memory hierarchy has been investigated to reduce its energy consumption because nonvolatile memories consume zero leakage power in memory cells. One of the difficulties is, however, that the endurance of most nonvolatile memory technologies is much shorter than the conventional SRAM and DRAM technology. This has limited its usage to only the low levels of a memory hierarchy, e.g., disks, that is far from the CPU. Bo Zhao 0007, Jun Yang 0002, Youtao Zhang |
ISCA | 4 |
| 2009 | Variation-tolerant non-uniform 3D cache management in die stacked multicore processorabstractProcess variations in integrated circuits have significant impact on their performance, leakage and stability. This is particularly evident in large, regular and dense structures such as DRAMs. DRAMs are built using minimized transistors with presumably uniform speed in an organized array structure. Process variation can introduce latency disparity among different memory arrays. With the proliferation of 3D stacking technology, DRAMs become a favorable choice for stacking on top of a multicore processor as a last level cache for large capacity, high bandwidth, and low power. Hence, variations in bank speed creates a unique problem of non-uniform cache accesses in 3D space. Bo Zhao 0007, Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
MICRO | 3 |
| 2009 | SDC: Secure Data Collection for Time Based Queries in Tiered Wireless Sensor NetworksabstractTiered wireless sensor networks (WSNs) have many advantages over traditional WSNs. However they are vulnerable to security attacks, especially those targeting at the storage nodes that bu er and process the data readings from sensors. In this paper, we propose a Secure Data Collection protocol - SDC to support time-based queries in tiered WSNs. With small overhead introduced to data communication, SDC protects both data confidentiality and data integrity. In particular it employs data co-commitment such that it can detect and evaluate the message dropping attacks in the network. Zhiguang Qin, Youtao Zhang, Taieb Znati |
RTCSA | 3 |
| 2009 | Towards update-conscious compilation for energy-efficient code dissemination in WSNsabstractPostdeployment code dissemination in wireless sensor networks (WSN) is challenging, as the code has to be transmitted via energy-expensive wireless communication. In this article, we propose novel update-conscious compilation (UCC) techniques to achieve energy efficiency. By integrating the compilation decisions in generating the old binary, an update-conscious compiler strives to match the old decisions, which improves the binary code similarity, reduces the amount of transmitted data to remote sensors, and thus, consumes less energy. In this article, we develop update-conscious register allocation and data layout algorithms. Our experimental results show great improvements over the traditional, update-oblivious approaches. Youtao Zhang, Jun Yang 0002, Jiang Zheng 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2008 | Adaptive Buffer Management for Efficient Code Dissemination in Multi-Application Wireless Sensor NetworksabstractFuture wireless sensor networks (WSNs) are projected to run multiple applications in the same network infrastructure. While such multi-application WSNs (MA-WSNs) are economically more efficient and adapt better to the changing environments than traditional single-application WSNs, they usually require frequent code redistribution on wireless sensors, making it critical to design energy efficient post-deployment code dissemination protocols in MA-WSNs. Different applications in MA-WSNs often share some common code segments. Therefore when there is a need to disseminate a new application from the sink node, it is possible to disseminate its shared code segments from peer sensors instead of disseminating everything from the sink node. While dissemination protocols have been proposed to handle code of each single type, it is challenging to achieve energy efficiency when the code contains both types and needs simultaneous dissemination. In this paper we utilize an adaptive buffer management approach to achieve efficient code dissemination in MA-WSNs. Our experimental results show that adaptive buffer management can reduce the completion time and the message overhead up to 10% and 20% respectively. Yu Du 0002, Youtao Zhang, Bruce R. Childers, Jun Yang 0002 |
EUC (1) | 3 |
| 2008 | Thermal Management for 3D Processors via Task SchedulingabstractA rising horizon in chip fabrication is the 3D integration technology. It stacks two or more dies vertically with a dense, high-speed interface to increase the device density and reduce the delay of interconnects across the dies. However, a major challenge in 3D technology is the increased power density which brings the concern of heat dissipation within the processor. High temperatures trigger voltage and frequency throttlings in hardware which degrade the chip performance. Moreover, high temperatures impair the processorpsilas reliability and reduce its lifetime. To alleviate this problem, we propose in this paper an OS-level scheduling algorithm that performs thermal-aware task scheduling on a 3D chip. Our algorithm leverages the inherent thermal variations within and across different tasks, and schedules them to keep the chip temperature low. We observed that vertically adjacent dies have strong thermal correlations, and the scheduler should consider them jointly. Our proposed algorithm can remove on average 54% of hardware DTMs and result in 7.2% performance improvement over the base case. Xiuyi Zhou, Yi Xu 0010, Yu Du 0002, Youtao Zhang, Jun Yang 0002 |
ICPP | 4 |
| 2008 | Towards energy-efficient code dissemination in wireless sensor networksabstractPost-deployment code dissemination has become an important design issue for many applications in wireless sensor networks (WSNs). While several dissemination protocols have been proposed, challenges still exist due to the high energy consumption of transmitting wireless signals. In this paper, we present update-conscious compilation (UCC) techniques for energy-efficient code dissemination in WSNs. An update-conscious compiler, when compiling the modified code, includes the compilation decisions that were made when generating the old binary. In most cases, matching the previous decisions improves the binary code similarity, reduces the amount of data to be transmitted to remote sensors, and thus, consumes less energy. In this paper, we focus on the development of update-conscious register allocation (UCC-RA) algorithms. Our experimental results show that UCC-RA can achieve great improvements over the traditional, update-oblivious approaches. Youtao Zhang, Jun Yang 0002 |
IPDPS | 1 |
| 2008 | Dynamic Thermal Management through Task SchedulingabstractThe evolution of microprocessors has been hindered by their increasing power consumption and the heat generation speed on-die. High temperature impairs the processor's reliability and reduces its lifetime. While hardware level dynamic thermal management (DTM) techniques, such as voltage and frequency scaling, can effectively lower the chip temperature when it surpasses the thermal threshold, they inevitably come at the cost of performance degradation. We propose an OS level technique that performs thermal- aware job scheduling to reduce the number of thermal trespasses. Our scheduler reduces the amount of hardware DTMs and achieves higher performance while keeping the temperature low. Our methods leverage the natural discrepancies in thermal behavior among different workloads, and schedule them to keep the chip temperature below a given budget. We develop a heuristic algorithm based on the observation that there is a difference in the resulting temperature when a hot and a cool job are executed in a different order. To evaluate our scheduling algorithms, we developed a lightweight runtime temperature monitor to enable informed scheduling decisions. We have implemented our scheduling algorithm and the entire temperature monitoring framework in the Linux kernel. Our proposed scheduler can remove 10.5-73.6% of the hardware DTMs in various combinations of workloads in a medium thermal environment. As a result, the CPU throughput was improved by up to 7.6% (4.1% on average) even under a severe thermal environment. Jun Yang 0002, Xiuyi Zhou, Marek Chrobak, Youtao Zhang, Lingling Jin |
ISPASS | 4 |
| 2007 | UCC: update-conscious compilation for energy efficiency in wireless sensor networksabstractWireless sensor networks (WSN), composed of a large number of low-cost, battery-powered sensors, have recently emerged as promising computing platforms for many non-traditional applications. The preloaded code on remote sensors often needs to be updated after deployment in order for the WSN to adapt to the changing demands from the users. Post-deployment code dissemination is challenging as the data are transmitted via battery-powered wireless communication. Recent studies show that the energy for sending a single bit is about the same as executing 1000 instructions in aWSN. Therefore it is important to achieve energy efficiency in code dissemination. Youtao Zhang, Jun Yang 0002, Jiang Zheng 0002 |
PLDI | 2 |
| 2007 | The design and evaluation of path matching schemes on compressed control flow traces
Yongjing Lin, Youtao Zhang, Rajiv Gupta 0001 |
J. Syst. Softw. | 2 |
| 2006 | A low-cost memory remapping scheme for address bus protectionabstractThe address sequence on the processor-memory bus can reveal abundant information about the control flow of a program. This can lead to critical information leakage such as encryption keys or proprietary algorithms. Addresses can be observed by attaching a hardware device on the bus that passively monitors the bus transaction. Such side-channel attacks should be given rising attention especially in a distributed computing environment, where remote servers running sensitive programs are not within the physical control of the client.Two previously proposed hardware techniques tackled this problem through randomizing address patterns on the bus. One proposal permutes a set of contiguous memory blocks under certain conditions, while the other approach randomly swaps two blocks when necessary. In this paper, we present an anatomy of these attempts and show that they impose great pressure on both the memory and the disk. This leaves them less scalable in high-performance systems where the bandwidth of the bus and memory are critical resources. We propose a lightweight solution to alleviating the pressure without compromising the security strength. The results show that our technique can reduce the memory traffic by a factor of 10 compared with the prior scheme, while keeping almost the same page fault rate as a baseline system with no security protection. Lan Gao 0002, Jun Yang 0002, Marek Chrobak, Youtao Zhang, San Nguyen, Hsien-Hsin S. Lee |
PACT | 4 |
| 2006 | Efficient Group KeyManagement with Tamper-resistant ISA ExtensionsabstractWe present a tamper-resistant architectural enhancement for secure group key management in group communication applications. Using specially designed four cryptographic instructions, we show that the hardware assisted design can greatly reduce the management overhead to the order of O(1) in terms of rekey messages, storage cost, and the encryption computation cost Youtao Zhang, Jun Yang 0002, Lan Gao 0002 |
ASAP | 1 |
| 2006 | Locating Compromised Sensor Nodes Through Incremental Hashing Authentication
Youtao Zhang, Jun Yang 0002, Lingling Jin |
DCOSS | 1 |
| 2006 | InfoShield: a security architecture for protecting information usage in memoryabstractCyber theft is a serious threat to Internet security. It is one of the major security concerns by both network service providers and Internet users. Though sensitive information can be encrypted when stored in non-volatile memory such as hard disks, for many e-commerce and network applications, sensitive information is often stored as plaintext in main memory. Documented and reported exploits facilitate an adversary stealing sensitive information from an application's memory. These exploits include illegitimate memory scan, information theft oriented buffer overflow, invalid pointer manipulation, integer overflow, password stealing Trojans and so forth. Today's computing system and its hardware cannot address these exploits effectively in a coherent way. This paper presents a unified and lightweight solution, called InfoShield that can strengthen application protection against theft of sensitive information such as passwords, encryption keys, and other private data with a minimal performance impact. Unlike prior whole memory encryption and information flow based efforts, InfoShield protects the usage of information. InfoShield ensures that sensitive data are used only as defined by application semantics, preventing misuse of information. Comparing with prior art, InfoShield handles a broader range of information theft scenarios in a unified framework with less overhead. Evaluation using popular network client-server applications shows that InfoShield is sound for practical use and incurs little performance loss because InfoShield only protects absolute, critical sensitive information. Based on the profiling results, only 0.3% of memory accesses and 0.2% of executed codes are affected by InfoShield. Joshua B. Fryman, Guofei Gu, Hsien-Hsin S. Lee, Youtao Zhang, Jun Yang 0002 |
HPCA | 5 |
| 2006 | Reduce Register Files Leakage Through Discharging CellsabstractWe propose a low-leakage register file cell design based on the observation that the physical registers in a superscalar processor have very short life cycles. When a register is dead, we discharge its cells to '0' to greatly reduce the leakage current from the read bitlines to the ground. Our design has no impact to critical register read access path. Projected to future 45 nm technology, our design yields additional 38% and 47% leakage power savings on top of the existing low-leakage cell designs for 64-bit and 32-bit datapath, respectively. Taking into the account of dynamic energy savings due to the elimination of write '0' operations, our design saves nearly 20% of total energy. Lingling Jin, Wei Wu 0024, Jun Yang 0002, Chuanjun Zhang, Youtao Zhang |
ICCD | 5 |
| 2006 | The interleaved authentication for filtering false reports in multipath routing based sensor networksabstractIn this paper, we consider filtering false reports in braided multipath routing sensor networks. While multi-path routing provides better resilience to various faults in sensor networks, it has two problems regarding the authentication design. One is that, due to the large number of partially overlapped routing paths between the source and sink nodes, the authentication overhead could be very high if these paths are authenticated individually; the other is that false reports may escape the authentication check through the newly identified node association attack. In this paper we propose enhancements to solve both problems such that secure and efficient authentication can be achieved in multi-path routing. The proposed scheme is (t + 1)-resilient, i.e. it is secure with up to t compromised nodes. The upper bound number of hops that a false report may be forwarded in the network is O(t/sup 2/). Youtao Zhang, Jun Yang 0002, Hai Trong Vu |
IPDPS | 1 |
| 2006 | Dynamic Authentication-Key Re-assignment for Reliable Report DeliveryabstractSensor networks deployed in hostile environments are subject to various types of attacks. While multipath routing and en-route authentication schemes have been proposed to defend packet dropping and injection attacks respectively, it is challenging to defend both at the same time. The paper addresses this problem through an annulus based authentication-key reassignment scheme in multipath routing with en-route authentication. Our experimental results show that the proposed scheme achieves better trade-offs in true report delivery, false report filtering and energy consumption Youtao Zhang, Jun Yang 0002 |
MASS | 2 |
| 2006 | Compressing heap data for improved memory performanceabstractAbstract We introduce a class of transformations that modify the representation of dynamic data structures used in programs with the objective of compressing their sizes. Based upon a profiling study of data value characteristics, we have developed the common‐prefix and narrow‐data transformations that respectively compress a 32 bit address pointer and a 32 bit integer field into 15 bit entities. A pair of fields that have been compressed by the above compression transformations are packed together into a single 32 bit word. The above transformations are designed to apply to data structures that are partially compressible, that is, they compress portions of data structures to which transformations apply and provide a mechanism to handle the data that is not compressible. The accesses to compressed data are efficiently implemented by designing data compression extensions (DCX) to the processor's instruction set. We have observed average reductions in heap allocated storage of 25% and average reductions in execution time and power consumption of 30%. If DCX support is not provided the reductions in execution times fall from 30% to 18%. Copyright © 2006 John Wiley & Sons, Ltd. Youtao Zhang, Rajiv Gupta 0001 |
Softw. Pract. Exp. | 1 |
| 2005 | Performance Comparison of Path Matching Algorithms over Compressed Control Flow TracesabstractA control flow trace captures the complete sequence of dynamically executed basic blocks and function calls. It is usually stored in compressed form due to its large size. Matching an intraprocedural path in a control flow trace faces path interruption and path context problems and therefore requires the extension of traditional pattern matching algorithms. In this paper we evaluate different path matching schemes including those matching in the compressed data directly and those matching after the decompression. We design simple indices for the compressed data and show that they can greatly improve the performance. Our experimental results show that these schemes are useful and can be adapted to environments with different hardware settings and path matching requests. Yongjing Lin, Youtao Zhang |
DCC | 2 |
| 2005 | SENSS: Security Enhancement to Symmetric Shared Memory MultiprocessorsabstractWith the increasing concern of the security on high performance multiprocessor enterprise servers, more and more effort is being invested into defending against various kinds of attacks. This paper proposes a security enhancement model called SENSS, that allows programs to run securely on a symmetric shared memory multiprocessor (SMP) environment. In SENSS, a program, including both code and data, is stored in the shared memory in encrypted form but is decrypted once it is fetched into any of the processors. In contrast to the traditional uniprocessor XOM model (Lie et al., 2000), the main challenge in developing SENSS lies in the necessity for guarding the clear text communication between processors in a multiprocessor environment. In this paper we propose an inexpensive solution that can effectively protect the shared bus communication. The proposed schemes include both encryption and authentication for bus transactions. We develop a scheme that utilizes the cipher block chaining mode of the advanced encryption standard (CBC-AES) to achieve ultra low latency for the shared bus encryption and decryption. In addition, CBC-AES can generate integrity checking code for the bus communication over time, achieving bus authentication. Further, we develop techniques to ensure the cryptographic computation throughput meets the high bandwidth of gigabyte buses. We performed full system simulation using Simics to measure the overhead of the security features on a SMP system with a snooping write invalidate cache coherence protocol. Overall, only a slight performance degradation of 2.03% on average was observed when the security is provided at the highest level. Youtao Zhang, Lan Gao 0002, Jun Yang 0002, Xiangyu Zhang 0001, Rajiv Gupta 0001 |
HPCA | 1 |
| 2005 | A low energy cache design for multimedia applications exploiting set access locality
Jun Yang 0002, Jia Yu 0008, Youtao Zhang |
J. Syst. Archit. | 3 |
| 2005 | Improving Memory Encryption Performance in Secure ProcessorsabstractDue to the widespread software piracy and virus attacks, significant efforts have been made to improve security for computer systems. For stand-alone computers, a key observation is that, other than the processor, any component is vulnerable to security attacks. Recently, an execution only memory (XOM) architecture has been proposed to support copy and tamper resistant software. In this design, the program and data are stored in an encrypted format outside the CPU boundary. The decryption is carried out after they are fetched from memory and before they are used by the CPU. As a result, the lengthened critical path causes a serious performance degradation. We present an innovative technique in which the cryptography computation is shifted off from the memory access critical path. We propose using a different encryption scheme, namely, "pseudo-one-time pad" encryption, to produce the instructions and data ciphertext. With some additional on-chip storage, cryptography computations are carried in parallel with memory accesses, minimizing the performance penalty. We performed experiments to study the trade-off between storage size and performance penalty. Our technique reduces the performance overhead from 20.79 percent to 1.28 percent on average for reasonably sized (64 KB) on-chip storage. Jun Yang 0002, Lan Gao 0002, Youtao Zhang |
IEEE Trans. Computers | 3 |
| 2005 | Cost and precision tradeoffs of dynamic data slicing algorithmsabstractDynamic slicing algorithms are used to narrow the attention of the user or an algorithm to a relevant subset of executed program statements. Although dynamic slicing was first introduced to aid in user level debugging, increasingly applications aimed at improving software quality, reliability, security, and performance are finding opportunities to make automated use of dynamic slicing. In this paper we present the design and evaluation of three precise dynamic data slicing algorithms called the full preprocessing (FP), no preprocessing (NP) and limited preprocessing (LP) algorithms. The algorithms differ in the relative timing of constructing the dynamic data dependence graph and its traversal for computing requested dynamic data slices. Our experiments show that the LP algorithm is a fast and practical precise data slicing algorithm. In fact we show that while precise data slices can be orders of magnitude smaller than imprecise dynamic data slices, for small number of data slicing requests, the LP algorithm is faster than an imprecise dynamic data slicing algorithm proposed by Agrawal and Horgan. Xiangyu Zhang 0001, Rajiv Gupta 0001, Youtao Zhang |
ACM Trans. Program. Lang. Syst. | 3 |
| 2004 | Scalable Duplication Strategy with Bounded Availability of Processors
Youtao Zhang, Yongjing Lin, Yaochun Huang |
ICPADS | 2 |
| 2004 | Efficient Forward Computation of Dynamic Slices Using Reduced Ordered Binary Decision Diagrams
Xiangyu Zhang 0001, Rajiv Gupta 0001, Youtao Zhang |
ICSE | 3 |
| 2003 | Enabling Partial Cache Line Prefetching Through Data CompressionabstractHardware prefetching is a simple and effective technique for hiding cache miss latency and thus improving the overall performance. However, it comes with addition of prefetch buffers and causes significant memory traffic increase. We propose a new prefetching scheme which improves performance without increasing memory traffic or requiring prefetch buffers. We observe that a significant percentage of dynamically appearing values exhibit characteristics that enable their compression using a very simple compression scheme. The bandwidth freed by transferring values from lower levels in memory hierarchy to upper levels in compressed form is used to prefetch additional compressible values. These prefetched values are held in vacant space created in the data cache by storing values in compressed form. Thus, in comparison to other prefetching schemes, our scheme does not introduce prefetch buffers or increase the memory traffic. In comparison to a baseline cache that does not support prefetching, on average, our cache design reduces the memory traffic by 10%, reduces the data cache miss rate by 14%, and speeds up program execution by 7%. Youtao Zhang, Rajiv Gupta 0001 |
ICPP | 1 |
| 2003 | Procedural Level Address Offset Assignment of DSP Applications with LoopsabstractAutomatic optimization of address offset assignment for DSP applications, which reduces the number of address arithmetic instructions to meet the tight memory size restrictions and performance requirements, received a lot of attention in recent years. However, most of current research focuses at the basic block level and does not distinguish different program structures, especially loops. Moreover, the effectiveness of modify register (MR) is not fully exploited since it is used only in the post optimization step. A novel address offset assignment approach is proposed at the procedural level. The MR is effectively used in the address assignment for loop structures. By taking advantage of MR, variables accessed in sequence within a loop are assigned to memory words of equal distances. Both static and dynamic addressing instruction counts are greatly reduced. For DSPSTONE benchmarks and on average, 9.9%, 17.1% and 21.8% improvements are achieved over address offset assignment [R. Leupers et al., (1996)] together with MR optimization when there is 1, 2 and 4 address registers respectively. Youtao Zhang, Jun Yang 0002 |
ICPP | 1 |
| 2003 | Precise Dynamic Slicing AlgorithmsabstractDynamic slicing algorithms can greatly reduce the debugging effort by focusing the attention of the user on a relevant subset of program statements. In this paper we present the design and evaluation of three precise dynamic slicing algorithms called the full preprocessing (FP), no preprocessing (NP) and limited preprocessing (LP) algorithms. The algorithms differ in the relative timing of constructing the dynamic data dependence graph and its traversal for computing requested dynamic slices. Our experiments show that the LP algorithm is a fast and practical precise slicing algorithm. In fact we show that while precise slices can be orders of magnitude smaller than imprecise dynamic slices, for small number of slicing requests, the LP algorithm is faster than an imprecise dynamic slicing algorithm proposed by Agrawal and Horgan. Xiangyu Zhang 0001, Rajiv Gupta 0001, Youtao Zhang |
ICSE | 3 |
| 2003 | Lightweight set buffer: low power data cache for multimedia applicationabstractA new architectural technique to reduce power dissipation in data caches is proposed. In multimedia applications, a major portion of data cache accesses hit in the same cache set continuously before going to a different set. This feature allows us to remove unnecessary driving power in data arrays as long as the same cache set is accessed incessantly. Power saving is achieved through buffering and accessing the cache set instead of the main data array. The proposed technique does not incur performance degradation and accomplishes up to 57% of power reduction for data caches. Jun Yang 0002, Youtao Zhang |
ISLPED | 2 |
| 2003 | Low cost instruction cache designs for tag comparison eliminationabstractTag comparison elimination (TCE) is an effective approach to reduce I-cache energy. Current research focuses on finding good tradeoffs between hardware cost and percentage of comparisons that can be removed. For this purpose, two low cost innovations are proposed in this paper. We design a small dedicated TCE table whose size is flexible both horizontally (entry size) and vertically (number of entries). The design also minimizes interactions with the I-cache. For a 64-way 16K cache, the new design reduces the tag comparisons to 4.0% with a fraction only 20% of the hardware cost of the way memoization technique [5]. The result is 40% better compared to a recent proposed low cost design [2] of comparable hardware cost. Youtao Zhang, Jun Yang 0002 |
ISLPED | 1 |
| 2003 | Fast Secure Processor for Inhibiting Software Piracy and TamperingabstractDue to the widespread software piracy and virus attacks, significant efforts have been made to improve security for computer systems. For stand-alone computers, a key observation is that other than the processor, any component is vulnerable to security attacks. Recently, an execution only memory (XOM) architecture has been proposed to support copy and tamper resistant software by D. Lie et al. (2000), D. Lie et al. (2003) and T. Gilmont et al. (1999). In this design, the program and data are stored in encrypted format outside the CPU boundary. The decryption is carried after they are fetched from memory, and before they are used by the CPU. As a result, the lengthened critical path causes a serious performance degradation. In this paper, we present an innovative technique in which the cryptography computation is shifted off from the memory access critical path. We propose to use a different encryption scheme, namely "one-time pad" encryption, to produce the instructions and data ciphertext. With some additional on-chip storage, cryptography computations are carried in parallel with memory accesses, minimizing performance penalty. We performed experiments to study the trade-off between storage size and performance penalty. Our technique improves the execution speed of the XOM architecture by 34% at maximum. Jun Yang 0002, Youtao Zhang, Lan Gao 0002 |
MICRO | 2 |
| 2002 | A Representation for Bit Section Based Analysis and Optimization
Rajiv Gupta 0001, Eduard Mehofer, Youtao Zhang |
CC | 3 |
| 2002 | Data Compression Transformations for Dynamically Allocated Data Structures
Youtao Zhang, Rajiv Gupta 0001 |
CC | 1 |
| 2002 | Path Matching in Compressed Control Flow TraceabstractDue to its large size, a whole program path (WPP) is stored in compressed form generated using SEQUITUR. The occurrence of an intraprocedural path in the WPP cannot be carried out using existing algorithms due to the path interruption and path context problems. We present an algorithm that addresses the above problems. The complexity of the algorithm is analyzed and experimental data is presented to demonstrate that our algorithm is efficient in practice. Youtao Zhang, Rajiv Gupta 0001 |
DCC | 1 |
| 2001 | Timestamped Whole Program Path Representation and its ApplicationsabstractA whole program path (WPP) is a complete control flow trace of a program's execution. Recently Larus [18] showed that although WPP is expected to be very large (100's of MBytes), it can be greatly compressed (to 10's of MBytes) and therefore saved for future analysis. While the compression algorithm proposed by Larus is highly effective, the compression is accompanied with a loss in the ease with which subsets of information can be accessed. In particular, path traces pertaining to a particular function cannot generally be obtained without examining the entire compressed WPP representation. To solve this problem we advocate the application of compaction techniques aimed at providing easy access to path traces on a per function basis. Youtao Zhang, Rajiv Gupta 0001 |
PLDI | 1 |
| 2000 | Frequent Value Locality and Value-Centric Data Cache DesignabstractBy studying the behavior of programs in the SPECint95 suite we observed that six out of eight programs exhibit a new kind of value locality, the frequent value locality, according to which a few values appear very frequently in memory locations and are therefore involved in a large fraction of memory accesses. In these six programs ten distinct values occupy over 50% of all memory locations and on an average account for nearly 50% of all memory accesses during program execution. This observation holds for smaller blocks of consecutive memory locations and the set of frequent values remains quite stable over the execution of the program.In the six benchmarks with frequent value locality, on an average 50% of all cache misses occur during the reading or writing of the ten most frequently accessed values. We propose a new data cache structure, the frequent value cache (FVC), which employs a value-centric approach to caching data locations for exploiting the frequent value locality phenomenon. FVC is a small direct-mapped cache which is dedicated to holding only frequently occurring values. The value-centric nature of FVC enables us to store data in a compressed form where the compression is achieved by encoding the frequent values using a few bits. Moreover this simple compression scheme preserves the random access to data values in a cache line.Our experiments demonstrate that by augmenting a direct mapped cache (DMC) with a direct mapped FVC of size no more than 3 Kbytes we can obtain reductions in miss rates ranging from 1% to 68%. In fact we observed that higher reductions in miss rates can he achieved by augmenting a DMC with a small FVC as opposed to doubling the size of DMC for the 124.m88ksim and 134.perl benchmarks. Youtao Zhang, Jun Yang 0002, Rajiv Gupta 0001 |
ASPLOS | 1 |
| 2000 | Frequent value compression in data cachesabstractSince the area occupied by cache memories on processor chips continues to grow, an increasing percentage of power is consumed by memory. We present the design and evaluation of the compression cache (CC) which is a first level cache that has been designed so that each cache line can either hold one uncompressed line or two cache lines which have been compressed to at least half their lengths. We use a novel data compression scheme based upon encoding of a small number of valves that appear frequently during memory accesses. This compression scheme preserves the ability to randomly access individual data items. We observed that the contents of 40%, 52% and 51% of the memory blocks of size 4, 8, and 16 words respectively in SPECint95 benchmarks can be compressed to at least half their sizes by encoding the top 2, 4, and 8 frequent valves respectively. Compression allows greater amounts of data to be stored leading to substantial reductions in miss rates (0-36.4%), off-chip traffic (3.9-48.1%), and energy consumed (1-27%). Traffic and energy reductions are in part derived by transferring data over external buses in compressed form. Jun Yang 0002, Youtao Zhang, Rajiv Gupta 0001 |
MICRO | 2 |