Dong Tong 0001

dblp:56/464 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0003-0987-6351ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8Software engineering, systems software and programming languages · 3
YearPublicationVenuePosition
2023 FlexPointer: Fast Address Translation Based on Range TLB and Tagged Pointers
abstract
Page-based virtual memory relies on TLBs to accelerate the address translation. Nowadays, the gap between application workloads and the capacity of TLB continues to grow, bringing many costly TLB misses and making the TLB a performance bottleneck. Previous studies seek to narrow the gap by exploiting the contiguity of physical pages. One promising solution is to group pages that are both virtually and physically contiguous into a memory range. Recording range translations can greatly increase the TLB reach, but ranges are also hard to index because they have arbitrary bounds. The processor has to compare against all the boundaries to determine which range an address falls in, which restricts the usage of memory ranges. In this article, we propose a tagged-pointer-based scheme, FlexPointer, to solve the range indexing problem. The core insight of FlexPointer is that large memory objects are rare, so we can create memory ranges based on such objects and assign each of them a unique ID. With the range ID integrated into pointers, we can index the range TLB with IDs and greatly simplify its structure. Moreover, because the ID is stored in the unused bits of a pointer and is not manipulated by the address generation, we can shift the range lookup to an earlier stage, working in parallel with the address generation. According to our trace-based simulation results, FlexPointer can reduce nearly all the L1 TLB misses, and page walks for a variety of memory-intensive workloads. Compared with a 4K-page baseline system, FlexPointer shows a 14% performance improvement on average and up to 2.8x speedup in the best case. For other workloads, FlexPointer shows no performance degradation.
Dongwei Chen, Dong Tong 0001, Jiangfang Yi, Xu Cheng 0001
ACM Trans. Archit. Code Optim.2
2022 FlexPointer: Fast Address Translation Based on Range TLB and Tagged Pointers
abstract
Page-based virtual memory suffers from costly page walks because of the gap between application workload sizes and TLB capacity. In this paper, we propose a tagged-pointer-based system, FlexPointer, to solve this problem. FlexPointer creates memory ranges from objects larger than a certain threshold by allocating contiguous physical pages for them. Virtual addresses within a range share a common translation entry, thus greatly expanding the TLB capacity. Because such objects are rare, we can assign each of them a unique ID and pass it to the hardware in pointer tags. According to our trace-based simulation results, FlexPointer can reduce nearly all the L1 TLB misses and page walks for a variety of memory-intensive workloads, providing a 14% performance improvement on average.
Dongwei Chen, Dong Tong 0001, Jiangfang Yi, Xu Cheng 0001
PACT2
2021 Prediction of Register Instance Usage and Time-sharing Register for Extended Register Reuse Scheme
Shuxin Zhou, Huandong Wang, Dong Tong 0001
ASP-DAC3
2021 MetaTableLite: An Efficient Metadata Management Scheme for Tagged-Pointer-Based Spatial Safety
abstract
A tagged-pointer-based memory spatial safety protection system utilizes the unused bits in a pointer to store the boundary information of an object. This paper proposed a hybrid metadata management scheme, MetaTableLite, for tagged-pointer-based protections. We observed that objects of a large size only take a minority part of all the objects in a program. However, recording their boundary metadata with traditional pointer tags will incur large memory overheads. Based on this observation, we introduce a small supplementary table to maintain metadata for these few large objects. For small ones, MetaTableLite represents their boundaries with a 14-bit pointer tag, well utilizing the unused 16 bits in a conventional 64-bit pointer. MetaTableLite can achieve a 6% average memory overhead without alternating the conventional pointer representation.
Dongwei Chen, Dong Tong 0001, Xu Cheng 0001
ICCD2
2017 A Staged Memory Resource Management Method for CMP systems
abstract
Memory interference is a critical impediment to system performance in CMP systems. To address this problem, we first propose a Dynamically Proportional Bandwidth Throttling policy (DPBT), which dynamically throttles back memory-intensive applications based on their memory access behavior. DPBT achieves a more balance memory bandwidth partitioning. Moreover, we improve the previous memory channel partitioning scheme by integrating it with a bank partitioning. We further integrate DPBT with the improved memory channel partitioning scheme and a memory scheduling policy to leverage the architecture advantages, and present a Stage Memory Resource Management Method (SRM). Experimental results show that DPBT improves system throughput/fairness by 13.5%/31.1%. SRM provides 27.1% better system throughput and 34.8% better system fairness.
Yangguo Liu, Junlin Lu, Dong Tong 0001, Xu Cheng 0001
ASAP3
2017 Locality-aware bank partitioning for shared DRAM MPSoCs
abstract
Memory interference is a critical impediment to system performance in MPSoCs. To address this problem, we first propose a Locality-Aware Bank Partitioning (LABP), which partitions memory banks according to applications' memory access behavior. The key idea is to separate memory intensive applications with high row-buffer locality from the other applications. Moreover, we integrate LABP with a bandwidth allocation scheme to leverage the architecture advantages, and present a comprehensive approach named Integrated Bandwidth and Bank Partitioning (IBBP) to further alleviate the interference. Experimental results show LABP improves system throughput/fairness by 10.8%/26.4%. IBBP provides 14.1% better system throughput and 34.2% better system fairness. Our methods are better than other recent work, including bandwidth throttling, DBP and DBP-TCM.
Yangguo Liu, Junlin Lu, Dong Tong 0001, Xu Cheng 0001
ASP-DAC3
2016 MFAP: Fair Allocation between fully backlogged and non-fully backlogged applications
abstract
In this paper, we consider the problem of ensuring fairness in systems serving a mixture of fully backlogged applications, which continuously demand resources, and non-fully backlogged applications. We introduce a fairness metric, called interference fairness, the basic idea underlying which is that the interference caused by application A for another application B should be equal to that caused by B for A. To effectively and efficiently guarantee this fairness metric, we propose Mutual Fair Allocation Policy (MFAP), a simple and powerful resource sharing policy, and show how it guarantees interference fairness between any pair of applications. We also show that MFAP, unlike other viable policies, satisfies several highly desirable properties, including some from game theory, as well as common sense intuitions. As a use case, we implemented MFAP on a disk scheduling framework. The experimental results based on synthetic and real workloads show how our implementation achieved interference fairness and improved non-fully backlogged applications performance.
Yan Sui, Dong Tong 0001, Xianhua Liu 0001, Xu Cheng 0001
ICCD3
2014 Improving system throughput and fairness simultaneously in shared memory CMP systems via Dynamic Bank Partitioning
abstract
Applications running concurrently in CMP systems interfere with each other at DRAM memory, leading to poor system performance and fairness. Memory access scheduling reorders memory requests to improve system throughput and fairness. However, it cannot resolve the interference issue effectively. To reduce interference, memory partitioning divides memory resource among threads. Memory channel partitioning maps the data of threads that are likely to severely interfere with each other to different channels. However, it allocates memory resource unfairly and physically exacerbates memory contention of intensive threads, thus ultimately resulting in the increased slowdown of these threads and high system unfairness. Bank partitioning divides memory banks among cores and eliminates interference. However, previous equal bank partitioning restricts the number of banks available to individual thread and reduces bank level parallelism. In this paper, we first propose a Dynamic Bank Partitioning (DBP), which partitions memory banks according to threads' requirements for bank amounts. DBP compensates for the reduced bank level parallelism caused by equal bank partitioning. The key principle is to profile threads' memory characteristics at run-time and estimate their demands for bank amount, then use the estimation to direct our bank partitioning. Second, we observe that bank partitioning and memory scheduling are orthogonal in the sense; both methods can be illuminated when they are applied together. Therefore, we present a comprehensive approach which integrates Dynamic Bank Partitioning and Thread Cluster Memory scheduling (DBP-TCM, TCM is one of the best memory scheduling) to further improve system performance. Experimental results show that the proposed DBP improves system performance by 4.3% and improves system fairness by 16% over equal bank partitioning. Compared to TCM, DBP-TCM improves system throughput by 6.2% and fairness by 16.7%. When compared with MCP, DBP-TCM provides 5.3% better system throughput and 37% better system fairness. We conclude that our methods are effective in improving both system throughput and fairness.
Mingli Xie, Dong Tong 0001, Kan Huang, Xu Cheng 0001
HPCA2
2014 A General Low-Cost Indirect Branch Prediction Using Target Address Pointers
Zichao Xie, Dong Tong 0001, Mingkai Huang
J. Comput. Sci. Technol.2
2013 An adaptive filtering mechanism for energy efficient data prefetching
abstract
As data prefetching is used in embedded processors, it is crucial to reduce the wasted energy for improving the energy efficiency. In this paper, we propose an adaptive prefetch filtering (APF) mechanism to reduce the wasted bandwidth and energy as well as the cache pollution caused by useless prefetches. APF records the prefetch-victim address pairs of issued prefetches and collects information about which address in each pair is first accessed by the processor to guide the filtering of new generated useless prefetches. Meanwhile, filtered prefetches are recorded for building the feedback mechanism to avoid filtering useful prefetches. Experimental results demonstrate that APF reduces useless prefetches by an average of 53.81% with a mere 5.28% reduction of useful prefetches, thus reducing the memory access bandwidth consumption by 59.92% and the L2 cache energy by 6.19%. APF also improves the performance of several programs by reducing the cache pollution incurred by useless prefetches, thus gaining an average performance improvement of 2.12%.
Xianglei Dang, Xiaoyin Wang, Dong Tong 0001, Zichao Xie, Lingda Li
ASP-DAC3
2013 VFCC: A verification framework of cache coherence using parallel simulation
abstract
A cache coherence protocol is a vital component of a multiprocessor to maintain the data consistency. In this paper, we proposed VFCC, which is a simulation framework to validate a cache-coherence protocol implementation of a commercial 64-bit superscalar multiprocessor. It exploits multiple-level parallelism to accelerate validation without overheads among threads. Our experimental results demonstrate VFCC has a 5.0× speedup than a traditional simulator on a conventional 16-core host machine.
Qiaoli Xiong, Jiangfang Yi, Tianbao Song, Zichao Xie, Dong Tong 0001
ASP-DAC5
2013 An energy-efficient branch prediction technique via global-history noise reduction
abstract
Accurate branch prediction can improve processor performance, while reducing energy waste. Though some existing branch predictors have been proved effective, they usually require large amount of storage or complicate the processor front-end. This paper proposes a novel branch prediction technique called History Artificially Selected (HAS) prediction. It is a hardware technique that bases on the existing branch predictors to detect history noises and avoid noise interferences when predicting branches. It separates the original branch predictor into sub-predictors, each of which performs differently in branch history updating. With the help of some history stacks, one sub-predictor saves and restores the branch history at the entrance and the exit of loops and program subroutines where history noise usually exists. Through using a tournament mechanism, HAS prediction selectively uses the modified branch history to eliminate the history noise interferences and retain those useful history correlations at the same time. Our experimental results show that for three representative branch predictors, gshare, perceptron, and TAGE, it reduces the MPKI by 1.49, 2.85, and 1.10 respectively, resulting in 4.55%, 10.16%, and 4.45% performance improvement. It also reduces energy consumption by 4.02%, 7.78%, and 3.91%, respectively.
Zichao Xie, Dong Tong 0001, Xu Cheng 0001
ISLPED2
2013 Page policy control with memory partitioning for DRAM performance and power efficiency
abstract
DRAM performance and power efficiency considerations are becoming increasingly important. Bank partitioning partitions memory banks among cores and eliminates inter-thread interference, thus improving system performance of shared memory CMP systems. However, it doesn't take into account DRAM power consumption. We propose an application-aware page policy, which exploits potential benefits of page policy to optimize DRAM performance or minimize power consumption. The key idea is to dynamically assign page policy to applications according to their memory characteristics. As an improvement, we propose a power-aware bank partitioning to balance DRAM performance and power consumption. Experimental results show that our proposal increases system performance and significantly improves DRAM power efficiency.
Mingli Xie, Dong Tong 0001, Yi Feng 0003, Kan Huang, Xu Cheng 0001
ISLPED2
2013 SPIRE: improving dynamic binary translation through SPC-indexed indirect branch redirecting
abstract
Dynamic binary translation system must perform an address translation for every execution of indirect branch instructions. The procedure to convert Source binary Program Counter (SPC) address to Translated Program Counter (TPC) address always takes more than 10 instructions, becoming a major source of performance overhead. This paper proposes a novel mechanism called SPc-Indexed REdirecting (SPIRE), which can significantly reduce the indirect branch handling overhead. SPIRE doesn't rely on hash lookup and address mapping table to perform address translation. It reuses the source binary code space to build a SPC-indexed redirecting table. This table can be indexed directly by SPC address without hashing. With SPIRE, the indirect branch can jump to the originally SPC address without address translation. The trampoline residing in the SPC address will redirect the control flow to related code cache. Only 2-6 instructions are needed to handle an indirect branch execution. As part of the source binary would be overwritten, a shadow page mechanism is explored to keep transparency of the corrupt source binary code page. Online profiling is adopted to reduce the memory overhead.
Ning Jia 0004, Jing Wang 0055, Dong Tong 0001
VEE4
2012 Optimal bypass monitor for high performance last-level caches
abstract
In the last-level cache, large amounts of blocks have reuse distances greater than the available cache capacity. Cache performance and efficiency can be improved if some subset of these distant reuse blocks can reside in the cache longer. The bypass technique is an effective and attractive solution that prevents the insertion of harmful blocks.
Lingda Li, Dong Tong 0001, Zichao Xie, Junlin Lu, Xu Cheng 0001
PACT2
2012 S/DC: A storage and energy efficient data prefetcher
abstract
Energy efficiency is becoming a major constraint in processor designs. Every component of the processor should be reconsidered to reduce wasted energy and area. Prefetching is an important technique for tolerating memory latency. Prefetcher designs have important impact on the energy efficiency of the memory hierarchy. Stride prefetchers require little storage, but cannot handle irregular access patterns. Delta correlation (DC) prefetchers can handle complicated access patterns, but waste storage because of storing multiple miss addresses for a stride pattern. Moreover, DC prefetchers waste the bandwidth and energy of the memory hierarchy because they cannot identify whether an address has been prefetched and generate a large number of redundant prefetches. In this paper, we propose a storage and energy efficient data prefetcher called stride/DC (S/DC) to combine the advantages of stride and DC prefetchers. S/DC uses a pattern prediction table (PPT) which stores two recent miss addresses in each entry to capture stride patterns. PPT avoids recording multiple miss addresses for a stride pattern, and thus improves the storage efficiency. When handling stride patterns, each PPT entry maintains a counter for obtaining the last prefetched address to avoid generating redundant prefetches. When handling other patterns, S/DC compares the new predicted address with earlier generated addresses in the prefetch queue and filters the redundant ones. In addition, to expand the filtering scope, S/DC uses a prefetch filter to store addresses evicted from the prefetch queue. In this way, S/DC reduces the bandwidth requirements and energy consumption of prefetching. Experimental results demonstrate that S/DC achieves comparable performance with only 24% of the storage and reduces 11.46% of the L2 cache energy, as compared to the CZone/DC prefetcher.
Xianglei Dang, Xiaoyin Wang, Dong Tong 0001, Junlin Lu, Jiangfang Yi
DATE3
2012 Energy-efficient branch prediction with Compiler-guided History Stack
abstract
Branch prediction is critical in exploring instruction level parallelism for modern processors. Previous aggressive branch predictors generally require significant amount of hardware storage and complexity to pursue high prediction accuracy. This paper proposes the Compiler-guided History Stack (CHS), an energy-efficient compiler-microarchitecture cooperative technique for branch prediction. The key idea is to track very-long-distance branch correlation using a low-cost compiler-guided history stack. It relies on the compiler to identify branch correlation based on two program substructures: loop and procedure, and feed the information to the predictor by inserting guiding instructions. At runtime, the processor dynamically saves and restores the global history using a low-cost history stack structure according to the compiler-guided information. The modification on the global history enables the predictor to track very-long-distance branch correlation and thus improves the prediction accuracy. We show that CHS can be combined with most of existing branch predictors and it is especially effective with small and simple predictors. Our evaluations show that the CHS technique can reduce the average branch mispredictions by 28.7% over gshare predictor, resulting in average performance improvement of 10.4%. Furthermore, it can also improve those aggressive perceptron, OGEHL and TAGE predictors.
Mingxing Tan, Xianhua Liu 0001, Zichao Xie, Dong Tong 0001, Xu Cheng 0001
DATE4
2012 Improving inclusive cache performance with two-level eviction priority
abstract
Inclusive cache hierarchies are widely adopted in modern processors, since they can simplify the implementation of cache coherence. However, it sacrifices some performance to guarantee inclusion. Many recent intelligent management policies are proposed to improve the last-level cache (LLC) performance by evicting blocks with poor locality earlier. Unfortunately, they are inapplicable in inclusive LLCs. In this paper, we propose Two-level Eviction Priority (TEP) policy. Besides the eviction priority provided by the baseline replacement policy, TEP appends an additional high level of eviction priority to LLC blocks, which is decided at the insertion time and cannot be changed during their lifetime in the LLC. When blocks with high eviction priority are not in inner caches anymore, they get evicted from the LLC preferentially. Thus, the LLC can retain more useful blocks to improve performance. TEP can cooperate well with various baseline replacement policies. Our evaluation shows that TEP with NRU can improve the performance of inclusive LLCs significantly while requiring negligible extra storage. It also outperforms other recent proposals including QBS, DIP, and DRRIP.
Lingda Li, Dong Tong 0001, Zichao Xie, Junlin Lu, Xu Cheng 0001
ICCD2
2012 SOLE: Speculative one-cycle load execution with scalability, high-performance and energy-efficiency
abstract
Conventional superscalar processors usually contain large CAM-based LSQ (load/store queue) with poor scalability and high energy consumption. Recently proposals only focus on improving the LSQ scalability to increase the in-flight instruction capacity, but with poor performance improvement and energy efficiency. This paper presents a novel speculative store-load forwarding mechanism, named SOLE (speculative one-cycle load execution)1. Firstly, SOLE uses address identifiers to determine the memory disambiguation, rather than the exact memory addresses as the traditional LSQ does. Since the address identifier is just simple hash from the address base and offset, the speculative store-load forwarding could be advanced earlier to reduce the load execution latency and avoid unnecessary energy consumption by filtering unnecessary accesses to the data cache. Secondly, SOLE enlarges the forwarding communication range by using SSN (store sequential number) to determine the age order between stores, which further improves the performance. Finally, the implementation of SOLE all uses set-associative structures that avoid the non-scalable problem of CAM-based LSQ. Experiments show that performance of SOLE outperforms the traditional LSQ by 13.57% in terms of performance, with only 75.2% execution energy consumption of the loads and stores.
Zhen-Hao Zhang, Dong Tong 0001, Xiaoyin Wang, Jiangfang Yi
ICCD2
2012 SWIP Prediction: Complexity-Effective Indirect-Branch Prediction Using Pointers
Zichao Xie, Dong Tong 0001, Mingkai Huang, Qinqing Shi, Xu Cheng 0001
J. Comput. Sci. Technol.2
2012 Active Store Window: Enabling Far Store-Load Forwarding with Scalability and Complexity-Efficiency
Zhen-Hao Zhang, Xiaoyin Wang, Dong Tong 0001, Jiangfang Yi, Junlin Lu
J. Comput. Sci. Technol.3
2011 TAP prediction: Reusing conditional branch predictor for indirect branches with Target Address Pointers
abstract
Indirect-branch prediction is becoming more important for modern processors as more programs are written in object-oriented languages. Previous hardware-based indirect-branch predictors generally require significant hardware storage or use aggressive algorithms which make the processor front-end more complex. In this paper, we propose a fast and cost-efficient indirect-branch prediction strategy, called Target Address Pointer (TAP) Prediction. TAP Prediction reuses the history-based branch direction predictor to detect occurrences of indirect branches, and then stores indirect-branch targets in the Branch Target Buffer (BTB). The key idea of TAP Prediction is to predict the Target Address Pointers, which generate virtual addresses to index the targets stored in the BTB, rather than to predict the indirect-branch targets directly. TAP Prediction also reuses the branch direction predictor to construct several small predictors. When fetching an indirect branch, these small predictors work in parallel to generate the target address pointer. Then TAP prediction accesses the BTB to fetch the predicted indirect-branch target using the generated virtual address. This mechanism could achieve time cost comparable to that of dedicated-storage-predictors, without requiring additional large amounts of storage. Our evaluation shows that for three representative direction predictors-Hybrid, Perceptrons, and O-GEHL-TAP schemes improve performance by 18.19%, 21.52%, and 20.59%, respectively, over the baseline processor with the most commonly-used BTB prediction. Compared with previous hardware-based indirect-branch predictors, the TAP-Perceptrons scheme achieves performance improvement equivalent to that provided by a 48KB TTC predictor, and it also outperforms the VPC predictor by 14.02%.
Zichao Xie, Dong Tong 0001, Mingkai Huang, Xiaoyin Wang, Qinqing Shi, Xu Cheng 0001
ICCD2
2010 FPGA prototyping of an amba-based windows-compatible SoC
abstract
For the increasing market of smart phones, mobile internet devices, and ultra-mobile PCs, mainstream vendors propose two approaches: one is based on ARM SoC, and the other is based on power-efficient x86 processor. However, either approach has its own limitation. The ARM-based approach lacks application software while the x86-based approach does not support flexible SoC extension. To overcome the limitations, we propose the PKUnity86 SoC architecture, which is based on AMBA bus architecture to support fast IP integration. Furthermore, it contains a reduced AMD Geode GX2 processor and several specific designs to support Microsoft Windows and exploit the massive PC software resources.
Kan Huang, Junlin Lu, Jiufeng Pang, Yansong Zheng, Dong Tong 0001, Xu Cheng 0001
FPGA6
2010 Research Progress of UniCore CPUs and PKUnity SoCs
Xu Cheng 0001, Xiaoyin Wang, Junlin Lu, Jiangfang Yi, Dong Tong 0001, Xuetao Guan, Xianhua Liu 0001, Yi Feng 0003
J. Comput. Sci. Technol.5
2009 WHOLE: A low energy I-Cache with separate way history
abstract
Set-associative instruction caches achieve low miss rates at the expense of significant energy dissipation. Previous energy-efficient approaches usually suffer from performance degradation and redundant extension bits. In this paper, we propose a Way History Oriented Low Energy Instruction Cache (WHOLE-Cache) design for single issue and in-order execution processors. The WHOLE-Cache design not only achieves a significant portion of energy reduction by effectively reducing dynamic energy dissipation of set-associative instruction cache, but also leads to no additional cycle penalties. Tag comparison results are stored into either the Branch Target Buffer (BTB) or the Instruction Cache (I-Cache) to avoid tag checks and unnecessary way activation for subsequent accesses to visited cache lines. The extended BTB uses way history bits for branch instructions, while the I-Cache extension bits are used in case of fetching consecutive instructions resided in different cache lines. A valid flag is associated with each stored tag comparison result to indicate whether the instruction to be fetched is resided in the recorded location. A simple invalidation scheme is implemented in the cache miss replacement operation. Whenever a cache line is replaced, the pointers to it, which reside in the BTB or other I-cache lines, will be invalidated accordingly. We model the WHOLE-Cache design in Verilog. By deriving basic parameters from TSMC 65nm technology, we use Wattch simulator to evaluate the performance and energy reduction of the WHOLE-Cache in the instruction fetch stage. We use SPEC2000 and Mediabench as benchmarks. It is observed that compared with a conventional 4-way set-associative I-Cache, the energy consumption of the WHOLE-Cache is reduced by 65% without any performance penalty.
Zichao Xie, Dong Tong 0001, Xu Cheng 0001
ICCD2
2008 CASA: A New IFU Architecture for Power-Efficient Instruction Cache and TLB Designs
Han-Xin Sun, Kun-Peng Yang, Yulai Zhao 0003, Dong Tong 0001, Xu Cheng 0001
J. Comput. Sci. Technol.4
2007 Clock domain crossing fault model and coverage metric for validation of SoC design
Yi Feng 0003, Dong Tong 0001, Xu Cheng 0001
DATE3
2007 Reuse Distance Based Cache Leakage Control
Yulai Zhao 0003, Dong Tong 0001, Xu Cheng 0001
HiPC3
2007 An Energy-Efficient Instruction Scheduler Design with Two-Level Shelving and Adaptive Banking
Yulai Zhao 0003, Dong Tong 0001, Xu Cheng 0001
J. Comput. Sci. Technol.3