EDBT 2026 Demo / reviewers in the wild / expert
Xu Cheng 0001
dblp:30/828-1
· DBLP profile ↗
57ranked-venue papers
2as first author
19since 2021 · last 2024
0000-0002-5544-8852ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Oblivious Demand Paging with Ring ORAM in RISC-V Trusted Execution EnvironmentsabstractTrusted execution environments based on RISC-V architecture like Keystone remain susceptible to leaking page access patterns of applications via simple demand paging, in which a malicious Operating System (OS) deduces sensitive information from it. To address this issue, Keystone requires protecting sensitive access patterns from being revealed to the malicious OS by implementing oblivious demand paging. In this paper, we use Oblivious RAM (ORAM) techniques that obfuscate access patterns while simultaneously making demand paging oblivious for Keystone. Furthermore, we present customized optimizations to Ring ORAM, aimed at minimizing the performance overhead incurred by applications during both secure and unsecure demand paging in Keystone. These optimizations encompass strategies such as encoding the position map within the page table, utilizing a resizable tree structure and selective eviction of only the root bucket. These improvements collectively contribute to minimizing performance slowdown. We implement and evaluate our optimized Ring ORAM for oblivious demand paging, which shows the average performance slowdown of 7.1x in comparison to the simple Ring ORAM slowdown of 26.2x. Wenjing Cai, Yusha Zhang, Xu Cheng 0001 |
CSCWD | 5 |
| 2024 | Multi: Reduce Energy Overhead of Criticality-Aware Dynamic Instruction Scheduling for Energy EfficiencyabstractCriticality-aware dynamic instruction scheduling (DIS) focuses on prioritizing the execution of critical instructions, thereby significantly improving performance. However, criticality-aware DIS is energy-intensive. Energy overhead mainly comes from the design of DIS strategy and the multi-port tables used for extracting critical instruction slices. In this study, Multi is proposed, which develops an energy-efficient DIS strategy and an adaptive multi-port table coordinated with the DIS strategy to reduce the energy overhead of criticality-aware DIS. Evaluation on the SPEC CPU2017 benchmark shows that Multi outperforms the state-of-the-art work in terms of core energy consumption with a geometric mean reduction of 5.9%. Moreover, post-layout simulation results demonstrate a 61.2% reduction in energy consumption of the proposed adaptive multi-port table. It also proves the applicability of Multi in ultra-wide processors with unified or distributed schedulers, contributing to enhanced energy efficiency. Honglan Zhan, Xianhua Liu 0001, Xu Cheng 0001 |
ICCD | 6 |
| 2024 | Hyperion: A Highly Effective Page and PC Based Delta PrefetcherabstractHardware prefetching plays an important role in modern processors for hiding memory access latency. Delta prefetchers show great potential at the L1D cache level, as they can impose small storage overhead by recording deltas. Furthermore, local delta prefetchers, such as Berti, have been shown to achieve high L1D accuracy. However, there is still room for improving the L1D coverage of existing delta prefetchers. Our goal is to develop a delta prefetcher capable of achieving both high L1D coverage and accuracy. We explore delta prefetchers trained on various types of contextual information, ranging from coarse-grained to fine-grained, and analyze their L1D coverage and accuracy. Our findings indicate that training deltas based on the access histories of both PCs and memory pages for individual PCs and memory pages can lead to increased L1D coverage alongside high accuracy. Therefore, we introduce Hyperion, a highly efficient Page and PC-based delta prefetcher. In terms of the vital component of recording access histories, we implement three different structures and engage in a detailed discussion about them. Furthermore, Hyperion utilizes micro-architecture information (e.g., L1D hits or misses, PQ occupancy) and real-time L1D accuracy to dynamically adjust its issuing mechanism, further enhancing performance and L1D accuracy. Our results show that Hyperion achieves an L1D accuracy of 92.4% and an L1D coverage of 51.9%, along with an L2C coverage of 63.0% and an LLC coverage of 67.5% across a diverse range of applications, including SPEC CPU2006, SPEC CPU2017, GAP, and PARSEC, with a baseline of no prefetching. Regarding performance, Hyperion achieves a 50.1% performance gain, outperforming the state-of-the-art delta prefetcher Berti by 5.0% over baseline across all memory-intensive traces from the four benchmark suites. Wei Chen 0168, Xu Cheng 0001, Jiangfang Yi |
ACM Trans. Archit. Code Optim. | 3 |
| 2023 | MBAPIS: Multi-Level Behavior Analysis Guided Program Interval Selection for Microarchitecture StudiesabstractUnderstanding program behavior is crucial in computer architecture research, but the growing size of benchmarks makes analyzing and simulating entire programs increasingly challenging. In practice, researchers often select representative program intervals for analysis and testing. These intervals are different sections of continuous execution of a program. SimPoint is a well-known method for selecting representative intervals using hardware-independent information. However, when focusing on a specific microarchitecture study, it is desirable to select intervals that are more relevant to that study. For instance, intervals with more branch mispredictions are more appropriate for branch prediction studies. We refer to these intervals as “tailored intervals” for branch prediction studies. This paper presents a Multi-level Behavior Analysis guided Program Interval Selection (MBAPIS) for selecting tailored intervals. For a given microarchitecture study, the first level of MBAPIS uses hardware performance counters to prioritize selecting the intervals that exhibit clearer microarchitectural characteristics relevant to that study. The second level analyzes the processor performance bottlenecks to further select the intervals where the concerned microarchitecture design more strongly impacts performance. Finally, MBAPIS performs clustering analysis with the basic block information of each interval selected by the first two levels, and selects the representative intervals among them while preserving the diverse software behavior. Additionally, we present a general and extensible interval-replaying design to accurately re-execute selected intervals. The SPEC CPU2006 and CPU2017 benchmarks are used for evaluation. The results demonstrate that MBAPIS can select representative and tailored intervals for two typical microarchi-tecture studies and deliver accurate estimates of the concerned hardware events for all tailored intervals in each benchmark, with an average error rate of less than 1.5%. Moreover, the interval-replaying design effectively restores the hardware behavior of the intervals selected by MBAPIS, with an average relative error rate of annroximately 1.4%. Hongwei Cui, Honglan Zhan, Shuhao Liang, Xianhua Liu 0001, Xu Cheng 0001 |
PACT | 7 |
| 2023 | Detecting and Mitigating Cache Side Channel Threats on Intel SGXabstractIntel Software Guard Extensions (SGX) protect sensitive content of applications on the cloud platform by creating an isolated environment on an untrusted operating system. However, resent works have shown that the SGX is vulnerable to a variety of side channel attacks which could be severely damage the data confidentiality provided by SGX, such as the cache side channel attack. Unfortunately, existing defense mechanisms either provide an incomplete protection or incur too much performance costs. In this paper, we propose a defense countermeasure against cache side channel attacks for SGX by detecting abnormal each level cache use behaviors. We create auxiliary threads for each enclave thread and detect when asynchronous enclave exits (AEX) occur, which defeats the condition of L1/L2 cache side channel attacks that attacker and victim threads execute in the same physical core. We put some guard data to the cache lines and inspect access time, which detects last level cache eviction set behaviors. More importantly, we utilize optimizations to reduce the performance overhead caused by AEX detection. In comparison to existing approaches, our design is secure against any cache level side channel attacks and its performance loss increases less. Wenjing Cai, Yusha Zhang, Xu Cheng 0001 |
CSCWD | 5 |
| 2023 | A Hardware-Software Cooperative Interval-Replaying for FPGA-based Architecture EvaluationabstractOpen-source processors and FPGA provide more real and accurate results of the new microarchitecture design, but the long execution time for running large benchmarks on FPGA boards still hinders researchers. This paper proposes a hardware-software cooperative interval-replaying. It uses simula-tors to create checkpoints for arbitrary program intervals and provides an extensible and portable checkpoint loader to re-execute selected intervals. In addition, this paper extends RISC-VISA and proposes an event-based sampling design to find hot program intervals with more representative microarchitecture characteristics. By using checkpoints in hot regions, researchers can quickly verify the effectiveness of microarchitecture designs on FPGA and alleviate the speed bottleneck of FPGA. The correctness and effectiveness of the checkpoint scheme and the event-based sampling design are evaluated on FPGA. The experimental results show that the solution is effective. Hongwei Cui, Shuhao Liang, Honglan Zhan, Xianhua Liu 0001, Xu Cheng 0001 |
DATE | 8 |
| 2023 | High-Speed and Energy-Efficient Single-Port Content Addressable Memory to Achieve Dual-Port OperationabstractHigh-speed and energy-efficient multi-port content addressable memory (CAM) is very important to modern superscalar processors. In order to overcome the disadvantages of multi-port CAM and improve the performance of searching stage, a high-speed and energy-efficient single-port (SP) CAM is introduced to achieve dual-port (DP) operation. For different bit cell topologies - the traditional 9T CAM cell and 6T SRAM cell, two novel peripheral schemes - CShare and VClamp are proposed. The proposed schemes are verified using all possible corners, a wide range of temperature and detailed Monte-Carlo variation analysis. With 65-nm process and 1.2 V supply, the search delay of CShare and VClamp is 0.55 ns and 0.6 ns, respectively, a reduction of approximately 87% compared to the state-of-the-art works. In addition, compared with the recently proposed IOT BCAM, CShare and VClamp can provide 84.9% and 85.1% energy reduction in the TT corner, respectively. Experimental results in an 8 Kb CAM at 1.2 V supply and across different corners show that the energy efficiency is improved by 45.56% (CShare) and 45.64% (VClamp) on average in comparison with DP CAM. Honglan Zhan, Hongwei Cui, Xianhua Liu 0001, Xu Cheng 0001 |
DATE | 6 |
| 2023 | HUND: Enhancing Hardware Performance Counter Based Malware Detection Under System Resource Competition Using Explanation MethodabstractHardware performance counter (HPC) has been widely used in malware detection because of its low access overhead and the ability of revealing dynamic behavior during program's execution. However, HPC based malware detection (HMD) suffers from performance decline due to HPC's non- determinism caused by resource competition. Current work enables malware detection under resource competition but still leaves misclassifications. In this paper, we propose HUND, a framework for improving the detection ability of HMD models under resource competition. To this end, we first introduce an explanation module to make the program's prediction interpretable and accurate on the whole. We then design a rectification module for troubleshooting HMDMs' errors by generating modified samples and lowering the effects of false classified instances on model decision. We evaluate HUND by performing HMD models two datasets of HPC-level behaviors. The experimental results show HUND explains HMDMs with high fidelity and HUND's effectiveness in troubleshooting the errors of HMDMs. Yanfei Hu, Shuailou Li, Xu Cheng 0001, Yu Wen 0001 |
ISCC | 4 |
| 2023 | Towards Dynamic Backdoor Attacks against LiDAR Semantic Segmentation in Autonomous DrivingabstractLiDAR perception is widely deployed in high-level autonomous vehicles (AVs) to gain accurate information about the driving environment, where 3D semantic segmentation plays a critical role as it can provide more fine-grained scene understanding than other tasks. The mainstream LiDAR perception systems mainly adopt deep neural networks (DNNs) to achieve good performance. However, densely annotating LiDAR point clouds and training complex LiDAR segmentation models are time-consuming and resource-intensive. It is common for ordinary self-driving developers to outsource the data annotation or model training task to third-party platforms, which could inevitably expose a natural attack injection point. In this paper, we propose BadLiSeg, the first work to investigate backdoor attacks against LiDAR semantic segmentation. Specifically, we present a general attack strategy based on which the attacker can inject a dynamic backdoor into the victim model by constructing a trigger pattern pool and a location pool. Afterward, the attacker can choose a common physical object (e.g., drone and traffic sign) as the trigger and place it around an arbitrary location in a selected area to easily fool the backdoored LiDAR segmentation model. Our extensive experiments on five representative segmentation models and one public dataset demonstrate that BadLiSeg can not only achieve a high attack success rate but also maintain the normal segmentation performance of the backdoored model. We further show the effectiveness of BadLiSeg on some practical attack scenarios collected from a high-fidelity simulator. Yu Wen 0001, Xu Cheng 0001 |
TrustCom | 3 |
| 2023 | BadLiDet: A Simple Backdoor Attack against LiDAR Object Detection in Autonomous DrivingabstractAutonomous vehicles (AVs) widely deploy LiDAR-based 3D object detection to accurately perceive and understand the surrounding environment. The mainstream LiDAR detection systems primarily adopt deep neural networks (DNNs) to achieve satisfactory performance. However, annotating large amounts of LiDAR data and training complex LiDAR detection models are time-consuming and resource-intensive. A common practice for some individual developers and self-driving companies is to outsource data annotation or model training tasks to third-party platforms, which could expose a natural attack injection point. In this paper, we propose BadLiDet, a simple yet effective backdoor attack against LiDAR object detection in autonomous driving. Specifically, we present a model-agnostic attack strategy that enables an attacker to inject a shape-independent backdoor into the victim model by poisoning its training data. Afterward, the attacker can choose an ordinary object in different shapes as the trigger and place it around a specific location to easily deceive the LiDAR detection system. Our extensive experiments on five representative models and a public dataset demonstrate that BadLiDet can achieve a high attack success rate while preserving the utility of victim models. We further show the effectiveness of BadLiDet on the continuous attack scenario collected from a high-fidelity simulator. Moreover, the end-to-end simulation evaluation on an open-source self-driving platform shows that BadLiDet can cause a 100% vehicle collision rate. Yu Wen 0001, Huiying Wang, Xu Cheng 0001 |
TrustCom | 4 |
| 2023 | Secure Speculation via Speculative Secret Flow Tracking
Hongwei Cui, Xu Cheng 0001 |
J. Comput. Sci. Technol. | 3 |
| 2023 | FlexPointer: Fast Address Translation Based on Range TLB and Tagged PointersabstractPage-based virtual memory relies on TLBs to accelerate the address translation. Nowadays, the gap between application workloads and the capacity of TLB continues to grow, bringing many costly TLB misses and making the TLB a performance bottleneck. Previous studies seek to narrow the gap by exploiting the contiguity of physical pages. One promising solution is to group pages that are both virtually and physically contiguous into a memory range. Recording range translations can greatly increase the TLB reach, but ranges are also hard to index because they have arbitrary bounds. The processor has to compare against all the boundaries to determine which range an address falls in, which restricts the usage of memory ranges. In this article, we propose a tagged-pointer-based scheme, FlexPointer, to solve the range indexing problem. The core insight of FlexPointer is that large memory objects are rare, so we can create memory ranges based on such objects and assign each of them a unique ID. With the range ID integrated into pointers, we can index the range TLB with IDs and greatly simplify its structure. Moreover, because the ID is stored in the unused bits of a pointer and is not manipulated by the address generation, we can shift the range lookup to an earlier stage, working in parallel with the address generation. According to our trace-based simulation results, FlexPointer can reduce nearly all the L1 TLB misses, and page walks for a variety of memory-intensive workloads. Compared with a 4K-page baseline system, FlexPointer shows a 14% performance improvement on average and up to 2.8x speedup in the best case. For other workloads, FlexPointer shows no performance degradation. Dongwei Chen, Dong Tong 0001, Jiangfang Yi, Xu Cheng 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | Information Leakage Attacks Exploiting Cache Replacement in Commercial ProcessorsabstractCaches have been used to construct various covert and side channels. Most existing cache channels exploit the timing difference between cache hits and cache misses. We highlight that cache misses in different states may have more significant time differences. This paper presents in detail how replacement latency differences can be used to construct timing-based channels (called WB channels) to leak information. Any modification to a cache line by a sender will set it to the dirty state, and the receiver can observe this through measuring the latency of replacing this cache set. This paper evaluates WB channels from implementation complexity, stability, scalability, bandwidth, and stealthiness. Experimental results show that WB channel can not only transmit information covertly with high bandwidth but also has the strong anti-interference ability. Moreover, many previous cache defense and detection mechanisms target attacks exploiting the timing difference between cache hits and cache misses. This paper discusses the effectiveness of the WB channels against some such strategies. Moreover, this paper shows how to use our WB channel to mount a side-channel attack against a real-world security-sensitive application, such as AES implemented in OpenSSL-1.0.1e. Hongwei Cui, Xu Cheng 0001 |
IEEE Trans. Computers | 3 |
| 2023 | A High-Coverage and Efficient Instruction-Level Testing Approach for x86 ProcessorsabstractThe processors have long been treated as trusted black boxes for running software. However, processors may have undocumented instructions and instruction flaws, which increase the attack surface of the computing system. Hardware-related attack surfaces can bypass malware detection tools, resulting in undefined system behavior, instability, and insecurity. Unfortunately, the existing testing methods for undocumented instructions and instruction flaws have issues of insufficient test coverage and low test efficiency. We proposed an approach Skipscan to address these issues, which tests both the legal instructions and the reserved instructions. For the first time, to improve the test coverage, Skipscan leverages anoptimized combination algorithmto generate instruction prefix combinations, which covers the entire types of legal prefix combinations. To improve the test efficiency, Skipscan skips a considerable number of redundant legal instructions by leveraging theminimal test setof immediate and displacement operands. We evaluated Skipscan on eight x86 processors from Intel and AMD. The number of legal instructions and reserved instructions tested by Skipscan are 121.4 and 259.55 times that of Sandsifter on average, respectively. The test efficiency of Skipscan is on average 4 times that of Sandsifter. The ratio of legal instructions is reduced from 78.2% to 20.1% on average. Furthermore, we found more undocumented instructions on x86 processors and instruction flaws in x86 disassemblers. Guang Wang 0005, Xu Cheng 0001, Dan Meng 0002 |
IEEE Trans. Computers | 3 |
| 2022 | FlexPointer: Fast Address Translation Based on Range TLB and Tagged PointersabstractPage-based virtual memory suffers from costly page walks because of the gap between application workload sizes and TLB capacity. In this paper, we propose a tagged-pointer-based system, FlexPointer, to solve this problem. FlexPointer creates memory ranges from objects larger than a certain threshold by allocating contiguous physical pages for them. Virtual addresses within a range share a common translation entry, thus greatly expanding the TLB capacity. Because such objects are rare, we can assign each of them a unique ID and pass it to the hardware in pointer tags. According to our trace-based simulation results, FlexPointer can reduce nearly all the L1 TLB misses and page walks for a variety of memory-intensive workloads, providing a 14% performance improvement on average. Dongwei Chen, Dong Tong 0001, Jiangfang Yi, Xu Cheng 0001 |
PACT | 5 |
| 2022 | Abusing Cache Line Dirty States to Leak Information in Commercial ProcessorsabstractCaches have been used to construct various types of covert and side channels to leak information. Most existing cache channels exploit the timing difference between cache hits and cache misses. However, we introduce a new and broader classification of cache covert channel attacks: Hit+Miss, Hit+Hit, and Miss+Miss. We highlight that cache misses (or cache hits) for cache lines in different states may have more significant time differences, and these can be used as timing channels. Based on this classification, we propose a new stable and stealthy Miss+Miss cache channel. Write-back caches are widely deployed in modern processors. This paper presents in detail a way in which replacement latency differences can be used to construct timing-based channels (called WB channels) to leak information in a write-back cache. Any modification to a cache line by a sender will set it to the dirty state, and the receiver can observe this through measuring the latency of replacing this cache set. We also demonstrate how senders could exploit a different number of dirty cache lines in a cache set to improve transmission bandwidth with symbols encoding multiple bits. The peak transmission bandwidths of the WB channels in commercial systems can vary between 1300 and 4400 kbps per cache set in a hyper-threaded setting without shared memory between the sender and the receiver. In contrast to most existing cache channels, which always target specific memory addresses, the new WB channels focus on the cache set and cache line states, making it difficult for the channel to be disturbed by other processes on the core, and they can still work in a cache using a random replacement policy. We also analyzed the stealthiness of WB channels from the perspective of the number of cache loads and cache miss rates. We discuss and evaluate possible defenses. The paper finishes by discussing various forms of side-channel attack. Xu Cheng 0001 |
HPCA | 3 |
| 2022 | In-depth Testing of x86 Instruction Disassemblers with Feedback Controlled DFS AlgorithmabstractInstruction disassemblers can be used for software reverse engineering, malware analysis, and undocumented instructions detection. However, flaws in the disassemblers directly affect the accuracy of its related applications. For example, if the disassembler fails to decode or misdecodes the binary code of malware, the reverse engineers may misinterpret the functionality of the malware. Therefore, it is necessary to systematically test the disassemblers. Existing works leverage the depth-first search (DFS) algorithm to search the x86 instruction space. However, they cannot cover all x86 instruction opcodes and register operands. The root cause is that existing DFS algorithms cannot guarantee the search depth for some instruction space. We proposed an approach, named FedDFS, to improve the search depth of DFS algorithm. We analyzed the x86 instruction formats and summarized the essential search depth for each instruction format. We leveraged a feedback controlled DFS algorithm, which is controlled by comparing its search depth with essential search depth. If FedDFS detects that search depth is smaller than essential search depth, the feedback mechanism promptly increases the search depth until it reaches the proper search depth. We evaluated FedDFS on disassembler Capstone and processors from Intel and AMD. The experimental results proved that, after increasing the search depth, FedDFS does improve the coverage of x86 instruction opcodes and register operands. FedDFS tested trillions of instructions and found more instruction flaws in Capstone, which of them can only be found by FedDFS. Guang Wang 0005, Xu Cheng 0001, Dan Meng 0002 |
ICCD | 3 |
| 2021 | MetaTableLite: An Efficient Metadata Management Scheme for Tagged-Pointer-Based Spatial SafetyabstractA tagged-pointer-based memory spatial safety protection system utilizes the unused bits in a pointer to store the boundary information of an object. This paper proposed a hybrid metadata management scheme, MetaTableLite, for tagged-pointer-based protections. We observed that objects of a large size only take a minority part of all the objects in a program. However, recording their boundary metadata with traditional pointer tags will incur large memory overheads. Based on this observation, we introduce a small supplementary table to maintain metadata for these few large objects. For small ones, MetaTableLite represents their boundaries with a 14-bit pointer tag, well utilizing the unused 16 bits in a conventional 64-bit pointer. MetaTableLite can achieve a 6% average memory overhead without alternating the conventional pointer representation. Dongwei Chen, Dong Tong 0001, Xu Cheng 0001 |
ICCD | 4 |
| 2021 | Differential Testing of x86 Instruction Decoders with Instruction Operand Inferring AlgorithmabstractThe instruction decoders are tools for software analysis, sandboxing, malware detection, and undocumented instructions detection. The decoders must be accurate and consistent with the instruction set architecture manuals. The existing testing methods for instruction decoders are based on random and instruction structure mutation. Moreover, the methods are mainly aimed at the legal instruction space. However, there is little research on whether the instructions in the reserved instruction space can be accurately identified as invalid instructions. We propose an instruction operand inferring algorithm, based on the depth-first search algorithm, to skip considerable redundant legal instruction space. The algorithm keeps the types of instructions in the legal instruction space unchanged and guarantees the traversal of the reserved instruction space. In addition, we propose a differential testing method that discovers decoding discrepancies between instruction decoders. We applied the method to XED and Capstone and found four million inconsistent instructions between them. Compared with the existing instruction generation method based on the depth-first search algorithm, the efficiency of our method is improved by about four times. Guang Wang 0005, Shuan Li, Xu Cheng 0001, Dan Meng 0002 |
ICCD | 4 |
| 2017 | A Staged Memory Resource Management Method for CMP systemsabstractMemory interference is a critical impediment to system performance in CMP systems. To address this problem, we first propose a Dynamically Proportional Bandwidth Throttling policy (DPBT), which dynamically throttles back memory-intensive applications based on their memory access behavior. DPBT achieves a more balance memory bandwidth partitioning. Moreover, we improve the previous memory channel partitioning scheme by integrating it with a bank partitioning. We further integrate DPBT with the improved memory channel partitioning scheme and a memory scheduling policy to leverage the architecture advantages, and present a Stage Memory Resource Management Method (SRM). Experimental results show that DPBT improves system throughput/fairness by 13.5%/31.1%. SRM provides 27.1% better system throughput and 34.8% better system fairness. Yangguo Liu, Junlin Lu, Dong Tong 0001, Xu Cheng 0001 |
ASAP | 4 |
| 2017 | Locality-aware bank partitioning for shared DRAM MPSoCsabstractMemory interference is a critical impediment to system performance in MPSoCs. To address this problem, we first propose a Locality-Aware Bank Partitioning (LABP), which partitions memory banks according to applications' memory access behavior. The key idea is to separate memory intensive applications with high row-buffer locality from the other applications. Moreover, we integrate LABP with a bandwidth allocation scheme to leverage the architecture advantages, and present a comprehensive approach named Integrated Bandwidth and Bank Partitioning (IBBP) to further alleviate the interference. Experimental results show LABP improves system throughput/fairness by 10.8%/26.4%. IBBP provides 14.1% better system throughput and 34.2% better system fairness. Our methods are better than other recent work, including bandwidth throttling, DBP and DBP-TCM. Yangguo Liu, Junlin Lu, Dong Tong 0001, Xu Cheng 0001 |
ASP-DAC | 4 |
| 2017 | Content Look-Aside Buffer for Redundancy-Free Virtual Disk I/O and CachingabstractStorage consolidation in a virtualized environment introduces numerous duplications in virtual disks and imposes considerable pressure on disk I/O and caching. In this paper, we present a content look-aside buffer (CLB) approach for simultaneously providing redundancy-free virtual disk I/O and caching. CLB attaches persistent fingerprints to virtual disk blocks, which enables detection of I/O redundancy before disk access. At run time, CLB exploits content pages already present in the guest disk caches to service the redundant reads through page sharing, thus eliminating both redundant I/O requests and redundant disk cache copies. For write requests, CLB uses a group invalidating writeback protocol for updating fingerprints to support crash consistency while minimizing disk write overhead. By implementing and evaluating a CLB prototype on KVM hypervisor, we demonstrate that CLB delivers considerably improved I/O performance with realistic workloads. Our CLB prototype improves the throughput of sequential and random read on duplicate data by 4.1x and 26.2x, respectively. For typical read-intensive workloads, such as booting VM and launching application, CLB's I/O deduplication and cache deduplication eliminates 94.9%--98.5% of read requests and saves 50%--100% cache memory in each VM, respectively. Compared with the QEMU's raw virtual disk format, CLB improves the per-disk VM density by 8x--16x. For mixed read-write workloads, the cost of on-line fingerprint updating offsets the read benefit; nevertheless, CLB substantially improves overall performance. Xianhua Liu 0001, Xu Cheng 0001 |
VEE | 3 |
| 2016 | MFAP: Fair Allocation between fully backlogged and non-fully backlogged applicationsabstractIn this paper, we consider the problem of ensuring fairness in systems serving a mixture of fully backlogged applications, which continuously demand resources, and non-fully backlogged applications. We introduce a fairness metric, called interference fairness, the basic idea underlying which is that the interference caused by application A for another application B should be equal to that caused by B for A. To effectively and efficiently guarantee this fairness metric, we propose Mutual Fair Allocation Policy (MFAP), a simple and powerful resource sharing policy, and show how it guarantees interference fairness between any pair of applications. We also show that MFAP, unlike other viable policies, satisfies several highly desirable properties, including some from game theory, as well as common sense intuitions. As a use case, we implemented MFAP on a disk scheduling framework. The experimental results based on synthetic and real workloads show how our implementation achieved interference fairness and improved non-fully backlogged applications performance. Yan Sui, Dong Tong 0001, Xianhua Liu 0001, Xu Cheng 0001 |
ICCD | 5 |
| 2015 | Exploration of the Relationship Between Just-in-Time Compilation Policy and Number of Cores
Mingkai Huang, Xianhua Liu 0001, Xu Cheng 0001 |
ICA3PP (4) | 4 |
| 2015 | An Energy-Efficient Branch Prediction with Grouped Global HistoryabstractBranch prediction has been playing an increasingly important role in improving the performance and energy efficiency for modern microprocessors. The state-of-the-art branch predictors, such as the perceptron and TAGE predictors, leverage novel prediction algorithms to explore longer branch history for higher prediction accuracy. We observe that as the branch history is becoming longer, the efficiency of global history is degraded by the interference of different branch instructions. In order to mitigate the excessive influence of the branch history interference, we propose the Grouped Global History (GGH) based branch predictor, a lightweight yet efficient branch predictor. Unlike existing branch predictors that make use of a unified global history for prediction, GGH divides the global history into a set of subgroups such that the interference resulted by frequently executed branch instructions could be restricted. With subgroups of global history, GGH also enables us to track even longer effective branch correlation without introducing hardware storage overhead. Our experimental results based on SPEC CINT 2006 workloads demonstrate that our approach can significantly reduce the branch mispredictions per kilo instructions (MPKI) by 4.76 over the baseline perceptron predictor, with a simple control logic extension. Mingkai Huang, Xianhua Liu 0001, Mingxing Tan, Xu Cheng 0001 |
ICPP | 5 |
| 2014 | Improving system throughput and fairness simultaneously in shared memory CMP systems via Dynamic Bank PartitioningabstractApplications running concurrently in CMP systems interfere with each other at DRAM memory, leading to poor system performance and fairness. Memory access scheduling reorders memory requests to improve system throughput and fairness. However, it cannot resolve the interference issue effectively. To reduce interference, memory partitioning divides memory resource among threads. Memory channel partitioning maps the data of threads that are likely to severely interfere with each other to different channels. However, it allocates memory resource unfairly and physically exacerbates memory contention of intensive threads, thus ultimately resulting in the increased slowdown of these threads and high system unfairness. Bank partitioning divides memory banks among cores and eliminates interference. However, previous equal bank partitioning restricts the number of banks available to individual thread and reduces bank level parallelism. In this paper, we first propose a Dynamic Bank Partitioning (DBP), which partitions memory banks according to threads' requirements for bank amounts. DBP compensates for the reduced bank level parallelism caused by equal bank partitioning. The key principle is to profile threads' memory characteristics at run-time and estimate their demands for bank amount, then use the estimation to direct our bank partitioning. Second, we observe that bank partitioning and memory scheduling are orthogonal in the sense; both methods can be illuminated when they are applied together. Therefore, we present a comprehensive approach which integrates Dynamic Bank Partitioning and Thread Cluster Memory scheduling (DBP-TCM, TCM is one of the best memory scheduling) to further improve system performance. Experimental results show that the proposed DBP improves system performance by 4.3% and improves system fairness by 16% over equal bank partitioning. Compared to TCM, DBP-TCM improves system throughput by 6.2% and fairness by 16.7%. When compared with MCP, DBP-TCM provides 5.3% better system throughput and 37% better system fairness. We conclude that our methods are effective in improving both system throughput and fairness. Mingli Xie, Dong Tong 0001, Kan Huang, Xu Cheng 0001 |
HPCA | 4 |
| 2014 | Block value based insertion policy for high performance last-level cachesabstractLast-level cache performance has been proved to be crucial to the system performance. Essentially, any cache management policy improves performance by retaining blocks that it believes to have higher values preferentially. Most cache management policies use the access time or reuse distance of a block as its value to minimize total miss count. However, cache miss penalty is variable in modern systems due to i) variable memory access latency and ii) the disparity in latency toleration ability across different misses. Some recently proposed policies thus take into account the miss penalty as the block value. However, only considering miss penalty is not enough. In fact, the value of a block includes not only the penalty on its misses, but also the reduction of processor stall cycles on its hits, i.e., hit benefit. Therefore, we propose a method to compute both miss penalty and hit benefit. Then, the value of a block is calculated by accumulating all the miss penalty and hit benefits of its requests. Using our notion of block value, we propose Value based Insertion Policy (VIP) which aims to reserve more blocks with higher values in the cache. VIP keeps track of a small number of incoming and victim block pairs to learn the relationship between the value of the incoming block and that of the victim. On a miss, if the value of the incoming block is learned to be lower than that of the victim block in the past, VIP will predict that the incoming block is valueless and insert it with a high eviction priority. The evaluation shows that VIP can improve cache performance significantly in both single-core and multi-core environment while requiring a low storage overhead. Lingda Li, Junlin Lu, Xu Cheng 0001 |
ICS | 3 |
| 2014 | SPTU: Improving Dynamic Binary Translation through Software Prediction with Target UpdatingabstractIn dynamic translation system, handling indirect branch is a major source of performance overhead, because it must perform an on-the-fly address translation at each indirect branch execution. The translation systems usually adopt software prediction to reduce the overhead of address translation, but the low prediction accuracy restricts the performance improvement. Ning Jia 0004, Xu Cheng 0001 |
SYSTOR | 4 |
| 2014 | Retention Benefit Based Intelligent Cache Replacement
Lingda Li, Junlin Lu, Xu Cheng 0001 |
J. Comput. Sci. Technol. | 3 |
| 2013 | An energy-efficient branch prediction technique via global-history noise reductionabstractAccurate branch prediction can improve processor performance, while reducing energy waste. Though some existing branch predictors have been proved effective, they usually require large amount of storage or complicate the processor front-end. This paper proposes a novel branch prediction technique called History Artificially Selected (HAS) prediction. It is a hardware technique that bases on the existing branch predictors to detect history noises and avoid noise interferences when predicting branches. It separates the original branch predictor into sub-predictors, each of which performs differently in branch history updating. With the help of some history stacks, one sub-predictor saves and restores the branch history at the entrance and the exit of loops and program subroutines where history noise usually exists. Through using a tournament mechanism, HAS prediction selectively uses the modified branch history to eliminate the history noise interferences and retain those useful history correlations at the same time. Our experimental results show that for three representative branch predictors, gshare, perceptron, and TAGE, it reduces the MPKI by 1.49, 2.85, and 1.10 respectively, resulting in 4.55%, 10.16%, and 4.45% performance improvement. It also reduces energy consumption by 4.02%, 7.78%, and 3.91%, respectively. Zichao Xie, Dong Tong 0001, Xu Cheng 0001 |
ISLPED | 3 |
| 2013 | Page policy control with memory partitioning for DRAM performance and power efficiencyabstractDRAM performance and power efficiency considerations are becoming increasingly important. Bank partitioning partitions memory banks among cores and eliminates inter-thread interference, thus improving system performance of shared memory CMP systems. However, it doesn't take into account DRAM power consumption. We propose an application-aware page policy, which exploits potential benefits of page policy to optimize DRAM performance or minimize power consumption. The key idea is to dynamically assign page policy to applications according to their memory characteristics. As an improvement, we propose a power-aware bank partitioning to balance DRAM performance and power consumption. Experimental results show that our proposal increases system performance and significantly improves DRAM power efficiency. Mingli Xie, Dong Tong 0001, Yi Feng 0003, Kan Huang, Xu Cheng 0001 |
ISLPED | 5 |
| 2012 | Optimal bypass monitor for high performance last-level cachesabstractIn the last-level cache, large amounts of blocks have reuse distances greater than the available cache capacity. Cache performance and efficiency can be improved if some subset of these distant reuse blocks can reside in the cache longer. The bypass technique is an effective and attractive solution that prevents the insertion of harmful blocks. Lingda Li, Dong Tong 0001, Zichao Xie, Junlin Lu, Xu Cheng 0001 |
PACT | 5 |
| 2012 | An integrated and automated memory optimization flow for FPGA behavioral synthesisabstractBehavioral synthesis tools have made significant progress in compiling high-level programs into register-transfer level (RTL) specifications. But manually rewriting code is still necessary in order to obtain better quality of results in memory system optimization. In recent years different automated memory optimization techniques have been proposed and implemented, such as data reuse and memory partitioning, but the problem of integrating these techniques into an applicable flow to obtain a better performance has become a challenge. In this paper we integrate data reuse, loop pipelining, memory partitioning, and memory merging into an automated optimization flow (AMO) for FPGA behavioral synthesis. We develop memory padding to help in the memory partitioning of indices with modulo operations. Experimental results on Xilinx Virtex-6 FPGAs show that our integrated approach can gain an average 5.8× throughput and 4.55× latency improvement compared to the approach without memory partitioning. Moreover, memory merging saves up to 44.32% of block RAM (BRAM). Peng Zhang 0007, Xu Cheng 0001, Jason Cong |
ASP-DAC | 3 |
| 2012 | Energy-efficient branch prediction with Compiler-guided History StackabstractBranch prediction is critical in exploring instruction level parallelism for modern processors. Previous aggressive branch predictors generally require significant amount of hardware storage and complexity to pursue high prediction accuracy. This paper proposes the Compiler-guided History Stack (CHS), an energy-efficient compiler-microarchitecture cooperative technique for branch prediction. The key idea is to track very-long-distance branch correlation using a low-cost compiler-guided history stack. It relies on the compiler to identify branch correlation based on two program substructures: loop and procedure, and feed the information to the predictor by inserting guiding instructions. At runtime, the processor dynamically saves and restores the global history using a low-cost history stack structure according to the compiler-guided information. The modification on the global history enables the predictor to track very-long-distance branch correlation and thus improves the prediction accuracy. We show that CHS can be combined with most of existing branch predictors and it is especially effective with small and simple predictors. Our evaluations show that the CHS technique can reduce the average branch mispredictions by 28.7% over gshare predictor, resulting in average performance improvement of 10.4%. Furthermore, it can also improve those aggressive perceptron, OGEHL and TAGE predictors. Mingxing Tan, Xianhua Liu 0001, Zichao Xie, Dong Tong 0001, Xu Cheng 0001 |
DATE | 5 |
| 2012 | Improving inclusive cache performance with two-level eviction priorityabstractInclusive cache hierarchies are widely adopted in modern processors, since they can simplify the implementation of cache coherence. However, it sacrifices some performance to guarantee inclusion. Many recent intelligent management policies are proposed to improve the last-level cache (LLC) performance by evicting blocks with poor locality earlier. Unfortunately, they are inapplicable in inclusive LLCs. In this paper, we propose Two-level Eviction Priority (TEP) policy. Besides the eviction priority provided by the baseline replacement policy, TEP appends an additional high level of eviction priority to LLC blocks, which is decided at the insertion time and cannot be changed during their lifetime in the LLC. When blocks with high eviction priority are not in inner caches anymore, they get evicted from the LLC preferentially. Thus, the LLC can retain more useful blocks to improve performance. TEP can cooperate well with various baseline replacement policies. Our evaluation shows that TEP with NRU can improve the performance of inclusive LLCs significantly while requiring negligible extra storage. It also outperforms other recent proposals including QBS, DIP, and DRRIP. Lingda Li, Dong Tong 0001, Zichao Xie, Junlin Lu, Xu Cheng 0001 |
ICCD | 5 |
| 2012 | CVP: an energy-efficient indirect branch prediction with compiler-guided value patternabstractIndirect branch prediction is becoming increasingly important in modern high-performance processors. However, previous indirect branch predictors either require a significant amount of hardware storage and complexity, or heavily rely on the expensive manual profiling. Mingxing Tan, Xianhua Liu 0001, Xu Cheng 0001 |
ICS | 4 |
| 2012 | SWIP Prediction: Complexity-Effective Indirect-Branch Prediction Using Pointers
Zichao Xie, Dong Tong 0001, Mingkai Huang, Qinqing Shi, Xu Cheng 0001 |
J. Comput. Sci. Technol. | 5 |
| 2011 | TAP prediction: Reusing conditional branch predictor for indirect branches with Target Address PointersabstractIndirect-branch prediction is becoming more important for modern processors as more programs are written in object-oriented languages. Previous hardware-based indirect-branch predictors generally require significant hardware storage or use aggressive algorithms which make the processor front-end more complex. In this paper, we propose a fast and cost-efficient indirect-branch prediction strategy, called Target Address Pointer (TAP) Prediction. TAP Prediction reuses the history-based branch direction predictor to detect occurrences of indirect branches, and then stores indirect-branch targets in the Branch Target Buffer (BTB). The key idea of TAP Prediction is to predict the Target Address Pointers, which generate virtual addresses to index the targets stored in the BTB, rather than to predict the indirect-branch targets directly. TAP Prediction also reuses the branch direction predictor to construct several small predictors. When fetching an indirect branch, these small predictors work in parallel to generate the target address pointer. Then TAP prediction accesses the BTB to fetch the predicted indirect-branch target using the generated virtual address. This mechanism could achieve time cost comparable to that of dedicated-storage-predictors, without requiring additional large amounts of storage. Our evaluation shows that for three representative direction predictors-Hybrid, Perceptrons, and O-GEHL-TAP schemes improve performance by 18.19%, 21.52%, and 20.59%, respectively, over the baseline processor with the most commonly-used BTB prediction. Compared with previous hardware-based indirect-branch predictors, the TAP-Perceptrons scheme achieves performance improvement equivalent to that provided by a 48KB TTC predictor, and it also outperforms the VPC predictor by 14.02%. Zichao Xie, Dong Tong 0001, Mingkai Huang, Xiaoyin Wang, Qinqing Shi, Xu Cheng 0001 |
ICCD | 6 |
| 2011 | Dynamic Memory Demand Estimating Based on the Guest Operating System Behaviors for Virtual MachinesabstractIn the virtualized environment, memory can be efficiently utilized if the dynamic memory demands of virtual machines can be estimated at runtime. An efficient memory estimator should report the appropriate size of the memory which can be made full use of by the virtual machine while keeping reasonable performance. However, the appropriate size is hard to be estimated accurately with low overhead. This paper presents a memory demand estimator based on the guest operating system behaviors architecturally visible to the virtual machine monitor, and it can accurately reports the expected appropriate memory size with negligible overhead. The estimator consists of two components which respectively, track the amount of the memory residing in virtual address space, and the memory used as page cache only accessible in kernel mode. The experimental results show that the estimation error is only 0.4%~2.1%, and the runtime overhead is only 0.8% on average due to no additional memory protection traps are introduced. Yan Niu, Xu Cheng 0001 |
ISPA | 3 |
| 2010 | FPGA prototyping of an amba-based windows-compatible SoCabstractFor the increasing market of smart phones, mobile internet devices, and ultra-mobile PCs, mainstream vendors propose two approaches: one is based on ARM SoC, and the other is based on power-efficient x86 processor. However, either approach has its own limitation. The ARM-based approach lacks application software while the x86-based approach does not support flexible SoC extension. To overcome the limitations, we propose the PKUnity86 SoC architecture, which is based on AMBA bus architecture to support fast IP integration. Furthermore, it contains a reduced AMD Geode GX2 processor and several specific designs to support Microsoft Windows and exploit the massive PC software resources. Kan Huang, Junlin Lu, Jiufeng Pang, Yansong Zheng, Dong Tong 0001, Xu Cheng 0001 |
FPGA | 7 |
| 2010 | Bit-level optimization for high-level synthesis and FPGA-based accelerationabstractAutomated hardware design from behavior-level abstraction has drawn wide interest in FPGA-based acceleration and configurable computing research field. However, for many high-level programming languages, such as C/C++, the description of bitwise access and computation is not as direct as hardware description languages, and high-level synthesis of algorithmic descriptions may generate suboptimal implementations for bitwise computation-intensive applications. In this paper we introduce a bit-level transformation and optimization approach to assisting high-level synthesis of algorithmic descriptions. We introduce a bit-flow graph to capture bit-value information. Analysis and optimizing transformations can be performed on this representation, and the optimized results are transformed back to the standard data-flow graphs extended with a few instructions representing bitwise access. This allows high-level synthesis tools to automatically generate circuits with higher quality. Experiments show that our algorithm can reduce slice usage by 29.8% on average for a set of real-life benchmarks on Xilinx Virtex-4 FPGAs. In the meantime, the clock period is reduced by 13.6% on average, with an 11.4% latency reduction. Jiyu Zhang, Zhiru Zhang, Mingxing Tan, Xianhua Liu 0001, Xu Cheng 0001, Jason Cong |
FPGA | 6 |
| 2010 | Research Progress of UniCore CPUs and PKUnity SoCs
Xu Cheng 0001, Xiaoyin Wang, Junlin Lu, Jiangfang Yi, Dong Tong 0001, Xuetao Guan, Xianhua Liu 0001, Yi Feng 0003 |
J. Comput. Sci. Technol. | 1 |
| 2009 | WHOLE: A low energy I-Cache with separate way historyabstractSet-associative instruction caches achieve low miss rates at the expense of significant energy dissipation. Previous energy-efficient approaches usually suffer from performance degradation and redundant extension bits. In this paper, we propose a Way History Oriented Low Energy Instruction Cache (WHOLE-Cache) design for single issue and in-order execution processors. The WHOLE-Cache design not only achieves a significant portion of energy reduction by effectively reducing dynamic energy dissipation of set-associative instruction cache, but also leads to no additional cycle penalties. Tag comparison results are stored into either the Branch Target Buffer (BTB) or the Instruction Cache (I-Cache) to avoid tag checks and unnecessary way activation for subsequent accesses to visited cache lines. The extended BTB uses way history bits for branch instructions, while the I-Cache extension bits are used in case of fetching consecutive instructions resided in different cache lines. A valid flag is associated with each stored tag comparison result to indicate whether the instruction to be fetched is resided in the recorded location. A simple invalidation scheme is implemented in the cache miss replacement operation. Whenever a cache line is replaced, the pointers to it, which reside in the BTB or other I-cache lines, will be invalidated accordingly. We model the WHOLE-Cache design in Verilog. By deriving basic parameters from TSMC 65nm technology, we use Wattch simulator to evaluate the performance and energy reduction of the WHOLE-Cache in the instruction fetch stage. We use SPEC2000 and Mediabench as benchmarks. It is observed that compared with a conventional 4-way set-associative I-Cache, the energy consumption of the WHOLE-Cache is reduced by 65% without any performance penalty. Zichao Xie, Dong Tong 0001, Xu Cheng 0001 |
ICCD | 3 |
| 2009 | A Heterogeneous Auto-offloading Framework Based on Web Browser for Resource-Constrained DevicesabstractWeb applications are becoming increasingly popular. More and more types of multimedia data have been embedded into Web pages for providing abundant information and user interface. However, these complex Web pages could not be smoothly presented on devices without adequate functionalities or resource. An alternative approach is shifting tasks from resource-constrained device to surrogate with rich resource. In this paper, we present a heterogeneous auto-offloading framework based on Web browser, and implement plug-in proxy to offload multimedia processing tasks to heterogeneous surrogate automatically. The plug-in proxy works transparently with existing browser, and does not require modifying, recompiling or relinking. We perform several plug-in proxies of Firefox by offloading multimedia data and execution logic to surrogate, and merging remote display seamlessly with Web browser interface. The experimental results show that the proposed framework and plug-in proxies could enrich browser functionalities, reduce resource consumed by client, or improve performance of multimedia tasks. Xuetao Guan, Xu Cheng 0001 |
ICIW | 4 |
| 2009 | PaS: A Preemption-aware Scheduling Interface for Improving Interactive Performance in Consolidated Virtual Machine EnvironmentabstractAs virtualization technology is used widely in cloud computing, there are more and more interactive workloads being deployed on virtual machine (VM) environment. Although improving interactive performance has been heavily studied in operating system area, in consolidated VM environment, the improvements of guest OS are usually offset by the more coarse-grained VM scheduler, which may cause poor interactive performance. The guest OS scheduler and VM scheduler are totally independent with each other, which leads to the so called 'semantic gap'. To reduce this semantic gap, this paper presents PaS (Preemption-aware Scheduling) as an extension of VM scheduling interface. PaS introduces only two interfaces: one to register VM preemption conditions, the other to check if a VM is preempting. Thanks to the sophisticated techniques of interactive-process identification and optimization in traditional OS, it is trivial for guest OS to use the new interfaces: only 10 lines of code are added into Linux 2.6.18.8. The evaluation results show that PaS can significantly improve the interactive performance of consolidated VMs while keeping the fairness and performance isolation. Yubin Xia, Xu Cheng 0001 |
ICPADS | 3 |
| 2008 | Super-K: A SoC for single-chip ultra mobile computerabstractIn this panel discussion, I will make a brief introduction to the on-going Super-K project at the Microprocessor Research and Development Center (MPRC). The background of this project is the convergence of computers, consumer electronics and communication, so called 3C. MPRC has a ten-year history in microprocessor and system-on-chip (SoC) research and design. In the past, we have developed our own 32-bit RISC processor named UniCore, and two generations of SoC named PKUnity-863 (in acknowledgement of the support from China National High-Tech Program 863, abbr. 863 Project). From 2003 onwards, network computers based on UniCore have been used in commercial context. Xu Cheng 0001 |
ASP-DAC | 1 |
| 2008 | CASA: A New IFU Architecture for Power-Efficient Instruction Cache and TLB Designs
Han-Xin Sun, Kun-Peng Yang, Yulai Zhao 0003, Dong Tong 0001, Xu Cheng 0001 |
J. Comput. Sci. Technol. | 5 |
| 2007 | An Efficient SSA-Based Algorithm for Complete Global Value Numbering
Jiu-Tao Nie, Xu Cheng 0001 |
APLAS | 2 |
| 2007 | GISP: A Transparent Superpage Support Framework for LinuxabstractThough all of the current main-stream OSs have supported superpage to some extent, most of them need runtime information provided by applications, simulator or other tools. Transparent superpage support is a further step, while so far there are only very few primitive attempts for Linux. In this paper, we propose the design of GISP (global information based super page support), a transparent superpage support framework in Linux kernel. GISP adopts the basic idea of the reservation-based policy, and uses LMO (lightweight memory object) and POPMAP (population map) to manage the page allocation for applications. GISP could provide the core functions of superpage support while keeping the memory continuity by dynamic pages recycling. We implement it in Linux 2.4.17 on PKUnity SoC, and evaluate it for real workloads and benchmarks. We obtain substantial performance benefits from 8.1% to 24.0%. Compared with the best transparent superpage support in Linux up to now, we achieve better performance from 0.4% to 4.6% in most cases; even the worst case has comparable performance improvements within 2.6%. Otherwise, it keeps low management cost during system running which is suitable for not only scientific applications but also commodity applications. Ning Qu, Yansong Zheng, Xu Cheng 0001 |
ASAP | 4 |
| 2007 | A Retargetable Software Timing Analyzer Using Architecture Description LanguageabstractWorst case execution time (WCET) is an essential input for performance and schedulability analysis of real-time systems. Static WCET analysis requires program path analysis and microarchitecture modeling. Despite almost two decades of research, WCET analysis has not enjoyed wide acceptance in industry. This is in part due to the difficulty in microarchitecture modeling of modern processors. Given the large number of embedded processors available in the market, retargetability of the WCET analysis framework is a serious issue. In this paper, we address it using architecture description language (ADL). Starting with the ADL of a target processor, the proposed framework automatically generates graph-based execution models to capture timing effects of instructions in the pipeline. This pipeline model coupled with parameterized models of cache and branch prediction lead to a WCET framework that is safe, accurate and retargetable. Abhik Roychoudhury, Tulika Mitra, Prabhat Mishra 0001, Xu Cheng 0001 |
ASP-DAC | 5 |
| 2007 | Clock domain crossing fault model and coverage metric for validation of SoC design
Yi Feng 0003, Dong Tong 0001, Xu Cheng 0001 |
DATE | 4 |
| 2007 | A Fast Lossless Codec of Continuous-Tone Images for Thin Client ComputingabstractSummary form only given. We propose a fast and efficient lossless codec of continuous-tone images, SPEDIC (simple predictor and edge detector based image codec), which is uniquely suitable for coding screen updates generated by multimedia applications in thin client computing systems. A codec for thin client computing should take account of the tradeoff between compression ratio and coding complexity, because screen update images in thin client computing should be sent from server to client in a timely fashion. We argue that by combining similar but simpler building blocks of the state-of-the-art methods, a little inferior compression ratio of JPEG-LS (the standard for lossless image compression today) can be attained with much lower coding complexity. Yan Niu, Yubin Xia, Xu Cheng 0001 |
DCC | 4 |
| 2007 | Reuse Distance Based Cache Leakage Control
Yulai Zhao 0003, Dong Tong 0001, Xu Cheng 0001 |
HiPC | 4 |
| 2007 | A Fast and Efficient Codec for Multimedia Applications in Wireless Thin-Client ComputingabstractThin-client computing is uniquely suitable for mobile environments, where resource-poor devices may need to access critical applications over wireless networks. In thin-client computing, applications run on a powerful server, which sends screen updates to the client in real time. However, fast and efficient coding methods for screen updates of multimedia applications are very challenging in wireless environments, for the limited bandwidth and the real-time constraint. In this paper, we present SPEDIC (Simple Predictor and Edge Detector based Image Codec), a fast and efficient codec of continuous-tone images, which is appropriate for multimedia applications in wireless thin-client computing. SPEDIC adopts predictive coding, edge coding and run coding, depending on the characteristics of image blocks, and trades off between compression ratio and coding complexity. The experimental results show that SPEDIC provides good compression ratio with very low coding complexity, especially for continuous-tone images, and achieves the best end-to-end latency over wireless networks. Compared with JPEG-LS, the standard of lossless image compression, SPEDIC compresses 2.5 times faster and decompresses 3.2 times faster than it does for a series of continuous-tone images, and is only 6.2% inferior to it in compression ratio. Yan Niu, Yubin Xia, Xu Cheng 0001 |
WOWMOM | 4 |
| 2007 | An Energy-Efficient Instruction Scheduler Design with Two-Level Shelving and Adaptive Banking
Yulai Zhao 0003, Dong Tong 0001, Xu Cheng 0001 |
J. Comput. Sci. Technol. | 4 |
| 2005 | Bitwidth-aware scheduling and binding in high-level synthesisabstractMany high-level description languages, such as C/C++ or Java, lack the capability to specify the bitwidth information for variables and operations. Synthesis from these specifications without bitwidth analysis may introduce wasted resources. Furthermore, conventional high-level synthesis techniques usually focus on uniform-width resources, thus they cannot obtain the full resource savings even with bitwidth information. This work develops a bitwidth-aware synthesis flow, including bitwidth analysis, scheduling and binding, and register allocation and binding, to exploit the multi-bitwidth nature of operations and variables for area-efficient designs. We also develop lower bound estimation to evaluate the efficiency of our proposed solutions for register allocation and binding. The flow is implemented in the MCAS synthesis system [11]. Experimental results show that our proposed bitwidth-aware synthesis flow reduces area by 36% and wire-length by 52% on average compared to the uniform-width MCAS flow, while achieving the same performance. Jason Cong, Yiping Fan, Guoling Han, Yizhou Lin, Junjuan Xu, Zhiru Zhang, Xu Cheng 0001 |
ASP-DAC | 7 |
| 2005 | BluePower - A New Distributed Multihop Scatternet Formation Protocol for Bluetooth NetworksabstractBluetooth is a promising local area wireless technology designed to establish both personal area and multihop ad hoc networks. In this paper, we present BluePower as a novel and practical distributed scheme for building large multihop scatternets based on device transmitting power. The protocol is executed at each node with no prior knowledge of network topology and constructs the topology simply by enabling each node to alternate between scatternet formation and communication. Different from existing solutions, our design integrates device mobility and network self-healing from partitions with topology formation. Besides, partial loop detection is applied to minimize the redundant links in the network. The simulation results show that BluePower has low scatternet formation latency, decent number of slaves per piconet, short average route length and small connection delay for nodes to establish the first communication links in dynamic environments. Weijia Jia 0001, Xu Cheng 0001 |
ICPP | 4 |