VLDB 2026 Research / reviewers in the wild / expert
Mingyu Chen 0001
dblp:62/5558
· DBLP profile ↗
100ranked-venue papers
0as first author
30since 2021 · last 2026
0000-0003-4469-1037ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 74 · 26 since 2021Software engineering, systems software and programming languages · 10 · 2 since 2021Computer networks · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ANIMo: Accelerating Nested Isolation with Monitor-free Domain Transition
Yibin Xu, Tianyi Huang, Tianyue Lu, Mingyu Chen 0001 |
ASP-DAC | 6 |
| 2026 | S-MSHR: A Scalable MSHR Architecture Using Cache Tag Data-Ready Bits and Index Queues
Xu Zhang 0033, Yibin Xu, Tianyue Lu, Mingyu Chen 0001 |
CCGrid | 6 |
| 2026 | LightSerial: Accelerating In-Process Isolation via Implicit Dependency ExposureabstractWith the frequent emergence of in-process security threats, in-process isolation has become essential for software security. While hardware security primitives enable lightweight isolation, an implicit dependency persists between permission-setting instructions and their subsequent checked instructions. To enforce this implicit dependency, modern out-of-order processors resort to enforcing strict serialization by flushing the entire pipeline during permission switches. This approach severely degrades instruction-level parallelism. Moreover, serialization leaves an exploitable security window for side-channel attacks by failing to revert speculative microarchitectural side effects. Yibin Xu, Tianyue Lu, Mingyu Chen 0001 |
CF | 5 |
| 2026 | AExec: Asynchronous Multi-accelerator Execution and Management Mechanism
Xiaokun Pei, Zhuolun Jiang, Mingyu Chen 0001, Songyue Wang, Tianyue Lu |
CF | 3 |
| 2026 | RaidenSwap: A Multi-Swap Remote System for Multi-core ApplicationsabstractKernel-based remote memory systems are gaining traction in datacenters due to their significant improvement in memory utilization and their ability to transparently provide applications with unlimited memory capacity. However, the high degree of parallelism in contemporary applications leads to a significant demand for remote access throughput, which mismatches with the state-of-the-art kernel swap path due to its inherently limited parallelism. As a result, it cannot scale up the multi-core applications. We dive into the implementation of the swap path and identify the root cause behind it - significant lock contentions and inefficient swap tasks offloading. Kefan Liu, Ke Liu 0004, Xu Zhang 0033, Ning Liu 0031, Sa Wang, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
EuroSys | 10 |
| 2026 | Chrono-Fabric: A Decoupled Hierarchical Framework for Cycle-Accurate Coordination in Multi-FPGA Systems
Congwu Zhang, Panyu Wang, Bibo Yang, Mingyu Chen 0001, Yungang Bao, Ke Zhang 0017 |
FPGA | 5 |
| 2025 | CoroAMU: Unleashing Memory-Driven Coroutines through Latency-Aware Decoupled OperationsabstractModern data-intensive applications face memory latency challenges exacerbated by disaggregated memory systems. Recent work shows that coroutines are promising in effectively interleaving tasks and hiding memory latency, but they struggle to balance latency-hiding efficiency with runtime overhead. We present CoroAMU, a hardware-software co-designed system for memory-centric coroutines. It introduces compiler procedures that optimize coroutine code generation, minimize context, and coalesce requests, paired with a simple interface. With hardware support of decoupled memory operations, we enhance the Asynchronous Memory Unit to further exploit dynamic coroutine schedulers by coroutine-specific memory operations and a novel memory-guided branch prediction mechanism. It is implemented with LLVM and open-source XiangShan RISCV processor over the FPGA platform. Experiments demonstrate that the CoroAMU compiler achieves a $\mathbf{1. 5 1} \boldsymbol{\times}$ speedup over state-of-the-art coroutine methods on Intel server processors. When combined with optimized hardware of decoupled memory access, it delivers $3.39 \times$ and $4.87 \times$ average performance improvements over the baseline processor on FPGA-emulated disaggregated systems under 200 ns and 800 ns latency respectively. Zhuolun Jiang, Songyue Wang, Xiaokun Pei, Tianyue, Mingyu Chen 0001 |
PACT | 5 |
| 2025 | A Fast, Iterative Clock Skew Scheduling Algorithm with Dynamic Sequential Graph ExtractionabstractClock skew scheduling (CSS) is a well-known technique that improves design timing slack by adjusting clock latency to flipflops. CSS requires obtaining timing path information between sequential elements (including flip-flops and I/O ports), known as sequential graph extraction, which is the most time-consuming part of advanced CSS. In this paper, to quickly identify the potential of clock skew in slack optimization, we propose an iterative CSS algorithm that leverages timing propagation to facilitate sequential graph extraction. Then, we provide a comprehensive skew calculation method that considers multiple clock latency constraints, obtaining the target latency of each flip-flop. Finally, we present slack optimization techniques to achieve the target latencies. Our algorithm achieves a $49.11 \times$ speedup compared to the advanced CSS algorithm based on partial graph extraction, reducing 90.05% of the extracted edges. Compared to a state-of-the-art CSS-based slack optimization methodology, our algorithm delivers a $27.01 \times$ speedup with superior slack improvement. Shijian Chen, Yihang Qiu, Biwei Xie, Mingyu Chen 0001 |
DAC | 4 |
| 2025 | DASICS: Efficient In-Process Protection with Hardware-Assisted Dynamic CompartmentalizationabstractHardware-assisted in-process compartmentalization reduces attack surface at low cost, but existing methods face practical challenges: inefficient dynamic permission management, weak metadata/instruction protection, and limited resource isolation. To tackle these problems, this paper proposes DASICS, a lightweight and efficient design of hardware-assisted in-process compartmentalization. DASICS partitions the process code segments into trusted and untrusted compartments and implements a user-mode protection runtime in the trusted compartment for dynamic permission management. It employs boundary registers to enforce dynamic access-control restrictions on instructions within different untrusted compartments. Additionally, it applies metadata access restriction, control-flow checks, and systemcall filtering for the untrusted compartments to achieve comprehensive protection. We implemented a hardware prototype of DASICS on the RISC-V XiangShan superscalar out-of-order processor and validated its effectiveness on FPGA. Our prototype increases less than 5% LUTs cost, and experimental results show that DASICS isolation incurs an average overhead of${6.02 \%}$on Memcached key-value store and 8.18% on NGINX webserver. DASICS project is publicly available at github.com/DASICS-ICT. Yibin Xu, Tianyi Huang, Tianyue Lu, Mingyu Chen 0001 |
ICCD | 7 |
| 2025 | A Unified Framework for DRL-Based Congestion Control to Optimize QoS Over Mobile NetworksabstractDeep Reinforcement Learning-based Congestion Control Algorithms (DRL-based CCA) have shown their great potential to adapt to various environments automatically (e.g., Orca). However, it is difficult for existing DRL-based CCAs to achieve superior QoS consistently, particularly over mobile networks with rapid network fluctuations. The fundamental problem stems from the training of a single model that encompasses a wide range of network conditions. The resulting model can be overly generalized, rendering it less accurately or optimally tailored for a specific network condition. To tackle this challenge, we develop Network-segmented Model Specialization (NMS), a framework that automatically maximizes QoS for any DRLbased CCA under different network conditions. Specifically, NMS generates a set of models offline, each trained for a specific network segment. Online, it selects the model based on the current segment. We showed that NMS not only improves the QoSs of existing DRL-based CCAs consistently, but also opens a new way for the exploration of a clean-slate approach, as opposed to following the hybrid TCP/DRL approach. Hence, we designed Galaxy, a novel clean-slate DRL-based CCA that addresses the inherent limitations in clean-slate approaches by incorporating network segments' knowledge. Extensive evaluations show NMSoptimized Galaxy further explores NMS to achieve superior QoS. Ke Liu 0004, Jack Y. B. Lee, Theophilus Benson, Yungang Bao, Mingyu Chen 0001 |
IWQoS | 7 |
| 2025 | DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack Communication
Xu Zhang 0033, Ke Liu 0004, Yuan Hui 0001, Yisong Chang, Yizhou Shan, Ke Zhang 0017, Yungang Bao, Mingyu Chen 0001, Chenxi Wang 0005 |
USENIX ATC | 10 |
| 2025 | A Survey of Hardware-Assisted Intra-Address Space Protections
Tian-Yi Huang, Si-Yuan Zeng, Yi-Bin Xu, Tian-Yue Lu, Mingyu Chen 0001 |
J. Comput. Sci. Technol. | 7 |
| 2024 | iEDA: An Open-source infrastructure of EDAabstractBy leveraging the power of open-source software, the EDA tool offers a cost-effective and flexible solution for designers, researchers, and hobbyists alike. Open-source EDA promotes collaboration, innovation, and knowledge sharing within the EDA community. It emphasizes the role of the toolchain in accelerating the development of electronic systems, reducing design costs, and improving design quality. This paper presents an open-source EDA project, iEDA, aiming to build a basic infrastructure for EDA technology evolution and closing the industrial-academic gap in the EDA area. As the foundation for developing EDA tools and researching EDA algorithms and technologies, iEDA is mainly composed of file system, database, manager, operator and interface. To demonstrate the effectiveness of iEDA, we implement and tape out four chips of different scales (from 700k to 500M gates) on different process nodes (110nm and 28nm) with iEDA. iEDA is publicly available on the project home page https://github.com/OSCC-Project/iEDA. Zengrong Huang, Simin Tao, Zhipeng Huang 0009, Chunan Zhuang, Yihang Qiu, Guojie Luo, Huawei Li 0001, Haihua Shen, Mingyu Chen 0001, Dongbo Bu, Wenxing Zhu, Ye Cai 0001, Xiaoming Xiong, Yi Heng, Peng Zhang 0007, Bei Yu 0001, Biwei Xie, Yungang Bao |
ASPDAC | 12 |
| 2024 | Planaria: Pattern Directed Cross-page Composite PrefetcherabstractGiven the memory wall, the performance of the memory system significantly influences the user experience of mobile phones. The system cache (SC), located on the memory side, is shared among all CPUs and GPUs within the mobile phone, serving as the last line of defense before resorting to time-consuming off-chip memory access. Managing SC is a challenge due to its large working set and irregular access patterns. Despite occupying a substantial on-chip area, SC's effectiveness in terms of hit rate is relatively low. It has been observed that neither state-of-the-art cache replacement policies nor increasing cache size significantly improve SC performance. Prefetchers designed for higher-level caches cannot be seamlessly applied to SC due to the absence of the required program counter (PC) on the memory side and the violation of stringent power constraints in mobile phones by aggressive prefetch traffic. Yuhang Liu 0001, Mingyu Chen 0001 |
DAC | 2 |
| 2024 | XUNI: Virtual Machine Abstraction for Self-contained and Multi-tenant Cloud FPGAsabstractFPGAs have become essential infrastructural components as well as publicly rentable resources in cloud and datacenters. Although one tenant is equipped with a couple of individual physical FPGA devices to solve a single problem, cloud FPGAs are still underutilized in many scenarios. Considering the economic cost of such heterogeneous computing resources, it is essential to explore opportunities for FPGA virtualization such that multiple tenants can share one physical device. Unlike the conventional host-centric virtualization approaches considering FPGAs as I/O peripherals, we propose XUNI, an FPGA-centric and self-contained virtual machine (VM) abstraction without involving the host-side virtualization techniques. Specifically, we design a hardware-software co-designed hypervisor for resource management and provisioning of various FPGA VMs. First, XUNI partitions the FPGA fabric into a series of reconfigurable regions that can be flexibly assembled for the deployment of large-scale designs. Second, we introduce both static partition and dynamic allocation schemes for FPGA-side DRAM sharing in XUNI. Last but not least, a hierarchical multi-tenant hardware network stack is built to provide an I/O interface for each FPGA VM. We implement XUNI and conduct infrastructural evaluations on a custom cloud FPGA node populated with an AMD/Xilinx Zynq MPSoC chip. Preliminary results demonstrate that XUNI is capable of handling hundreds of thousands of FPGA-VM-initiated memory requests per second. The hardware network stack exhibits line rate (~100Gbps) when receiving packets with the default 1500-byte MTU size. Moreover, each FPGA VM boots up within hundreds of milliseconds, which is comparable to emerging lightweight host-side VMs. Guiyuan Zhu, Yunhai Liu, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001 |
FPGA | 6 |
| 2024 | Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory AccessabstractThe growing memory demands of modern applications have driven the adoption of far memory technologies in data centers to provide cost-effective, high-capacity memory solutions. However, far memory presents new performance challenges because its access latencies are significantly longer and more variable than local DRAM. For applications to achieve acceptable performance on far memory, a high degree of memory-level parallelism (MLP) is needed to tolerate the long access latency. While modern out-of-order processors are capable of exploiting a certain degree of MLP, they are constrained by resource limitations and hardware complexity. The key obstacle is the synchronous memory access semantics of traditional load/store instructions, which occupy critical hardware resources for a long time. The longer far memory latencies exacerbate this limitation. This article proposes a set of Asynchronous Memory Access Instructions (AMI) and its supporting function unit, Asynchronous Memory Access Unit (AMU), inside contemporary Out-of-Order Core. AMI separates memory request issuing from response handling to reduce resource occupation. Additionally, AMU architecture supports up to several hundreds of asynchronous memory requests through re-purposing a portion of L2 Cache as scratchpad memory (SPM) to provide sufficient temporal storage. Together with a coroutine-based programming framework, this scheme can achieve significantly higher MLP for hiding far memory latencies. Evaluation with a cycle-accurate simulation shows AMI achieves 2.42× speedup on average for memory-bound benchmarks with 1μs additional far memory latency. Over 130 outstanding requests are supported with 26.86× speedup for GUPS (random access) with 5 μs latency. These demonstrate how the techniques tackle far memory performance impacts through explicit MLP expression and latency adaptation. Luming Wang, Xu Zhang 0033, Songyue Wang, Zhuolun Jiang, Tianyue Lu, Mingyu Chen 0001, Siwei Luo, Keji Huang |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | Suppressing the Interference Within a Datacenter: Theorems, Metric and StrategyabstractAs the paradigm of cloud computing, a datacenter accommodates many co-running applications sharing system resources. Although highly concurrent applications improve resource utilization, the resulting resource contention can increase the uncertainty of quality of services (QoS). Previous studies have shown that achieving high resource utilization and high QoS simultaneously is challenging. Moreover, quantifying the intensity of interference across multiple concurrent applications in a datacenter, where applications can be either latency-critical (LC) or best-effort (BE), poses a significant challenge. To address these issues, we propose Ah-Q, which comprises a series of theorems, a quantification theory and a scheduling strategy. Firstly, we present the necessary and sufficient conditions to precisely test whether a datacenter is both QoS guaranteed and high-throughput. We also present and prove a theorem that reveals the relationship between tail latency and throughput. Our theoretical results are insightful and useful for building datacenters that have desirable performance. By applying our theoretical results, datacenter architects can more effectively balance the trade-off between resource utilization and QoS, leading to improved performance for co-running applications. Secondly, we propose the “System Entropy” (E$\rm {_{S}}$) theory to quantitatively and analytically measure interference in a datacenter. Interference arises due to resource scarcity or irrational scheduling, and effective scheduling can alleviate resource scarcity. To assess the effectiveness of a resource scheduling strategy, we introduce the concept of “resource equivalence”. We evaluate various resource scheduling strategies to demonstrate the correctness and effectiveness of the proposed theory. Thirdly, we introduce a new resource scheduling strategy, ARQ, that leverages both isolation and sharing of resources. Our evaluations show that ARQ significantly outperforms state-of-the-art strategies PARTIES and CLITE in reducing the tail latency of LC applications and increasing the IPC of BE applications. Yuhang Liu 0001, Jiapeng Zhou, Mingyu Chen 0001, Yungang Bao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Rethinking Design Paradigm of Graph Processing System with a CXL-like Memory Semantic FabricabstractWith the evolution of network fabrics, message-passing clusters have been promising solutions for large-scale graph processing. Alternatively, the shared-memory model is also introduced to avoid redundant copies and extra storage space of graph data. Compared to conventional network fabrics, with the capability of fine-grained, byte-addressable remote memory access, emerging memory semantic interconnects and fabrics, e.g., Intel's Compute Express Link (CXL), are intuitively more appropriate for adoption in shared-memory clusters. However, due to the latency gap between local and remote memory, it is still challenging to take advantage of the shared-memory graph processing with memory semantic fabrics. To tackle this problem, in this paper, we first investigate memory access characterizations of graph vertex propagation based on the shared-memory model. Then we propose GraCXL, a series of design paradigms to address high-frequency and long-latency of remote memory access potentially incurred in CXL-based clusters. For system adaptiveness, we elaborate GraCXL towards the general-purpose CPU cluster and the domain-specific FPGA accelerator array, respectively. We design a custom fabric with the CXL.mem protocol and leverage a couple of ARM SoC-equipped FPGAs to build an evaluation prototype in the absence of commodity CXL hardware and platforms. Experimental results show that the proposed GraCXL CPU and FPGA clusters achieve 1.33x-8.92x and 2.48x-5.01x performance improvement, respectively. Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Zhang 0017, Mingyu Chen 0001 |
CCGrid | 5 |
| 2023 | MARB: Bridge the Semantic Gap between Operating System and Application Memory Access BehaviorabstractThe virtual memory subsystem (VMS) is a long-standing and integral part of an operating system (OS). It plays a vital role in enabling remote memory systems over fast data center networks and is promising in terms of transparency and generality. Specifically, these systems use three VMS mechanisms: demand paging, page swapping, and page prefetching. However, the VMS inherent data path is costly, which takes a huge toll on performance. Despite prior efforts to propose page swapping and prefetching algorithms to minimize the occurrences of the data path, they still fall short due to the semantic gap between the OS and applications - the VMS has limited knowledge of its running applications' memory access behaviors. In this paper, orthogonal to prior efforts, we take a fundamen-tally different approach by building an efficient framework to collect full memory access traces at the local bus, and make them available to the OS through CPU cache. Consequently, the page swapping and page prefetching can use this trace to make better decisions, thereby improving the overall performance of systems. We implement a proof-of-concept prototype on commodity x86 servers using a hardware-based memory tracking tool. To show-case our framework's benefits, we integrate it with a state-of-the-art remote memory system and the default kernel page eviction subsystem. Our evaluation shows promising improvements. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yisong Chang, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
DATE | 10 |
| 2023 | HoPP: Hardware-Software Co-Designed Page Prefetching for Disaggregated MemoryabstractMemory disaggregation is a promising direction to mitigate memory contention in datacenters. To make memory disaggregation practical, prior efforts expose remote memory to applications transparently via virtual memory subsystem’s swapping interface. However, due to the semantic gap between OS and applications – OS cannot know the memory accessing sequences of an application but via page faults. This approach has two limitations. First, it learns little from page faults’ access history, which leads to sub-optimal prefetching predictions. Second, a page fault can still occur even if there is a prefetch-hit which leads to a large kernel overhead.To address such limitations, our key insight is to decouple the address capturing from page faults by collecting full memory access traces in the memory controller. Using this idea, we buildHoPP– a hardware-software co-designed prefetching framework.HoPPadds hardware modules to the memory controller to feed sufficient hot pages to OS in real-time, which has three benefits inHoPP’s software design: 1) it improves existing prefetching algorithms with simple revamps, also offers more insights to build better policies; 2) the prefetch algorithm can run as a separate data path alongside the normal remote data path via page faults, potentially hiding the swap latency from applications, and enabling fine-grained control over prefetching behaviors; 3) the prefetch-hit overhead can be eliminated by early page table entry (PTE) injection, i.e., inject PTE for the prefetched page as soon as it returns. We implemented a proof-of-concept prototype using commodity servers along with a hardware-based memory tracking tool calledHMTTto emulate a modified memory controller. Results show that compared to Fastswap and Leap,HoPP-optimized prefetching algorithm achieves over 90% accuracy and coverage, which leads to up to 59% completion time improvement for various datacenter applications. Ke Liu 0004, Ting Liang, Zuojun Li, Tianyue Lu, Yinben Xia, Yungang Bao, Mingyu Chen 0001, Yizhou Shan |
HPCA | 9 |
| 2023 | Ah-Q: Quantifying and Handling the Interference within a Datacenter from a System PerspectiveabstractInterference among applications frequently occurs in a datacenter and significantly influences the cost-efficiency and the user experience. However, it is challenging for us to quantify the exact intensity of the interference that occurred in the overall system of a datacenter, because there are many concurrent applications in a datacenter, and their type can be either latency-critical (LC) and best-effort (BE). To address this issue, we present the Ah-Q which includes a theory and a strategy.First, we propose the "system entropy" (ES) theory to holistically and analytically quantify the interference in a datacenter to address this vital issue. The interference is caused by the scarcity of resources or/and the irrationality of scheduling. As more appropriate scheduling can compensate for resource scarcity, we derive the concept of "resource equivalence" to quantify the effectiveness of a resource scheduling strategy. We evaluate different resource scheduling strategies to validate the correctness and effectiveness of the proposed theory.Moreover, using the theory to eliminate interference, we propose a new resource scheduling strategy; i.e., ARQ, which dynamically allocates the isolated resources and the shared resources to simultaneously harvest the benefits of isolation and sharing. Our results show that compared to the state-of-the-art strategies (PARTIES and CLITE), ARQ is more effective to reduce the tail latency of the LC applications and to increase the IPC of the BE applications. Compared with PARTIES and CLITE, ARQ increases the yield (the ratio of satisfactory LC applications) by 25% and 20%, respectively; when the load is low, ARQ increases IPC of BE applications by 63.8% and 37.1%, respectively; ARQ reduces ESby 36.4% and 33.3%, respectively. The effectiveness of ARQ has saved resources significantly to achieve the same satisfactory overall user experience. Yuhang Liu 0001, Jiapeng Zhou, Mingyu Chen 0001, Yungang Bao |
HPCA | 4 |
| 2023 | REMU: Enabling Cost-Effective Checkpointing and Deterministic Replay in FPGA-based EmulationabstractLeft-shifted integration and evaluation of hardware and software design are increasingly crucial in pre-silicon validation of processor-centric computing systems. With the inherent cycle-accurate deployment of target processor design in programmable logic fabrics, FPGA-based emulation has attracted academic attention for early-stage performance evaluations. However, it is difficult to conduct system-level inspection and debugging within open-source academic FPGA-based emulation frameworks due to the limited HW-SW visibility at run-time.To fill such a gap, we present REMU, an FPGA-based emulation framework enabling hardware checkpointing and deterministic replay to acquire bit-accurate visibility of target processors and other system components. Specially, we first employ a cost-effective scan-chain insertion method and related implementation strategies within an optimized open-source synthesis tool for status capturing of the emulated circuit primitives. Then, we introduce mechanisms in the design of emulated memory and I/O peripheral components to precisely describe behaviors and ensure deterministic replay of system-level interactions. Experimental results show that REMU drastically speeds up the scan-chain insertion flow by 1.6x-32.7x, and the proposed mechanisms for deterministic replay in the emulated external memory introduce negligible overhead in FPGA resource utilization. Yuxiao Chen 0009, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001, Yungang Bao |
ICCD | 4 |
| 2023 | Morpheus: An Adaptive DRAM Cache with Online Granularity Adjustment for Disaggregated MemoryabstractDisaggregated memory introduces a cost-effective solution for improving the memory utilization rate of data centers, by sharing a distributed memory pool among several individual servers. However, latency penalty in the existing connection between a computing node and the memory pool introduces performance degradation due to frequent far memory accesses. Based on our observation, page caching in the local DRAM, despite its reductions in the number of far memory accesses, still faces severe data over-fetching problem.With a detailed analysis of far memory access traces collected via several representative real-world applications, we argue that exploiting the various page-specific preferences of caching granularity is the key point of solving the data over-fetching problem in the DRAM cache. Consequently, in this paper, we present that it is influential to enable 1) dynamic selection of caching granularity for each page to not only guarantee sufficient spatial localities compared to the conventional fine-grained cache lines but also avoid data over-fetching caused by the coarse-grained pages, as well as 2) adaptive adjustment of cache capacity during execution for each granularity to accommodate the varying proportion of pages with different granularity preferences. Specifically, we propose Morpheus, an adaptive DRAM cache architecture that determines an optimal page-specific caching granularity at run-time and dynamically adjusts capacity occupations of different caching granularities. Based on our modeling and evaluations within the DRAMSim3 simulator, Morpheus exhibits 1.17-1.34x performance speedup for a wide range of workloads against the state-of-the-art DRAM cache design. Xu Zhang 0033, Tianyue Lu, Yisong Chang, Ke Zhang 0017, Mingyu Chen 0001 |
ICCD | 5 |
| 2023 | A Data-Driven Framework for TCP to Achieve Flexible QoS Control in Mobile Data NetworksabstractLearning-based approaches have shown their great potential to adapt themselves to various environments (e.g., PCC and Sprout). Unfortunately, they do not consistently achieve superior QoS across different network conditions and configurations in mobile networks. Furthermore, although they can offer multiple application objectives by adjusting a preference weight vector, it is challenging for users to accurately express an application objective with a weight vector. In this work, we argue that, if configured correctly, the delay-based TCP scheme can outperform learned ones, and allow users to directly specify their objectives. To this end, we propose Post-QoS Analysis (PQSA), a data-driven framework that trains the key QoS-impacting parameters of the scheme to capture the statistical correlations between QoS objectives, network conditions, and configurations, thereby determining the optimal parameter-set that meets the user-defined QoS objective under different network conditions and configurations. To support this, we enhance conventional delay-based TCP design to develop a Generalized TCP-like Rate controller (GR) by exporting three key parameters. Extensive evaluations show that PQSA-optimized GR outperforms existing schemes in different scenarios consistently, and enables service providers to control the QoS flexibly. Ke Liu 0004, Ting Liang, Theophilus Benson, Jack Y. B. Lee, Vaneet Aggarwal, Yungang Bao, Mingyu Chen 0001 |
IWQoS | 8 |
| 2022 | Increasing Flexibility of Cloud FPGA VirtualizationabstractFPGA virtualization enables multiple tenants to share programmable hardware resources for application accelerations in cloud. However, such technique is still of limited usage in commercial FPGA cloud platforms, which mainly lies in: 1) absence of direct programming interfaces of the virtualized FPGA accelerators (vFPGAs) in tenants' virtual machines (VMs), 2) a fixed VM-vFPGA data movement scheme that is inadaptive to a wide range of data sizes among different applications, and 3) performance degradation due to unregulated inter-vFPGA competitions for limited shareable external resources (e.g., off-chip DRAM bandwidth). To tackle all the above issues, we propose a flexible FPGA virtualization framework and prototype an open cloud platform with ARM SoC-equipped FPGAs. Under such framework, tenants are allowed to directly initiate FPGA partial reconfiguration in isolated VMs via a direct I/O-like vFPGA device driver with as low as 20ms overhead. A hybrid data movement approach that leverages both memory-mapped I/O and DMA is also introduced in our framework to adaptively guarantee moderate VM-vFPGA bandwidth towards various data sizes. Moreover, a lightweight priority-based hardware scheduler is elaborated to monitor and dynamically allocate off-chip DRAM bandwidth among vFPGAs. Based on our preliminary infrastructure-level evaluation results, the proposed framework and the open prototyping are of significant interests to researchers looking forward to conducting further explorations in FPGA virtualization. Jinjie Ruan, Yisong Chang, Ke Zhang 0017, Kan Shi, Mingyu Chen 0001, Yungang Bao |
FPL | 5 |
| 2022 | FPL Demo: SERVE: Agile Hardware Development Platform with Cloud IDE and Cloud FPGAsabstractWe introduce SERVE, a cloud platform for agile hardware software co-design, with cloud IDE and cloud FPGAs integrated. SERVE enables users to focus on logic designs, without facing the hassle of setting up FPGA tools and development environment. Users can write and simulate hardware logic in the cloud IDE and then generate bitstream files through a Continuous Integration (CI) pipeline. Finally, the bitstream files are deployed on an FPGA board. A great amount of testbenches will be executed to ensure the correctness of the hardware logic. We will demo a workflow of modifying a RISC- V processor and getting the design change quickly evaluated using SERVE. Ke Zhang 0017, Yisong Chang, Yanlong Yin, Yuxiao Chen 0009, Songyue Wang, Mingyu Chen 0001, Yungang Bao |
FPL | 8 |
| 2022 | GraFF: A Multi-FPGA System with Memory Semantic Fabric for Scalable Graph ProcessingabstractFPGA has been a promising solution for graph processing in many scenarios. With a rapid growth in graph size, the on/off-chip memory capacity of a single FPGA is insufficient to hold large-scale graphs. To tackle such problem, in this position paper, we introduce GraFF, a Graph processing system with multiple FPGAs interconnected via a custom memory semantic Fabric. In order to efficiently exploit system parallelism, we first split the traversal of graph data into a series of independent fine-grained flits that are concurrently delivered among FPGAs as sheer memory semantic transactions. Then we relax FPGAs' synchronization from strict barrier boundaries between adjacent supersteps to fully parallelize graph traversing and computing. We build a prototype of GraFF with four custom FPGA nodes. Preliminary evaluation result based on the Breadth First Search (BFS) algorithm shows that the peak performance of GraFF reaches up to 6.23 GTEPS. Moreover, GraFF exhibits linear scalability when the number of FPGAs rises from one to four. Xu Zhang 0033, Yisong Chang, Tianyue Lu, Ke Liu 0004, Ke Zhang 0017, Mingyu Chen 0001 |
FPT | 6 |
| 2022 | HCMonitor: An accurate measurement system for high concurrent network servicesabstractAbstract This article aims to enhance the monitoring accuracy of high concurrent network services. As modern network services grow rapidly in data centers, tail latency has become one of the most crucial deciding factors on user experience. Latency measurement and anomaly detection are essential in evaluating service performance. Existing monitoring tools can be divided into two categories according to estimation methods. First, approaches based on sample traffic sample network packets to unburden the measurement. Second, approaches based on full traffic like wrk, analyze all of the packets from the kernel network stack and load the client‐side overhead into response delay. Therefore, we propose a high‐performance monitor system named HCMonitor, which computes the server‐side response latency and the round‐trip time of per‐request. It can afford full traffic monitoring on the basis of userspace, “zero copy” and pipeline. By switch mirroring, the measured latency eliminates the kernel network stack overhead and the queuing delay of the client‐side. Such measurement results in improved accuracy, online analysis, anomaly detection, real‐time display and transparent to network services. Our evaluations show HCMonitor obtains a higher throughput compared with tcpdump by over 200 times. Compared with wrk, the tail latency accuracy shows an increase by up to 72%–76% in high concurrent networks. Ke Liu 0004, Yifan Shen 0002, Mingyu Chen 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | LSP: Collective Cross-Page Prefetching for NVMabstractAs an emerging technique, non-volatile memory (NVM) provides valuable opportunities for boosting the memory system, which is vital for the computing system performance. However, one challenge preventing NVM from replacing DRAM as the main memory is that NVM row activation's latency is much longer (by approximately 10x) than that of DRAM. To address this issue, we present a collective cross-page prefetching scheme that can accurately open an NVM row in advance and then prefetch the data blocks from the opened row with low overhead. We identify a memory access pattern (referred to as a ladder stream) to facilitate prefetching that can cross page boundary, and propose the ladder stream prefetcher (LSP) for NVM. In LSP, two crucial components have been well designed. Collective Prefetch Table is proposed to reduce the interference with demand requests caused by prefetching through speculatively scheduling the prefetching according to the states of the memory queue. It is implemented with low overhead by using single entry to track multiple prefetches. Memory Mapping Table is proposed to accurately prefetch future pages by maintaining the mapping between physical and virtual addresses. Experimental evaluations show that LSP improves the memory system performance with no prefetching by 66%, and the improvement over the state-of-the-art prefetchers, Access Map Pattern Matching Prefetcher (AMPM), Best-Offset Prefetcher (BOP) and Signature Path Prefetcher (SPP) is 26.6%. 21.7% and 27.4%. respectively. Haiyang Pan, Yuhang Liu 0001, Tianyue Lu, Mingyu Chen 0001 |
DATE | 4 |
| 2021 | EdUCAS: An In-house CI/CD Platform with Cloud FPGAs for Agilely Conducting Computer Systems Course ProjectsabstractIn recent years, there has been a rapidly growing recognition of the importance of conducting hands-on HW-SW co-design labs with real hardware (e.g., programmable logic chips named FPGAs) while studying computing curricula, especially the computer systems (CSys) courses. However, using FPGA is quite atime-consuming and error-prone process for students, and manipulating FPGA development tools and boards also distracts students and instructors. To overcome these obstacles and improve agility, we introduce an in-house platform named EdUCAS, in combination with the previously designed cloud FPGA servers for students to automatically conduct computer systems course projects. EdUCAS aims at enabling students to concentrate on their logic designs using hardware description language (e.g., Verilog HDL), without wasting useless time in FPGA tools and experimental environment. Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002 |
ITiCSE (2) | 5 |
| 2020 | Freeway: an order-less user-space framework for non-real-time applicationsabstractThe demand for high network capacity has been rapidly increasing, such as inter-DC WANs. But the transport over the network with high bandwidth and delay cannot fully utilize the bandwidth because of the inevitable packet loss and the flow control bottleneck caused by the out-of-order data blocked in the receive buffer. We further found that a lot of applications over inter-DC WANs are non-real-time, which are insensitive to data arriving sequence. Thus, we design and implement Freeway, a user-space bulk-data network transfer framework, to improve the bandwidth of these non-real-time applications over inter-DC WANs. Experimental results show that Freeway achieves 100% more bandwidth utilization than Linux TCP stack, and significantly reduces memory cost. Yifan Shen 0002, Ke Liu 0004, Ziting Guo, Vaneet Aggarwal, Mingyu Chen 0001 |
CF | 7 |
| 2020 | Labeled Network Stack: A High-Concurrency and Low-Tail Latency Cloud Server Framework for Massive IoT Devices
Ke Liu 0004, Yifan Shen 0002, Yazhu Lan, Mingyu Chen 0001, Yuan-Fei Chen |
J. Comput. Sci. Technol. | 6 |
| 2020 | IMPULP: A Hardware Approach for In-Process Memory Protection via User-Level Partitioning
Mingyu Chen 0001, Yuhang Liu 0001, Zong-Hao Yang, Zonghui Hong, Yunge Guo |
J. Comput. Sci. Technol. | 2 |
| 2020 | Optimizing TCP Loss Recovery Performance Over Mobile Data NetworksabstractRecent advances in high-speed mobile networks have revealed new bottlenecks in ubiquitous TCP protocol deployed in the Internet. In addition to differentiating non-congestive loss from congestive loss, our experiments revealed two significant performance bottlenecks during the loss recovery phase: flow control bottleneck and application stall, resulting in degradation in QoS performance. To tackle these two problems, we first develop a novel opportunistic retransmission algorithm to eliminate the flow control bottleneck, which enables TCP sender to transmit new packets even if receiver's receiving window is exhausted. Second, application stall can be significantly alleviated by carefully monitoring and tuning the TCP sending buffer growth mechanism. We implemented and modularized the proposed algorithms in the Linux kernel thus they can plug-and-play with the existing TCP loss recovery algorithms easily. We evaluated our proposed algorithms over emulated and real experiments and showed that, compared to existing TCP loss recovery algorithms, the proposed optimization algorithms improve the bandwidth efficiency by up to 133 percent and completely mitigate RTT spikes, i.e., over 50 percent RTT reduction, over the loss recovery phase. Ke Liu 0004, Zhongbin Zha, Wenkai Wan, Vaneet Aggarwal, Binzhang Fu, Mingyu Chen 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2019 | Engaging Heterogeneous FPGAs in the CloudabstractFPGA has become an essential infrastructural component in commercial cloud and datacenter for improving system performance and efficiency. Meanwhile, a heterogeneous FPGA chip (Hetero-FPGA) in which a multi-core System-on-Chip (SoC) is tightly integrated with an FPGA fabric has been successfully pioneered. Given its hardware-software co-programmability, Hetero-FPGA is supposed to become an independent and first-class cloud computing resource with networking capabilities in order to avoid involving brawny commodity x86 servers as carriers for FPGA fabrics which are usually the cases in current commercial FPGA clouds from several web vendors. Following this design paradigm, we present HeFA, a self-contained Hetero-FPGA Array architecture in cloud. We construct a high-level hardware template as well as a software stack for the Hetero-FPGA node, enabling the SoC as a primary engine to manage, coordinate and incorporate with the dominant FPGA fabric. We also propose a fully scripted design flow to make HeFA as an easy-to-use cloud infrastructure. Based on these techniques, we implement an academia prototype chassis of HeFA that includes 32 Hetero-FPGA nodes with Xilinx's Zynq MPSoC chips. By a customized cloud resource manager, the prototype is flexibly provisioned as either 32 individual FPGA nodes or multiple scalable sub-clusters to abstract arbitrary volume of reconfigurable fabrics as on-demand cloud services. In this manner, a versatile research and educational platform is delivered for agile hardware-software co-design in scenarios such as domain-specific accelerator development, open instruction set architecture-based chip design, computer system-related experimental project, and so on. Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002 |
FPGA | 3 |
| 2019 | Make Page Coloring more Efficient on Slice-Based Three-Level CacheabstractOn modern multi-core machines, page coloring has been used to alleviate the competition at Last Level Cache (LLC). However, the latest development of CPU architecture has brought new issues to page coloring. Firstly, in the case of three-level cache, previous works about page coloring did not discuss the impact on L2 cache of color allocation and the competition for L2 cache is not considered concurrently under hyper-threading. In addition, as the last level cache structure is changed from shared to slice-based and undocumented hash function is applied, page coloring is more complex and slice information is also not fully utilized. This paper presents solutions to these issues. Firstly, by making small changes to the traditional page coloring, the problem that page coloring may waste L2 cache is alleviated. At the same time, we rethink the vertical allocation of L2 cache and LLC in page coloring under hyper-threading, and discuss the impact of color allocation on programs, especially those with different sensitivity to L2 cache and LLC. Finally, we make full use of slice information and propose Partial Conflict Color (PCC). At the same time, we also propose a fast method to obtain PCC. Experiments show that using PCC can improve system performance when the number of colors is insufficient. Tianyue Lu, Yuhang Liu 0001, Mingyu Chen 0001 |
ICPADS | 4 |
| 2019 | HCMonitor: An Accurate Measurement System for High Concurrent Network ServicesabstractAs user-interactive services grow explosively in datacenters, latency has become one of the most deciding factors on user experience. Therefore, estimating the latency and detecting anomalies from the expected latency is essential to evaluate services' performance. Although many existing tools have been used widely, their estimation methods can be divided into two categories. First, the traffic-sample-based approaches sample the network traffic for accelerating the estimation rather than measure every response time. Second, the full-traffic-based approaches, such as tcpdump and wrk, analyze data from kernel and leave the latency computation to the client-side. In this paper, we attempt to compute the applications' server-side latency for every request in real-time, and eliminate kernel processing delay. We propose a system named HCMonitor. It monitors all the traffic by switch mirroring, which results in high throughput and more accuracy in server-side latency estimation. The latency measurement is transparent to network services and can be displayed in real time. Our evaluations show HCMonitor obtains higher throughput than tcpdump by over 1000 times. Compared to wrk, the tail latency accuracy estimated by HCMonitor shows a promotion by up to 72%~76% in high concurrent network, by eliminating delay produced by packet transfer, kernel network stack and packets queuing on client side. Ke Liu 0004, Yifan Shen 0002, Mingyu Chen 0001 |
NAS | 5 |
| 2019 | HCMA: Supporting High Concurrency of Memory Accesses with Scratchpad Memory in FPGAsabstractCurrently many researches focus on new methods of accelerating memory accesses between memory controller and memory modules. However, the absence of an accelerator for memory accesses between CPU and memory controller wastes the performance benefits of new methods. Therefore, we propose a coordinated batch method to support high concurrency of memory accesses (HCMA). Compared to the conventional method of holding outstanding memory access requests in miss status handling registers (MSHRs), HCMA method takes advantage of scratchpad memory in FPGAs or SoCs to circumvent the limitation of MSHR entries. The concurrency of requests is only limited by the capacity of scratchpad memory. Moreover, to avoid the higher latency when searching more entries, we design an efficient coordinating mechanism based on circular queues.We evaluate the performance of HCMA method on an MP-SoC FPGA platform. Compared to conventional methods based on MSHRs, HCMA method supports ten times of concurrent memory accesses (from 10 to 128 entries on our evaluation platform). HCMA method achieves up to 2.72× memory bandwidth utilization for applications that access memory with massive fine-grained random requests, and to 3.46× memory bandwidth utilization for stream-based memory accesses. For real applications like CG, our method improves speedup performance by 29.87%. Yuhang Liu 0001, Mingyu Chen 0001 |
NAS | 4 |
| 2019 | Computer Organization and Design Course with FPGA CloudabstractComputer Organization and Design (COD) is a fundamentally required early-stage undergraduate course in most computer science and engineering curricula. During the two sessions (lecture and project part) of one COD course, educational platforms play an important role in cultivating students' computational thinking, especially the ability of viewing the hardware and software in a computer system as a whole (computer system thinking ability for short in this paper). In order to improve teaching quality, in this paper, we discuss the deployment of an inexpensive in-house Field Programmable Gate Array (FPGA) cloud platform, which can provide students with hardware-software co-design methodology and practice. The platform includes 32 FPGA nodes and the scale can be dynamically changed. Each cloud node is heterogeneously composed of an ARM processor and a tightly-coupled reconfigurable fabric to provide students with hands-on hardware and software programming experiences. We illustrate our efforts to make the FPGA cloud as an easy-to-use resource pool to elastically support a class with 92 undergrads via Internet access and to monitor students' experimental behaviors. We also present key insights in our teaching activities that indicate such appliance is feasible to provide practice of both basic principles and emerging co-design techniques for students. We believe that our cost-effective FPGA cloud is of significant interests to educators looking forward to improving computer system-related courses. Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002 |
SIGCSE | 3 |
| 2019 | RAGuard: An Efficient and User-Transparent Hardware Mechanism against ROP AttacksabstractControl-flow integrity (CFI) is a general method for preventing code-reuse attacks, which utilize benign code sequences to achieve arbitrary code execution. CFI ensures that the execution of a program follows the edges of its predefined static Control-Flow Graph: any deviation that constitutes a CFI violation terminates the application. Despite decades of research effort, there are still several implementation challenges in efficiently protecting the control flow of function returns (Return-Oriented Programming attacks). The set of valid return addresses of frequently called functions can be large and thus an attacker could bend the backward-edge CFI by modifying an indirect branch target to another within the valid return set. This article proposes RAGuard, an efficient and user-transparent hardware-based approach to prevent Return-Oreiented Programming attacks. RAGuard binds a message authentication code (MAC) to each return address to protect its integrity. To guarantee the security of the MAC and reduce runtime overhead: RAGuard (1) computes the MAC by encrypting the signature of a return address with AES-128, (2) develops a key management module based on a Physical Unclonable Function (PUF) and a True Random Number Generator (TRNG), and (3) uses a dedicated register to reduce MACs’ load and store operations of leaf functions. We have evaluated our mechanism based on the open-source LEON3 processor and the results show that RAGuard incurs acceptable performance overhead and occupies reasonable area. Rui Hou 0001, Wei Song 0002, Sally A. McKee, Zhen Jia 0001, Chen Zheng 0001, Mingyu Chen 0001, Lixin Zhang 0002, Dan Meng 0002 |
ACM Trans. Archit. Code Optim. | 7 |
| 2018 | Labeled Network Stack: A Co-designed Stack for Low Tail-Latency and High Concurrency in Datacenter Services
Ke Liu 0004, Lan Yu, Mingyu Chen 0001 |
NPC | 5 |
| 2018 | ZyForce: An FPGA-based Cloud Platform for Experimental Curriculum of Computer System in University of Chinese Academy of Sciences (Abstract Only)abstractTo cultivate students" capability of computer system thinking and software/hardware programming, experimental curriculum of computer system is regarded as one of the most effective methods. Some universities have set up hardware labs equipped with several or dozens of FPGA (Field Programmable Gate Array) boards for these courses. However, these lab kits are always in a relatively low utilization rate and how the students" capability is improved by these assets is not easy to be evaluated. Inspired by the merits of FPGA public cloud (e.g. Amazon AWS F1 instance), an in-house-designed FPGA-based online cloud platform (named ZyForce) is proposed to deploy in UCAS. This platform is equipped with 40 custom designed boards using Xilinx Zynq UltraScale+ MPSoC FPGAs and the utilization rate of these education resources is boosted by means of advanced cloud computing technology. With ZyForce, students remotely carry out lab assignments (e.g. MIPS, RISC-V or domain-specific architecture processor design with cache/memory, DMA, accelerator and performance counter) as using local FPGA boards; instructors can analyze the downloaded operation log file for each student and know how these kits are being used. It"s believed that this kind of online hardware lab appliances provides a novel pay-as-you-go service model for those universities in remote regions who cannot afford to set up their own hardware laboratories, and also facilitates our students, the future scientists and engineers, with this promising cloud development approach. Ke Zhang 0017, Mingyu Chen 0001, Yungang Bao |
SIGCSE | 2 |
| 2018 | PTAT: An Efficient and Precise Tool for Tracing and Profiling Detailed TLB MissesabstractAs the memory access footprints of applications in areas like data analytics increase, the latency overhead of translation lookaside buffer (TLB) misses increases. Thus, the efficiency of TLB becomes increasingly critical for overall system performance. Analyzing TLB miss traces is useful for hardware architecture design and software application optimization. Utilizing cycle-accurate simulators or instrumentation tools is very time-consuming and/or inaccurate for tracing and profiling TLB misses. In this article, we propose an efficient and precise tool to collect and profile last-level TLB misses. This tool utilizes a novel software method called Page Table Access Tracing (PTAT), storing last-level page table entries of certain workload processes into a reserved uncached memory region. Therefore, each last-level TLB miss incurred by user process corresponds to one uncached page table access to main memory, which can be captured and recorded by a hardware memory bus monitor. The detected information is then dumped into offline storage. In this manner, full TLB miss traces are collected and can be analyzed flexibly. Compared to previous software-based methods, this method achieves higher performance. Experiments show that, compared with a state-of-the-art kernel instrumentation method (BadgerTrap), which lacks complete dumping trace function, the speedup is still up to 3.88-fold for memory-intensive benchmarks. Due to the improved efficiency and completeness of tracing, case studies validate that more flexible profiling can be conducted, which is of great significance for TLB performance optimization. The accuracy of PTAT is verified by both dedicated sequence and performance counters. Jiutian Zhang, Yuhang Liu 0001, Mingyu Chen 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | SMEFF: A scalable memory extension fabric for FPGAabstractIn resource-constrained FPGA systems, off-chip memory plays an important role in both prototype verification and acceleration systems for big data. As the scale of applications become increasingly large and complex, the data to be processed grows exponentially. In contrast, FPGAs provide limited memory capacity and bandwidth, severely limiting the scale and performance of prototype verification systems and acceleration systems. Furthermore, data movement is expected to be a dominant consumer of energy, thus inefficient data movement between different DRAM modules also incurs significant performance and energy penalties. This paper proposes a practical design: A Scalable Memory Extension Fabric for FPGA (SMEFF), which is an asynchronous memory access mechanism and exploits cascaded technology to solve the problem of memory capacity and bandwidth. SMEFF uses two key technologies to achieve memory capacity and bandwidth improvements, and shrink the latency and overhead of data movement-the first is an FPGA-based high-speed serial bus to build a multi-level memory fabric instead of the traditional parallel bus mechanism to solve the signal integrity problem. The second is a module to module (M-To-M) DMA data movement technology, which reduces the latency and overhead of data movement between memory modules. We implement SMEFF on an FPGA-based prototype to demonstrate the feasibility of our approach. Experimental results show that SMEFF provides 5x memory capacity increase and up to 3.6x memory bandwidth improvement compared to state-of-the-art FPGA-based memory systems, and outperforms PCIe-based systems. The data movement of M-TO-M's DMA technology obtains up to 3x latency reduction, and average of 21.1% to 61.1% energy reduction compared to state-of-the-art FPGA-based memory systems. SMEFF thus increases FPGA-based prototype memory systems capacity and bandwidth. More importantly, our architecture provides opportunities for the design of scalable, cost-effective FPGA-based memory subsystems. Yuhang Liu 0001, Mingyu Chen 0001 |
FPT | 4 |
| 2017 | TDV Cache: Organizing Off-Chip DRAM Cache of NVMM from a Fusion PerspectiveabstractEmerging Non-Volatile Memory (NVM) provides both larger memory capacity and higher energy efficiency, but has much longer access latency than traditional DRAM, thus DRAM can be used as an efficient cache to hide the long latency of Non-Volatile Main Memory (NVMM) system. Transparent Off-chip DRAM cache (TOD cache) is a new DRAM cache structure where off-chip DRAM module is used as L4 cache and managed by hardware. The capacity and latency ratio of TOD cache over NVM are both quite different from those of traditional on-chip SRAM or die-stacked DRAM cache over off-chip DRAM memory. All the factors including hit latency, miss latency and hit rate need to be re-considered for TOD cache design. In this study, we first point out that three types of traditional cache schemes cannot be used directly for TOD cache, since set-associative cache suffers from extra tag lookup latency, direct-mapped cache has low hit rate and tag cache is too small to efficiently hold the working sets of tags for DRAM cache. Based on these observations, we propose a novel cache scheme, TDV, that fuses these three different types of cache together to take their advantages. In TDV, a direct-mapped cache is used as the first-level cache to achieve short access latency, a set-associative victim cache is taken as the second-level cache to obtain extra high hit rate, and a SRAM tag cache only serves for the victim cache rather than the whole DRAM cache and thus improves the hit rate of tag cache significantly. The simulation results show that, TDV cache has a performance improvement of 6.3% and 8.3% on average than state-of-the-art direct-mapped (Alloy cache) and set-associative cache (ATCache) with same DRAM and SRAM capacity. Tianyue Lu, Yuhang Liu 0001, Haiyang Pan, Mingyu Chen 0001 |
ICCD | 4 |
| 2017 | PTAT: An efficient and precise tool for collecting detailed TLB miss tracesabstractIt is well known that the TLB performance impacts the memory system performance, which is critical for overall system performance. Similar to multi-level caches, multilevel TLBs have become an important leverage for boosting data access performance. Applications have increasingly large working sets. Servers targeting such applications have thus been built with ever larger main memory capacities, but there has been no commensurate growth in TLB sizes. Designing high performance and energy efficient memory hierarchies require insight into the behavior of current designs: when do they work well, and when do they fall short of expectations. Profiling the TLB misses is the prerequisite for memory system optimization. Both designing efficient TLB architecture and TLB-friendly applications require analysis of TLB miss behavior. Although researchers have extensively studied TLB behavior, current approaches have some issues in either efficiency or precision. Jiutian Zhang, Yuhang Liu 0001, Yuan Ruan, Mingyu Chen 0001 |
ISPASS | 5 |
| 2017 | Stem: A Table-Based Congestion Control Framework for Virtualized Data Center Networks
Binzhang Fu, Mingyu Chen 0001 |
NPC | 3 |
| 2017 | Efficient Regional Congestion Awareness (ERCA) for Load Balance with Aggregated Congestion InformationabstractIn this paper, we propose Efficient Regional Congestion Awareness (ERCA), a novel adaptive routing technique which utilizes both local and non-local/aggregated congestion information to estimate network congestion. To reduce the interference caused by the noises of the congestion information, two optimizations are performed: first, to minimize the noises of the congestion information, ERCA only considers the congestion status of the links adjacent to the nodes defined from current to the corresponding boundary, second, to minimize the influence of the noises, ERCA exploits dynamic instead of static weights to evaluate port's congestion. Furthermore, ERCA uses one wire per dimension direction or quadrant to transmit congestion information between two adjacent nodes, so the wiring overhead of ERCA is minimal. Compared with the DBAR, ERCA can maximally improve the saturation throughput by 11.7% and averagely improve the saturation throughput by 6.02%. Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
PDP | 4 |
| 2017 | Understanding the GPU Microarchitecture to Achieve Bare-Metal Performance TuningabstractIn this paper, we present a methodology to understand GPU microarchitectural features and improve performance for compute-intensive kernels. The methodology relies on a reverse engineering approach to crack the GPU ISA encodings in order to build a GPU assembler. An assembly microbenchmark suite correlates microarchitectural features with their performance factors to uncover instruction-level and memory hierarchy preferences. We use SGEMM as a running example to show the ways to achieve bare-metal performance tuning. The performance boost is achieved by tuning FFMA throughput by activating dual-issue, eliminating register bank conflicts, adding non-FFMA instructions with little penalty, and choosing proper width of global/shared load instructions. On NVIDIA Kepler K20m, we develop a faster SGEMM with 3.1Tflop/s performance and 88% efficiency; the performance is 15% higher than cuBLAS7.0. Applying these optimizations to convolution, the implementation gains 39%-62% performance improvement compared with cuDNN4.0. The toolchain is an attempt to automatically crack different GPU ISA encodings and build an assembler adaptively for the purpose of performance enhancements to applications on GPUs. Xiuxia Zhang, Guangming Tan, Shuangbai Xue, Jiajia Li 0001, Keren Zhou 0001, Mingyu Chen 0001 |
PPoPP | 6 |
| 2017 | Joint Upload-Download TCP Acceleration over Mobile Data NetworksabstractUpload and download traffic often coexist in mobile networks. However, TCP download throughput could be substantially degraded by upload traffic even if the downlink is not the bottleneck. Previous works such as RSFC and TCP-RRE can substantially improve TCP download throughput in the presence of concurrent TCP upload flows, albeit at the expense of significantly degraded upload throughput performance. This work addresses this limitation by developing a novel Aggregate Transmission Rate Controller with Upload and Download flows aggregations (ATRC-UD) to jointly accelerate concurrent TCP upload and download flows from/to the same mobile device. The insight is that existing TCP as well as other flow-based approaches all suffer from ACK packets delayed by data packets from TCP flows in the opposite direction, resulting in significant errors in bandwidth estimation. By contrast, ATRC-UD exploits data packets of the opposite direction to enable continuously estimation of the downlink bandwidth and queueing delay even when ACK packets are significantly delayed. This allows ATRC-UD to track the bandwidth and delay variations more closely to maintain a shorter queue length at the downlink, thus jointly improve the download-upload throughput. Extensive emulated and real-world experiments showed that ATRC-UD enables TCP to achieve 96% downlink bandwidth utilization while improving uplink bandwidth utilization by over 115% compared to existing approaches, such as TCP-RRE and RSFC. Ke Liu 0004, Vaneet Aggarwal, Ziyu Shao, Mingyu Chen 0001 |
SECON | 4 |
| 2017 | HAP: Hybrid-Memory-Aware Partition in Shared Last-Level CacheabstractData-center servers benefit from large-capacity memory systems to run multiple processes simultaneously. Hybrid DRAM-NVM memory is attractive for increasing memory capacity by exploiting the scalability of Non-Volatile Memory (NVM). However, current LLC policies are unaware of hybrid memory. Cache misses to NVM introduce high cost due to long NVM latency. Moreover, evicting dirty NVM data suffer from long write latency. We propose hybrid memory aware cache partitioning to dynamically adjust cache spaces and give NVM dirty data more chances to reside in LLC. Experimental results show Hybrid-memory-Aware Partition (HAP) improves performance by 46.7% and reduces energy consumption by 21.9% on average against LRU management. Moreover, HAP averagely improves performance by 9.3% and reduces energy consumption by 6.4% against a state-of-the-art cache mechanism. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2016 | sAXI: A High-Efficient Hardware Inter-Node Link in ARM Server for Remote Memory AccessabstractThe ever-growing need for fast big-data operations has made in-memory processing increasingly important in modern datacenters. To mitigate the capacity limitation of a single server node, techniques of inner-rack cross-node memory access have drawn attention recently. However, existing proposals exhibit inefficiency in remote memory access among server nodes due to inter-protocol conversions and non-transparent coarse-grained accesses. In this study, we propose the high-performance and efficient serialized AXI (sAXI) link and its associated cross-node memory access mechanism for emerging ARM-based servers. The key idea behind sAXI is directly extending the on-chip AMBA AXI-4.0 interconnection of the SoC in a local server node to the outside, and then bringing into remote server nodes via high-speed serial lanes. As a result, natively accessing remote memory in adjacent nodes in the same manner of local assets is supported by purely using existing software. Experimental results show that, using the sAXI data-path, performance of remote memory access in the user-level micro-benchmark is very promising (min. latency: 1.16µs, max. bandwidth: 1.52GB/s on our in-house FPGA prototype). In addition, through this efficient hardware inter-node link, performance of an in-memory key-value framework, Redis, can be improved up to 1.72x and large latency overhead of database query can be effectively hidden. Ke Zhang 0017, Yisong Chang, Lixin Zhang 0002, Mingyu Chen 0001, Zhiwei Xu 0002 |
CCGrid | 4 |
| 2016 | Intra-host Rate Control with Centralized ApproachabstractToday's datacenter is shared among various applications with different QoS requirements, which poses a great challenge to deliver low delay transport with high throughput. Most of works address this challenge by reducing the in-network delay, but assumes a negligible local delay. However, we show that this assumption does not hold for a multi-tenant datacenter that a physical machine is shared by multiple tenants with virtual machines running different applications. As measured, we found that VMs in a PM competing for bandwidth resources introduce delays as high as 13 ms, resulted from the packet queueing at QDisc layer of that PM, because current VMs' rate control still operates in a distributed manner without exploiting knowledge of the QoS requirements of applications running in VMs. This work addresses this problem by proposing a centralized rate adaptation (CERA) that operates in the host PM, dynamically schedules the flows from all VMs in a centralized manner. We implemented a CERA prototype and evaluated CERA through testbed experiments. Our results show that CERA reduces the local delay significantly thus reduces the average request latency of delay sensitive applications, e.g., memcached, by a factor of 6.3, without sacrificing the throughput performance of throughput intensive applications, e.g., iperf. Ke Liu 0004, Yifan Shen 0002, Jack Y. B. Lee, Mingyu Chen 0001, Lixin Zhang 0002 |
CLUSTER | 5 |
| 2016 | Extending On-chip Interconnects for rack-level remote resource accessabstractThe need to perform data analytics on exploding data volumes coupled with the rapidly changing workloads in cloud computing places great pressure on data-center servers. To improve hardware resource utilization across servers within a rack, we propose Direct Extension of On-chip Interconnects (DEOI), a high-performance and efficient architecture for remote resource access among server nodes. DEOI extends an SoC server node's on-chip interconnect to access resources in adjacent nodes with no protocol changes, allowing remote memory and network resources to be used as if they were local. Our results on a four-node FPGA prototype show that the latency of user-level, cross-node, random reads to DEOI-connected remote memory is as low as 1.16µs, which beats current commercial technologies. We exploit DEOI remote access to improve performance of the Redis in-memory key-value framework by 47%. When using DEOI to access remote network resources, we observe an 8.4% average performance degradation and only a 2.52µs ping-pong latency disparity compared to using local assets. These results suggest that DEOI can be a promising mechanism for increasing both performance and efficiency in next-generation data-center servers. Yisong Chang, Ke Zhang 0017, Sally A. McKee, Lixin Zhang 0002, Mingyu Chen 0001, Liqiang Ren, Zhiwei Xu 0002 |
ICCD | 5 |
| 2016 | Isolating bandwidth guarantees from work conservation in the cloudabstractTo predict lower bounds on the performance of applications, the cloud should provide guarantees on bandwidth that each virtual machine can obtain. By competing for spare network bandwidth, current solutions that provide bandwidth guarantees can achieve work conservation as well. However, they usually fail to provide accurate bandwidth guarantees, for the interference between traffic for achieving the two objectives respectively. In order to eliminate the interference, they reserve sufficient bandwidth headroom for every link, which cannot be allocated to tenants as guarantees, incurring a decrease in the total of guarantees that each link can offer and thus a decline in the revenue of the cloud provider. To address the problem, this paper proposes a new mechanism, namely DFlow, which achieves bandwidth guarantees and work conservation simultaneously. Specifically, DFlow isolates its solutions for achieving bandwidth guarantees and work conservation from each other by splitting every flow into two subflows with distinct priorities. They are then used to achieve the two objectives respectively. Our evaluations show that DFlow can provide accurate bandwidth guarantees without reserving any bandwidth headroom while achieving work conservation to effectively utilize spare network bandwidth. Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
ISCC | 4 |
| 2016 | Adaptive rate control over mobile data networks with heuristic rate compensationsabstractMobile data networks exhibit highly variable data rates and stochastic non-congestion-related packet loss. These challenges result in key performance bottlenecks in current Transmission Control Protocol (TCP) implementations: bandwidth inefficiency and large end-to-end delay. This work addresses these challenges by first developing a Sliding Interval based Rate Adaptation (SIRA) that tracks bandwidths with a fixed time interval and applies them to its transmission rate periodically. Extensive experiments confirmed that SIRA achieves 96.3% bandwidth utilization and reduces the average queueing delay by a factor of 1.37, compared to TCP CUBIC, the preferred variant for Internet servers. However, the resultant end-to-end delay is still much larger for interactive applications, thus we complement SIRA with two heuristic rate compensation algorithms (SIRA-H) given that the bandwidth does not vary significantly in long time scales. Specifically, SIRA-H first reduces the transmission rate of SIRA if the estimated RTT is above a prefigured threshold. Meanwhile, it computes the amount of unsent data that would be transmitted if SIRA were used, and compensates the rate reduction with those unsent data as if their ACKs were received, when the queue is detected to be empty. We evaluated SIRA-H through a combination of trace-driven emulations and real-world experiments, and showed that it reduces the 95thpercentile queueing delay by a factor of over 3.9, while maintains a similar throughput compared to the original SIRA. In comparison to state of the art protocols such as Sprout and Verus, SIRA-H also reduces the 95thpercentile queueing delay by a factor of over 0.8. Ke Liu 0004, Jack Y. B. Lee, Mingyu Chen 0001, Lixin Zhang 0002 |
IWQoS | 4 |
| 2015 | Exploiting Program Semantics to Place Data in Hybrid MemoryabstractLarge-memory applications like data analytics and graph processing benefit from extended memory hierarchies, and hybrid DRAM/NVM (non-volatile memory) systems represent an attractive means by which to increase capacity at reasonable performance/energy tradeoffs. Compared to DRAM, NVMs generally have longer latencies and higher energies for writes, which makes careful data placement essential for efficient system operation. Data placement strategies that resort to monitoring all data accesses and migrating objects to dynamically adjust data locations incur high monitoring overhead and unnecessary memory copies due to mispredicted migrations. We find that program semantics (specifically, global access characteristics) can effectively guide initial data placement with respect to memory types, which, in turn, makes run-time migration more efficient. We study a combined offline/online placement scheme that uses access profiling information to place objects statically and then selectively monitors run-time behaviors to optimize placements dynamically. We present a software/hardware cooperative framework, 2PP, and evaluate it with respect to state-of-the-art migratory placement, finding that it improves performance by an average of 12.1%. Furthermore, 2PP improves energy efficiency by up to 51.8%, and by an average of 18.4%. It does so by reducing run-time monitoring and migration overheads. Dejun Jiang 0001, Sally A. McKee, Jin Xiong, Mingyu Chen 0001 |
PACT | 5 |
| 2015 | Improving Memory Access Performance of In-Memory Key-Value Store Using Data Prefetching Techniques
Guangyu Sun 0003, Peng Wang 0025, Mingyu Chen 0001 |
APPT | 4 |
| 2015 | A Reliable Distributed Convolutional Neural Network for Biology Image SegmentationabstractMany modern advanced biology experiments are carried on by Electron Microscope(EM) image analysis. Segmentation is one of the most important and complex steps in the process of image analysis. Previous ISBI contest results and related research show that Convolution Neural Network(CNN)has high classification accuracy in EM image segmentation. Besides it eliminates the pain of extracting complex features which's indispensable for traditional classification algorithms. HoweverCNN's extremely time-consuming and fault vulnerability due to long time execution prevent it from being widely used in practice. In this paper, we try to address these problems by providing reliable high performance CNN framework for medial image segmentation. Our CNN has light weighted user level checkpoint, which costs seconds when doing one checkpoint and restart. On the fact of lacking in platform diversity in current parallel CNN framework, our CNN system tries to make it general by providing distributed cross-platform parallelism implementation. Currently we have integrated Theano's GPU implementation in our CNNsystem, and we explore parallelism potential on multi-core CPUs and many-core Intel Phi by testing performance of main kernel functions of CNN. In the future, we will integrate implementation son other two platforms into our CNN framework. Xiuxia Zhang, Guangming Tan, Mingyu Chen 0001 |
CCGRID | 3 |
| 2015 | Optimizing TCP loss recovery performance over mobile data networksabstractRecent advances in high-speed mobile networks have revealed new bottlenecks in ubiquitous TCP protocol deployed in the Internet. In addition to differentiating non-congestive loss from congestive loss, our experiments revealed two significant performance bottlenecks during loss recovery phase: flow control bottleneck and application stall, resulting in degradation in QoS performance. To tackle these two problems we firstly develop a novel opportunistic retransmission algorithm to eliminate the flow control bottleneck, which enables TCP sender to transmit new packets even if receiver's receiving window is exhausted. Secondly, the application stall can be significantly alleviated by carefully monitoring and tuning the TCP send buffer growth mechanism. We implemented and modularized the proposed algorithms in the Linux kernel thus they can plug-and-play with the existing TCP loss recovery algorithms easily. Using emulated experiments we showed that with the proposed optimization techniques the existing loss recovery algorithms can at most achieve 98.3% bandwidth utilization during loss recovery phase, and reduce RTT by at most 80% after loss recovery phase. Zhongbin Zha, Ke Liu 0004, Binzhang Fu, Mingyu Chen 0001 |
SECON | 4 |
| 2014 | A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring OptimizationsabstractWe propose a novel cache set index scheme called SWAP (swap-based cache set index). SWAP introduces a pseudo-physical address space that is used by the operating system. The real physical address used for cache and main memory access is obtained by simply swapping some of superpage number bits with cache set index bits from the pseudo-physical address. By adding a level of indirection to the physical memory management, we simultaneously support both page coloring and superpage optimizations. These work together to improve TLB and shared LLC performance with negligible cost. Our results show that SWAP can improve performance by an average of 15.1% (by up to 25.2%) compared to 7.34% and 8.26% for superpage and page coloring, respectively. Zehan Cui, Licheng Chen, Yungang Bao, Mingyu Chen 0001 |
DAC | 4 |
| 2014 | Achieving efficient packet-based memory system by exploiting correlation of memory requestsabstractPacket-based interface is a trend for future memory system to alleviate memory capacity and bandwidth bottlenecks. On the other hand fine-grained memory access has been proven to efficiently reduce memory power. However leveraging both these two technologies will result in high packet overhead, because previous implementations of packet-based interface all adopt a simple design that a single packet is dedicated to a single request (SPSR). In this paper, we propose three optimizations to overcome the problem by exploiting correlations of memory requests. First, we propose a novel single packet multiple requests (SPMR) interface that encapsulates multiple requests into a packet to share packet header and tail. Second, we propose an adaptive address compression mechanism within a packet by adopting a base-difference algorithm. Third, we propose a mechanism to merge multiple memory requests with continuous access addresses into a single request before packing. By this way, the granularity constraint of cache line size is broken to enable efficiently row buffer scheduling. The experimental results show that, for certain memory-intensive workloads, the optimizations can effectively reduce packet overhead by about 53.9% and improve system performance by about 63.6% in average. Tianyue Lu, Licheng Chen, Mingyu Chen 0001 |
DATE | 3 |
| 2014 | Dandelion: A locally-high-performance and globally-high-scalability hierarchical data center networkabstractThe increasing customer demand is driving modern data centers to embrace the freely-expandable network architecture. Unfortunately, state-of-the-art freely-expandable networks suffer from either the large granularity of expansion or the prohibitive implementation cost. Furthermore, a recent research showed that data center traffic tends to be highly clustered. Based on above observations, this paper proposes a freely-expandable network architecture, namely the dandelion. Dandelion is a two-level hierarchical network, where the first level aims at “high performance” and the second level aims at “high scalability”. The resulting network has two distinct advantages. First, it could arbitrarily expand with a reasonable granularity. Second, the router architecture is efficient as well as highly scalable since 1) the routing table is significantly compressed and 2) a fixed number of virtual channels per physical channel are required regardless of the network size. Finally, the traffic characteristics of four typical cloud applications are analyzed, and the generated traffic patterns are used to evaluate the proposed network architecture. Simulation results prove that the dandelion is a promising network architecture for future data centers. Binzhang Fu, Wentao Bao, Guolong Jiang, Mingyu Chen 0001, Lixin Zhang 0002, Yidong Tao, Junfeng Zhao 0003 |
ICCCN | 5 |
| 2014 | HAP: Hybrid-memory-Aware Partition in shared Last-Level CacheabstractData-center servers require large capacity main memory to run multiple workloads simultaneously. However, the scalability and power consumption of DRAM limit its capability of constructing large capacity memory. Emerging non-volatile memories (e.g. PCM and STT-RAM) provide better scalability and lower power leakage than DRAM. Especially, hybrid memory consisting of DRAM and NVM is able to exploit advantages of different memory medias. However, NVMs have a few drawbacks, such as relatively longer read and write latency. Cache miss at the shared last level cache (LLC) suffers from longer latency if the missing data resides in NVM. Current LLC policies manage the cache space without being aware of the underlying heterogeneous medias. This results in cache performance degradation if a large number of missing data come from NVM. Taking the asymmetric cache miss cost into account, we first propose a new performance metric -TMPKI, which can exactly reflect the LLC performance on the top of hybrid memories. Then we propose a hybrid memory aware cache partitioning technique (HAP) to dynamically adjust the cache spaces for DRAM and NVM data based on TMPKI. Experimental results show that HAP improves performance against the traditional LRU policy by up to 54.3% (19.6% on average) while it incurs a little storage overhead (0.2%). Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001 |
ICCD | 4 |
| 2014 | DTail: a flexible approach to DRAM refresh managementabstractDRAM cells must be refreshed (or rewritten) periodically to maintain data integrity, and as DRAM density grows, so does the refresh time and energy. Not all data need to be refreshed with the same frequency, though, and thus some refresh operations can safely be delayed. Tracking such information allows the memory controller to reduce refresh costs by judiciously choosing when to refresh different rows Zehan Cui, Sally A. McKee, Zhongbin Zha, Yungang Bao, Mingyu Chen 0001 |
ICS | 5 |
| 2014 | DWC: dynamic write consolidation for phase change memory systemsabstractPhase change memory (PCM) is promising to become an alternative main memory thanks to its better scalability and lower leakage than DRAM. However, the long write latency of PCM puts it at a severe disadvantage against DRAM. In this paper, we propose a Dynamic Write Consolidation (DWC) scheme to improve PCM memory system performance while reducing energy consumption. This paper is motivated by the observation that a large fraction of a cache line being written back to memory is not actually modified. DWC exploits the unnecessary burst writes of unmodified data to consolidate multiple writes targeting the same row into one write. By doing so, DWC enables multiple writes to be send within one. DWC incurs low implementation overhead and shows significant efficiency. The evaluation results show that DWC achieves up to 35.7% performance improvement, and 17.9% on average. The effective write latency are reduced by up to 27.7%, and 16.0% on average. Moreover, DWC reduces the energy consumption by up to 35.3%, and 13.9% on average. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
ICS | 4 |
| 2014 | Pipelined Compaction for the LSM-TreeabstractWrite-optimized data structures like Log-Structured Merge-tree (LSM-tree) and its variants are widely used in key-value storage systems like Big Table and Cassandra. Due to deferral and batching, the LSM-tree based storage systems need background compactions to merge key-value entries and keep them sorted for future queries and scans. Background compactions play a key role on the performance of the LSM-tree based storage systems. Existing studies about the background compaction focus on decreasing the compaction frequency, reducing I/Os or confining compactions on hot data key-ranges. They do not pay much attention to the computation time in background compactions. However, the computation time is no longer negligible, and even the computation takes more than 60% of the total compaction time in storage systems using flash based SSDs. Therefore, an alternative method to speedup the compaction is to make good use of the parallelism of underlying hardware including CPUs and I/O devices. In this paper, we analyze the compaction procedure, recognize the performance bottleneck, and propose the Pipelined Compaction Procedure (PCP) to better utilize the parallelism of CPUs and I/O devices. Theoretical analysis proves that PCP can improve the compaction bandwidth. Furthermore, we implement PCP in real system and conduct extensive experiments. The experimental results show that the pipelined compaction procedure can increase the compaction bandwidth and storage system throughput by 77% and 62% respectively. Zigang Zhang, Yinliang Yue, Bingsheng He, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
IPDPS | 5 |
| 2014 | Going vertical in memory management: Handling multiplicity by multi-policyabstractMany emerging applications from various domains often exhibit heterogeneous memory characteristics. When running in combination on parallel platforms, these applications present a daunting variety of workload behaviors that challenge the effectiveness of any memory allocation strategy. Prior partitioning-based or random memory allocation schemes typically manage only one level of the memory hierarchy and often target specific workloads. To handle diverse and dynamically changing memory and cache allocation needs, we augment existing “horizontal” cache/DRAM bank partitioning with vertical partitioning and explore the resulting multi-policy space. We study the performance of these policies for over 2000 workloads and correlate the results with application characteristics via a data mining approach. Based on this correlation we derive several practical memory allocation rules that we integrate into a unified multi-policy framework to guide resources partitioning and coalescing for dynamic and diverse multi-programmed/threaded workloads. We implement our approach in Linux kernel 2.6.32 as a restructured page indexing system plus a series of kernel modules. Extensive experiments show that, in practice, our framework can select proper memory allocation policy and consistently outperforms the unmodified Linux kernel, achieving up to 11% performance gains compared to prior techniques. Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ISCA | 5 |
| 2014 | Intelligent frame refresh for energy-aware display subsystems in mobile devicesabstractFrame refreshes, that are used to retain frame images from frame buffers for display subsystems in mobile devices, waste energy and memory bandwidth. In this paper, we propose an intelligent frame refresh mechanism to reduce redundant frame refreshes and useless data accesses to frame buffers, which bridges the semantic gap between frame buffers and frame refreshes, and exploits the knowledge of frame buffers to guide frame refreshes. Based on this mechanism, we introduce two detailed schemes to optimize refreshes by utilizing different information. The flipping-aware frame refresh scheme uses the frame buffer switching operations to detect frame image updates and triggers useful refreshes. The row-level frame refresh scheme supports to refresh only modified rows instead of the whole frame, under the guidance of pixel status information of frame buffers. Our evaluation results show that our proposed mechanism can reduce memory requests by nearly 50% and memory power consumption up to 30%, compared to conventional fixed frame refresh mechanism. Yongbing Huang, Mingyu Chen 0001, Lixin Zhang 0002, Shihai Xiao, Junfeng Zhao 0003, Zhulin Wei |
ISLPED | 2 |
| 2014 | Moby: A mobile benchmark suite for architectural simulatorsabstractMobile devices such as smartphones and tablets have become the primary consumer computing devices, and their rate of adoption continues to grow. The applications that run on these mobile platforms vary in how they use hardware resources, and their diversity is increasing. Performance and power limitations also vary widely across mobile platforms. Thus there is a growing need for tools to help computer architects design systems to meet the needs of mobile workloads. Full-system simulators are invaluable tools for designing new architectures, but we still need appropriate benchmark suites that capture the behaviors of emerging mobile applications. Current benchmark suites cover only a small range of mobile applications, and many cannot run directly in simulators due to their user interaction requirements. In this paper, we introduce and characterize Moby, a benchmark suite designed to make it easier to use full-system architectural simulators to evaluate microarchitectures for mobile processors. Moby contains popular Android applications, including a web browser, a social networking application, an email client, a music player, a video player, a document processing application, and a map program. To facilitate microarchitectural exploration, we port the Moby benchmark suite to the popular gem5 simulator. We characterize the architecture-independent features of Moby applications on the simulator and analyze the architecture-dependent features on a current-generation mobile platform. Our results show that mobile applications exhibit complex instruction execution behaviors and poor code locality, but current mobile platforms especially instruction-related components cannot meet their requirements. Yongbing Huang, Zhongbin Zha, Mingyu Chen 0001, Lixin Zhang 0002 |
ISPASS | 3 |
| 2014 | CMD: classification-based memory deduplication through page access characteristicsabstractLimited main memory size is considered as one of the major bottlenecks in virtualization environments. Content-Based Page Sharing (CBPS) is an efficient memory deduplication technique to reduce server memory requirements, in which pages with same content are detected and shared into a single copy. As the widely used implementation of CBPS, Kernel Samepage Merging (KSM) maintains the whole memory pages into two global comparison trees (a stable tree and an unstable tree). To detect page sharing opportunities, each tracked page needs to be compared with pages already in these two large global trees. However since the vast majority of compared pages have different content with it, that will induce massive futility comparisons and thus heavy overhead. Licheng Chen, Zehan Cui, Mingyu Chen 0001, Haiyang Pan, Yungang Bao |
VEE | 4 |
| 2014 | A High-Performance and Cost-Efficient Interconnection Network for High-Density Servers
Wentao Bao, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002 |
J. Comput. Sci. Technol. | 3 |
| 2014 | MIMS: Towards a Message Interface Based Memory System
Licheng Chen, Mingyu Chen 0001, Yuan Ruan, Yongbing Huang, Zehan Cui, Tianyue Lu, Yungang Bao |
J. Comput. Sci. Technol. | 2 |
| 2014 | HMTT: A hybrid hardware/software tracing system for bridging the DRAM access trace's semantic gapabstractDRAM access traces (i.e., off-chip memory references) can be extremely valuable for the design of memory subsystems and performance tuning of software. Hardware snooping on the off-chip memory interface is an effective and nonintrusive approach to monitoring and collecting real-life DRAM accesses. However, compared with software-based approaches, hardware snooping approaches typically lack semantic information, such as process/function/object identifiers, virtual addresses, and lock contexts, that is essential to the complete understanding of the systems and software under investigation. In this article, we propose a hybrid hardware/software mechanism that is able to collect off-chip memory reference traces with semantic information. We have designed and implemented a prototype system called HMTT (Hybrid Memory Trace Tool), which uses a custom-made DIMM connector to collect off-chip memory references and a high-level event-encoding scheme to correlate semantic information with memory references. In addition to providing complete, undistorted DRAM access traces, the proposed system is also able to perform various types of low-overhead profiling, such as object-relative accesses and multithread lock accesses. Yongbing Huang, Licheng Chen, Zehan Cui, Yuan Ruan, Yungang Bao, Mingyu Chen 0001, Ninghui Sun |
ACM Trans. Archit. Code Optim. | 6 |
| 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systemsabstractThe main memory system is a shared resource in modern multicore machines that can result in serious interference leading to reduced throughput and unfairness. Many new memory scheduling mechanisms have been proposed to address the interference problem. However, these mechanisms usually employ relative complex scheduling logic and need modifications to Memory Controllers (MCs), which incur expensive hardware design and manufacturing overheads. This article presents a practical software approach to effectively eliminate the interference without any hardware modifications. The key idea is to modify the OS memory management system and adopt a page-coloring-based Bank-level Partitioning Mechanism (BPM) that allocates dedicated DRAM banks to each core (or thread). By using BPM, memory requests from distinct programs are segregated across multiple memory banks to promote locality/fairness and reduce interference. We further extend BPM to BPM+ by incorporating channel-level partitioning, on which we demonstrate additional gain over BPM in many cases. To achieve benefits in the presence of diverse application memory needs and avoid performance degradation due to resource underutilization, we propose a dynamic mechanism upon BPM/BPM+ that assigns appropriate bank/channel resources based on application memory/bandwidth demands monitored through PMU (performance-monitoring unit) and a low-overhead OS page table scanning process. We implement BPM/BPM+ in Linux 2.6.32.15 kernel and evaluate the technique on four-core and eight-core real machines by running a large amount of randomly generated multiprogrammed and multithreaded workloads. Experimental results show that BPM/BPM+ can improve the overall system throughput by 4.7%/5.9%, on average, (up to 8.6%/9.5%) and reduce the unfairness by an average of 4.2%/6.1% (up to 15.8%/13.9%). Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ACM Trans. Archit. Code Optim. | 5 |
| 2013 | Scattered superpage: A case for bridging the gap between superpage and page coloringabstractSuperpage and page coloring are two important practical techniques to improve the performance of Translation Lookaside Buffers (TLBs) and shared Last Level Cache (LLC) respectively. However, there exists a gap between these two techniques in current hardware-architecture design, resulting in the contradiction in adopting these two optimizations simultaneously: a superpage requires hundreds of contiguous (e.g. a power of two) base pages in both virtual and physical memory, which would compulsorily occupy all available page colors (or cache sets), thus making page coloring failed to work. This is because most contemporary architecture adopts the design with cache set indexes placed in the least significant part of block address. In this paper, we propose a lightweight approach named Scattered Superpage to bridge this gap. Scattered Superpage decouples a superpage from the limitation of occupying multiple contiguous physical base pages. A superpage is still contiguous in virtual memory, but it is scattered mapping into multiple physical superpages, and it just occupies specified partial page colors in each physical superpage, thus it allows us to configure page color for each superpage. The huge TLB is slightly modified to store page color configuration for each superpage and to calculate target physical address based on this configuration when doing address translation. The experimental results show that the Scattered Superpage can improve system performance by 20.51% and reduce unfairness by 27.77% in our 4-core simulation system (with multi-program memory-intensive workloads). It achieves this by reducing last level cache miss by 17.05% and reducing TLB miss by 86.02% simultaneously. Licheng Chen, Zehan Cui, Yongbing Huang, Yungang Bao, Mingyu Chen 0001 |
ICCD | 6 |
| 2013 | SMAT: an input adaptive auto-tuner for sparse matrix-vector multiplicationabstractSparse Matrix Vector multiplication (SpMV) is an important kernel in both traditional high performance computing and emerging data-intensive applications. By far, SpMV libraries are optimized by either application-specific or architecture-specific approaches, making the libraries become too complicated to be used extensively in real applications. In this work we develop a Sparse Matrix-vector multiplication Auto-Tuning system (SMAT) to bridge the gap between specific optimizations and general-purpose usage. SMAT provides users with a unified programming interface in compressed sparse row (CSR) format and automatically determines the optimal format and implementation for any input sparse matrix at runtime. For this purpose, SMAT leverages a learning model, which is generated in an off-line stage by a machine learning method with a training set of more than 2000 matrices from the UF sparse matrix collection, to quickly predict the best combination of the matrix feature parameters. Our experiments show that SMAT achieves impressive performance of up to 51GFLOPS in single-precision and 37GFLOPS in double-precision on mainstream x86 multi-core processors, which are both more than 3 times faster than the Intel MKL library. We also demonstrate its adaptability in an algebraic multigrid solver from Hypre library with above 20% performance improvement reported. Jiajia Li 0001, Guangming Tan, Mingyu Chen 0001, Ninghui Sun |
PLDI | 3 |
| 2012 | HaLock: hardware-assisted lock contention detection in multithreaded applicationsabstractMultithreaded programming relies on locks to ensure the consistency of shared data. Lock contention is the main reason of low parallel efficiency and poor scalability of multithreaded programs. Lock profiling is the primary approach to detect lock contention. Prior lock profiling tools are able to track lock behaviors but directly store profiling data into local memory regardless of the memory interference on targeted programs. Yongbing Huang, Zehan Cui, Licheng Chen, Yungang Bao, Mingyu Chen 0001 |
PACT | 6 |
| 2012 | A software memory partition approach for eliminating bank-level interference in multicore systemsabstractMain memory system is a shared resource in modern multicore machines, resulting in serious interference, which causes performance degradation in terms of throughput slowdown and unfairness. Numerous new memory scheduling algorithms have been proposed to address the interference problem. However, these algorithms usually employ complex scheduling logic and need hardware modification to memory controllers, as a result, industrial venders seem to have some hesitation in adopting them. Lei Liu 0030, Zehan Cui, Mingjie Xing, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
PACT | 5 |
| 2012 | Evaluation and Optimization of Breadth-First Search on NUMA ClusterabstractGraph is widely used in many areas. Breadth-First Search (BFS), a key subroutine for many graph analysis algorithms, has become the primary benchmark for Graph500 ranking. Due to the high communication cost of BFS, multi-socket nodes with large memory capacity (NUMA) are supposed to reduce network pressure. However, the longer latency to remote memory may cause problem if not treated well. In this work, we first demonstrate that simply spawning and binding one MPI process for each socket can achieve the best performance for MPI/OpenMP hybrid programmed BFS algorithm, resulting in 1.53X of performance on 16 nodes. Nevertheless, we notice that one MPI process per socket may exacerbate the communication cost. We propose to share some communication data structure among the processes inside the same node, to eliminate most of the intra-node communication. To fully utilize the network bandwidth, we make all the processes in a node to perform communication simultaneously. We further adjust the granularity of a key bitmap for better cache locality to speed up the computation. With all the optimizations for NUMA, communication and computation together, 2.44X of performance is achieved on 16 nodes, which is 39.2 Billion Traversed Edges per Second for an R-MAT graph of scale 32 (4 billion vertices and 64 billion edges). Zehan Cui, Licheng Chen, Mingyu Chen 0001, Yungang Bao, Yongbing Huang, Huiwei Lv |
CLUSTER | 3 |
| 2012 | Supporting User-directed Fault Tolerance over Standard MPIabstractUser-directed means the process of carrying out fault tolerance is dynamic and the fault tolerance mode is chosen by users based on application requirements. In this paper, we introduce a general scheme based on standard MPI to provide the user directed support for application level algorithmic fault tolerance. The user-directed fault tolerance plays the role as a connection between applications and algorithmic fault tolerance. As a case study, our scheme has been incorporated to HPL combined with a non-blocking ABFT technique. We have tested the functional availability of our scheme for fault tolerance in real circumstance. We also evaluated that when there is no failure occurring, our support only brings 2.5 percent overhead. When failure occurs, with our scheme, the scalability of algorithmic fault tolerance maintains well. Zhimin Wu, Weizhi Xu 0001, Mingyu Chen 0001, Erlin Yao |
ICPADS | 4 |
| 2012 | An optimized large-scale hybrid DGEMM design for CPUs and ATI GPUsabstractIn heterogeneous systems that include CPUs and GPUs, the data transfers between these components play a critical role in determining the performance of applications. Software pipelining is a common approach to mitigate the overheads of those transfers. In this paper we investigate advanced software-pipelining optimizations for the double-precision general matrix multiplication (DGEMM) algorithm running on a heterogeneous system that includes ATI GPUs. Our approach decomposes the DGEMM workload to a finer detail and hides the latency of CPU-GPU data transfers to a higher degree than previous approaches in literature. We implement our approach in a five-stage software pipelined DGEMM and analyze its performance on a platform including x86 multi-core CPUs and an ATI Radeon™ HD5970 GPU that has two Cypress GPU chips on board. Our implementation delivers 758 GFLOPS (82% floating-point efficiency) when it uses only the GPU, and 844 GFLOPS (80% efficiency) when it distributes the workload on both CPU and GPU. We analyze the performance of our optimized DGEMM as the number of GPU chips employed grows from one to two, and the results show that resource contention on the PCIe bus and on the host memory are limiting factors. Jiajia Li 0001, Xingjian Li 0002, Guangming Tan, Mingyu Chen 0001, Ninghui Sun |
ICS | 4 |
| 2012 | A Case Study of Designing Efficient Algorithm-based Fault Tolerant Application for Exascale ParallelismabstractFault tolerance overhead of high performance computing (HPC) applications is becoming critical to the efficient utilization of HPC systems at large scale. Today's HPC applications typically tolerate fail-stop failures by check pointing. However, check pointing will lose its efficiency when system becoming very large. An alternative method is algorithm-based fault recovery which has been proved to be more efficient than check pointing. In this paper, we first point out by theoretical analysis that algorithm-based fault recovery will also lose its efficiency when systems scale up to Exa flops. Then, a more efficient algorithm-based fault tolerance scheme for HPC applications at large scale is presented. The new method has two novel skills. One is algorithm-based hot replacement, which avoids the stop-and-wait time after failure. Second is background accelerated recovery, which guarantees the system to endure multiple failures in succession. As a case study, this method is incorporated to High Performance Lin pack (HPL). Theoretical analysis shows that the fault tolerance overhead can be reduced to 2/log(p, 2) of that of algorithm-based fault recovery method (p is the number of computation processes), so that the new method will still be efficient in Exascale. Experimental results for up to 1800 processes show that the overhead of the new method is about 25% of that of algorithm-based fault recovery method, which is close to the theoretical prediction. Erlin Yao, Mingyu Chen 0001, Guangming Tan, Ninghui Sun |
IPDPS | 3 |
| 2012 | A lightweight hybrid hardware/software approach for object-relative memory profilingabstractMemory profiling is the process of collecting memory address traces during the execution of a program, then analyzing and characterizing the memory behavior of the program offline. With the trend that there will be more and more cores integrated in a processor chip, the “Memory Wall” problem will become more serious in the chip multiprocessor (CMP) system. Thus accurate and effective memory profiling is becoming one of the keys to identify the source of memory system bottlenecks. A large body of work has been contributed to memory profiling, however, most adopts instrumentation, simulator which suffers heavy overhead, or hardware performance counter which is lack of detail trace information. Furthermore, correlating the raw memory address traces with object-relative information allows us to separate regular pattern for certain object from the irregular mixed, thus helps the optimization. In this paper, we propose a lightweight hybrid hardware/software approach for object-relative memory profiling. We monitor physical memory addresses through hardware snooping with negligible overhead; meanwhile we dump Linux kernel page tables of processes, as well as object-relative memory allocation information. Our approach supports not only to collect applications' full memory traces with detail object relative information, but also to identify hardware-generated memory accesses such as page memory walks due to TLB miss at object level. The experimental results on real system show that our approach is highly accurate (the largest error is 2.04%) and low overhead (the average overhead is 1.60%). Furthermore, we profile two multi-thread applications in detail, and successfully identity hot TLB-miss objects. With object-targeted optimization, we can improve applications' performance by nearly 6.86%. Licheng Chen, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Yongbing Huang, Guangming Tan |
ISPASS | 4 |
| 2011 | Building algorithmically nonstop fault tolerant MPI programsabstractWith the growing scale of high-performance computing (HPC) systems, today and more so tomorrow, faults are a norm rather than an exception. HPC applications typically tolerate fail-stop failures under the stop-and-wait scheme, where even if only one processor fails, the whole system has to stop and wait for the recovery of the corrupted data. It is now a more-or-less accepted fact that the stop-and-wait scheme will not scale to the next generation of HPC systems. Inspired by the previous stop-and-wait algorithm-based fault tolerance (ABFT) recovery technique, we propose in this paper a nonstop fault tolerance scheme at the application level and describe its implementation. When failure occurs during the execution of applications, we do not stop to wait for the recovery of the corrupted node; instead, we replace it with the corresponding redundant node and continue the execution. At the end of execution, the correct solution can be recovered algorithmically at a very low cost. In order to implement the scheme, some new fault-tolerant features of the Message Passing Interface (MPI) have been investigated and utilized in the MPICH implementation of MPI. We also describe a case study using High Performance Linpack (HPL) with these new features and evaluate the performance of both our new scheme and ABFT recovery. Experimental results show the advantage of our new scheme over ABFT recovery even in a small scale. Erlin Yao, Mingyu Chen 0001, Guangming Tan, Pavan Balaji, Darius Buntinas |
HiPC | 3 |
| 2011 | Experience of parallelizing cryo-EM 3D reconstruction on a CPU-GPU heterogeneous systemabstractHeterogeneous architecture is becoming an important way to build a massive parallel computer system, i.e. the CPU-GPU heterogeneous systems ranked in Top500 list. However, it is a challenge to efficiently utilize massive parallelism of both applications and architectures on such heterogeneous systems. In this paper we present a practice on how to exploit and orchestrate parallelism at algorithm level to take advantage of underlying parallelism at architecture level. A potential Petaflops application -- cryo-EM 3D reconstruction is selected as an example. We exploit all possible parallelism in cryo-EM 3D reconstruction, and leverage a self-adaptive dynamic scheduling algorithm to create a proper parallelism mapping between the application and architecture. The parallelized programs are evaluated on a subsystem of Dawning Nebulae supercomputer, whose node is composed of two Intel six-core Xeon CPUs and one Nvidia Fermi GPU. The experiment confirms that hierarchical parallelism is an efficient pattern of parallel programming to utilize capabilities of both CPU and GPU in a heterogeneous system. The CUDA kernels run more than 3 times faster than the OpenMP parallelized ones using 12 cores (threads). Based on the GPU-only version, the hybrid CPU-GPU program further improves the whole application's performance by 30% on the average. Linchuan Li, Xingjian Li 0002, Guangming Tan, Mingyu Chen 0001, Peiheng Zhang |
HPDC | 4 |
| 2011 | Poster: revisiting virtual channel memory for performance and fairness on multi-core architectureabstractIn modern multi-core chip architecture, the DRAM system is shared by more and more cores and high bandwidth I/O devices. This trend would make the problem of request contention and un-fairness more serious. Previous research focused on memory sche-duling mechanisms to efficiently and fairly serve memory requests generated by multiple cores. However, the performance is mod-erately improved due to the limited bank-level parallelism in preva-lent DRAM chips. Based on the observation that virtual channel memory (VCM) provides more opportunities for exploiting MLP because it has more channel buffers than banks in conventional DRAM chip, we evaluate VCM technology as an alternative to DRAM for addressing the issues of contention, unfairness and MLP. In this work we implement VCM and leverage the state of art scheduling mechanism on a multi-core architecture. The experi-mental results show that (i) VCM with 32 channels improves ho-mogeneous workloads' IPC by 2.08X on a 16-core system compared to the system with conventional DRAM chips, causing extra area cost by 0.5%, and dynamic and background power pe-nalties by only 5.8% and 0.03% respectively. (ii) For heterogene-ous workloads, VCM significantly reduces unfairness by 82.0% as well as improves the workloads' performance by 1.86X in term of system throughput. Licheng Chen, Yongbing Huang, Yungang Bao, Onur Mutlu, Guangming Tan, Mingyu Chen 0001 |
ICS | 6 |
| 2011 | What Hill-Marty model learn from and break through Amdahlʼs law?
Erlin Yao, Yungang Bao, Mingyu Chen 0001 |
Inf. Process. Lett. | 3 |
| 2010 | DMA cache: Using on-chip storage to architecturally separate I/O data from CPU data for improving I/O performanceabstractAs technology advances both in increasing bandwidth and in reducing latency for I/O buses and devices, moving I/O data in/out memory has become critical. In this paper, we have observed the different characteristics of I/O and CPU memory reference behavior, and found the potential benefits of separating I/O data from CPU data. We propose a DMA cache technique to store I/O data in dedicated on-chip storage and present two DMA cache designs. The first design, Decoupled DMA Cache (DDC), adopts additional on-chip storage as the DMA cache to buffer I/O data. The second design, Partition-Based DMA Cache (PBDC), does not require additional on-chip storage, but can dynamically use some ways of the processor's last level cache (LLC) as the DMA cache. We have implemented and evaluated the two DMA cache designs by using an FPGA-based emulation platform and the memory reference traces of real-world applications. Experimental results show that, compared with the existing snooping-cache scheme, DDC can reduce memory access latency (in bus cycles) by 34.8% on average (up to 58.4%), while PBDC can achieve about 80% of DDC's performance improvements despite no additional on-chip storage. Dan Tang 0002, Yungang Bao, Weiwu Hu, Mingyu Chen 0001 |
HPCA | 4 |
| 2010 | Automatically Tuned Dynamic Programming with an Algorithm-by-BlocksabstractAs the complexity of current computer architecture increases, domain-specific program generators are extensively used to implement performance portable libraries. Dynamic programming is a performance-critical kernel in many applications including engineering operations and bioinformatics. In this paper, we propose an Automatically Tuned Dynamic Programming (ATDP) to optimize performance of dynamic programming algorithm across various architectures. First, an algorithm-by-blocks for dynamic programming is designed to facilitate optimizing with well-known techniques including cache and register tiling. Further, the parameterized algorithm-by-blocks is cooperative with an auto-tuning framework and leverages a hill climbing algorithm to search the possible best program on a given platform. The experiments on two ×86 processors demonstrate that (i) the generated scalar programs improve performance by over 10 times, (ii) the vector programs further speedup the scalar ones by a factor of 4 and 2 for single-precision and double-precision, respectively. Jiajia Li 0001, Guangming Tan, Mingyu Chen 0001 |
ICPADS | 3 |
| 2010 | GenerOS: An asymmetric operating system kernel for multi-core systemsabstractDue to complex abstractions implemented over shared data structures protected by locks, conventional symmetric multithreaded operating system kernel such as Linux is hard to achieve high scalability on the emerging multi-core architectures, which integrate more and more cores on a single die. This paper presents GenerOS - a general asymmetric operating system kernel for multi-core systems. In principal, GenerOS partitions processing cores into application core, kernel core and interrupt core, each of which is dedicated to a specified function. In implementation, we conduct a delicate modification to Linux kernel and provide the same interface as Linux kernel so that GenerOS is compatible with legacy applications. The better performance of GenerOS mainly benefits from: (1) Applications run on their own cores with minimal interrupt and kernel support; (2) Every kernel service is encapsulated in to a serial process so that there will be fewer contentions than conventional symmetric kernel; (3) A slim schedule policy is used in the kernel core to support schedule between system calls with low overhead. Experiments with two typical workloads on 16-core AMD machine show that GenerOS behaves better than original Linux kernel when there are more processing cores (19.6% for TPC-H using oracle database management system and 42.8% for httperf using apache web server). Qingbo Yuan, Mingyu Chen 0001, Ninghui Sun |
IPDPS | 3 |
| 2010 | QTL: An efficient scheduling policy for 10Gbps network intrusion detection systemabstractBroad network bandwidth and deep inspection impose great challenge for the capability of 10Gpbs network security monitoring. Proper scheduling policies can improve system capability without requiring additional resources. LAS, a size-based scheduling policy which can achieve optimal mean response time by giving preferential analysis to short flows, is widely used in various aspects of network field. Due to the high variability property of Internet traffic, LAS favors short flows without penalizing large flows very much. Unfortunately, the inspection of large flows can not be guaranteed in those network intrusion detection systems on 10Gbps links, which are usually heavily loaded, or even overloaded. Although tiny in percentage, large flows comprise more than 50% of the total load, and therefore can not be ignored, especially when specified by users as critical. How to avoid starving large flows while still giving higher priority to short flows is a dilemma we have to face in practice. In this paper, we propose a QoS-supported three-level scheduling policy (QTL), which can remedy LAS' defect. The experimental results show that our QTL scheduling policy has approximately the same performance as LAS for short flows, and meanwhile exhibits greatly enhanced processing capability for large flows. Weibing Yang, Mingyu Chen 0001, Jianping Fan 0002 |
ISCC | 3 |
| 2010 | Robust TCP Reassembly with a Hardware-Based Solution for Backbone TrafficabstractThere is a growing interest in designing high-speed network devices to perform packet processing at stream layer. However, variety kinds of out-of- sequence packets in real traffic will make trouble for hardware-based TCP reassembly system which is less flexible for exceptional processing. In this paper, we present a detailed analysis of behavior characteristic of out-of-sequence packets in real backbone traffic and propose a hardware-based solution for TCP reassembly based on the results of analysis. This solution could reassemble a TCP stream with concurrent multi discontinuous out-of-order data segments and provide a robust buffer management strategy for out-of-order data. We have also assessed the memory size and bandwidth required for reassembling real 10G traffic. The simulation result shows that the system can process over 99% of the 10G backbone traffic using reasonable storage resources. A FPGA-based prototype is also implemented for evaluation. Yuan Ruan, Weibing Yang, Mingyu Chen 0001, Jianping Fan 0002 |
NAS | 3 |
| 2010 | Achieving Flow-Level Controllability in Network Intrusion Detection SystemabstractCurrent network intrusion detection systems are lack of controllability, manifested as significant packet loss due to the long-term resources occupation by a single flow. The reasons can be classified into two kinds. The first kind is known as normal reasons, that is, the processing of mass arriving packets of a large flow can not be limited to a determinable period of time and thus makes other flows starved. The second kind, in which the CPU is trapped in a dead-loop like state due to processing some packets with particular content of a flow, is considered as abnormal reasons. In fact, it is a kind of software crashes. In this paper, we discuss the innate defects of traditional packet-driven NIDS, and implement a flow-driven framework which can achieve fine-grained controllability. An Active Two-threshold scheme based on ideal Exit-Point (ATEP) is proposed in order to diminish data preserving overhead during flow switches and to detect crash in time. A quick crash recovery mechanism is also given which can recover the trapped thread from 90% crashes in 0.2 ms. The experimental results show that our flow-driven framework with ATEP scheme can achieve higher throughput and less packet loss ratio than the uncontrollable packet-driven systems with less than 1% of extra CPU overhead. What's more, in the case of crash occurrence, the ATEP scheme is still able to maintain rather steady throughput without sudden decrease. Weibing Yang, Mingyu Chen 0001, Jianping Fan 0002 |
SNPD | 3 |
| 2009 | Single-particle 3d reconstruction from cryo-electron microscopy images on GPUabstractSingle-particle 3D reconstruction from cryo-electron microscopy (cryo-EM) images is a kernel application of biological molecules analysis, as the computational requirement of which is now beyond PetaFlop for a high-resolution 3D structure. In this paper, we quantitatively analyze the workload, computational intensity and memory performance of the application, parallelize it on an emerging multicore architecture GPU-CUDA. Further we apply a percolation technique to decouple computation with memory operations and orchestrate thread-data mapping to reduce the overhead off-chip memory operations. Finally we tested our optimization strategy on a popular open-source package EMAN to GPU-CUDA, which achieves a relative speedup of about 10X to the original CPU-only EMAN. The experimental results also show that the proposed percolation programming greatly improves utilization of memory bandwidth and floating-point units. Guangming Tan, Mingyu Chen 0001, Dan Meng 0002 |
ICS | 3 |
| 2009 | A Scalability Analysis of the Symmetric Multiprocessing Architecture in Multi-Core SystemabstractThe quickly development of the multi-core technology brings plenty of logical processors to the symmetric multiprocessing (SMP) system. As all of cores share the same system bus and memory bandwidth, the additional computing resources canpsilat fully play their roles. It is the basic restrict to the scalability of such a system. Furthermore, the operating system which runs in this system typically provides complex abstractions implemented over shared data structures protected by locks. More contentions come along with the increase of cores in such type of kernel. After several detailed experiments to 5 different types of benchmarks, we recognize these problems in the multi-core SMP system. At last, reasons causing the problems are analyzed and corresponding solutions are raised briefly. Qingbo Yuan, Yungang Bao, Mingyu Chen 0001, Ninghui Sun |
NAS | 3 |
| 2009 | SimK: A Large-Scale Parallel Simulation Engine
Mingyu Chen 0001, Gui Zheng, Zheng Cao 0003, Huiwei Lv, Ninghui Sun |
J. Comput. Sci. Technol. | 2 |
| 2008 | HMTT: a platform independent full-system memory trace monitoring systemabstractMemory trace analysis is an important technology for architecture research, system software (i.e., OS, compiler) optimization, and application performance improvements. Many approaches have been used to track memory trace, such as simulation, binary instrumentation and hardware snooping. However, they usually have limitations of time, accuracy and capacity.In this paper we propose a platform independent memory trace monitoring system, which is able to track virtual memory reference trace of full systems (including OS, VMMs, libraries, and applications). The system adopts a DIMM-snooping mechanism that uses hardware boards plugged in DIMM slots to snoop. There are several advantages in this approach, such as fast, complete, undistorted, and portable. Three key techniques are proposed to address the system design challenges with this mechanism: (1) To keep up with memory speeds, the DDR protocol state machine is simplified, and large FIFOs are added between the state machine and the trace transmitting logic to handle burst memory accesses; (2) To reconstruct physical-tovirtual mapping and distinguish one process' address space from others, an OS kernel module, which collects page table information, and a synchronization mechanism, which synchronizes the page table information with the memory race, are developed; (3) To dump massive trace data, we employ a straightforward method to compress the trace and use Gigabit Ethernet and RAID to send and receive the compressed trace.We present our implementation of an initial monitoring system, named HMTT (Hyper Memory Trace Tracker). Using HMTT, we have observed that burst bandwidth utilization is much larger than average bandwidth utilization, by up to 5X in desktop applications. We have also confirmed that the stream memory accesses of many applications contribute even more than 40% of L2 Cache misses and OS virtual memory management may decrease stream accesses in view of memory controller (or L2 Cache), by up to 30.2%. Moreover, we have evaluated OS impact on memory performance in real systems. The evaluations and case studies show the feasibility and effectiveness of our proposed monitoring mechanism and techniques. Yungang Bao, Mingyu Chen 0001, Yuan Ruan, Li Liu 0038, Jianping Fan 0002, Qingbo Yuan |
SIGMETRICS | 2 |
| 2005 | A Reconfigurable Optical Interconnect System for DSAGabstractHigh performance computing research is facing challenges and innovation on architecture is urgent. DSAG architecture is proposed and delivers "Architecture on Demand" feature. In DSAG, components in different catalogs are parted, while the ones in same catalog are congregated. This architecture can be enabled by optical interconnect and reconfigurable computing technology. Using advanced optical devices and enhanced reconfigurable computing devices (FPGA), we build a prototype system for DSAG. Optical interconnect can reach 16Gbps bandwidth; DDRAM interface is selected as host communication interface to match the bandwidth of optical channel; Reconfigurable logic and embedded processors are employed for flexible reconfiguration. The system is featured by high bandwidth, owerful, flexible. Lei Li 0005, Zheng Cao 0003, Mingyu Chen 0001, Jianping Fan 0002 |
PDCAT | 3 |
| 2004 | HPL Performance Prevision to Intending System Improvement
Mingyu Chen 0001, Jianping Fan 0002 |
ISPA | 2 |