Teng Ma 0006

dblp:24/5931-6 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-7104-1526ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 6 first-author · 15 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TArCS: Trusted and Attack-Resilient Clock Source with TEE and RDMA
Yunpeng Xu, Yu Jin 0010, Shuangjie Yao, Teng Ma 0006, Shuwen Deng
APPT6
2026 Bridging the GPU Utilization Gap: Predictive Multi-Dimensional Resource Scheduling for AI Workloads
abstract
Modern AI data centers face a critical paradox: while machine learning workloads dominate infrastructure demands, actual GPU utilization remains consistently low. Existing schedulers fail to coordinate heterogeneous resources effectively, lack predictive capabilities for dynamic workloads, and cannot balance isolation requirements with sharing optimization in multi-tenant clusters. This paper presents Wind, a novel resource scheduler that bridges the GPU utilization gap through predictive scheduling and geometric resource coordination. Wind introduces three key innovations: (1) a resource prediction framework that leverages historical execution patterns to forecast task requirements and completion times with high accuracy;(2) a unified scheduling architecture supporting isolation, sharing, preemption, and prioritization policies that eliminate resource fragmentation while maintaining performance guarantees; and (3) a Hilbert curve-based multi-dimensional scheduling algorithm that maps CPU-memory-GPU resource space to preserve spatial locality while achieving linear computational complexity.
Yilei Lu 0002, Dongbiao He, Teng Ma 0006, Letian Ruan, Jinlei Jiang, Yongwei Wu 0001
EuroSys3
2026 LightDSA: Enabling Efficient DSA Through Hardware-Aware Transparent Optimization
abstract
Data streaming operations consume a significant portion of CPU resources in data centers. The Data Streaming Accelerator (DSA), integrated into modern Intel CPUs in datacenter, offers promising acceleration for these operations with user-friendly features. However, previous studies have overlooked DSA's internal mechanisms and the performance implications of these features, leaving key performance issues unresolved in real-world usage.
Yuansen Wang, Teng Ma 0006, Yuanhui Luo, Dongbiao He, Zheng Liu 0022, Yunpeng Chai
EuroSys2
2026 HybridSkipList+: Rethinking Distributed Skiplist With Hybrid RDMA and Caching
abstract
Remote Direct Memory Access (RDMA) offers high performance through OS kernel bypass and has become a key technology in modern data centers. By exploiting memorysemantic operations, RDMA-based data structures can achieve high scalability and significantly reduce CPU utilization compared with traditional Ethernet-based systems. However, classic sorted indexes such asSkiplistsuffer from low throughput under RDMA memory semantics due to frequent and costly remote accesses. To achieve both scalability and high throughput for a lock-based concurrent Skiplist, this paper introduces a hybrid paradigm that combines one-sided and two-sided RDMA operations. This proposal is built upon a re-evaluation of core design choices in RDMA, including transport modes, caching strategies, and memory management. Building on this paradigm, we designHybridSkipList+, a distributed Skiplist system that integrates a client-side coherent cache with pull-based synchronization to reduce expensive network round trips, and a semi-continuous memory allocator to enhance RDMA access locality. We implementHybridSkipList+ on an eight-machine RDMA cluster and conduct extensive evaluations. Results show that HybridSkipList+ outperforms two baseline systems by up to 4.41× and 3.45× under typical workload conditions.
Yilei Lu 0002, Teng Ma 0006, Dongbiao He, Zhe Wang 0015, Cédric Westphal, Linghe Kong
IEEE Trans. Computers2
2025 CXL-INTERPLAY: Unraveling and Characterizing CXL Interference in Modern Computer Systems
abstract
Compute Express Link (CXL) is a promising technology that addresses memory and storage challenges. Despite its advantages, CXL faces performance threats from external interference when coexisting with current memory and storage systems. This interference is under-explored in existing research. To address this, we develop CXL-Interplay, systematically characterizing and analyzing interference from memory and storage systems. To the best of our knowledge, we are the first to characterize CXL interference on real CXL hardware. We also provide reverse-reasoning analysis with performance counters and kernel functions. In the end, we propose and evaluate mitigation solutions.
Shunyu Mao, Jiajun Luo, Jiapeng Zhou, Zheng Liu 0022, Teng Ma 0006, Shuwen Deng
DAC7
2025 Ragnar: Exploring Volatile-Channel Vulnerabilities on RDMA NIC
abstract
With the surge in data computation, Remote Direct Memory Access (RDMA) becomes crucial to offering low-latency and highthroughput communication for data centers, but it faces new security threats. This paper presents RAGNAR, a comprehensive suite of hardware-contention-based volatile-channel attacks leveraging the underexplored security vulnerabilities in RDMA hardware. Through comprehensive microbenchmark reverse engineering, we analyze RDMA NICs at multiple granularity levels and then construct covert-channel attacks, achieving 3.2x the bandwidth of state-of-the-art RDMA-targeted attacks on CX-5. We apply side-channel attacks on real-world distributed databases and disaggregated memory, where we successfully fingerprint operations and recover sensitive address data with 95.6% accuracy.
Yunpeng Xu, Yuchen Fan 0002, Teng Ma 0006, Shuwen Deng
DAC3
2025 Utilizing Contrastive Learning for Locating Network Anomalies in Real-time Conferencing Applications
abstract
Real-time conferencing applications (RCA) are crucial for online learning and e-commerce. However, they can be affected by network fluctuations because they are heavily dependent on cloud network connections. However, there is a dearth of systematic studies that aim to pinpoint the specific network links where these fluctuations occur. We introduce a contrastive learning approach for locating anomalies, based on actual traffic from real-time conferencing applications. This method is trained on unlabeled data, which means that it does not require the creation of a large-scale training dataset. The results illustrate the robust localization ability, achieving an accuracy rate of more than 95%, demonstrating its adaptability to commonly used real-time conferencing applications.
Teng Ma 0006, Dongbiao He, Zhongxing Ming, Laizhong Cui, Yunpeng Chai
ICME1
2025 DSA-2LM: A CPU-Free Tiered Memory Architecture with Intel DSA
Ruili Liu, Teng Ma 0006, Yingdi Shan, Zheng Liu 0022, Lingfeng Xiang, Hui Lu 0001, Jia Rao, Kang Chen 0001, Yongwei Wu 0001
USENIX ATC2
2025 Shard: A Scalable and Resize-optimized Hash Index on Disaggregated Memory
Hantian Zha, Teng Ma 0006, Baotong Lu, Yuansen Wang, Dongbiao He, Yuanhui Luo, Yunpeng Chai, Yuxing Chen 0003, Anqun Pan
Proc. VLDB Endow.2
2025 MemTunnel: A CXL-Based Rack-Scale Host Memory Pooling Architecture for Cloud Service
abstract
Memory underutilization poses a significant challenge in cloud services, leading to performance inefficiencies and resource wastage. The tightly coupled computing and memory resources in cloud servers are identified as the root cause of this problem. To address this issue, memory pooling has been the subject of extensive research for decades, providing centralized or distributed shared memory pools as flexible memory resources for various applications running on different servers. However, existing memory disaggregation solutions sacrifice memory resources, add extra hardware (such as memory boxes/blades/drives), and degrade memory performance to achieve flexibility. To overcome these limitations, this paper proposes MemTunnel, a rack-scale host memory pooling architecture that provides a low-cost memory pooling solution based on Compute Express Link (CXL). MemTunnel is the first hardware and software architecture to offer symmetric, memory-semantic memory pooling over CXL, with an FPGA-based platform to demonstrate its feasibility in a real implementation. MemTunnel is orthogonal to the existing CXL-based memory pool and provides an additional layer of abstraction for memory disaggregation. Evaluation results show that MemTunnel achieves comparable performance to the existing CXL-based memory pool for a single machine and provides better rack-scale performance with minor hardware overheads.
Tianchan Guan, Yijin Guan, Zhaoyang Du, Jiacheng Ma 0001, Boyu Tian, Teng Ma 0006, Zheng Liu 0022, Yuan Xie 0001, Mingyu Gao 0001, Guangyu Sun 0003, Hongzhong Zheng, Dimin Niu
IEEE Trans. Parallel Distributed Syst.7
2024 LogGenius: An Unsupervised Log Parsing Framework with Zero-shot Prompt Engineering
abstract
Efficient and accurate parsing of unstructured logs is crucial for anomaly detection, root cause localization, and log compression. Although many existing works have made good progress relying on Large Language Models (LLMs) and prompt engineering techniques, most of them require a certain degree of labeling or few-shot prompts, which limits their applicability in large-scale real-time heterogeneous log environments. To tackle this issue, we develop LogGenius, a novel unsupervised log parsing framework. It initially enriches the diversity of the parsed logs by leveraging generative LLMs with zero-shot prompts. It then employs an unsupervised parsing model on the augmented log data to accomplish log parsing. In order to alleviate the impact of potential hallucination issues caused by generative LLMs, we conduct a meticulous analysis and summarize the biases inherent in LLMs when directly applying them to generate diversified logs. Building upon these insights, we propose an effective log diversity augmentation algorithm to mitigate the aforementioned concerns.We thoroughly evaluate LogGenius based on various open-source system runtime log datasets and a new alarm log dataset from a commercial cloud production environment. The experimental results demonstrate that LogGenius can improve the parsing accuracy by up to about 30%, and the parsing accuracy in unseen logs by up to about 100%, compared to the state-of-the-art unsupervised-based methods.
Shengxi Nong, Dongbiao He, Weijie Zheng 0001, Teng Ma 0006, Ning Liu 0014, Gaogang Xie
ICWS5
2024 TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
abstract
Serverless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv, a co-designed integration of the serverless platform with the operating system and CXL/RDMA-based remote memory pools in two key areas. Firstly, TrEnv introduces repurposable sandboxes, which can be shared across different functions and hence, substantially decrease the overhead associated with creating isolation sandboxes. Secondly, it augments the OS with "memory templates" that enable rapid restoration of function states stored on remote memory. These innovations allow TrEnv to facilitate rapid transitions between instances of different functions and enable memory sharing across multiple nodes. Our evaluations using a variety of representative and real-world workloads demonstrate that TrEnv can initiate a container within 10 milliseconds, achieving up to a 7× speedup in P99 end-to-end latency and reducing memory usage by 48% on average compared to state-of-the-art on-demand restoring systems.
Teng Ma 0006, Zheng Liu 0022, Sixing Lin, Kang Chen 0001, Jinlei Jiang, Xia Liao, Yingdi Shan, Mengting Lu, Tao Ma 0006, Haifeng Gong, Yongwei Wu 0001
SOSP3
2024 HydraRPC: RPC in the CXL Era
Teng Ma 0006, Zheng Liu 0022, Chengkun Wei, Youwei Zhuo, Yijin Guan, Dimin Niu, Tao Ma 0006
USENIX ATC1
2024 Zero+: Monitoring Large-Scale Cloud-Native Infrastructure Using One-Sided RDMA
abstract
Cloud services have shifted from monolithic designs to microservices running on cloud-native infrastructure with monitoring systems to ensure service level agreements (SLAs). However, traditional monitoring systems no longer meet the demands of cloud-native monitoring. In Alibaba’s “double eleven” shopping festival, it is observed that the monitor occupies resources of the monitored infrastructure and even disrupts services. In this paper, we propose a novel monitoring system named for cloud-native monitoring. achieves zero overhead in collecting raw metrics using one-sided remote direct memory access (RDMA) and remedies network congestion by adopting a receiver-driven flow control scheme. also features a priority queue mechanism to meet different quality of service requirements and an efficient batch processing design to relieve CPU occupation. has been deployed and evaluated in four different clusters with heterogeneous RDMA NIC devices and architectures in Alibaba Cloud. Results show that achieves no CPU occupation at the monitored host and supports$1\sim10k$hosts with$0.1\sim1s$sampling interval using a single thread for network I/O. significantly relieves the incast issue and maintains$80\sim95\%$of bandwidth utilization in several clusters when monitoring$1k$hosts. also ensures services with high priority accomplish collecting metrics earlier than low priority ones by at least$400 \mu s$when monitoring$1k$hosts.
Jiejian Wu, Teng Ma 0006, Zhe Wang 0015, Linghe Kong, Zhenzao Wen, Yong Yang 0013, Tao Ma 0006, Zheng Liu 0022, Guihai Chen
IEEE/ACM Trans. Netw.3
2023 Efficient Scheduler Live Update for Linux Kernel with Modularization
abstract
The scheduler is a critical component of the operating system (OS)and is tightly coupled with Linux. Production-level clouds often host various workloads, and these workloads require different schedulers to achieve high performance. Thus the capability of updating the scheduler lively without rebooting the OS is crucial for the production environments. However, emerging live update techniques only apply for the fine-grained function-level updates or require extra constraints such as microkernel. It fails to update the entire heavy process scheduler subsystem lively. We therefore propose Plugsched to enable scheduler live update, and there are two key novelties. First of all, with the idea of modularization, Plugsched decouples the scheduler from the Linux kernel to be an independent module; Secondly, Plugsched uses the data rebuild technique to migrate the state from the old scheduler to the new one. This scheme can be directly applied to the Linux kernel scheduler in production environments without modifying kernel code. Unlike current function-level live update solutions, Plugsched allows developers to update the entire scheduler subsystem and modify internal scheduler data via the rebuilding technique. Moreover, an optimized stack inspection method is introduced to further effectively reduce the downtime due to the update. Experimental and production results show that Plugsched can effectively update kernel scheduler lively and the downtime is less than tens of milliseconds.
Teng Ma 0006, Shanpei Chen, Erwei Deng, Quan Chen 0002, Minyi Guo
ASPLOS (3)1
2023 LigBee: Symbol-Level Cross-Technology Communication from LoRa to ZigBee
abstract
Low-power wide-area networks (LPWAN) evolve rapidly with advanced communication primitives (e.g., coding, modulation) being continuously invented. This rapid iteration on LPWAN, however, forms a communication barrier between legacy wireless sensor nodes deployed years ago (e.g., ZigBee-based sensor node) with their latest competitor running a different communication protocol (e.g., LoRa-based IoT node): they work on the same frequency band but share different MAC- and PHY-layer regulations and thus cannot talk to each other directly. To break this barrier, we propose LigBee, a cross-technology communication (CTC) solution that enables symbol-level communication from the latest LPWAN LoRa node to legacy ZIGBEE node. We have implemented LigBee on both software-defined radios and commercial-off-the-shelf (COTS) LoRa and ZigBee nodes, and demonstrated that LigBee builds a reliable CTC link from LoRa node to ZigBee node on both platforms. Our experimental results show that i) LigBee achieves a bit error rate (BER) in the order of 10−3with 70 ∼ 80% frame reception ratio (FRR), ii) the range of LigBee link is over 300m, which is 6 ∼ 7.5× the typical range of legacy ZigBee and state-of-the-art solution, and iii) the throughput of LigBee link is maintained on the order of kbps, which is close to the LoRa’s throughput.
Zhe Wang 0015, Linghe Kong, Longfei Shangguan, Liang He 0002, Kangjie Xu, Yifeng Cao, Qiao Xiang, Jiadi Yu, Teng Ma 0006, Zheng Liu 0022, Guihai Chen
INFOCOM10
2023 Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared Memory
abstract
The efficiency of distributed shared memory (DSM) has been greatly improved by recent hardware technologies. But, the difficulty of distributed memory management can still be a major obstacle to the democratization of DSM, especially when a partial failure of the participating clients (e.g., due to crashed processes or machines) should be tolerated.
Teng Ma 0006, Jinqi Hua, Zheng Liu 0022, Kang Chen 0001, Fan Du, Jinlei Jiang, Tao Ma 0006, Yongwei Wu 0001
SOSP2
2022 Log-ROC: Log Structured RAID on Open-Channel SSD
abstract
With the development of high-density flash memory, the price decreases, and the program/erase capability for each storage cell also decreases. Error correcting should be applied at all levels in SSD (Solid State Drive) based storage systems. This paper focuses on how to build a redundant array of Open-Channel SSDs, and proposes Log-ROC to enforce reliability at the drive level. Log-ROC adopts log structure, which means each write will be firstly buffered in the internal cache, and then the data in the cache are encoded in corresponding RAID level and flushed out in new locations as logs. Such mechanism is integrated with the host-side FTL (flash translation layer), effectively eliminating the parity update issue existing in RAID5 systems. We have built Log-ROC on 4 Open-Channel SSDs and evaluated Log-ROC with FIO and Filebench workloads. Compared with RAID5 on conventional SSDs, Log-ROC can achieve 1.91∼5.29× speedup in throughput on different workloads.
Teng Ma 0006, Ning Liu 0014
ICCD1
2022 SeqDLM: A Sequencer-Based Distributed Lock Manager for Efficient Shared File Access in a Parallel File System
abstract
Distributed locks are used to guarantee the distributed client-cache coherence in parallel file systems. However, they lead to poor performance in the case of parallel writes under high-contention workloads. We analyze the distributed lock manager and find out that lock conflict resolution is the root cause of the poor performance, which involves frequent lock revocations and slow data flushing from client caches to data servers. We design a distributed lock manager named SeqDLM by exploiting the sequencer mechanism. SeqDLM mitigates the lock conflict resolution overhead using early grant and early revocation while keeping the same semantics as traditional distributed locks. To evaluate SeqDLM, we have implemented a parallel file system called ccPFS using both SeqDLM and traditional distributed locks. Evaluations on 96 nodes show SeqDLM outperforms the traditional distributed locks by up to$\boldsymbol{10.3}\times$for high-contention parallel writes on a shared file with multiple stripes.
Shaonan Ma, Kang Chen 0001, Teng Ma 0006, Xin Liu 0081, Dexun Chen, Yongwei Wu 0001, Zuoning Chen
SC4
2022 Zero Overhead Monitoring for Cloud-native Infrastructure using RDMA
Zhe Wang 0015, Teng Ma 0006, Linghe Kong, Zhenzao Wen, Guihai Chen, Wei Cao 0006
USENIX ATC2
2022 A Survey of Storage Systems in the RDMA Era
abstract
Remote Direct Memory Access (RDMA) based network devices are increasingly being deployed in modern data centers. RDMA brings significant performance improvements over traditional network devices such as Ethernet due to its unique features:protocol offloadingandmemory semantics. In particular, it can achieve microsecond level latency, which is about 2$\sim$3 orders of magnitude improvement. With such improvement in hardware, the software stack, including device drivers and programming libraries, is becoming a new performance bottleneck. Developers need to use new programming libraries to take full advantage of the performance of the underlying hardware. Storage systems are very important in modern data centers. This article surveys the current efforts to use RDMA for optimizing storage systems. We first present five classes of RDMA-based storage systems, including key-value stores, file systems, distributed memory systems, database systems, and systems using smart NICs, to demonstrate different design choices. Then, we examine the core modules of storage systems from different perspectives: communication mode, concurrency control, fault tolerance, caching, and resource management. Finally, we provide some design guidelines for new RDMA-based storage systems, as well as a discussion of opportunities and challenges.
Shaonan Ma, Teng Ma 0006, Kang Chen 0001, Yongwei Wu 0001
IEEE Trans. Parallel Distributed Syst.2
2021 Thinking More about RDMA Memory Semantics
abstract
RDMA (Remote Direct Memory Access) provides memory semantics to access the remote memory directly bypassing remote CPUs. It can provide low latency and high throughput that can benefit many data center applications. Though a lot of efforts had been made in the literature, this paper tries to find more opportunities to boost the performance of memory semantic operations in the RDMA network. Similar to the optimizations for local memory operations, we find that the performance can be improved in the RDMA network after considering the vector IO mechanism, the performance asymmetry between sequential and random access, IO consolidation, NUMA effects, as well as the atomic operations (such as compare and swap) provided by the underlying hardware. We have done a comprehensive empirical study on the influences from these factors for the memory semantic operations in RDMA network and provide guidelines to improve applications. Experimental results show that four typical applications, disaggregated hashtable, distributed shuffle, distributed join, and distributed log are improved by 2.7×/5.8×/5.3×/9.1× respectively after considering memory semantics related optimizations.
Teng Ma 0006, Kang Chen 0001, Shaonan Ma, Yongwei Wu 0001
CLUSTER1
2021 HybridSkipList: A Case Study of Designing Distributed Data Structure with Hybrid RDMA
abstract
RDMA (Remote Direct Memory Access) delivers both high performance and OS kernel bypass thus triggers the revolutions of the modern data center. Especially, compared with traditional Ethernet, by exploiting memory semantics operations, RDMA-based data structure has shown great potential for high scalability and CPU utilization reduction. SkipList is an elegant sorted index data structure, and yet presents low throughput upon memory semantics RDMA due to the heavy remote access.By using SkipList as a case study, we revisit the current design paradigm of RDMA such as operation type, transport type, and caching. Accordingly, we propose a one-sided/two-sided hybrid paradigm to gain high scalability and at the same time achieve high throughput based on lock-based concurrent SkipList. To reduce the number of expensive network round trips, a client-sided cache is presented. With an in-depth analysis of the design choices, we implement HybridSkipList and deploy at an eight-machine RDMA cluster. The evaluations show it can outperform two baselines by 4.41× and 3.45× respectively.
Teng Ma 0006, Dongbiao He, Gordon Ning Liu
COMPSAC1
2021 CUBIST: High-Quality 360-Degree Video Streaming Services via Tile-based Edge Caching and FoV-Adaptive Prefetching
abstract
360-degree video streaming, which is becoming more and more popular as the fast development of VR/AR applications nowadays due to the immersive viewing experience it can offer, poses enormous challenges to the current network infrastructure in terms of high bandwidth and low latency requirements. To address this problem and to ensure the QoE (quality of experience) of end-users, this paper presents CUBIST, a method and system for high-quality 360-degree video streaming in networks with cache nodes at the edge. To the best of our knowledge, it is the first tile-based edge caching solution that incorporates proactive tile prefetching and hierarchical cache organization into reactive caching to maximize the caching benefit while reducing the cost of 360-degree video streaming. Experimental results show that CUBIST can achieve a cache hit ratio of 87 % and improve the effective video bitrate by 12.9 % with most rate transitions being small when compared with the latest FoV-aware edge caching scheme.
Dongbiao He, Jinlei Jiang, Teng Ma 0006, Guangwen Yang 0002, Cédric Westphal, J. J. Garcia-Luna-Aceves, Shutao Xia
ICWS3
2020 AsymNVM: An Efficient Framework for Implementing Persistent Data Structures on Asymmetric NVM Architecture
abstract
The byte-addressable non-volatile memory (NVM) is a promising technology since it simultaneously provides DRAM-like performance, disk-like capacity, and persistency. The current NVM deployment with byte-addressability is \em symmetric, where NVM devices are directly attached to servers. Due to the higher density, NVM provides much larger capacity and should be shared among servers. Unfortunately, in the symmetric setting, the availability of NVM devices is affected by the specific machine it is attached to. High availability can be achieved by replicating data to NVM on a remote machine. However, it requires full replication of data structure in local memory --- limiting the size of the working set. This paper rethinks NVM deployment and makes a case for the \em asymmetric byte-addressable non-volatile memory architecture, which decouples servers from persistent data storage. In the proposed \em \anvm architecture, NVM devices (i.e., back-end nodes) can be shared by multiple servers (i.e., front-end nodes) and provide recoverable persistent data structures. The asymmetric architecture, which follows the industry trend of \em resource disaggregation, is made possible due to the high-performance network (e.g., RDMA). At the same time, \anvm leads to a number of key problems such as, still relatively long network latency, persistency bottleneck, and simple interface of the back-end NVM nodes. We build \em \anvm framework based on \anvm architecture that implements: 1) high performance persistent data structure update; 2) NVM data management; 3) concurrency control; and 4) crash-consistency and replication. The key idea to remove persistency bottleneck is the use of \em operation log that reduces stall time due to RDMA writes and enables efficient batching and caching in front-end nodes. To evaluate performance, we construct eight widely used data structures and two transaction applications based on \anvm framework. In a 10-node cluster equipped with real NVM devices, results show that \anvm achieves similar or better performance compared to the best possible symmetric architecture while enjoying the benefits of disaggregation. We found the speedup brought by the proposed optimizations is drastic, --- 5$\sim$12× among all benchmarks.
Teng Ma 0006, Kang Chen 0001, Yongwei Wu 0001, Xuehai Qian
ASPLOS1
2019 X-RDMA: Effective RDMA Middleware in Large-scale Production Environments
abstract
X-RDMA is a communication middleware deployed and heavily used in Alibaba's large-scale cluster hosting cloud storage and database systems. Unlike recent research projects which purely focus on squeezing out the raw hardware performance, it puts emphasis on robustness, scalability and maintainability of large-scale production clusters. X-RDMA integrates necessary features, not available in current RDMA ecosystem, to release the developers from complex and imperfect details. X-RDMA simplifies the programming model, extends RDMA protocols for application awareness, and proposes mechanisms for resource management with thousands of connections per machine. It also reduces the work for administration and performance tuning with built-in tracing, tuning and monitoring tools. X-RDMA has been deployed in several large-scale clusters with over 4000 servers in Alibaba cloud since 2016. It can save at least 70% development and maintenance time over RDMA, effectively improve performance and reduce network jitter especially when production servers are under pressure. It also helped locate over 30 issues in different layers of productions with over 5000 connections for each server on average.
Teng Ma 0006, Tao Ma 0006, Huaixin Chang, Kang Chen 0001, Hai Jiang 0003, Yongwei Wu 0001
CLUSTER1
2019 RF-RPC: Remote Fetching RPC Paradigm for RDMA-Enabled Network
abstract
Remote Direct Memory Access (RDMA) devices are widely deployed in modern data centers. However, existing RDMA usages lead to a dilemma between performance and redesign cost. Server-reply mode directly replaces socket-based send/receive primitives with corresponding RDMA counterparts. It can not fully harness the power of RDMA devices. Server-bypass is distinct mode provided by RDMA for reaching the hardware limits. This mode can totally bypass the server by using one-sided RDMA operations but at the cost of redesigning software. This paper provides a different approach, called RF-RPC (Remote Fetching RPC Paradigm). Different from server-reply, RF-RPC makes client fetch the results from server using one-sided RDMA instead of waiting the results pushed by server. Different from server-bypass, RF-RPC demands server to process client's requests for supporting programming paradigm like RPC. Thus, RF-RPC can achieve high performance without abandoning traditional programming models. An RF-RPC supported in-memory key-value store shows that the performance can be improved by 1.6× comparing to server-reply paradigm and 4× comparing to server-bypass paradigm.
Yongwei Wu 0001, Teng Ma 0006, Maomeng Su, Kang Chen 0001
IEEE Trans. Parallel Distributed Syst.2
2016 Measuring and Optimizing Distributed Array Programs
abstract
Nowadays, there is a rising trend of building array-based distributed computing frameworks, which are suitable for implementing many machine learning and data mining algorithms. However, most of these frameworks only execute each primitive in an isolated manner and in the exact order defined by programmers, which implies a huge space for optimization. In this paper, we propose a novel array-based programming model, named K asen , which distinguishes itself from models in the existing literature by defining a strict computation and communication model. This model makes it easy to analyze programs' behavior and measure their performance, with which we design a corresponding optimizer that can automatically apply high-level optimizations to the original programs written by programmers. According to our evaluation, the optimizer of K asen can achieve a significant reduction on memory read/write, buffer allocation and network traffic, which leads to a speedup up to 5.82x.
Yongwei Wu 0001, Kang Chen 0001, Teng Ma 0006
Proc. VLDB Endow.4