EDBT 2026 Demo / reviewers in the wild / expert
Yufei Ren
dblp:118/5479
· DBLP profile ↗
18ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 24% Distributed systems · 21% Cloud and datacenter computing · 18% | |
| Artificial intelligence
3 papers |
Trustworthy machine learning · 77% Vision and language · 20% Efficient and distributed learning · 3% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | CGS-Mask: Making Time Series Predictions Intuitive for All · AAAI 2024 |
Machine learning › Trustworthy machine learning › interpretability
post-hoc explanation |
0.8 | 1 | 2024 | CGS-Mask: Making Time Series Predictions Intuitive for All · AAAI 2024 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
saliency methods |
0.8 | 1 | 2024 | CGS-Mask: Making Time Series Predictions Intuitive for All · AAAI 2024 |
Computer vision › Vision and language
multimodal fusion |
0.7 | 1 | 2023 | A Composite Multi-Attention Framework for Intraoperative Hypotension Early Warning · AAAI 2023 |
Medical and health informatics › clinical prediction
clinical event prediction |
0.7 | 1 | 2023 | A Composite Multi-Attention Framework for Intraoperative Hypotension Early Warning · AAAI 2023 |
Distributed systems
fault tolerance |
0.6 | 2 | 2026 | Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure · NSDI 2026 GaDei: On Scale-Up Training as a Service for Deep Learning · ICDM 2017 |
Memory systems
cache management |
0.5 | 2 | 2016 | zExpander: a key-value cache with both high performance and fewer misses · EuroSys 2016 Design, Implementation, and Evaluation of a NUMA-Aware Cache for iSCSI Storage Servers · IEEE Trans. Parallel Distributed Syst. 2015 |
High-performance computing
data-intensive computing |
0.3 | 2 | 2013 | Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systems · SC 2013 Protocols for wide-area data-intensive applications: design and performance issues · SC 2012 |
Distributed systems › fault tolerance › high availability
failover |
0.3 | 1 | 2026 | Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure · NSDI 2026 |
High-performance computing
data transfer |
0.3 | 1 | 2017 | RAMSYS: Resource-Aware Asynchronous Data Transfer with Multicore SYStems · IEEE Trans. Parallel Distributed Syst. 2017 |
Distributed systems › distributed machine learning
parameter server |
0.3 | 1 | 2017 | GaDei: On Scale-Up Training as a Service for Deep Learning · ICDM 2017 |
Parallel and multicore computing › parallel scheduling
resource-aware scheduling |
0.3 | 1 | 2017 | RAMSYS: Resource-Aware Asynchronous Data Transfer with Multicore SYStems · IEEE Trans. Parallel Distributed Syst. 2017 |
Parallel and multicore computing
task scheduling |
0.3 | 1 | 2017 | RAMSYS: Resource-Aware Asynchronous Data Transfer with Multicore SYStems · IEEE Trans. Parallel Distributed Syst. 2017 |
Memory systems › memory compression
cache compression |
0.2 | 1 | 2016 | zExpander: a key-value cache with both high performance and fewer misses · EuroSys 2016 |
Memory systems › cache management
cache replacement |
0.2 | 1 | 2016 | zExpander: a key-value cache with both high performance and fewer misses · EuroSys 2016 |
Memory systems › cache
key-value cache |
0.2 | 1 | 2016 | zExpander: a key-value cache with both high performance and fewer misses · EuroSys 2016 |
Storage systems
key-value storage |
0.2 | 1 | 2016 | zExpander: a key-value cache with both high performance and fewer misses · EuroSys 2016 |
Storage systems
networked storage |
0.2 | 1 | 2015 | Design, Implementation, and Evaluation of a NUMA-Aware Cache for iSCSI Storage Servers · IEEE Trans. Parallel Distributed Syst. 2015 |
Storage systems › networked storage › storage networking
iSCSI |
0.2 | 1 | 2013 | Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systems · SC 2013 |
Storage systems › networked storage
storage area network |
0.2 | 1 | 2013 | Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systems · SC 2013 |
Transport protocols and congestion control
transport protocols |
0.1 | 1 | 2012 | Protocols for wide-area data-intensive applications: design and performance issues · SC 2012 |
High-performance computing › data transfer
wide-area data transfer |
0.1 | 1 | 2012 | Protocols for wide-area data-intensive applications: design and performance issues · SC 2012 |
Memory systems
non-uniform memory access |
0.1 | 2 | 2017 | RAMSYS: Resource-Aware Asynchronous Data Transfer with Multicore SYStems · IEEE Trans. Parallel Distributed Syst. 2017 Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systems · SC 2013 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2017 | GaDei: On Scale-Up Training as a Service for Deep Learning · ICDM 2017 |
Operating systems › i/o › i/o subsystem
i/o scheduling |
0.1 | 1 | 2015 | Design, Implementation, and Evaluation of a NUMA-Aware Cache for iSCSI Storage Servers · IEEE Trans. Parallel Distributed Syst. 2015 |
Operating systems › resource management › process management › CPU scheduling
NUMA-aware scheduling |
0.1 | 1 | 2015 | Design, Implementation, and Evaluation of a NUMA-Aware Cache for iSCSI Storage Servers · IEEE Trans. Parallel Distributed Syst. 2015 |
Distributed systems › middleware
communication middleware |
0.0 | 1 | 2012 | Protocols for wide-area data-intensive applications: design and performance issues · SC 2012 |
Methods — techniques the papers use, named apart from their topics
multimodal fusion network · 1.3multi-attention mechanism · 1.3mask-based saliency · 0.8cellular genetic algorithm · 0.8mini-batch size tuning · 0.6hyperparameter tuning · 0.6i/o request scheduling · 0.4cache alignment · 0.4thread affinity · 0.3pipelining · 0.3asynchronous processing · 0.3data compression · 0.2compact data organization · 0.2parallel data transfer · 0.2task synchronization · 0.1flow control · 0.1connection management · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
Mayank Bansal, Milind Chabbi, Kenneth Bogh, Srikanth Prodduturi, Kevin Xu, David Bell, Ranjib Dey, Yufei Ren, Juan Marcano, Shriniket Kale, Subhav Pradhan, Ivan Beschastnikh, Miguel Covarrubias, Chien-Chih Liao, Sandeep Koushik Sheshadri, Ashish Samant, Sahil Rihan, Nimish Sheth, Albert G. Greenberg, Uday Kiran Medisetty |
NSDI | 9 |
| 2026 | Design and analysis of a novel time-variable parameter zeroing neural network approach for synchronization control of chaotic systems
Yufei Ren, Lin Xiao 0002 |
Neurocomputing | 2 |
| 2025 | Practical Universal Designated Verifier Transitive Signature Proof Scheme for Graph-Based Data SystemsabstractABSTRACT Transitive signatures are a special type of homomorphic signature proposed by Turing Award winners Micali and Rivest, which are highly suitable for authenticating dynamically growing graph‐based data systems. In such a signature scheme, anyone with the signer's public key is allowed to generate a signature for a composed edge , from two signatures on adjacent edges and . To prevent the problem of malicious dissemination of signatures by verifiers leading to data privacy leakage, researchers have proposed a series of universal designated verifier transitive signature (UDVTS) schemes. However, existing work requires that the designated verifier create its own secret‐public key pair using the public key parameters provided by the signer. Besides, these schemes suffer from significant performance defects due to expensive pairing or exponentiation operations. In this work, we design a pairing‐free and exponentiation‐free UDVTS proof scheme based on the SM2 digital signature algorithm and a zero‐knowledge proof scheme. We prove the security of our construction based on rigorous cryptographic assumptions. The performance comparison with related work shows that our UDVTS proof scheme has an optimal computational cost and desirable communication cost. For example, compared to the state‐of‐the‐art work, we reduce the signing cost by and the designated verification cost by . Yucan Xu, Yufei Ren, Wei Wu 0001, Shao-Jun Yang |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | CGS-Mask: Making Time Series Predictions Intuitive for AllabstractArtificial intelligence (AI) has immense potential in time series prediction, but most explainable tools have limited capabilities in providing a systematic understanding of important features over time. These tools typically rely on evaluating a single time point, overlook the time ordering of inputs, and neglect the time-sensitive nature of time series applications. These factors make it difficult for users, particularly those without domain knowledge, to comprehend AI model decisions and obtain meaningful explanations. We propose CGS-Mask, a post-hoc and model-agnostic cellular genetic strip mask-based saliency approach to address these challenges. CGS-Mask uses consecutive time steps as a cohesive entity to evaluate the impact of features on the final prediction, providing binary and sustained feature importance scores over time. Our algorithm optimizes the mask population iteratively to obtain the optimal mask in a reasonable time. We evaluated CGS-Mask on synthetic and real-world datasets, and it outperformed state-of-the-art methods in elucidating the importance of features over time. According to our pilot user study via a questionnaire survey, CGS-Mask is the most effective approach in presenting easily understandable time series prediction results, enabling users to comprehend the decision-making process of AI models with ease. Feng Lu 0003, Wei Li 0058, Cheng Song, Yufei Ren, Albert Y. Zomaya |
AAAI | 5 |
| 2024 | Fine-Grained Lesion Classification Framework for Early Auxiliary DiagnosisabstractThe deep neural networks are envisaged for the early disease diagnosis from medical images. However, in the early stage of the disease, the medical images of patients and healthy people have only subtle visual differences. Distinguishing the medical images for early diagnosis belongs to the Fine-Grained Visual Classification (FGVC) task. Many recent works are based on a standard FGVC learning paradigm: locate the discriminative regions first and then classify by fusing the information of these regions. However, it is still not enough for medical images. Because the shape and size of the lesions are variable, and the relationship between lesions and the background is complex. In order to solve these problems, we propose a fine-grained lesion classification framework for early auxiliary diagnosis. We first locate and extract multiple lesions with different sizes and shapes from the original image and then fuse the feature of lesion and background based on attention mechanism. As shown by experiment results in two real-world clinical data sets, our model can locate accurately and perform better. Feng Lu 0003, Wei Li 0058, Canyu Li, Minghao Fang, Xiaojing Zou, Yufei Ren, Xiaofei Liao, Hai Jin 0001, Albert Y. Zomaya |
IEEE Trans. Comput. Biol. Bioinform. | 10 |
| 2023 | A Composite Multi-Attention Framework for Intraoperative Hypotension Early WarningabstractIntraoperative hypotension (IOH) events warning plays a crucial role in preventing postoperative complications, such as postoperative delirium and mortality. Despite significant efforts, two fundamental problems limit its wide clinical use. The well-established IOH event warning systems are often built on proprietary medical devices that may not be available in all hospitals. The warnings are also triggered mainly through a predefined IOH event that might not be suitable for all patients. This work proposes a composite multi-attention (CMA) framework to tackle these problems by conducting short-term predictions on user-definable IOH events using vital signals in a low sampling rate with demographic characteristics. Our framework leverages a multi-modal fusion network to make four vital signals and three demographic characteristics as input modalities. For each modality, a multi-attention mechanism is used for feature extraction for better model training. Experiments on two large-scale real-world data sets show that our method can achieve up to 94.1% accuracy on IOH events early warning while the signals sampling rate is reduced by 3000 times. Our proposal CMA can achieve a mean absolute error of 4.50 mm Hg in the most challenging 15-minute mean arterial pressure prediction task and the error reduction by 42.9% compared to existing solutions. Feng Lu 0003, Wei Li 0058, Cheng Song, Yufei Ren, Xiaofei Liao, Hai Jin 0001, Ailin Luo, Albert Y. Zomaya |
AAAI | 7 |
| 2017 | GaDei: On Scale-Up Training as a Service for Deep LearningabstractDeep learning (DL) training-as-a-service (TaaS) is an important emerging industrial workload. TaaS must satisfy a wide range of customers who have no experience and/or resources to tune DL hyper-parameters (e.g., mini-batch size and learning rate), and meticulous tuning for each user's dataset is prohibitively expensive. Therefore, TaaS hyper-parameters must be fixed with values that are applicable to all users. Unfortunately, few research papers have studied how to design a system for TaaS workloads. By evaluating the IBM Watson Natural Language Classfier (NLC) workloads, the most popular IBM cognitive service used by thousands of enterprise-level clients globally, we provide empirical evidence that only the conservative hyper-parameter setup (e.g., small mini-batch size) can guarantee acceptable model accuracy for a wide range of customers. Unfortunately, smaller mini-batch size requires higher communication bandwidth in a parameter-server based DL training system. In this paper, we characterize the exceedingly high communication bandwidth requirement of TaaS using representative industrial deep learning workloads. We then present GaDei, a highly optimized shared-memory based scale-up parameter server design. We evaluate GaDei using both commercial benchmarks and public benchmarks and demonstrate that GaDei significantly outperforms the state-of-the-art parameter-server based implementation while maintaining the required accuracy. GaDei achieves near-best-possible runtime performance, constrained only by the hardware limitation. Furthermore, to the best of our knowledge, GaDei is the only scale-up DL system that provides fault-tolerance. Wei Zhang 0057, Minwei Feng, Yunhui Zheng, Yufei Ren, Yandong Wang 0001, Peng Liu 0010, Bing Xiang, Li Zhang 0002, Bowen Zhou 0002, Fei Wang 0001 |
ICDM | 4 |
| 2017 | Lightweight Replication Through Remote Backup Memory Sharing for In-memory Key-Value StoresabstractMemory price will continue dropping in the next few years according to Gartner. Such trend renders it affordable for in-memory key-value stores (IMKVs) to maintain redundant memory-resident copies of each key-value pair to provision enhanced reliability and high availability services. Though contemporary IMKVs have reached unprecedented performance, delivering single-digit microsecond-scale latency with up to tens of millions queries per second throughput, existing replication protocols are unable to keep pace with such an advancement of IMKVs, either incurring unbearable latency overhead or demanding intensive resource usage. Consequently, the adoption of those replication techniques always results in substantial performance degradation.In this paper, we propose MacR, a RDMA-based high-performance and lightweight replication protocol for IMKVs. The design of MacR centers around sharing the remote backup memory to enable RDMA-based replication protocol, and synthesizes a collection of optimizations, including memory allocator cooperative replication and adaptive bulk data synchronization to control the number of network operations and to enhance the recovery performance. Performance evaluations with a variety of YCSB workloads demonstrate that MacR can efficiently outperform alternative replication methods in terms of the throughput while preserving sufficiently low latency overhead. It can also efficiently speed up the recovery process. Yandong Wang 0001, Li Zhang 0002, Michel Hack, Yufei Ren |
MASCOTS | 4 |
| 2017 | Nexus: Bringing Efficient and Scalable Training to Deep Learning FrameworksabstractDemand is mounting in the industry for scalable GPU-based deep learning systems. Unfortunately, existing training applications built atop popular deep learning frameworks, including Caffe, Theano, and Torch, etc, are incapable of conducting distributed GPU training over large-scale clusters. To remedy such a situation, this paper presents Nexus, a platform that allows existing deep learning frameworks to easily scale out to multiple machines without sacrificing model accuracy. Nexus leverages recently proposed distributed parameter management architecture to orchestrate distributed training by a large number of learners spread across the cluster. Through characterizing the run-time behavior of existing single-node based applications, Nexus is equipped with a suite of optimization schemes, including hierarchical and hybrid parameter aggregation, enhanced network and computation layer, and quality-guided communication adjustment, etc, to strengthen the communication channels and resource utilization. Empirical evaluations with a diverse set of deep learning applications demonstrate that Nexus is easy to integrate and can deliver efficient distributed training services to major deep learning frameworks. In addition, Nexus's optimization schemes are highly effective to shorten the training time with targeted accuracy bounds. Yandong Wang 0001, Li Zhang 0002, Yufei Ren, Wei Zhang 0057 |
MASCOTS | 3 |
| 2017 | Analysis of NUMA effects in modern multicore systems for the design of high-performance data transfer applications
Yufei Ren, Dantong Yu, Shudong Jin |
Future Gener. Comput. Syst. | 2 |
| 2017 | RAMSYS: Resource-Aware Asynchronous Data Transfer with Multicore SYStemsabstractHigh-speed data transfer is vital to data-intensive computing that often requires moving large data volumes efficiently within a local data center and among geographically dispersed facilities. Effective utilization of the abundant resources in modern multicore environments for data transfer remains a persistent challenge, particularly, for Non-Uniform Memory Access (NUMA) systems wherein the locality of data accessing is an important factor. This requires rethinking how to exploit parallel access to data and to optimize the storage and network I/Os. We address this challenge and present a novel design of asynchronous processing and resource-aware task scheduling in the context of high-throughput data replication. Our software allocates multiple sets of threads to different stages of the processing pipeline, including storage I/O and network communication, based on their capacities. Threads belonging to each stage follow an asynchronous model, and attain high performance via multiple locality-aware and peer-aware mechanisms, such as task grouping, buffer sharing, affinity control and communication protocols. Our design also integrates high performance features to enhance the scalability of data transfer in several scenarios, e.g., file-level sorting, block-level asynchrony, and thread-level pipelining. Our experiments confirm the advantages of our software under different types of workloads and dynamic environments with contention for shared resources, including a 28-160 percent increase in bandwidth for transferring large files, 1.7-66 times speed-up for small files, and up to 108 percent larger throughput for mixed workloads compared with three state of the art alternatives, GridFTP , BBCP and Aspera. Yufei Ren, Dantong Yu, Shudong Jin |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | zExpander: a key-value cache with both high performance and fewer missesabstractWhile key-value (KV) cache, such as memcached, dedicates a large volume of expensive memory to holding performance-critical data, it is important to improve memory efficiency, or to reduce cache miss ratio without adding more memory. As we find that optimizing replacement algorithms is of limited effect for this purpose, a promising approach is to use a compact data organization and data compression to increase effective cache size. However, this approach has the risk of degrading the cache's performance due to additional computation cost. A common perception is that a high-performance KV cache is not compatible with use of data compacting techniques. Xingbo Wu, Li Zhang 0002, Yandong Wang 0001, Yufei Ren, Michel Hack, Song Jiang 0001 |
EuroSys | 4 |
| 2015 | Resources-Conscious Asynchronous High-Speed Data Transfer in Multicore Systems: Design, Optimizations, and EvaluationabstractOne constant challenge in multicourse systems is to utilize fully the abundant resources, while assuring superior performance for individual tasks, particularly, in Non-uniform Memory Access (NUMA) systems where the locality of access is an important factor. To achieve this goal requires rethinking how to exploit parallel data access and I/O related optimizations. In the context of developing software for high-speed data transfer, we offer a novel design using asynchronous processing, and detail the advantages of resources-conscious task scheduling. In our design, multiple sets of threads are allocated to the different stages of the processing pipeline based on the capacity of resources, including storage I/O, and network communication operations. The threads in these stages are executed in an asynchronous mode, and they communicate efficiently via localized mechanisms in NUMA systems, e.g., task grouping, buffer memory, and locks. With this design, multiple effective optimizations are seamlessly integrated particularly for improving the performance and scalability of end-to-end data transfer. To validate the benefits of the design and optimizations therein, we conducted extensive experiments on the state-of-the-art multicourse systems. Our results highlighted the performance advantages of our software across different typical workloads, compared to the widely adopted data transfer tools, Graft and BBCP. Yufei Ren, Dantong Yu, Shudong Jin |
IPDPS | 2 |
| 2015 | Design, Implementation, and Evaluation of a NUMA-Aware Cache for iSCSI Storage ServersabstractIn an iSCSI based storage area network, target hosts serve concurrent I/O requests from initiators to achieve both high throughput and low latency. Existing iSCSI leverages the OS page cache to ensure data sharing and reuse. However, the non-uniform memory access (NUMA) architecture introduces another dimension of complexity, i.e., asymmetric memory access in multi-core and many-core platforms. Within a NUMA platform, an iSCSI target often dispatches an access request with a cache hit to an I/O thread remote to cached data, and thus cannot fully utilize multi-core systems. We encounter this problem in the context of ultra high-speed data transfer between two iSCSI storage systems, during which inferior NUMA remote memory access lags behind available high network bandwidth, and thereby becomes a bottleneck of the entire end-to-end data transfer path. We design a NUMA-aware cache mechanism to align cache memory with local NUMA nodes and threads, and then schedule I/O requests to those threads that are local to the data being accessed. This NUMA-aware solution results in lower access latency and higher system throughput. We implement a cache system within the Linux SCSI target framework, and evaluated it on our NUMA-based iSCSI testbed. Experimental results show the NUMA-aware cache can significantly improve the performance of iSCSI as measured by several benchmark tools and confirm its viability in data intensive applications and real-life workloads. Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Characterization of Input/Output Bandwidth Performance Models in NUMA Architecture for Data Intensive ApplicationsabstractData-intensive applications frequently rely on multicore computer systems, in which Non-Uniform Memory Access (NUMA) is a dominant architecture. To transfer data into and out from these high-performance computers becomes a bottleneck, and thus it is crucial to understand their I/O performance characteristics. However, the complexity in NUMA architecture presents a new challenge in modeling its I/O access cost, and thus lead to difficulties in configuring proper processor and memory affinity. In this paper, we show that existing NUMA experimental methods and metrics are inappropriate on contemporary high-end systems. We characterize a state-of-the-art NUMA host, and propose, to the best of our knowledge, the first methodology to simulate I/O operations using memory semantics, and model the I/O bandwidth performance. Our methodology is thoroughly tested and validated by mapping multiple parallel I/O streams to different sets of hardware components (CPU, memory, network cards, and SSDs) and by measuring the performance of each mapping. The experimental results and analysis reveal that our methodology can dramatically reduce characterization workload, accurately estimate the overall I/O performance, and effectively mitigate resource contention among I/O tasks. Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi |
ICPP | 2 |
| 2013 | Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systemsabstractData-intensive applications place stringent requirements on the performance of both back-end storage systems and front-end network interfaces. However, for ultra high-speed data transfer, for example, at 100 Gbps and higher, the effects of multiple bottlenecks along a full end-to-end path, have not been resolved efficiently. In this paper, we describe our implementation of an end-to-end data transfer software at such high-speeds. At the back-end, we construct a storage area network with the iSCSI protocols, and utilize efficient RDMA technology. At the front-end, we design network communication software to transfer data in parallel, and utilize NUMA techniques to maximize the performance of multiple network interfaces. We demonstrate that our system can deliver the full 100 Gbps end-to-end data transfer throughput. The software product is tested rigorously and demonstrated applicable to supporting various data-intensive applications that constantly move bulk data within and across data centers. Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi |
SC | 1 |
| 2013 | Design and testbed evaluation of RDMA-based middleware for high-performance data transfer applications
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi |
J. Syst. Softw. | 1 |
| 2012 | Protocols for wide-area data-intensive applications: design and performance issuesabstractProviding high-speed data transfer is vital to various data-intensive applications.While there have been remarkable technology advances to provide ultra-high-speed network bandwidth, existing protocols and applications may not be able to fully utilize the bare-metal bandwidth due to their inefficient design.We identify the same problem remains in the field of Remote Direct Memory Access (RDMA) networks.RDMA offloads TCP/IP protocols to hardware devices.However, its benefits have not been fully exploited due to the lack of efficient software and application protocols, in particular in wide-area networks.In this paper, we address the design choices to develop such protocols.We describe a protocol implemented as part of a communication middleware.The protocol has its flow control, connection management, and task synchronization.It maximizes the parallelism of RDMA operations.We demonstrate its performance benefit on various local and wide-area testbeds, including the DOE ANI testbed with RoCE links and InfiniBand links. Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi, Brian Tierney, Eric Pouyoul |
SC | 1 |