VLDB 2026 Research / reviewers in the wild / expert
Yanfei Guo
dblp:117/8957
· DBLP profile ↗
70ranked-venue papers
21as first author
43since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 14 · 7 first-author · 14 since 2021Security and privacy · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRANet: Pathological relationship perception and dual attention guided network for diabetic retinopathy grading
Yanfei Guo, Hangli Du, Yuncui Wang, Dengwang Li |
Appl. Intell. | 1 |
| 2026 | MAINet: Multi-scale attention interaction network for diabetic retinopathy lesion segmentation
Yanfei Guo, Chenglong Yang, Yuanke Zhang, Fei Ma 0004, Jing Meng 0001 |
Expert Syst. Appl. | 1 |
| 2026 | Fusion-from-zero network with information fusion for multimodal head and neck tumor segmentation
Yanjun Peng, Yanfei Guo, Hengzhong Li |
Expert Syst. Appl. | 3 |
| 2026 | Binocular dual attention interaction siamese network for diabetic retinopathy grading
Yanfei Guo, Yuncui Wang, Fei Ma 0004, Jing Meng 0001, Xiaofeng Zou |
Inf. Sci. | 1 |
| 2026 | TVFNet: text and visual attention feature fusion network for multi-lesion segmentation of diabetic retinopathy
Yanfei Guo, Yuanke Zhang, Fei Ma 0004, Jing Meng 0001, Shasha Yuan, Jindong Sun |
Neural Comput. Appl. | 1 |
| 2026 | DDNet: dual-domain network for OCT angiography retinal vessel segmentation
Fei Ma 0004, Zhaohui Zhang 0006, Fen Yan, Meirong Chen, Yuefeng Ma, Yanfei Guo, Jing Meng 0001, Ronghua Cheng |
J. Supercomput. | 6 |
| 2026 | Practical Machine Learning Autotuning for Large-Scale Collective Communication
Michael Wilkins, Yanfei Guo, Rajeev Thakur, Peter A. Dinda, Nikos Hardavellas |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | TMSurv: A Multi-task Assisted Multimodal and Multi-granularity Survival Prediction Network for Head and Neck CancerabstractSurvival prediction for head and neck cancer (HNC) is crucial for clinical treatment. However, current methods extract features only from the tumors and fail to fully use multimodal data. To address these limitations, we propose a multi-task assisted multimodal and multi-granularity survival prediction network. First, a multi-task assisted learning module performs tumor segmentation while capturing global contextual information and prognostic features from both 2D slices and 3D volumes. Second, a multi-graph feature aggregation module constructs distinct graph representations to fuse contextual information from imaging data and adaptively learn relationships within clinical records. Finally, a multimodal and multi-granularity information interaction module integrates 2D and 3D imaging features with clinical and radiomic data for final survival prediction. Experimental results on three HNC datasets demonstrate that the proposed method achieves superior prognostic performance. Yanjun Peng, Yanfei Guo |
BIBM | 3 |
| 2025 | Semi-Supervised Gaussian Mixture Variational Autoencoder with Graph Representation for Epileptic Seizure DetectionabstractAccurate electroencephalogram(EEG) annotation is essential for seizure detection but costly and error-prone, which can affect subsequent tasks. Moreover, the brain is a non-Euclidean topological structure, which contains spatial information for seizure detection. Based on these, this paper proposes a semi- supervised model based on Gaussian mixture variational autoencoder with graph representation, named GGMVAE. Firstly, for unlabeled EEG signals, we construct an adjacency matrix via Pearson correlation between channels. Then, matrix and EEG features are fed into a Gaussian Mixture VAE to learn the temporal and spatiall features through unsupervised training. Finally, the pre-trained encoder extracts low-dimensional features from partially labeled data for classification. The method is evaluated on the epilepsy dataset at the University of Helsinki, achieving the accuracy of 97.27 %, precision of 96.17 %, recall of 98.41 %, and F1-score of 97.28 %. The results indicate that this semi-supervised model can effectively learn the temporal and spatial features of EEG signals, improving the performance of seizure detection.. Shasha Yuan, Chenchen Jiang, Manman Yuan, Qianqian Ren, Yanfei Guo |
BIBM | 5 |
| 2025 | ghZCCL: Advancing GPU-aware Collective Communications with Homomorphic CompressionabstractIn the exascale computing era, collective communication has emerged as a significant bottleneck for GPU-based applications, as network bandwidth lags behind rapid GPU advancements.While traditional GPU-aware approaches employ error-bounded lossy compression to mitigate this issue, they incur substantial decompression-operation-compression (DOC) overhead.To overcome these limitations, we introduce ghZCCL, a first-ever GPU-aware homomorphic compression-accelerated collective communications library that enables direct computation and communication on compressed data, eliminating the DOC workflow.We design the first GPU homomorphic compressor, surpassing the fastest existing GPU lossy compressor, cuSZp2, by 3.47-3.89×for DOC workloads.We also propose co-design strategies to further optimize GPU-aware collective communications with homomorphic compression.Experiments on up to 512 NVIDIA A100 GPUs show that ghZCCL outperforms three state-of-the-art communication libraries-gZCCL, NCCL, and Cray MPI-by achieving speedups of up to 2.29×, 5.81×, and 188×, respectively, while maintaining high data accuracy. Jiajun Huang 0001, Sheng Di, Yafan Huang, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
ICS | 6 |
| 2025 | Research on Detection and Reconstruction of Multiple Types of Anomalies in Wind Speed-Power Data of Wind Farms
Shouyi Chen, Yiyi He, Yanfei Guo, Chung-Lun Wei |
KSEM (5) | 3 |
| 2025 | Examining MPI and its Extensions for Asynchronous Multithreaded Communication
Jiakun Yan, Marc Snir, Yanfei Guo |
EuroMPI | 3 |
| 2025 | Implementing True MPI Sessions and Evaluating MPI Initialization Scalability
Hui Zhou 0012, Kenneth Raffenetti, Yanfei Guo, Michael Wilkins, Rajeev Thakur |
EuroMPI | 3 |
| 2025 | Dynamic momentum contrastive learning network for diabetic retinopathy grading
Yanfei Guo, Chenglong Yang, Hangli Du, Yuanke Zhang, Fei Ma 0004, Shasha Yuan |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | DSCN-Net: domain-specific contrastive network for unsupervised low-dose CT denoising
Rui Zhang 0137, Yuanke Zhang, Yanfei Guo, Hanxiang Wang, Bingbing Wei, Fei Ma 0004, Jing Meng 0001, Jianlei Liu, Hongbing Lu |
Neurocomputing | 3 |
| 2025 | SMFDNet: spatial and multi-frequency domain network for OCT angiography retinal vessel segmentation
Sien Li, Fei Ma 0004, Fen Yan, Jing Meng 0001, Yanfei Guo, Hongjuan Liu, Ronghua Cheng |
J. Supercomput. | 5 |
| 2025 | WHANet: wavelet and hybrid attention network for vessel segmentation in OCTA fundus images
Shuxin Xue, Zhaohui Zhang 0006, Fen Yan, Fei Ma 0004, Guangmei Jia, Yanfei Guo, Yuefeng Ma, Xiaofei Ai, Jing Meng 0001 |
J. Supercomput. | 6 |
| 2025 | WS-SAM: self-prompting SAM with wavelet and spatial domain for OCTA retinal vessel segmentation
Zhaohui Zhang 0006, Fei Ma 0004, Hongjuan Liu, Xiwei Dong, Yanfei Guo, Jing Meng 0001 |
J. Supercomput. | 5 |
| 2024 | gZCCL: Compression-Accelerated Collective Communication Framework for GPU ClustersabstractGPU-aware collective communication has become a major bottleneck for modern computing platforms as GPU computing power rapidly rises. A traditional approach is to directly integrate lossy compression into GPU-aware collectives, which can lead to serious performance issues such as underutilized GPU devices and uncontrolled data distortion. In order to address these issues, in this paper, we propose gZCCL, a first-ever general framework that designs and optimizes GPU-aware, compression-enabled collectives with an accuracy-aware design to control error propagation. To validate our framework, we evaluate the performance on up to 512 NVIDIA A100 GPUs with real-world applications and datasets. Experimental results demonstrate that our gZCCL-accelerated collectives, including both collective computation (Allreduce) and collective data movement (Scatter), can outperform NCCL as well as Cray MPI by up to 4.5 × and 28.7 ×, respectively. Furthermore, our accuracy evaluation with an image-stacking application confirms the high reconstructed data quality of our accuracy-aware framework. Jiajun Huang 0001, Sheng Di, Xiaodong Yu 0001, Jinyang Liu 0003, Yafan Huang, Kenneth Raffenetti, Hui Zhou 0012, Kai Zhao 0008, Xiaoyi Lu 0001, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
ICS | 13 |
| 2024 | An Optimized Error-controlled MPI Collective Framework Integrated with Lossy CompressionabstractWith the ever-increasing computing power of supercomputers and the growing scale of scientific applications, the efficiency of MPI collective communications turns out to be a critical bottleneck in large-scale distributed and parallel processing. The large message size in MPI collectives is particularly concerning because it can significantly degrade the overall parallel performance. To address this issue, prior research simply applies the off-the-shelf fix-rate lossy compressors in the MPI collectives, leading to suboptimal performance, limited generalizability, and unbounded errors. In this paper, we propose a novel solution, called C-Coll, which leverages error-bounded lossy compression to significantly reduce the message size, resulting in a substantial reduction in communication cost. The key contributions are three-fold. (1) We develop two general, optimized lossy-compression-based frameworks for both types of MPI collectives (collective data movement as well as collective computation), based on their particular characteristics. Our framework not only reduces communication cost but also preserves data accuracy. (2) We customize SZx, an ultra-fast error-bounded lossy compressor, to meet the specific needs of collective communication. (3) We integrate C-Coll into multiple collectives, such as MPI Allreduce, MPI Scatter, and MPI Bcast, and perform a comprehensive evaluation based on real-world scientific datasets. Experiments show that our solution outperforms the original MPI collectives as well as multiple baselines and related efforts by 1.8–2.7×. Jiajun Huang 0001, Sheng Di, Xiaodong Yu 0001, Jinyang Liu 0003, Xiaoyi Lu 0001, Kenneth Raffenetti, Hui Zhou 0012, Kai Zhao 0008, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
IPDPS | 13 |
| 2024 | Accelerating Lossy and Lossless Compression on Emerging BlueField DPU ArchitecturesabstractData compression has become a crucial technique in addressing performance bottlenecks caused by increasing data volumes in High-Performance Computing (HPC), Big Data, and Deep Learning (DL). Despite its potential to boost system performance, recent studies have identified significant challenges with existing compression methods, mainly due to their high computational demands amidst continuously growing data sizes. Concurrently, the advent of Data Processing Units (DPUs), equipped with programmable System-on-Chip (SoC) and specialized compression accelerators, offers a promising opportunity to alter the landscape of data compression. This paper explores the complexities and potential of leveraging NVIDIA BlueField DPUs to accelerate lossy and lossless compression. Towards this, we introduce PEDAL, an innovative library that leverages the hardware capabilities of DPUs to unify and optimize data compression designs. Moreover, we seamlessly co-design PEDAL with the popular MPICH MPI library, demonstrating up to 101x speedup in compression time and 88x decrease in communication latency. Drawing on these achievements, we share our experience with various research communities about accelerating data compression on DPUs in communication-oriented HPC scenarios. Yuke Li 0003, Arjun Kashyap, Weicong Chen 0002, Yanfei Guo, Xiaoyi Lu 0001 |
IPDPS | 4 |
| 2024 | POSTER: Optimizing Collective Communications with Error-bounded Lossy Compression for GPU ClustersabstractGPU-aware collective communication has become a major bottleneck for modern computing platforms as GPU computing power rapidly rises. To address this issue, traditional approaches integrate lossy compression directly into GPU-aware collectives, which still suffer from serious issues such as underutilized GPU devices and uncontrolled data distortion. In this paper, we propose GPU-LCC, a general framework that designs and optimizes GPU-aware, compression-enabled collectives with well-controlled error propagation. To validate our framework, we evaluate the performance on up to 64 NVIDIA A100 GPUs with real-world applications and datasets. Experimental results demonstrate that our GPU-LCC-accelerated collective computation (Allreduce), can outperform NCCL as well as Cray MPI by up to 3.4× and 18.7×, respectively. Furthermore, our accuracy evaluation with an image-stacking application confirms the high reconstructed data quality of our accuracy-aware framework. Jiajun Huang 0001, Sheng Di, Xiaodong Yu 0001, Jinyang Liu 0003, Yafan Huang, Kenneth Raffenetti, Hui Zhou 0012, Kai Zhao 0008, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
PPoPP | 12 |
| 2024 | hZCCL: Accelerating Collective Communication with Co-Designed Homomorphic CompressionabstractAs network bandwidth struggles to keep up with rapidly growing computing capabilities, the efficiency of collective communication has become a critical challenge for exa-scale distributed and parallel applications. Traditional approaches directly utilize error-bounded lossy compression to accelerate collective computation operations, exposing unsatisfying performance due to the expensive decompression-operation-compression (DOC) workflow. To address this issue, we present a first-ever homomorphic compression-communication co-design, hZCCL, which enables operations to be performed directly on compressed data, saving the cost of time-consuming decompression and recompression. In addition to the co-design framework, we build a light-weight compressor, optimized specifically for multi-core CPU platforms. We also present a homomorphic compressor with a run-time heuristic to dynamically select efficient compression pipelines for reducing the cost of DOC handling. We evaluate $h \mathbf{Z C C L}$ with up to 512 nodes and across five application datasets. The experimental results demonstrate that our homomorphic compressor achieves a CPU throughput of up to $379.08 \mathrm{~GB} / \mathrm{s}$, surpassing the conventional DOC workflow by up to $36.53 \times$. Moreover, our $\boldsymbol{h Z C C L}$-accelerated collectives outperform two state-of-the-art baselines, delivering speedups of up to $2.12 \times$ and $6.77 \times$ compared to original MPI collectives in single-thread and multi-thread modes, respectively, while maintaining data accuracy. Jiajun Huang 0001, Sheng Di, Xiaodong Yu 0001, Jinyang Liu 0003, Zizhe Jian, Xin Liang 0001, Kai Zhao 0008, Xiaoyi Lu 0001, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur |
SC | 12 |
| 2023 | Exploring Wavelet Transform Usages for Error-bounded Scientific Data CompressionabstractTo address the challenges raised by the data management of exascale scientific data, error-bounded lossy compression has been proposed and well-researched as a prominent solution. Among the existing works, a recent trend leverages wavelet transforms in the error-bounded lossy compression task to effectively capture long-term data correlations within the inputs. Applying those transforms as data preprocessors and decorrelators, wavelet-based lossy compressors have achieved optimized compression rate-distortion on several datasets. However, certain significant limitations of wavelet-based compressors have also been observed: On one hand, attributed to the high computational cost of wavelet transforms, wavelet-based compressors suffer from relatively low computational efficiencies compared to other state-of-the-art compressors. On the other hand, one certain type of wavelet transform cannot perform well on all variations of scientific data. Consequently, to further fine-tune the wavelet-based scientific data lossy compression, more in-depth and systematic research and analysis needs to be conducted. In this paper, based on the FAZ auto-tuning-based modular compression framework, we have integrated a great number of wavelet transforms into the framework and evaluated them with various real-world scientific datasets and fields. From the analysis of those evaluations and the comparison to existing state-of-the-art wavelet-based and non-wavelet-based error-bounded lossy compressors, we conclude and present several essential takeaways for designing and optimizing the wavelet-based scientific error-bounded lossy compressor. Jiajun Huang 0001, Jinyang Liu 0003, Sheng Di, Zizhe Jian, Shixun Wu, Kai Zhao 0008, Zizhong Chen, Yanfei Guo, Franck Cappello |
IEEE Big Data | 9 |
| 2023 | PiP-MColl: Process-in-Process-based Multi-object MPI CollectivesabstractIn the era of exascale computing, the adoption of a large number of CPU cores and nodes by high-performance computing (HPC) applications has made MPI collective performance increasingly crucial. As the number of cores and nodes increases, the importance of optimizing MPI collective performance becomes more evident. Current collective algorithms, including kernel-assisted inter-process data exchange techniques and data sharing based shared-memory approaches, are prone to significant performance degradation due to the overhead of system calls and page faults or the cost of extra data-copy latency. These issues can negatively impact the efficiency and scalability of HPC applications. To address these issues, we propose PiP-MColl, a Process-in-Process-based Multi-object Interprocess MPI Collective design that maximizes small message MPI collective performance at scale. We also present specific designs to boost the performance for larger messages, such that we observe a comprehensive improvement for a series of message sizes beyond small messages. PiP-MColl features efficient multiple sender and receiver collective algorithms and leverages Process-in-Process shared memory techniques to eliminate unnecessary system call, page fault overhead and extra data copy, which results in improved intra- and inter-node message rate and throughput. Experimental results demonstrate that PiP-MColl significantly outperforms popular MPI libraries, including OpenMPI, MVAPICH2, and Intel MPI, by up to 4.6X for the MPI collectives MPI_Scatter, MPI_Allgather, and MPI_Allreduce. Jiajun Huang 0001, Kaiming Ouyang, Jinyang Liu 0003, Min Si, Kenneth Raffenetti, Hui Zhou 0012, Atsushi Hori, Zizhong Chen, Yanfei Guo, Rajeev Thakur |
CLUSTER | 10 |
| 2023 | Generalized Collective Algorithms for the Exascale EraabstractExascale supercomputers have renewed the exigence of improving distributed communication, specifically MPI collectives. Previous works accelerated collectives for specific scenarios by changing the radix of the collective algorithms. However, these approaches fail to explore the interplay between modern hardware features, such as multi-port networks, and software features, such as message size. In this paper, we present a novel approach that uses system-agnostic, generalized (i.e., variableradix) algorithms to capture relevant features and provide broad speedups for upcoming exascale-class supercomputers.We identify hardware commonalities found on announced exascale systems and three omnipresent communication kernels (binomial tree, ring, and recursive doubling) that can be generalized to better leverage these features, creating 10 total implementations. For each kernel, we develop analytical models to intuit algorithm performance with varying radix values.Experiments on the world’s first exascale supercomputer (Frontier at ORNL) and a pre-exascale system (Polaris at ANL) show that our generalized algorithms outperform the baseline open-source and proprietary vendor MPI implementations by a significant margin, up to over 4.5x. We empirically determine optimal algorithms and parameter values, identifying where the analytical models are accurate and where hardware features directly determine performance. Most notably, we show how a single, system-agnostic implementation of a generalized algorithm can optimize for multiple hardware/software features across multiple systems. Michael Wilkins, Hanming Wang, Peizhi Liu, Bangyen Pham, Yanfei Guo, Rajeev Thakur, Peter A. Dinda, Nikos Hardavellas |
CLUSTER | 5 |
| 2023 | Characterizing Lossy and Lossless Compression on Emerging BlueField DPU ArchitecturesabstractThe Data Processing Unit (DPU) (i.e., programmable SmartNICs with System-on-Chip or SoC cores) has emerged as a valuable supplementary resource to the host CPU. The DPU architecture has been attracting significant attention within High-Performance Computing (HPC) and data center clusters due to its advanced capabilities and accelerators, which include a hardware-based data compression engine. This positions the DPU as a prospective tool for accelerating and offloading compression workloads from the hosts, which can potentially speed up data-intensive applications. The convergence of Big Data, HPC, and Machine Learning (ML) systems has rendered large data volumes a major performance bottleneck in message communication and data storage. While compression can boost performance, recent studies reveal that compression techniques (e.g., lossy and lossless) are compute-intensive and time-consuming, particularly with larger data sizes. Consequently, this paper characterizes the performance of three lossy (SZ3) and lossless (DEFLATE and zlib) compression algorithms with seven real-world data sets on the popular NVIDIA’s BlueField DPUs to explore potential opportunities for offloading these workloads from the host. We find that compared to DPU’s SoC cores, DPU’s hardware compression engine can obtain up to 26.8x performance speedup. Furthermore, we discuss the challenges and opportunities associated with employing NVIDIA’s BlueField DPUs to accelerate lossy and lossless compression/decompression workloads. Our research discloses five important takeaways which shed light on future research directions for lossy and lossless compressions on DPUs. Yuke Li 0003, Arjun Kashyap, Yanfei Guo, Xiaoyi Lu 0001 |
HOTI | 3 |
| 2023 | Accelerating MPI Collectives with Process-in-Process-based Multi-object TechniquesabstractIn the exascale computing era, optimizing MPI collective performance in high-performance computing (HPC) applications is critical. Current algorithms face performance degradation due to system call overhead, page faults, or data-copy latency, affecting HPC applications' efficiency and scalability. To address these issues, we propose PiP-MColl, a Process-in-Process-based Multi-object Inter-process MPI Collective design that maximizes small message MPI collective performance at scale. PiP-MColl features efficient multiple sender and receiver collective algorithms and leverages Process-in-Process shared memory techniques to eliminate unnecessary system call, page fault overhead, and extra data copy, improving intra- and inter-node message rate and throughput. Our design also boosts performance for larger messages, resulting in comprehensive improvement for various message sizes. Experimental results show that PiP-MColl outperforms popular MPI libraries, including OpenMPI, MVAPICH2, and Intel MPI, by up to 4.6X for MPI collectives like MPI_Scatter and MPI_Allgather. Jiajun Huang 0001, Kaiming Ouyang, Jinyang Liu 0003, Min Si, Kenneth Raffenetti, Hui Zhou 0012, Atsushi Hori, Zizhong Chen, Yanfei Guo, Rajeev Thakur |
HPDC | 10 |
| 2023 | Quantifying the Performance Benefits of Partitioned Communication in MPIabstractPartitioned communication was introduced in MPI 4.0 as a user-friendly interface to support pipelined communication patterns, particularly common in the context of MPI+threads. It provides the user with the ability to divide a global buffer into smaller independent chunks, called partitions, which can then be communicated independently. In this work we first model the performance gain that can be expected when using partitioned communication. Next, we describe the improvements we made to MPICH to enable those gains and provide a high-quality implementation of MPI partitioned communication. We then evaluate partitioned communication in various common use cases and assess the performance in comparison with other MPI point-to-point and one-sided approaches. Specifically, we first investigate two scenarios commonly encountered for small partition sizes in a multithreaded environment: thread contention and overhead of using many partitions. We propose two solutions to alleviate the measured penalty and demonstrate their use. We then focus on large messages and the gain obtained when exploiting the delay resulting from computations or load imbalance. We conclude with our perspectives on the benefits of partitioned communication and the various results obtained. Thomas Gillis, Kenneth Raffenetti, Hui Zhou 0012, Yanfei Guo, Rajeev Thakur |
ICPP | 4 |
| 2023 | Frustrated With MPI+Threads? Try MPIxThreads!abstractMPI + Threads, embodied by the MPI/OpenMP hybrid programming model, is a parallel programming paradigm where threads are used for on-node shared-memory parallelization and MPI is used for multi-node distributed-memory parallelization. OpenMP provides an incremental approach to parallelize code, while MPI, with its isolated address space and explicit messaging API, affords straightforward paths to obtain good parallel performance. However, MPI + Threads is not an ideal solution. Since MPI is unaware of the thread context, it cannot be used for interthread communication. This results in duplicated efforts to create separate and sometimes nested solutions for similar parallel tasks. In addition, because the MPI library is required to obey message-ordering semantics, mixing threads and MPI via MPI_THREAD_MULTIPLE can easily result in miserable performance due to accidental serializations. Hui Zhou 0012, Kenneth Raffenetti, Junchao Zhang 0002, Yanfei Guo, Rajeev Thakur |
EuroMPI | 4 |
| 2023 | A histogram-driven generative adversarial network for brain MRI to CT synthesis
Yanjun Peng, Jindong Sun, Yande Ren, Dapeng Li 0001, Yanfei Guo |
Knowl. Based Syst. | 5 |
| 2023 | Near-Lossless MPI Tracing and Proxy Application AutogenerationabstractTraces of MPI communications are used by many performance analysis and visualization tools. Storing exhaustive traces of large-scale MPI applications is infeasible, however, because of their large volume. Aggregated or lossy MPI traces are smaller but provide much less information. In this paper we present Pilgrim, a near-lossless MPI tracing tool that, by using sophisticated compression techniques, generates small trace files at large scales and incurs only moderate overheads. We perform comprehensive studies of various compression techniques used for storing timestamps associated with each call. This timing information is essential for analysis purposes such as skews study. To demonstrate the usefulness of the detailed information stored by Pilgrim, we present a proxy application generator that can generate proxy apps that preserve original communication patterns from the Pilgrim traces. Chen Wang 0004, Yanfei Guo, Pavan Balaji, Marc Snir |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | ACCLAiM: Advancing the Practicality of MPI Collective Communication Autotuning Using Machine LearningabstractMPI collective communication is an omnipresent communication model for high-performance computing (HPC) systems. The performance of a collective operation depends strongly on the algorithm used to implement it. MPI libraries use inaccurate heuristics to select these algorithms, causing applications to suffer unnecessary slowdowns. Machine learning (ML)-based autotuners are a promising alternative. ML autotuners can intelligently select algorithms for individual jobs, resulting in near-optimal performance. However, these approaches currently spend more time training than they save by accelerating applications, rendering them impractical. We make the case that ML-based collective algorithm selection autotuners can be made practical and accelerate production applications on large-scale supercomputers. We identify multiple impracticalities in the existing work, such as inefficient training point selection and ignoring non-power-of-two feature values. We address these issues through variance-based point selection and model testing alongside topology-aware benchmark paral-lelization. Our approach minimizes training time by eliminating unnecessary training points and maximizing machine utilization. We incorporate our improvements in a prototype active learning system, ACCLAiM (Advancing Collective Communication (L) Autotuning using Machine Learning). We show that each of ACCLAiM's advancements significantly reduces training time compared with the best existing machine learning approach. Then we apply ACCLAiM on a leadership-class supercomputer and demonstrate the conditions where ACCLAiM can accelerate HPC applications, proving the advantage of ML autotuners in a production setting for the first time. Michael Wilkins, Yanfei Guo, Rajeev Thakur, Peter A. Dinda, Nikos Hardavellas |
CLUSTER | 2 |
| 2022 | MPIX Stream: An Explicit Solution to Hybrid MPI+X ProgrammingabstractThe hybrid MPI+X programming paradigm, where X refers to threads or GPUs, has gained prominence in the high-performance computing arena. This corresponds to a trend of system architectures growing more heterogeneous. The current MPI standard only specifies the compatibility levels between MPI and threading runtimes. No MPI concept or interface exists for applications to pass thread context or GPU stream context to MPI implementations explicitly. This lack has made performance optimization complicated in some cases and impossible in other cases. We propose a new concept in MPI, called MPIX stream, to represent the general serial execution context that exists in X runtimes. MPIX streams can be directly mapped to threads or GPU execution streams. Passing thread context into MPI allows implementations to precisely map the execution contexts to network endpoints. Passing GPU execution context into MPI allows implementations to directly operate on GPU streams, lowering the CPU/GPU synchronization cost. Hui Zhou 0012, Kenneth Raffenetti, Yanfei Guo, Rajeev Thakur |
EuroMPI | 3 |
| 2022 | Multiple lesion segmentation in diabetic retinopathy with dual-input attentive RefineNet
Yanfei Guo, Yanjun Peng |
Appl. Intell. | 1 |
| 2022 | MMNet: A multi-scale deep learning network for the left ventricular segmentation of cardiac MRI images
Yanjun Peng, Dapeng Li 0001, Yanfei Guo, Bin Zhang 0052 |
Appl. Intell. | 4 |
| 2022 | MFAUNet: Multiscale feature attentive U-Net for cardiac MRI structural segmentationabstractAbstract The accurate and robust automatic segmentation of cardiac structures in magnetic resonance imaging (MRI) is significant in calculating cardiac clinical functional indices, and diagnosing heart diseases. Most U‐Net based methods use pooling, transposed convolution, and skip connection operations to integrate the multiscale features for improved segmentation in cardiac MRI. However, this architecture lacks adequate semantic connection between the channel and spatial information, and robustness in segmenting objects with significant shape variations. In this paper, a new multiscale feature attentive U‐Net for cardiac MRI structural segmentation method is proposed. An attention mechanism is adopted after concatenating the multi‐level features to aggregate different scale features and determine on which features to focus. Cascade and parallel dilated convolution is also employed in the decoder blocks and skip connection is employed to enhance the ability of sensing receptive fields for multiscale context information. Furthermore, deep supervision approach with a loss function that combines the dice and cross‐entropy losses to reduce overfitting and ensure better prediction is introduced. The proposed method was evaluated on three public cardiac datasets. The experimental results indicate that the method achieved competitive segmentation performance with the three datasets, which verifies the robustness and generalisability of the proposed network. In comparison with conventional U‐Net methods, the model leverages attention mechanism and dilated convolution block, which increases the semantic connection between the channel and the spatial information, and improves the robustness of the right ventricle segmentation performance. From the view of the Dice scores and segmentation results, the multiscale feature attentive U‐Net method is one of effective methods in segmenting cardiac MRI structures. Dapeng Li 0001, Yanjun Peng, Yanfei Guo, Jindong Sun |
IET Image Process. | 3 |
| 2022 | DSLN: Dual-tutor student learning network for multiracial glaucoma detection
Yanfei Guo, Yanjun Peng, Jindong Sun, Dapeng Li 0001, Bin Zhang 0052 |
Neural Comput. Appl. | 1 |
| 2021 | RMACXX: An Efficient High-Level C++ Interface over MPI-3 RMAabstractParallel scientific applications can benefit from decoupling communication and synchronization. One-sided programming abstractions, which separate communication from synchronization, have in fact served as a motivation for partitioned global address space (PGAS) models. However, the use of PGAS models in application codes in a manner that fully exploits the benefit of these programming models requires significant development effort. Meanwhile, a vast majority of scientific codes already use the Message Passing Interface (MPI) and need convenient features to support application-specific one-sided communication scenarios. MPI Remote Memory Access (RMA) can be employed for this purpose. MPI is a low-level API, however, and developing applications with MPI RMA requires programmers to be well versed in its nuances. We present RMACXX, a compact set of C++ bindings to MPI-3 RMA, to ease the use of MPI RMA. Unlike other PGAS models, which may have interoperability issues with MPI, RMACXX is written on top of MPI and uses the same runtime as MPI. The basic functionality of RMACXX adds only a relatively small number of extra instructions (about 20) to the critical communication path. Moreover, RMACXX provides an intuitive API for building a wide variety of scientific applications while enjoying performance matching handwritten MPI-3 RMA codes. Yanfei Guo, Pavan Balaji, Assefaw Hadish Gebremedhin |
CCGRID | 2 |
| 2021 | In-situ workflow auto-tuning through combining component modelsabstractIn-situ parallel workflows couple multiple component applications via streaming data transfer to avoid data exchange via shared file systems. Such workflows are challenging to configure for optimal performance due to the huge space of possible configurations. Here, we propose an in-situ workflow auto-tuning method, ALIC, which integrates machine learning techniques with knowledge of in-situ workflow structures to enable automated workflow configuration with a limited number of performance measurements. Experiments with real applications show that ALIC identify better configurations than existing methods given a computer time budget. Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç |
PPoPP | 2 |
| 2021 | Bootstrapping in-situ workflow auto-tuning via combining performance models of component applicationsabstractIn an in-situ workflow, multiple components such as simulation and analysis applications are coupled with streaming data transfers. The multiplicity of possible configurations necessitates an auto-tuner for workflow optimization. Existing auto-tuning approaches are computationally expensive because many configurations must be sampled by running the whole workflow repeatedly in order to train the auto-tuner surrogate model or otherwise explore the configuration space. To reduce these costs, we instead combine the performance models of component applications by exploiting the analytical workflow structure, selectively generating test configurations to measure and guide the training of a machine learning workflow surrogate model. Because the training can focus on well-performing configurations, the resulting surrogate model can achieve high prediction accuracy for good configurations despite training with fewer total configurations. Experiments with real applications demonstrate that our approach can identify significantly better configurations than other approaches for a fixed computer time budget. Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç |
SC | 2 |
| 2021 | CAFR-CNN: coarse-to-fine adaptive faster R-CNN for cross-domain joint optic disc and cup segmentation
Yanfei Guo, Yanjun Peng |
Appl. Intell. | 1 |
| 2021 | Segmentation of the multimodal brain tumor image used the multi-pathway architecture method based on 3D FCN
Jindong Sun, Yanjun Peng, Yanfei Guo |
Neurocomputing | 3 |
| 2020 | Memory-Efficient and Skew-Tolerant MapReduce Over MPI for Supercomputing SystemsabstractData analytics has become an integral part of large-scale scientific computing. Among various data analytics frameworks, MapReduce has gained the most traction. Although some efforts have been made to enable efficient MapReduce for supercomputing systems, they are often limited to fairly homogeneous workloads where equal partitioning of input data across tasks results in essentially equal output or temporary data generated on each task. For workloads that are more skewed, however, current implementations can result in imbalance in memory usage and, consequently, can cause a slowdown in execution time and a loss in data scalability. To tackle this problem, we enhance a previously published memory-conscious MapReduce over MPI framework called Mimir. Our enhancements to Mimir include combiner and dynamic repartition optimizations to minimize and balance memory usage and to achieve close to optimal balance of the memory usage across processes and to reduce the execution time by up to 12 times. Experimental results show that Mimir can scale to at least 3072 processes on the Tianhe-2 supercomputer on skewed datasets. Yanfei Guo, Boyu Zhang 0002, Pietro Cicotti, Yutong Lu, Pavan Balaji, Michela Taufer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Optimized Execution of Parallel Loops via User-Defined Scheduling PoliciesabstractOn-node parallelism continues to increase in importance for high-performance computing and most newly deployed supercomputers have tens of processor cores per node. These higher levels of on-node parallelism exacerbate the impact of load imbalance and locality in parallel computations, and current programming systems notably lack features to enable efficient use of these large numbers of cores or require users to modify codes significantly. Our work is motivated by the need to address application-specific load balance and locality requirements with minimal changes to application codes. Seonmyeong Bak, Yanfei Guo, Pavan Balaji, Vivek Sarkar |
ICPP | 2 |
| 2019 | Software combining to mitigate multithreaded MPI contentionabstractEfforts to mitigate lock contention from concurrent threaded accesses to MPI have reduced contention through fine-grained locking, avoided locking altogether by offloading communication to dedicated threads, or alleviated negative side effects from contention by using better lock management protocols. The blocking nature of lock-based methods, however, wastes the asynchrony benefits of nonblocking MPI operations, and the offloading model sacrifices CPU resources and incurs unnecessary software offloading overheads under low contention. Abdelhalim Amer, Charles Archer, Michael Blocksome, Chongxiao Cao, Michael Chuvelev, Hajime Fujita 0002, María Jesús Garzarán, Yanfei Guo, Jeff R. Hammond, Shintaro Iwasaki, Kenneth Raffenetti, Mikhail Shiryaev, Min Si, Kenjiro Taura, Sagar Thapaliya, Pavan Balaji |
ICS | 8 |
| 2018 | On the Power of Combiner Optimizations in MapReduce Over MPI WorkflowsabstractAnalyzing large volumes of data is becoming more and more important in various scientific computing domains. MapReduce over MPI frameworks are an appealing solution to enable scalable big data analytics on supercomputing systems. These systems can further leverage features of MapReduce applications by merging (key/value) pairs before the reduce function in combiner optimizations. In this paper, we propose a pipeline combiner workflow and integrate it into Mimir, a cutting-edge implementation of Map Reduce over MPI. Our results with real datasets on the Tianhe-2 supercomputer prove that our pipeline combiner workflow can reduce memory usage up to 51% and improve the overall performance up to 61%. Yanfei Guo, Boyu Zhang 0002, Pietro Cicotti, Yutong Lu, Pavan Balaji, Michela Taufer |
ICPADS | 2 |
| 2017 | Bloomfish: A Highly Scalable Distributed K-mer Counting FrameworkabstractK-mer counting is a fundamental operation in DNA research and genome analytics; its application includes estimating genome assembly, understanding similarities in genomic samples, and merging a newly processed genome with a reference genome. As the genome dataset becomes larger and larger, designing a highly optimized distributed-memory implementation becomes more and more important. Current distributed-memory solutions have two limitations: they have a high memory footprint, and they do not provide advanced optimizations for loading enormous genome datasets into memory. Based on these observations, we present Bloomfish, a distributed, memory-efficient, scalable solution to the limits of current work. To keep a low memory footprint, Bloomfish leverages the compact hash array design of the single-node Jellyfish system and the optimized workflow of the high-performance MapReduce framework Mimir. We have also codesigned Mimir's I/O to efficiently load enormous datasets. We ran Bloomfish on the Tianhe-2 supercomputer with large sequence datasets (up to 24 TB). Our results show that Bloomfish achieves unprecedented scalability in genome analytics. Yanfei Guo, Yanjie Wei, Bingqiang Wang, Yutong Lu, Pietro Cicotti, Pavan Balaji, Michela Taufer |
ICPADS | 2 |
| 2017 | Mimir: Memory-Efficient and Scalable MapReduce for Large Supercomputing SystemsabstractIn this paper we present Mimir, a new implementation of MapReduce over MPI. Mimir inherits the core principles of existing MapReduce frameworks, such as MR-MPI, while redesigning the execution model to incorporate a number of sophisticated optimization techniques that achieve similar or better performance with significant reduction in the amount of memory used. Consequently, Mimir allows significantly larger problems to be executed in memory, achieving large performance gains. We evaluate Mimir with three benchmarks on two highend platforms to demonstrate its superiority compared with that of other frameworks. Yanfei Guo, Boyu Zhang 0002, Pietro Cicotti, Yutong Lu, Pavan Balaji, Michela Taufer |
IPDPS | 2 |
| 2017 | Memory Compression Techniques for Network Address Management in MPIabstractMPI allows applications to treat processes as a logical collection of integer ranks for each MPI communicator, while internally translating these logical ranks into actual network addresses. In current MPI implementations the management and lookup of such network addresses use memory sizes that are proportional to the number of processes in each communicator. In this paper, we propose a new mechanism, called AV-Rankmap, for managing such translation. AV-Rankmap takes advantage of logical patterns in rank-address mapping that most applications naturally tend to have, and it exploits the fact that some parts of network address structures are naturally more performance critical than others. It uses this information to compress the memory used for network address management. We demonstrate that AV-Rankmap can achieve performance similar to or better than that of other MPI implementations while using significantly less memory. Yanfei Guo, Charles Archer, Michael Blocksome, Scott Parker, Wesley Bland, Kenneth Raffenetti, Pavan Balaji |
IPDPS | 1 |
| 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1abstractThis paper provides an in-depth analysis of the software overheads in the MPI performance-critical path and exposes mandatory performance overheads that are unavoidable based on the MPI-3.1 specification. We first present a highly optimized implementation of the MPI-3.1 standard in which the communication stack---all the way from the application to the low-level network communication API---takes only a few tens of instructions. We carefully study these instructions and analyze the root cause of the overheads based on specific requirements from the MPI standard that are unavoidable under the current MPI standard. We recommend potential changes to the MPI standard that can minimize these overheads. Our experimental results on a variety of network architectures and applications demonstrate significant benefits from our proposed changes. Kenneth Raffenetti, Abdelhalim Amer, Lena Oden, Charles Archer, Wesley Bland, Hajime Fujita 0002, Yanfei Guo, Tomislav Janjusic, Dmitry Durnov, Michael Blocksome, Min Si, Akhil Langer, Gengbin Zheng, Masamichi Takagi, Paul K. Coffman, Sayantan Sur, Alexander Sannikov, Sergey Oblomov, Michael Chuvelev, Masayuki Hatanaka, Paul F. Fischer, Thilina Ratnayaka, Matthew Otten, Misun Min, Pavan Balaji |
SC | 7 |
| 2017 | Improving Performance of Heterogeneous MapReduce Clusters with Adaptive Task TuningabstractDatacenter-scale clusters are evolving toward heterogeneous hardware architectures due to continuous server replacement. Meanwhile, datacenters are commonly shared by many users for quite different uses. It often exhibits significant performance heterogeneity due to multi-tenant interferences. The deployment of MapReduce on such heterogeneous clusters presents significant challenges in achieving good application performance compared to in-house dedicated clusters. As most MapReduce implementations are originally designed for homogeneous environments, heterogeneity can cause significant performance deterioration in job execution despite existing optimizations on task scheduling and load balancing. In this paper, we observe that the homogeneous configuration of tasks on heterogeneous nodes can be an important source of load imbalance and thus cause poor performance. Tasks should be customized with different configurations to match the capabilities of heterogeneous nodes. To this end, we propose a self-adaptive task tuning approach, Ant, that automatically searches the optimal configurations for individual tasks running on different nodes. In a heterogeneous cluster, Ant first divides nodes into a number of homogeneous subclusters based on their hardware configurations. It then treats each subcluster as a homogeneous cluster and independently applies the self-tuning algorithm to them. Ant finally configures tasks with randomly selected configurations and gradually improves tasks configurations by reproducing the configurations from best performing tasks and discarding poor performing configurations. To accelerate task tuning and avoid trapping in local optimum, Ant uses genetic algorithm during adaptive task configuration. Experimental results on a heterogeneous physical cluster with varying hardware capabilities show that Ant improves the average job completion time by 31, 20, and 14 percent compared to stock Hadoop (Stock), customized Hadoop with industry recommendations (Heuristic), and a profilingbased configuration approach (Starfish), respectively. Furthermore, we extend Ant to virtual MapReduce clusters in a multi-tenant private cloud. Specifically, Ant characterizes a virtual node based on two measured performance statistics: I/O rate and CPU steal time. It uses k-means clustering algorithm to classify virtual nodes into configuration groups based on the measured dynamic interference. Experimental results on virtual clusters with varying interferences show that Ant improves the average job completion time by 20, 15, and 11 percent compared to Stock, Heuristic and Starfish, respectively. Dazhao Cheng, Jia Rao, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | iShuffle: Improving Hadoop Performance with Shuffle-on-WriteabstractHadoop is a popular implementation of the MapReduce framework for running data-intensive jobs on clusters of commodity servers.Shuffle, the all-to-all input data fetching phase between the map and reduce phase can significantly affect job performance. However, the shuffle phase and reduce phase are coupled together in Hadoop and the shuffle can only be performed by running the reduce tasks. This leaves the potential parallelism between multiple waves of map and reduce unexploited and resource wastage in multi-tenant Hadoop clusters, which significantly delays the completion of jobs in a multi-tenant Hadoop cluster. More importantly, Hadoop lacks the ability to schedule task efficiently and mitigate the data distribution skew among reduce tasks, which leads to further degradation of job performance. In this work, we propose to decouple shuffle from reduce tasks and convert it into a platform service provided by Hadoop. We presentiShuffle, a user-transparent shuffle service that pro-actively pushes map output data to nodes via a novelshuffle-on-writeoperation and flexibly schedules reduce tasks considering workload balance. Experimental results with representative workloads and Facebook workload trace show that iShuffle reduces job completion time by as much as 29.6 and 34 percent in single-user and multi-user clusters, respectively. Yanfei Guo, Jia Rao, Dazhao Cheng, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Moving Hadoop into the Cloud with Flexible Slot Management and Speculative ExecutionabstractLoad imbalance is a major source of overhead in parallel programs such as MapReduce. Due to the uneven distribution of input data, tasks with more data become stragglers and delay the overall job completion. Running Hadoop in a private cloud opens up opportunities for expediting stragglers with more resources but also introduces problems that often outweigh the performance gain: (1) performance interference from co-running jobs may create new stragglers; (2) there exists a semantic gap between the Hadoop task management and resource pool-based virtual cluster management preventing tasks from using resources efficiently. In this paper, we strive to make Hadoop more resilient to data skew and more efficient in cloud environments. We presentFlexSlot, a user-transparent task slot management scheme that automatically identifies map stragglers and resizes their slots accordingly to accelerate task execution. FlexSlot adaptively changes the number of slots on each virtual node to balance the resource usage so that the pool of resources can be efficiently utilized. FlexSlot further improves mitigation of data skew with an adaptive speculative execution strategy. Experimental results show that FlexSlot effectively reduces job completion time up to$47.2$percent compared to stock Hadoop and two recently proposed skew mitigation and speculative execution approaches. Yanfei Guo, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Authenticated key exchange with entities from different settings and varied groupsabstractAbstract Authenticated key exchange (AKE) is a very important primitive in cryptography. In the last decades, many AKE protocols appeared either in the certificate‐based (cert‐based) setting or in the identity‐based (id‐based) setting. In real applications, entities from different settings may also have the requirement to communicate with each other. Several papers have concentrated on supporting either multiple certification authorities or multiple key generation centers, but very few have considered the interoperability between the two settings. Furthermore, existing approaches are still inadequate in supporting parameters from different algebraic groups. In this paper, we consider AKE protocols integrating cert‐based and id‐based settings with varied groups. Based on two extract algorithms for id‐based entities, we present two AKE protocols where one entity is cert‐based and the other is id‐based, and the parameters of both entities come from different groups. An extended AKE security model of Chatterjee et al. and Ustaoǧlu [1, 2] is proposed to support multiple certification authorities and multiple key generation centers in which the proposed protocols are proved to be secure. Other variant protocols are also presented. Then, extensions to support forward secrecy and resistance to leakage of both ephemeral keys are provided. Finally, we present a more efficient integrating protocol than existing constructions for users who use the same group. Copyright © 2013 John Wiley & Sons, Ltd. Yanfei Guo, Zhenfeng Zhang |
Secur. Commun. Networks | 1 |
| 2016 | Autonomic Performance and Power Control for Co-Located Web Applications in Virtualized DatacentersabstractIn a datacenter, complex and time-varying interactions between various tiers and services of web applications, and the contention of shared resources among co-located virtual machines have significant impact on the user perceived performance and power consumption of the underlying system. We propose and develop APPLEware, an autonomic middleware for joint performance and power control of co-located web applications in virtualized datacenters. It features a distributed control structure that provides predictable performance and energy efficiency for large complex systems. It applies machine learning based self-adaptive modeling to capture the complex and time-varying relationship between the application performance and allocation of resources to various application components, in the face of highly dynamic and bursty workloads. The distributed controllers coordinate with each other and allocate resources to meet the service level agreements of applications in an agile and energy-efficient manner. Experimental results based on a testbed implementation with benchmark applications and large scale simulations demonstrate APPLEware's effectiveness, energy efficiency and scalability. Palden Lama, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | StoreApp: A shared storage appliance for efficient and scalable virtualized Hadoop clustersabstractVirtualizing Hadoop clusters provides many benefits, including rapid deployment, on-demand elasticity and secure multi-tenancy. However, a simple migration of Hadoop to a virtualized environment does not fully exploit these benefits. The dual role of a Hadoop worker, acting as both a compute node and a data node, makes it difficult to achieve efficient IO processing, maintain data locality, and exploit resource elasticity in the cloud. We find that decoupling per-node storage from its computation opens up opportunities for IO acceleration, locality improvement, and on-the-fly cluster resizing. To fully exploit these opportunities, we propose StoreApp, a shared storage appliance for virtual Hadoop worker nodes co-located on the same physical host. To completely separate storage from computation and prioritize IO processing, StoreApp pro-actively pushes intermediate data generated by map tasks to the storage node. StoreApp also implements late-binding task creation to take the advantage of prefetched data due to mis-aligned records. Experimental results show that StoreApp achieves up to 61% performance improvement compared to stock Hadoop and resizes the cluster to the (near) optimal degree of parallelism. Yanfei Guo, Jia Rao, Dazhao Cheng, Changjun Jiang 0002, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002 |
INFOCOM | 1 |
| 2015 | Fault tolerant MapReduce-MPI for HPC clustersabstractBuilding MapReduce applications using the Message-Passing Interface (MPI) enables us to exploit the performance of large HPC clusters for big data analytics. However, due to the lacking of native fault tolerance support in MPI and the incompatibility between the MapReduce fault tolerance model and HPC schedulers, it is very hard to provide a fault tolerant MapReduce runtime for HPC clusters. We propose and develop FT-MRMPI, the first fault tolerant MapReduce framework on MPI for HPC clusters. We discover a unique way to perform failure detection and recovery by exploiting the current MPI semantics and the new proposal of user-level failure mitigation. We design and develop the checkpoint/restart model for fault tolerant MapReduce in MPI. We further tailor the detect/resume model to conserve work for more efficient fault tolerance. The experimental results on a 256-node HPC cluster show that FT-MRMPI effectively masks failures and reduces the job completion time by 39%. Yanfei Guo, Wesley Bland, Pavan Balaji, Xiaobo Zhou 0002 |
SC | 1 |
| 2015 | Self-Tuning Batching with DVFS for Performance Improvement and Energy Efficiency in Internet ServersabstractPerformance improvement and energy efficiency are two important goals in provisioning Internet services in datacenter servers. In this article, we propose and develop a self-tuning request batching mechanism to simultaneously achieve the two correlated goals. The batching mechanism increases the cache hit rate at the front-tier Web server, which provides the opportunity to improve an application’s performance and the energy efficiency of the server system. The core of the batching mechanism is a novel and practical two-layer control system that adaptively adjusts the batching interval and frequency states of CPUs according to the service level agreement and the workload characteristics. The batching control adopts a self-tuning fuzzy model predictive control approach for application performance improvement. The power control dynamically adjusts the frequency of Central Processing Units (CPUs) with Dynamic Voltage and Frequency Scaling (DVFS) in response to workload fluctuations for energy efficiency. A coordinator between the two control loops achieves the desired performance and energy efficiency. We further extend the self-tuning batching with DVFS approach from a single-server system to a multiserver system. It relies on a MIMO expert fuzzy control to adjust the CPU frequencies of multiple servers and coordinate the frequency states of CPUs at different tiers. We implement the mechanism in a test bed. Experimental results demonstrate that the new approach significantly improves the application performance in terms of the system throughput and average response time. At the same time, the results also illustrate the mechanism can reduce the energy consumption of a single-server system by 13% and a multiserver system by 11%, respectively. Dazhao Cheng, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2014 | Black-Box Separations for One-More (Static) CDH and Its Generalization
Jiang Zhang 0001, Zhenfeng Zhang, Yu Chen 0003, Yanfei Guo, Zongyang Zhang |
ASIACRYPT (2) | 4 |
| 2014 | Security Analysis of EMV Channel Establishment Protocol in An Enhanced Security Model
Yanfei Guo, Zhenfeng Zhang, Jiang Zhang 0001, Xuexian Hu |
ICICS | 1 |
| 2014 | Improving MapReduce performance in heterogeneous environments with adaptive task tuningabstractThe deployment of MapReduce in datacenters and clouds present several challenges in achieving good job performance. Compared to in-house dedicated clusters, datacenters and clouds often exhibit significant hardware and performance heterogeneity due to continuous server replacement and multi-tenant interferences. As most Mapreduce implementations assume homogeneous clusters, heterogeneity can cause significant load imbalance in task execution, leading to poor performance and low cluster utilizations. Despite existing optimizations on task scheduling and load balancing, MapReduce still performs poorly on heterogeneous clusters. Dazhao Cheng, Jia Rao, Yanfei Guo, Xiaobo Zhou 0002 |
Middleware | 3 |
| 2014 | FlexSlot: Moving Hadoop Into the Cloud with Flexible Slot ManagementabstractLoad imbalance is a major source of overhead in Hadoop where the uneven distribution of input data among tasks can significantly delays the job completion. Running Hadoop in a private cloud opens up opportunities for mitigating data skew with elastic resource allocation, where stragglers are expedited with more resources, yet introduces problems that often cancel out the performance gain: (1) performance interference from co running jobs may create new stragglers, (2) there exist a semantic gap between Hadoop task management and resource pool-based virtual cluster management preventing efficient resource usage. We present Flex Slot, a user-transparent task slot management scheme that automatically identifies map stragglers and resizes their slots accordingly to accelerate task execution. Flex Slot adaptively changes the number of slots on each virtual node to promote efficient usage of resource pool. Experimental results with representative benchmarks show that Flex Slot effectively reduces job completion time by 46% and achieves better resource utilization. Yanfei Guo, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
SC | 1 |
| 2014 | Automated and Agile Server ParameterTuning by Coordinated Learning and ControlabstractAutomated server parameter tuning is crucial to performance and availability of Internet applications hosted in cloud environments. It is challenging due to high dynamics and burstiness of workloads, multi-tier service architecture, and virtualized server infrastructure. In this paper, we investigate automated and agile server parameter tuning for maximizing effective throughput of multi-tier Internet applications. A recent study proposed a reinforcement learning based server parameter tuning approach for minimizing average response time of multi-tier applications. Reinforcement learning is a decision making process determining the parameter tuning direction based on trial-and-error, instead of quantitative values for agile parameter tuning. It relies on a predefined adjustment value for each tuning action. However it is nontrivial or even infeasible to find an optimal value under highly dynamic and bursty workloads. We design a neural fuzzy control based approach that combines the strengths of fast online learning and self-adaptiveness of neural networks and fuzzy control. Due to the model independence, it is robust to highly dynamic and bursty workloads. It is agile in server parameter tuning due to its quantitative control outputs. We implemented the new approach on a testbed of virtualized data center hosting RUBiS and WikiBench benchmark applications. Experimental results demonstrate that the new approach significantly outperforms the reinforcement learning based approach for both improving effective system throughput and minimizing average response time. Yanfei Guo, Palden Lama, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | V-Cache: Towards Flexible Resource Provisioning for Multi-tier Applications in IaaS CloudsabstractAlthough the resource elasticity offered by Infrastructure-as-a-Service (IaaS) clouds opens up opportunities for elastic application performance, it also poses challenges to application management. Cluster applications, such as multi-tier websites, further complicates the management requiring not only accurate capacity planning but also proper partitioning of the resources into a number of virtual machines. Instead of burdening cloud users with complex management, we move the task of determining the optimal resource configuration for cluster applications to cloud providers. We find that a structural reorganization of multi-tier websites, by adding a caching tier which runs on resources debited from the original resource budget, significantly boosts application performance and reduces resource usage. We propose V-Cache, a machine learning based approach to flexible provisioning of resources for multi-tier applications in clouds. V-Cache transparently places a caching proxy in front of the application. It uses a genetic algorithm to identify the incoming requests that benefit most from caching and dynamically resizes the cache space to accommodate these requests. We develop a reinforcement learning algorithm to optimally allocate the remaining capacity to other tiers. We have implemented V-Cache on a VMware-based cloud testbed. Experiment results with the RUBiS and WikiBench benchmarks show that V-Cache outperforms a representative capacity management scheme and a cloud-cache based resource provisioning approach by at least 15% in performance, and achieves at least 11% and 21% savings on CPU and memory resources, respectively. Yanfei Guo, Palden Lama, Jia Rao, Xiaobo Zhou 0002 |
IPDPS | 1 |
| 2013 | Autonomic performance and power control for co-located Web applications on virtualized serversabstractIn a data center, various components of Web applications co-located on virtualized servers exhibit complex time-varying interactions and interference. It has a significant impact on the user perceived performance and power consumption of the underlying system. We propose and develop APPLEware, an autonomic middleware for joint performance and power control of co-located Web applications. It features a distributed control structure that provides performance assurance and energy efficiency for large complex systems. It applies machine learning based self-adaptive modeling to capture the complex and time-varying relationship between the application performance and allocation of resources to various application components, in the presence of highly dynamic and bursty workloads and inter-application performance interference. The distributed controllers perform coordinated resource allocation to meet the service level agreements of applications in an agile and energy-efficient manner. Experimental results based on a testbed implementation with benchmark applications demonstrate APPLEware's effectiveness and energy efficiency. Palden Lama, Yanfei Guo, Xiaobo Zhou 0002 |
IWQoS | 2 |
| 2013 | Self-Tuning Batching with DVFS for Improving Performance and Energy Efficiency in ServersabstractPerformance improvement and energy efficiency are two important goals in provisioning Internet services in data center servers. In this paper, we propose and develop a self-tuning request batching mechanism to simultaneously achieve the two correlated goals. The batching mechanism increases the cache hit rate at the front-tier Web server, which provides the opportunity to improve application's performance and energy efficiency of the server system. The core of the batching mechanism is a novel and practical two-layer control system that adaptively adjusts the batching interval and frequency states of CPUs according to the service level agreement and the workload characteristics. The batching control adopts a self-tuning fuzzy model predictive control approach for application performance improvement. The power control dynamically adjusts the frequency of CPUs with DVFS in response to workload fluctuations for energy efficiency. A coordinator between the two control loops achieves the desired performance and energy efficiency. We implement the mechanism in a test bed and experimental results demonstrate that the new approach significantly improves the application's performance in terms of the system throughput and average response time. The results also illustrate it can reduce the energy consumption of the server system by 13% at the same time. Dazhao Cheng, Yanfei Guo, Xiaobo Zhou 0002 |
MASCOTS | 2 |
| 2012 | Automated and Agile Server Parameter Tuning with Learning and ControlabstractServer parameter tuning in virtualized data centers is crucial to performance and availability of hosted Internet applications. It is challenging due to high dynamics and burstiness of workloads, multi-tier service architecture, and virtualized server infrastructure. In this paper, we investigate automated and agile server parameter tuning for maximizing effective throughput of multi-tier Internet applications. A recent study proposed a reinforcement learning based server parameter tuning approach for minimizing average response time of multi-tier applications. Reinforcement learning is a decision making process determining the parameter tuning direction based on trial-and-error, instead of quantitative values for agile parameter tuning. It relies on a predefined adjustment value for each tuning action. However it is nontrivial or even infeasible to find an optimal value under highly dynamic and bursty workloads. We design a neural fuzzy control based approach that combines the strengths of fast online learning and self-adaptive ness of neural networks and fuzzy control. Due to the model independence, it is robust to highly dynamic and bursty workloads. It is agile in server parameter tuning due to its quantitative control outputs. We implement the new approach on a test bed of virtualized HP Pro Liant blade servers hosting RUBiS benchmark applications. Experimental results demonstrate that the new approach significantly outperforms the reinforcement learning based approach for both improving effective system throughput and minimizing average response time. Yanfei Guo, Palden Lama, Xiaobo Zhou 0002 |
IPDPS | 1 |
| 2012 | Coordinated VM Resizing and Server Tuning: Throughput, Power Efficiency and ScalabilityabstractPerformance control and power management in virtualized machines (VM) are two major research issues in modern data centers. They are challenging due to complexities of hosted Internet applications, high dynamics in workloads and the shared virtualized infrastructure. Obtaining a model among VM capacity, server configuration, performance and power consumption is a very hard problem even for just one application. In this paper, we propose and develop GARL, a genetic algorithm with multi-agent reinforcement learning approach for coordinated VM resizing and server tuning. In GARL, model-independent reinforcement learning agents generate VM capacity and server configuration options and the genetic algorithm evaluates different combinations of those options for maximizing a global utilization function of system throughput and power efficiency. The multi-agent design makes GARL a scalable approach, which is important as more and more applications are hosted in data centers using cloud services. We build a testbed in a prototype data center and deploy multiple RUBiS benchmark applications. We apply a power budget in the testbed and observe superior system throughput and power efficiency of GARL. Experimental results also find that GARL significantly outperforms a representative reinforcement learning based approach in performance control. GARL shows better scalability when compared to a centralized approach. Yanfei Guo, Xiaobo Zhou 0002 |
MASCOTS | 1 |
| 2012 | Authenticated Key Exchange with Entities from Different Settings and Varied Groups
Yanfei Guo, Zhenfeng Zhang |
ProvSec | 1 |