EDBT 2026 Demo / reviewers in the wild / expert
Meng Hao 0002
dblp:184/7209-2
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-0043-4370ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 17 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Level Acceleration Scheme for AI Model Training on ARM Architecture ProcessorsabstractABSTRACT With the widespread application of reinforcement learning and deep learning on edge devices, training neural networks on ARM architecture processors has become an urgent demand. However, existing mainstream deep learning frameworks are not sufficiently optimized for training workloads on ARM CPUs, resulting in low training efficiency. To address this problem, this paper proposes and implements a multi‐level AI training acceleration scheme based on the open‐source C++ library mlpack, targeting the slow training speed and high resource consumption of Convolutional Neural Networks (CNNs) on ARM platforms. The scheme accelerates training by deeply optimizing the im2col algorithm to convert convolutions into efficient matrix multiplications, utilizing ARM NEON SIMD instructions to optimize linear operators, and integrating an FP64/FP16 mixed‐precision training strategy with dynamic loss scaling. Experimental results on LeNet‐5 and VGG11‐style CNNs show substantial performance gains over the original mlpack and mainstream frameworks. On an NVIDIA Jetson AGX Orin, our implementation achieves up to 7.3 speedup over the original mlpack baseline and up to 11.3 × and 5.69 × end‐to‐end speedups over PyTorch and TensorFlow, respectively, while still delivering multi‐fold reductions in training time on a low‐resource Raspberry Pi platform. In DQN‐based reinforcement learning for Atari Breakout, our solution attains a 4.98 × end‐to‐end speedup over the PyTorch single‐threaded baseline and maintains 2.37 × and 4.23 × advantages over 4‐threaded PyTorch and TensorFlow implementations. Ablation studies confirm the complementary nature of the proposed optimizations, with convolutional, linear, and mixed‐precision components jointly contributing to the overall speedup and enabling an attractive performance–accuracy trade‐off for ARM‐based edge computing. Mingdong Xie, Meng Hao 0002, Weizhe Zhang, NingCheng Wang |
Concurr. Comput. Pract. Exp. | 3 |
| 2026 | Performance Prediction of Concurrent DNN Training Tasks in GPU Spatial Sharing EnvironmentsabstractGPU sharing is commonly employed in GPU clusters to improve utilization, with spatial sharing being one of the most widely adopted techniques. However, spatial sharing can lead to resource interference, making task execution times difficult to predict. Predictable execution times for each task are crucial in GPU cluster management and task scheduling. In this article, we propose a performance predictor for multi-DNN training tasks in GPU spatial sharing environments. We first conduct experiments on spatial sharing for multiple DNN workloads on a single GPU, demonstrating that concurrent execution of multiple tasks improves overall performance and GPU resource utilization compared to serial execution. By analyzing warp stall reasons collected during task execution, we investigate the interference for computation and memory resources under MPS on GPUs. Finally, we design a performance predictor that predicts the execution time of a target DNN training task when it runs concurrently with other tasks under GPU spatial sharing via MPS. The predictor is capable of predicting the execution time of each task for previously unseen combinations of DNN training tasks. Extensive evaluations on modern GPUs show that compared to other baseline methods, our approach exhibits higher prediction accuracy, as well as improved stability and robustness. Experiments on multiple GPU architectures, as well as at higher concurrency levels, further demonstrate that our method possesses strong generalization and scalability. We also conducted a performance analysis under diverse workload pattern and a case study to validate the practical applicability of our predictor in real scheduling environments. Sichao Chen, Desheng Wang 0002, Weizhe Zhang, Meng Hao 0002, Yu-Chu Tian |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power ConstraintsabstractThe proliferation of deep learning inference services in power-constrained environments necessitates GPU management strategies that maximize throughput within strict power envelopes. Existing approaches often treat frequency scaling and resource partitioning as orthogonal problems or rely on static hardware assumptions, leading to suboptimal energy efficiency. This article presents PctoDL , a power-aware scheduling system that maximizes aggregate inference throughput by jointly optimizing spatial resource partitioning, batch size, and SM/memory frequency settings. To address the throughput–power tradeoff in power-constrained multi-tenant inference, PctoDL couples resource partitioning with coordinated frequency control under a fixed power cap. It combines a physics-informed iterative greedy partitioning algorithm, a thermodynamic model-predictive controller for runtime frequency regulation, and an online joint optimization mechanism for adaptive refinement. On the NVIDIA RTX 3080 Ti platform, PctoDL improves average throughput over BatchDVFS by 108.41%, with a peak gain of 262.74%. On the NVIDIA A100 platform, it delivers an average gain of 19.74% and a maximum gain of 57.03%. Compared with Morak’s coarse-grained partitioning approach, PctoDL achieves average/peak gains of 79.05%/137.93% on the RTX 3080 Ti and 26.33%/70.21% on the A100. Meng Hao 0002, Zikun Wu, Xueyang Tian, Siyu Yang 0002, Guotong Guo, Yiming Wang 0010, Farui Wang, Desheng Wang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 1 |
| 2026 | GreenDLS: An Energy-Efficient and SLO-Aware Deep Learning Serving SystemabstractThe growing demand for deploying deep learning (DL) models, particularly large language models (LLMs), has made it imperative to optimize GPU energy consumption while meeting service-level objectives (SLOs). Significant energy use and carbon dioxide (CO2) emissions from GPU-based inference tasks contribute substantially to the environmental footprint of the DL deployment. Existing approaches primarily rely on batching and dynamic voltage and frequency scaling (DVFS) to optimize service performance or throughput, but often overlook memory frequency adjustments and holistic energy optimization under dynamic workloads. This study presents GREENDLS, a DL serving system that integrates deep reinforcement learning (DRL) with offline prediction models to optimize energy consumption while adhering to inference latency SLOs, achieving significant energy savings. GREENDLS dynamically adjusts batch size, GPU streaming multiprocessor (SM) frequency, and GPU memory frequency based on inference request rates. It also accounts for GPU energy consumption during idle phases, such as batch filling, enabling multi-parameter and fine-grained energy optimization. Compared to the Clipper system, GREENDLS achieves energy savings of up to 45.57% on the RTX 3080Ti and 39.44% on the Tesla V100S. When compared to the EAIS system, which only combines batching with GPU SM frequency adjustment, GREENDLS achieves energy savings of up to 36.44% on the RTX 3080Ti and 15.44% on the Tesla V100S. Against the method proposed by Yu et al., GREENDLS attains energy savings of up to 46.34% on the RTX 3080Ti and 33.74% on the Tesla V100S. In LLM inference tasks using Qwen, GREENDLS reduces average energy consumption by 40.79% compared to Clipper, 10.42% compared to EAIS, and 42.42% compared to Yu et al. These results clearly demonstrate that GREENDLS more effectively optimizes energy consumption compared to traditional methods that rely primarily on batching or a combination of batching and DVFS, while still ensuring SLO compliance. Meng Hao 0002, Xueyang Tian, Siyu Yang 0002, Yiming Wang 0010, Desheng Wang 0002, Weizhe Zhang |
IEEE Trans. Computers | 1 |
| 2026 | Accelerating Secure Machine Learning Training on GPUs With Pipeline ParallelismabstractThe proliferation of data-driven machine learning (ML) applications makes privacy and data security risks increasingly prominent. Secure multi-party computation (MPC) offers a privacy-preserving ML method by enabling joint model training without data disclosure, but its reliance on complex cryptography adds significant delay and resource demands, challenging its practicality in large-scale applications. To address the low GPU utilization within the MPC framework, we propose an innovative privacy-preserving ML training optimization framework that exploits pipeline parallelism. Through a detailed analysis of the underlying principles of MPC-based training, we identify computation and communication as the primary bottlenecks for linear and non-linear computations, respectively. Drawing from traditional ML optimization strategies, we design a sub-network partitioning and pipeline parallelism method, specifically tailored for MPC training. This method not only allows for simultaneous training computations across different network layers, but also strategically overlaps computation with communication to enhance GPU utilization and reduce latency. Additionally, we develop a distributed communication mechanism to further improve communication efficiency. We integrate our framework into two distinct, state-of-the-art secure training frameworks: CryptGPU and Piranha. Compared to their original versions, our enhancements boost training speeds by up to 51%, significantly increasing GPU utilization, with negligible impact on model convergence speed and accuracy. Meng Hao 0002, Mingdong Xie, Weizhe Zhang, Linxuan Wang, Xueyang Tian, Desheng Wang 0002 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | DynGPU: A Dynamic GPU Sharing Framework for Enhanced Resource Utilization and Task Scheduling in Concurrent DNN TrainingabstractTraining deep neural networks (DNNs) is a common task in GPU clusters. However, in practical cluster environments, multiple concurrent DNN training tasks often fail to fully leverage GPU resources, resulting in suboptimal GPU utilization. Furthermore, existing GPU sharing frameworks primarily rely on static scheduling and frequently overlook task deadlines, leading to task delays and inefficient scheduling. To address these issues, we propose a dynamic GPU sharing framework (DynGPU) that intercepts GPU kernel executions to perform resource scheduling in multi-task environments. DynGPU incorporates a dynamic task priority adjustment mechanism that adapts task priorities in real time based on task progress, historical data, and remaining time to deadlines. By guaranteeing resources for high-priority tasks while maximizing resource allocation for low-priority tasks, DynGPU reduces resource contention and improves system throughput, enabling more timely task completions. Experiments show that, compared to dedicated GPU execution, DynGPU can reserve up to 97.5 % of throughput for high-priority tasks. Compared to state-of-the-art baselines, DynGPU achieves up to an 8.4 % improvement in task completion time. Zhiji Yu, Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Meng Hao 0002, Yu-Chu Tian |
ICPADS | 5 |
| 2025 | SEPPDL: A Secure and Efficient Privacy-Preserving Deep Learning Inference Framework for Autonomous DrivingabstractThe autonomous driving system necessitates using privacy-preserving deep learning (PPDL) technologies as the safety assurance for its extensive application. However, existing PPDL solutions depend on intricate protocol designs for robust security. Although leveraging advanced dedicated hardware platforms can significantly improve inference efficiency, the PPDL frameworks that make the best use of hardware platform computility are scarce. Thus, balancing efficiency and security in PPDL remains an open question. This study presents SEPPDL, a secure tripartite inference framework for deep learning based on secret-sharing to balance privacy security and computational efficiency. We reduce the communication and calculation time by designing a deep learning quantization representation scheme, two new computational protocols, and a computation library that utilizes the integer computation units of the GPU. The experimental results show that compared with state-of-the-art PPDL frameworks, the SEPPDL framework reduces the communication and computation delay in the model inference to 1/2 and 1/3 of the existing optimal frameworks while maintaining the accuracy of the model inference. Meanwhile, the SEPPDL framework achieves a 10-fold performance improvement in a lightweight model. As the model scale increases, the performance of the SEPPDL-based model even achieves an 86-fold improvement compared to VGG16. Wang Bobo, Meng Hao 0002, Weizhe Zhang |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2025 | Deep Learning Workload Mapping Optimization on Jetson PlatformsabstractTo improve the performance and energy efficiency of deep learning (DL) applications, recent edge computing platforms have built-in heterogeneous accelerators, such as general-purpose graphics processing units (GPUs) and neural processing units (NPUs). For example, widely used NVIDIA Jetson platforms contain CPU, GPU, and deep learning accelerator (DLA), a type of NPU. It is non-trivial to map DL workloads to suitable accelerators to improve performance, energy efficiency, or even both. This article presents JDIMO, 1 a Jetson-aware deep-learning inference workload mapping optimization framework, to simultaneously improve energy efficiency and performance. JDIMO first measures energy-performance data of the fundamental nodes and the sub-networks with energy-efficiency improvement potential according to the topology structure of a DL network. Then, under the guidance of an analytical energy-performance model, the framework exploits an algorithm based on the variable-length sliding window to find the optimal mapping configuration and the optimal number of CUDA streams. We evaluate JDIMO by applying it to seven DL applications on a Jetson Orin NX (16GB) platform. JDIMO saves 47.5% EDP (energy delay product) and 22.6% energy and improves 138.3% QPS (queries per second) on average compared to the DLA-possible configuration. JDIMO saves 22.5% EDP and 12.6% energy and improves 13.5% QPS on average compared to JEDI, the most similar work to ours. Meanwhile, JDIMO also reduces 93.8% optimization time on average compared to JEDI. Farui Wang, Meng Hao 0002, Siyu Yang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Dynamic Power Management Through Multi-agent Deep Reinforcement Learning for Heterogeneous SystemsabstractPower management and optimization play a significant role in modern computer systems, from battery-powered devices to servers running in data centers. Existing approaches for power capping fail to meet the requirements presented by dynamic workloads, and the situation becomes even more severe, given the divergent energy efficiency of workloads on heterogeneous hardware platforms. Adaptively optimizing energy consumption for dynamic workloads presents a great challenge to heterogeneous systems. To tackle this challenge, we present a machine learning based method to improve system-level power efficiency. We employ multi-agent deep reinforcement learning (MADRL) to automatically explore the relationship between long-term performance and the power budget for workloads of different types on classic CPU-GPU heterogeneous platforms. Our framework equips each device with an agent, enabling decentralized control over its power budget while maintaining centralized coordination to maximize the running time of applications within a power cap. We evaluate our approach against state-of-the-art methods on CPU-GPU platforms. Experimental results show that our method improves performance by an average of 8.5%. Additionally, our method is significantly more stable compared to the state-of-the-art heuristic approach. Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Weizhi Kong, Yuan Wen |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | HEngine: A High Performance Optimization Framework on a GPU for Homomorphic EncryptionabstractHomomorphic encryption (HE) represents an encryption technology that allows for direct computation on encrypted data without requiring decryption. However, the substantial computational complexity and significant latency associated with HE has impeded its broader adoption in practical applications. To address these challenges, we propose a GPU-based acceleration framework, namely HEngine, tailored for homomorphic encryption tasks. Specifically, we first propose a warp shuffle-based optimization method for two key phases, i.e., inverse Chinese Remainder Theorem (ICRT) and number theoretic transformation (NTT), to mitigate synchronization overhead in homomorphic encryption. Secondly, we propose to fuse the NTT kernel with the inner product kernel to address the imbalance between memory access and computation. Thirdly, considering the potential difference in the amount of tasks of users in the real-world, we design two different encoding methods for small batch and large batch inference tasks to improve computational efficiency. Finally, experiments demonstrate that our proposed framework achieves a 218× speedup on homomorphic multiplication tasks compared with the CPU-based SEAL library. In addition, for convolutional neural network inference tasks on shallow network structures, our proposed framework achieves amortized inference performance at the millisecond level and sub-millisecond level on small batch and large batch data, respectively. For convolutional neural network inference tasks on deeper network structures (i.e., ResNet-20), our proposed framework achieves second-level inference. Meng Hao 0002, Weizhe Zhang, Desheng Wang 0002 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Fast Memory Disaggregation with SwiftSwap
Xiangwei Zhang, Desheng Wang 0002, Weizhe Zhang, Zhiji Yu, Meng Hao 0002 |
NPC (1) | 5 |
| 2024 | Optimizing depthwise separable convolution on DCUabstractAbstract The integration of Large Language Models (LLMs) with Convolutional Neural Networks (CNNs) is significantly advancing the development of large models. However, the computational cost of large models is high, necessitating optimization for greater efficiency. One effective way to optimize the CNN is the use of depthwise separable convolution (DSC), which decouples spatial and channel convolutions to reduce the number of parameters and enhance efficiency. In this study, we focus on porting and optimizing DSC kernel functions from the GPU to the Deep Computing Unit (DCU), a computing accelerator developed in China. For depthwise convolution, we implement a row data reuse algorithm to minimize redundant data loading and memory access overhead. For pointwise convolution, we extend our dynamic tiling strategy to improve hardware utilization by balancing resource allocation among blocks and threads, and we enhance arithmetic intensity through a channel distribution algorithm. We implement depthwise and pointwise convolution kernel functions and integrate them into PyTorch as extension modules. Experiments demonstrate that our optimized kernel functions outperform the MIOpen library on the DCU, achieving up to a 3.59 $$\times$$ × speedup in depthwise convolution and up to a 3.54 $$\times$$ × speedup in pointwise convolution. These results highlight the effectiveness of our approach in leveraging the DCU’s architecture to accelerate deep learning operations. Meng Hao 0002, Weizhe Zhang, Gangzhao Lu, Xueyang Tian, Siyu Yang 0002, Mingdong Xie, Chenyu Yuan, Desheng Wang 0002 |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | DRLCAP: Runtime GPU Frequency Capping With Deep Reinforcement LearningabstractPower and energy consumption is the limiting factor of modern computing systems. As the GPU becomes a mainstream computing device, power management for GPUs becomes increasingly important. Current works focus on GPU kernel-level power management, with challenges in portability due to architecture-specific considerations. We presentDRLCap, a general runtime power management framework intended to support power management across various GPU architectures. It periodically monitors system-level information to dynamically detect program phase changes and model the workload and GPU system behavior. This elimination from kernel-specific constraints enhances adaptability and responsiveness. The framework leverages dynamic GPU frequency capping, which is the most widely used power knob, to control the power consumption.DRLCapemploys deep reinforcement learning (DRL) to adapt to the changing of program phases by automatically adjusting its power policy through online learning, aiming to reduce the GPU power consumption without significantly compromising the application performance. We evaluateDRLCapon three NVIDIA and one AMD GPU architectures. Experimental results show thatDRLCapimproves prior GPU power optimization strategies by a large margin. On average, it reduces the GPU energy consumption by 22% with less than 3% performance slowdown on NVIDIA GPUs. This translates to a 20% improvement in the energy efficiency measured by the energy-delay product (EDP) over the NVIDIA default GPU power management strategy. For the AMD GPU architecture,DRLCapsaves energy consumption by 10%, on average, with a 4% percentage loss, and improves energy efficiency by 8%. Yiming Wang 0010, Meng Hao 0002, Weizhe Zhang, Qiuyuan Tang, Zheng Wang 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Model-Free GPU Online Energy OptimizationabstractGPUs play a central and indispensable role as accelerators in modern high-performance computing (HPC) platforms, enabling a wide range of tasks to be performed efficiently. However, the use of GPUs also results in significant energy consumption and carbon dioxide (CO2) emissions. This article presents MF-GPOEO, a model-free GPU online energy efficiency optimization framework. MF-GPOEO leverages a synthetic performance index and a PID controller to dynamically determine the optimal clock frequency configuration for GPUs. It profiles GPU kernel activity information under different frequency configurations and then compares GPU kernel execution time and gap duration between kernels to derive the synthetic performance index. With the performance index and measured average power, MF-GPOEO can use the PID controller to try different frequency configurations and find the optimal frequency configuration under the guidance of user-defined objective functions. We evaluate the MF-GPOEO by running it with 74 applications on an NVIDIA RTX3080Ti GPU. MF-GPOEO delivers a mean energy saving of 26.2% with a slight average execution time increase of 3.4% compared with NVIDIA's default clock scheduling strategy. Farui Wang, Meng Hao 0002, Weizhe Zhang, Zheng Wang 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2022 | Online Power Management for Multi-Cores: A Reinforcement Learning Based ApproachabstractPower and energy is the first-class design constraint for multi-core processors and is a limiting factor for future-generation supercomputers. While modern processor design provides a wide range of mechanisms for power and energy optimization, it remains unclear how software can make the best use of them. This article presents a novel approach for runtime power optimization on modern multi-core systems. Our policy combines power capping and uncore frequency scaling to match the hardware power profile to the dynamically changing program behavior at runtime. We achieve this by employing reinforcement learning (RL) to automatically explore the energy-performance optimization space from training programs, learning the subtle relationships between the hardware power profile, the program characteristics, power consumption and program running times. Our RL framework then uses the learned knowledge to adapt the chip's power budget and uncore frequency to match the changing program phases for any new, previously unseen program. We evaluate our approach on two computing clusters by applying our techniques to 11 parallel programs that were not seen by our RL framework at the training stage. Experimental results show that our approach can reduce the system-level energy consumption by 12 percent, on average, with less than 3 percent of slowdown on the application performance. By lowering the uncore frequency to leave more energy budget to allow the processor cores to run at a higher frequency, our approach can reduce the energy consumption by up to 17 percent while improving the application performance by 5 percent for specific workloads. Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Dynamic GPU Energy Optimization for Machine Learning Training WorkloadsabstractGPUs are widely used to accelerate the training of machine learning workloads. As modern machine learning models become increasingly larger, they require a longer time to train, leading to higher GPU energy consumption. This paper presents GPOEO, an online GPU energy optimization framework for machine learning training workloads. GPOEO dynamically determines the optimal energy configuration by employing novel techniques for online measurement, multi-objective prediction modeling, and search optimization. To characterize the target workload behavior, GPOEO utilizes GPU performance counters. To reduce the performance counter profiling overhead, it uses an analytical model to detect the training iteration change and only collects performance counter data when an iteration shift is detected. GPOEO employs multi-objective models based on gradient boosting and a local search algorithm to find a trade-off between execution time and energy consumption. We evaluate the GPOEO by applying it to 71 machine learning workloads from two AI benchmark suites running on an NVIDIA RTX3080Ti GPU. Compared with the NVIDIA default scheduling strategy, GPOEO delivers a mean energy saving of 16.2% with a modest average execution time increase of 5.1%. Farui Wang, Weizhe Zhang, Shichao Lai, Meng Hao 0002, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Automatic translation of data parallel programs for heterogeneous parallelism through OpenMP offloading
Farui Wang, Weizhe Zhang, Meng Hao 0002, Gangzhao Lu, Zheng Wang 0001 |
J. Supercomput. | 4 |
| 2021 | Fine-Grained Powercap Allocation for Power-Constrained Systems Based on Multi-Objective Machine LearningabstractPower capping is an important solution to keep the system within a fixed power constraint. However, for the over-provisioned and power-constrained systems, especially the future exascale supercomputers, powercap needs to be reasonably allocated according to the workloads of compute nodes to achieve trade-offs among performance, energy and powercap. Thus it is necessary to model performance and energy and to predict the optimal powercap allocation strategies. Existing power allocation approaches have insufficient granularity within nodes. Modeling approaches usually model performance and energy separately, ignoring the correlation between objectives, and do not expose the Pareto-optimal powercap configurations. Therefore, this article combines the powercap with uncore frequency scaling and proposes an approach to predict the Pareto-optimal powercap configurations on the power-constrained system for input MPI and OpenMP parallel applications. Our approach first uses the elaborately designed micro-benchmarks and a small number of existing benchmarks to build the training set, and then applies a multi-objective machine learning algorithm which combines the stacked single-target method with extreme gradient boosting to build multi-objective models of performance and energy. The models can be used to predict the optimal processor and memory powercap settings, helping compute nodes perform fine-grained powercap allocation. When the optimal powercap configuration is determined, the uncore frequency scaling is used to further optimize the energy consumption. Compared with the reference powercap configuration, the predicted optimal configurations predicted by our method can achieve an average powercap reduction of 31.35 percent, an average energy reduction of 12.32 percent, and average performance degradation of only 2.43 percent. Meng Hao 0002, Weizhe Zhang, Yiming Wang 0010, Gangzhao Lu, Farui Wang, Athanasios V. Vasilakos |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Automatic generation of benchmarks for I/O-intensive parallel applications
Meng Hao 0002, Weizhe Zhang, Marc Snir, Laurence T. Yang |
J. Parallel Distributed Comput. | 1 |
| 2016 | Communication optimization for RDMA-based science data transmission tools
Weizhe Zhang, Meng Hao 0002 |
J. Supercomput. | 2 |