Yuchen Ma 0001

dblp:192/2001-1 · also Sam Ma 0001 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0001-8884-6278ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 27% Hardware accelerators and domain-specific architectures · 21% Memory systems · 12%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › numerical linear algebra
distributed linear algebra
1.012026
A Distributed Matrix-Block-Vector Multiplication in Presence of System Performance Variability · PPoPP 2026
Hardware accelerators and domain-specific architectures
matrix-vector multiplication
1.012026
A Distributed Matrix-Block-Vector Multiplication in Presence of System Performance Variability · PPoPP 2026
Storage systems
data placement
0.612022
LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications · PPoPP 2022
Memory systems
hybrid memory
0.612022
LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications · PPoPP 2022
Cloud and datacenter computing › resource management
memory management for HPC
0.612022
LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications · PPoPP 2022
Parallel and multicore computing › parallel programming models › task parallelism
task-parallel programs
0.612022
LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications · PPoPP 2022
High-performance computing
performance optimization
0.312026
A Distributed Matrix-Block-Vector Multiplication in Presence of System Performance Variability · PPoPP 2026
Performance modeling and evaluation › profiling
memory profiling with task semantics
0.212022
LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications · PPoPP 2022

Methods — techniques the papers use, named apart from their topics

static thread scheduling · 1.0simulation · 1.0performance modeling · 1.0cache-aware tiling · 1.0page migration · 0.6memory profiling · 0.6
YearPublicationVenuePosition
2026 A Distributed Matrix-Block-Vector Multiplication in Presence of System Performance Variability
abstract
Distributed matrix-block-vector multiplication (Matvec) algorithm is a critical component of many applications, but can be computationally challenging for dense matrices of dimension O(10^6–10^7) and blocks of O(10–100) vectors. We present performance analysis, implementation, and optimization of our SMatVec library for Matvec under the effect of system variability. Our modeling shows that 1D pipelining Matvec is as efficient as 2D algorithms at small to medium clusters, which are sufficient for these problem sizes. We develop a performance tracing framework and a simulator that reveal pipeline bubbles caused by modest ~5% system variability. To tolerate such variability, our SMatVec library, which combines on-the-fly kernel matrix generation and Matvec, integrates four optimizations: inter-process data preloading, unconventional static thread scheduling, cache-aware tiling, and multi-version unrolling. In our benchmarks on O(10^5) Matvec problems, SMatVec achieves up to 1.85× speedup over COSMA and 17× over ScaLAPACK. For O(10^6) problems, where COSMA and ScaLAPACK exceed memory capacity, SMatVec maintains linear strong scaling and achieves peak performance of 75% FMA Flop/s. Its static scheduling policy has a 2.27× speedup compared to the conventional work-stealing dynamic scheduler, and is predicted to withstand up to 108% performance variability under exponential distributed variability simulation.
Yuchen Ma 0001, Bin Ren 0002, Andreas Stathopoulos
PPoPP1
2022 LB-HM: load balance-aware data placement on heterogeneous memory for task-parallel HPC applications
abstract
The emergence of heterogeneous memory (HM) provides a cost-effective and high-performance solution to memory-consuming HPC applications. However, using HM, wisely migrating data objects on it is critical for high performance. In this work, we introduce a load balance-aware page management system, named LB-HM. LB-HM introduces task semantics during memory profiling, rather than being application-agnostic. Evaluating with a set of memory-consuming HPC applications, we show that we show that LB-HM reduces existing load imbalance and leads to an average of 17.1% and 15.4% (up to 26.0% and 23.2%) performance improvement, compared with a hardware-based solution and an industry-quality software-based solution on Optane-based HM.
Jie Liu 0096, Yuchen Ma 0001, Jiajia Li 0001, Dong Li 0001
PPoPP3
2019 PASTA: a parallel sparse tensor algorithm benchmark suite
Jiajia Li 0001, Yuchen Ma 0001, Ang Li 0006, Kevin J. Barker
CCF Trans. High Perform. Comput.2
2019 Optimizing sparse tensor times matrix on GPUs
Yuchen Ma 0001, Jiajia Li 0001, Chenggang Yan 0001, Jimeng Sun 0001, Richard W. Vuduc
J. Parallel Distributed Comput.1