Minyu Cui

dblp:191/6491 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-5983-1648ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Characterizing and Mitigating Performance Variability in Parallel Applications on Modern HPC multicore Systems
abstract
In high-performance computing (HPC), OpenMP has become the de facto programming model for shared-memory systems.However, running OpenMP-based parallel applications on multicore systems often faces the challenge of performance variability, particularly as core counts increase in modern HPC clusters.Factors spanning from Operating Systems (OS) and hardware feature to OpenMP implementation can significantly impact performance stability.This paper evaluates execution time variability across five multicore systems from multiple HPC clusters, covering two different ISAs and using five OpenMP benchmarks and a real-world mini-app compiled with both gcc and llvm/clang.We analyze the effects of various factors such as thread-pinning, OpenMP runtime implementations, OpenMP scalability, simultaneous multithreading (SMT), core resource reservation, frequency scaling, and platformspecific features such as hybrid architecture core configurations.Our findings highlight the complex interplay of these factors in performance variability and propose lightweight mitigation strategies to enhance the stability of OpenMP programs for developers and system users.
Minyu Cui, Miquel Pericàs
CF1
2025 Reliable and Energy Optimized Task Mapping for Heterogeneous Multicore NoC Based on Partial Task Duplication and Multipath Routing
abstract
The increasing integration of heterogeneous processors on a chip presents significant challenges for efficient management in Multi-Processor System-on-Chip (MPSoC) platforms. Network-on-Chip (NoC) architectures offer a flexible and scalable interconnection paradigm through router-based communication. However, mapping dependent, real-time tasks in NoC environments critically affects data processing and transmission efficiency. An optimized task mapping scheme must address constraints such as real-time deadlines, energy consumption, and reliability, which are key metrics for modern NoCs. Existing approaches often overlook the complex interplay between communication paths and their associated energy costs, resulting in suboptimal resource utilization. This paper proposes a comprehensive task mapping framework that jointly optimizes energy efficiency and reliability by integrating Dynamic Voltage and Frequency Scaling (DVFS), multi-path data routing, task allocation, scheduling, and partial task duplication. We formulate the problem as a complex combinatorial optimization task and transform it into a solvable form with reduced computational complexity. Simulation results demonstrate that the proposed method achieves superior energy efficiency by reducing energy consumption by up to 39.7%, reducing computation time, and improving task schedulability compared to existing state-of-the-art approaches.
Lei Mo, Tamim M. Al-Hasan, Minyu Cui, Xiaojun Zhai, Qing Gao 0001, Shibo He
IEEE Internet Things J.4
2023 Near-optimal energy-efficient partial-duplication task mapping of real-time parallel applications
Minyu Cui, Angeliki Kritikakou, Lei Mo, Emmanuel Casseau
J. Syst. Archit.1
2021 Fault-Tolerant Mapping of Real-Time Parallel Applications under multiple DVFS schemes
abstract
On multicore platforms, reliable task execution, as well as low energy consumption, are essential. Dynamic Voltage/Frequency Scaling (DVFS) is typically used for energy saving, but with a negative impact on reliability, especially when the frequency is low. Using high frequencies to meet reliability constraints is not always feasible, while multiple replicas increase energy consumption. To minimize energy consumption, enhancing reliability, without violating real-time constraints, we propose an approach that combines distinct reliability enhancement techniques, under task-level, processor-level and system-level DVFS. Our task mapping problem jointly optimizes task allocation, task frequency assignment, and task duplication, under multiple constraints. This is achieved by formulating the task mapping problem as a mixed integer non-linear programming problem and equivalently transforming it into a mixed integer linear programming, that is optimally solved. From the obtained results, the proposed approach achieves better energy consumption and finds solutions, when other approaches fail.
Minyu Cui, Angeliki Kritikakou, Lei Mo, Emmanuel Casseau
RTAS1