Juan Chen 0001

dblp:54/2717-1 · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-9941-9885ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 2 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AdaPolySI: Adaptive Polynomial Filtered Subspace Iteration for Hermitian Interior Eigenvalue Problems
Yuhui Ni, Shengguo Li, Juan Chen 0001, Jianchun Wang, José E. Román
ICS5
2026 CausalTuner: Feature-Aware Causal Guidance for Compiler Auto-tuning
abstract
Modern compilers like LLVM and GCC provide hundreds of optimization options (e.g., flags or passes), yet their fixed, predefined sequences (e.g., -O3) often fail to exploit the full performance potential of specific programs. Search-based auto-tuning has emerged to address this pass selection and ordering problem — known as the phase ordering problem. Existing approaches typically reduce search complexity by identifying critical flags or employing localized group-based mutations. However, these methods often remain feature-agnostic or rely on manually designed static mappings during the online search. Consequently, they fail to adaptively link program features to optimization logic. Furthermore, they are prone to being misled by spurious correlations and noise inherent in sparse performance data.
Jiaqing Zhong, Juan Chen 0001, Yichang Zhou, Kuan Li
LCTES2
2026 A survey of anomaly detection in HPC systems using machine learning
abstract
Abstract High-performance computing (HPC) systems must remain stable and reliable to consistently deliver robust computational power and ensure the proper execution of user jobs. Anomaly detection is a key means to ensure the stability and reliability of these systems. With the expansion of HPC systems and changes in their architecture, accurately identifying anomalies in dynamic environments has become increasingly challenging. Traditional detection methods rely on experience and rules, which could be inefficient and inaccurate. To address these issues, researchers have proposed machine learning-based methods to automatically process large amounts of complex data, improving the efficiency of anomaly identification and diagnosis. In this survey, we conduct a comprehensive and in-depth investigation of machine learning-based anomaly detection methods in HPC systems. Firstly, we summarize and introduce the background and challenges of anomaly detection in HPC systems. Secondly, we compare a series of machine learning-based anomaly detection works in detail and summarize their frameworks. We conclude their advantages and disadvantages and application scenarios. Finally, we discuss several promising development trends of machine learning-based HPC system anomaly detection.
Wei Zhang 0027, Yiqin Dai, Huijun Wu 0001, Zhenwei Wu, Hongyun Tian, Juan Chen 0001, Chubo Liu, Yong Dong
CCF Trans. High Perform. Comput.9
2026 CML-PowF: Data Clustering Matching Based Low-overhead Multiple CPU Real-time Power Forecasting
abstract
Efficient CPU power capping is essential for energy saving and fault tolerance in parallel computing clusters, but its effectiveness depends on accurate and timely processor power forecasting with minimal sampling overhead. Existing methods often struggle to balance these factors under scalability constraints, as hardware limitations tightly bound the available sampling resources. This article focuses on the issue of high-precision real-time processor power forecasting while maintaining (or minimally increasing) the total overhead of multiprocessor power forecasting, particularly when the parallelism scale ranges from P to 2P processors or when the problem size scales from M to 2M . We propose CML-PowF , a low-overhead multiprocessor real-time power forecasting approach based on data clustering. CML-PowF integrates two key algorithms: Alg-CEF , which conducts cluster matching on the runtime characteristics of the program at the P / M scale, and models the tradeoff among forecasting error, time span, and sampling overhead. Alg-MSF , which leverages execution patterns from smaller-scale runs to determine the optimal sampling overhead and forecasting time span at the 2P / 2M scale. We evaluate CML-PowF on x86 and ARM platforms with up to 32 computing nodes (2,048 cores). Results show that it achieves 3–6% forecasting error at large scales with only 0.2–0.5% degradation compared to the P / M scale, without increasing total sampling overhead. Integrated with the PowC control system, CML-PowF effectively maintains real-time processor power below target thresholds.
Rongyu Deng, Juan Chen 0001, Yuan Yuan 0034, Yong Dong, Aolin Cao, Yida Gu, Dingwen Tao
ACM Trans. Archit. Code Optim.2
2025 MIST: Towards MPI Instant Startup and Termination on Tianhe HPC Systems
abstract
As the size of MPI programs grows with expanding HPC resources and parallelism demands, the overhead of MPI startup and termination escalates due to the inclusion of less scalable global operations. Global operations involving extensive cross-machine communication and synchronization are crucial for ensuring semantic correctness. The current focus is on optimizing and accelerating these global operations rather than removing them, as the latter involves systematic changes to the system software stack and may impact program semantics. Given this background, we propose a systematic solution named MIST to safely eliminate global operations in MPI startup and termination. Through optimizing the generation of communication addresses, designing reliable communication protocols, and exploiting the resource release mechanism, MIST eliminates all global operations to achieve MPI instant startup and termination while ensuring correct program execution. Experiments on Tianhe-2A supercomputer demonstrate that MIST can reduce theMPI_Init()time by 32.5-77.6% and theMPI_Finalize()time by 28.9-85.0%.
Yiqin Dai, Ruibo Wang, Yong Dong, Juan Chen 0001, Huijun Wu 0001, Mingtian Shao, Kai Lu 0001
IEEE Trans. Parallel Distributed Syst.5
2024 EAtuner: Comparative Study of Evolutionary Algorithms for Compiler Auto-tuning
abstract
The manual adjustment of compilation flags by compiler users is impractical due to the exponential size of the search space. To address this, machine learning-based compiler auto-tuning methods, particularly evolutionary algorithms, have been proposed. However, existing works use different benchmarks and experimental setups, making it difficult to compare the strengths and weaknesses of various algorithms. To address this, we present EAtuner, an evolutionary algorithm-based framework for compiler auto-tuning, with the goal of benchmarking and identifying suitable algorithms for compiler auto-tuning. We implement ten discrete binary evolutionary algorithms and evaluate their effectiveness on the LLVM compiler through experiments. Notably, eight of these algorithms have not been previously applied to compiler flag optimization problems before our work. The results show that all ten algorithms can effectively achieve compiler auto-tuning, resulting in an average speedup of 1.204. However, there are notable differences in the effectiveness and efficiency of each algorithm, particularly in optimization efficiency, which is positively correlated with the number of program compilations. Based on this, we classify the algorithms into three levels, with Differential Evolution (DE) showing significant advantages in optimization effectiveness and efficiency. Additionally, we provide a comprehensive summary of the applicability of compiler flags, the correlation between them, and their relationship with programs.
Guojian Xiao, Siyuan Qin, Kuan Li, Juan Chen 0001, Jianping Yin
CSCWD4
2024 FIFO: Fuzzy Cluster Identification and High-dimensional Feature Clustering Optimization Based CPU Power Sampling Optimization
abstract
The high accuracy of processor power consumption modeling has consistently posed challenges in processor design and program power optimization. With the increasing complexity of processor architectures and the diversification of application types, enhancing the accuracy of processor power models has become increasingly difficult. In addition to model selection and feature selection for modeling parameters, the quantity and distribution of training set samples significantly affect the improvement of processor model accuracy. To reduce model complexity, high-dimensional feature spaces are often subjected to dimensionality reduction. However, this can sometimes lead to bias in sample point clustering within the feature space (fuzzy clustering), thereby impacting the accuracy of processor power models. Addressing this issue, this paper proposes a novel method called "FIFO: Fuzzy Cluster Identification and Feature Optimization for Processor Power Modeling Sample Optimization". The FIFO algorithm optimizes the distribution of training sample points for processor power models by implementing fuzzy cluster identification in low-dimensional feature space (FI), high-dimensional feature space restoration and clustering optimization (FO), and redundant point elimination based on mixed-dimension feature spaces, thereby enhancing the accuracy of processor modeling. Validation of the FIFO algorithm’s effectiveness was conducted through linear power modeling and neural network power modeling on both x86 and ARM processor platforms. Experimental results demonstrate that employing the FIFO algorithm reduces processor power model errors by an average of 13.69% on an ARMv8-based architecture processor platform, by 15.76% on the Intel Xeon Gold-6226R processor platform, and by 25.41% on the Intel Xeon E5-2660 processor platform.
Shaojun Feng, Juan Chen 0001, Yichang Zhou, Rongyu Deng, Xianyu Wu, Jiaqing Zhong
HPCC4
2024 CPU Power Modeling Through Training Data Selection
abstract
CPU power modeling aims to address issues related to power prediction and fine-grained power measurement during power management. However, existing studies have mainly focused on designing supervised learning methods for power modeling, neglecting the effect of training samples. Therefore, this paper presents a CPU power modeling method that adaptively selects high-quality sample data iteratively. The method is applied to linear regression (LR), neural networks (NN) and random forest models (RF) on x86 and ARM-based platforms. Experimental results show that the method can reduce the average MAPE by 13.33% for the LR model, 27.45% for the NN model and 28.30% for the RF model.
Jiaqing Zhong, Juan Chen 0001
HPCC3
2024 Energy-Aware Task Mapping for Multi-Core Processor: A Machine Learning Based Approach
abstract
When parallel workloads do not scale with the core, it is challenging to utilize parallelism effectively, especially on distributed cache/memory multi-core platforms. The mapping of processes to cores affects both the cache/memory latencies and the power consumption, determining the energy efficiency of computing subsystems. Consequently, performance and energy metrics usually cannot be simultaneously optimal, which demands a trade-off optimization mapping strategy. To this end, this paper presents a machine learning-based Energy-aware task Mapping Predictor (EMP). EMP adjusts memory-intensive and compute-intensive programs mapping on a single computing node by capturing hardware resources and program characteristics. For parallel programs on multiple nodes, EMP tries to add nodes and adjusts the mapping strategy to improve program performance while controlling energy consumption. We evaluate EMP on both ARM and x86 platforms. In single-node experiments, EMP reduces the Energy Delay Product (EDP) by 68.35% and 60.45% on ARM (A) and ARM (B), respectively, and in multi-node experiments, EMP reduces the EDP by 29.52% as compared to clustered task mapping.
Tao Xu 0052, Juan Chen 0001, Yong Dong
HPCC3
2024 Faster and Scalable MPI Applications Launching
abstract
Distributed parallel MPI applications are the dominant workload in many high-performance computing systems. While optimizing MPI application execution is a well-studied field, little work has considered optimizing the initial MPI application launching phase, which incurs extensive cross-machine communications and synchronization. The overhead of MPI application launching can be expensive, accounting for more than million core hours per 10K nodes annually on the production Tianhe-2A supercomputer, which will increase as the number of parallel machines used grows. Therefore, it is critical to optimize the MPI application launching process. This paper presents a novel approach to optimizing the MPI application launch. Our approach adopts a location-aware address generation rule to eliminate the need for address exchange and a topology-aware global communication scheme to optimize cross-machine synchronization. We then design a new application launch procedure to support the proposed optimizations to further reduce the pressure of the shared I/O system. Our techniques have been deployed to production in the Tianhe-2A supercomputer and the Next Generation Tianhe Supercomputer. Experimental results show that our approach scales well and outperforms alternative schemes, reducing the MPI application launching time by over 29% with 320K MPI processes.
Yong Dong, Yiqin Dai, Kai Lu 0001, Ruibo Wang, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001
IEEE Trans. Parallel Distributed Syst.6
2023 PowerDis: Fine-Grained Power Monitoring Through Power Disaggregation Model
Xinxin Qi, Juan Chen 0001, Rongyu Deng, Yuan Yuan 0034, Yonggang Che
ICA3PP (4)2
2023 HighRPM: Combining Integrated Measurement and Sofware Power Modeling for High-Resolution Power Monitoring
abstract
In an era where power and energy are the first-class constraints of computing systems, accurate power information is crucial for energy efficiency optimization in parallel computing systems. Existing power monitoring techniques rely on either software-centric power models that suffer from poor accuracy or integrated hardware measurement schemes that have a low reading update frequency and coarse granularity. These result in a low spatiotemporal resolution for power monitoring. This paper introduces HighRPM, a new method for accurately measuring power consumption on parallel computing systems. HighRPM combines coarse-grained power sensor readings and software power modeling techniques to improve temporal and spatial resolutions. To provide high-frequent power readings in the temporal domain, HighRPM employs statistical modeling and machine learning techniques to predict the long-term power trend and the short-term fluctuations in power consumption. To improve spatial coverage, HighRPM takes low-time resolution node-level power consumption and uses a neural network to distribute the power readings to lower-level computing components like CPUs and memory components. We evaluate HighRPM by applying it to both ARM-based and X86-based platforms. Experimental results show that HighRPM improves time resolution by 10 times, provides accurate readings for CPUs and memory, and reduces error by 7-24% compared to other power modeling methods.
Xinxin Qi, Juan Chen 0001, Yong Dong, Yuan Yuan 0034, Tao Xu 0052, Rongyu Deng, Kexing Zhou, Zheng Wang 0001
ICPP2
2023 Processor power forecasting through model sample analysis and clustering
Kexing Zhou, Yong Dong, Juan Chen 0001, Rongyu Deng, Yifei Guo, Zhixin Ou
CCF Trans. High Perform. Comput.3
2022 AOA: Adaptive Overclocking Algorithm on CPU-GPU Heterogeneous Platforms
abstract
Abstract Although GPUs have been used to accelerate various convolutional neural network algorithms with good performance, the demand for performance improvement is still continuously increasing. CPU/GPU overclocking technology brings opportunities for further performance improvement in CPU-GPU heterogeneous platforms. However, CPU/GPU overclocking inevitably increases the power of the CPU/GPU, which is not conducive to energy conservation, energy efficiency optimization, or even system stability. How to effectively constrain the total energy to remain roughly unchanged during the CPU/GPU overclocking is a key issue in designing adaptive overclocking algorithms. There are two key factors during solving this key issue. Firstly, the dynamic power upper bound must be set to reflect the real-time behavior characteristics of the program so that algorithm can better meet the total energy unchanging constraints; secondly, instead of independently overclocking at both CPU and GPU sides, coordinately overclocking on CPU-GPU must be considered to adapt to real-time load balance for higher performance improvement and better energy constraints. This paper proposes an Adaptive Overclocking Algorithm (AOA) on CPU-GPU heterogeneous platforms to achieve the goal of performance improvement while the total energy remains roughly unchanged. AOA uses the function $$F_k$$ F k to describe the variable power upper bound and introduces the load imbalance factor W to realize the CPU-GPU coordinated overclocking. Through the verification of several types convolutional neural network algorithms on two CPU-GPU heterogeneous platforms (Intel $$^\circledR $$ ® Xeon E5-2660 & NVIDIA $$^\circledR $$ ® Tesla K80; Intel $$^\circledR $$ ® Core™i9-10920X & NIVIDIA $$^\circledR $$ ® GeForce RTX 2080Ti), AOA achieves an average of 10.7% performance improvement and 4.4% energy savings. To verify the effectiveness of the AOA, we compare AOA with other methods including automatic boost, the highest overclocking and static optimal overclocking.
Zhixin Ou, Juan Chen 0001, Tao Xu 0052, Guodong Jiang, Zhengyuan Tan, Xinxin Qi
ICA3PP2
2022 CP3: Hierarchical Cross-Platform Power/Performance Prediction Using a Transfer Learning Approach
abstract
Abstract Cross-platform power/performance prediction is becoming increasingly important due to the rapid development and variety of software and hardware architectures in an era of heterogeneous multi-core. However, accurate power/performance prediction is faced with an obstacle caused by the large gap between architectures, which is often overcome by laborious and time-consuming fine-grained program profiling on the target platform. To overcome these problems, this paper introduces $$CP^3$$ C P 3 , a hierarchical Cross-platform Power/Performance Prediction framework, which focuses on utilizing architecture differences to migrate built models to target platforms. The core of $$CP^3$$ C P 3 is the three-step hierarchical transfer learning approach, hierarchical division, partial transfer learning, and model fusion, respectively. $$CP^3$$ C P 3 firstly builds a power/performance model on the source platform, then rebuilds it with the reduced training data on the target platform, and finally obtains a cross-platform model. We validate the effectiveness of $$CP^3$$ C P 3 using a group of benchmarks on X86- and ARM-based platforms that use three different types of commonly used processors. Evaluation results show that when applying $$CP^3$$ C P 3 , only 1% of the baseline training data is required to achieve high cross-platform prediction accuracy, with power prediction error being only 0.65%, and performance prediction error being only 4.64%.
Xinxin Qi, Juan Chen 0001
ICA3PP2
2022 The Fast and Scalable MPI Application Launch of the Tianhe HPC system
abstract
Fast and scalable MPI application launch helps achieve exascale performance and is becoming a common goal in high-performance computing. However, the traditional launch technique suffers from scalability deficiencies in the global information exchange and the global barrier operation. This drawback makes it challenging to launch MPI applications quickly in large-scale systems. In this paper, we propose a fast and scalable application launch technique and details its associated hardware and software support. The optimized launch technique includes a locality-aware static address generation rule for eliminating the need for address exchange and a topology-aware global communication scheme for improving global communication efficiency. We also propose an optimized application launch sequence for supporting the above launch technique. We implement and evaluate the proposed launch technique on the Tianhe-2A supercomputer and the Tianhe Exascale Prototype Upgrade System. Experimental results show that our technique can reduce the launch time by 26.1% when launching an application with 256K processes.
Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Mingtian Shao, Juan Chen 0001
IPDPS7
2022 Towards Scalable Resource Management for Supercomputers
abstract
Today's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin.
Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001
SC6
2021 VPC: Pruning connected components using vector-based path compression for Graph500
Xinbiao Gan, Tianjing Xu, Menghan Jia, Juan Chen 0001, Yiming Zhang 0003
CCF Trans. High Perform. Comput.6
2021 Correction to: VPC: Pruning connected components using vector-based path compression for Graph500
Xinbiao Gan, Tianjing Xu, Menghan Jia, Juan Chen 0001, Yiming Zhang 0003
CCF Trans. High Perform. Comput.6
2020 High-Performance Computing and Engineering Educational Development and Practice
abstract
This Innovate Practice Full Paper presents HPC and engineering educational development and practice. Educators and researchers are witnessing the emergence of parallel and high-performance computing (HPC) in many computing and engineering environments worldwide. Single-core processing is slowly becoming obsolete while parallel solutions to application development are now replacing serial solutions to achieve higher performance, particularly in many scientific, engineering, and computing fields. Some universities or research institutes worldwide have developed and have become centers of supercomputing. Most of these centers, particularly the world acclaimed Tianhe-2 supercomputer developed by China's National University of Defense Technology (NUDT), often have their own HPC educational approaches as well as their own curricula at both undergraduate and graduate levels. This paper presents some high-level HPC educational concepts, educational framework designs, and strategies for cultivating HPC talented graduates. In this work, the authors show unique perspectives and practical experiences on student capability, oriented toward HPC curricular development. They also show a step-by-step practical system in which real platform, real application, and research-teaching integration modes can immerse students into an HPC world for better understanding. The authors show how researchers with rich computing and engineering experience can become deeply involved in course design, in-class teaching, and in laboratory instruction. They also discuss the effectiveness of HPC courses at the NUDT and they show how HPC can become a stable element in computing and engineering education.
Juan Chen 0001, John Impagliazzo, Li Shen 0007
FIE1
2020 PMC-Based Dynamic Adaptive CPU and DRAM Power Modeling
Yunfang Zhang, Yong Dong, Juan Chen 0001, Zhixin Ou, Yuan Yuan 0034
ICA3PP (1)3
2020 Toward High Performance Computing Education
abstract
High Performance Computing (HPC) is the ability to process data and perform complex calculations at extremely high speeds. Current HPC platforms can achieve calculations on the order of quadrillions of calculations per second with quintillions on the horizon. The past three decades witnessed a vast increase in the use of HPC across different scientific, engineering and business communities, for example, sequencing the genome, predicting climate changes, designing modern aerodynamics, or establishing customer preferences. Although HPC has been well incorporated into science curricula such as bioinformatics, the same cannot be said for most computing programs. This working group will explore how HPC can make inroads into computer science education, from the undergraduate to postgraduate levels. The group will address research questions designed to investigate topics such as identifying and handling barriers that inhibit the adoption of HPC in educational environments, how to incorporate HPC into various curricula, and how HPC can be leveraged to enhance applied critical thinking and problem solving skills. Four deliverables include: (1) a catalog of core HPC educational concepts, (2) HPC curricula for contemporary computing needs, such as in artificial intelligence, cyberanalytics, data science and engineering, or internet of things, (3) possible infrastructures for implementing HPC coursework, and (4) HPC-related feedback to the CC2020 project.
Rajendra K. Raj, Carol J. Romanowski, Sherif G. Aly 0001, Brett A. Becker, Juan Chen 0001, Sheikh K. Ghafoor, Nasser Giacaman, Steven Gordon 0001, Cruz Izu, Nick Rahimi, Michael P. Robson, Neena Thota
ITiCSE5
2019 Analyzing time-dimension communication characterizations for representative scientific applications on supercomputer systems
Juan Chen 0001, Yong Dong, Feihao Wu, Enqiang Zhou, Yuhua Tang
Frontiers Comput. Sci.1
2018 Design of Practical Experiences to Improve Student Understanding of Efficiency and Scalability Issues in High Performance Computing: (Abstract Only)
abstract
With the increasing demand of big data technology, there has been a growing interest of introducing high performance computing in computer science curriculum. One challenge in helping students understand the nature of efficiency and scalability issues in high performance computing is the lack of opportunities for them to be engaged in large-scale applications that run on supercomputer system architecture. This poster presents a collection of example projects that have been used in a parallel computing course in multiple universities in China, including National University of Defense Technology, Sun Yat-sen University and Hunan University. These projects were adopted from a wide range of scientific computing applications such as CFD, text mining of biomedical literature and so on. The large-scale computing resource for courses is supported by two National Supercomputing Centers, one in Guangzhou and the other in Changsha. The poster describes the background, objective, structure, task, practice process and outcome for each project. It also discusses the impact on student understanding all kinds of key topics and major challenges related to computational efficiency and scalability. Such projects build a positive practical environment to make students indulge in doing all kinds of interesting and helpful trials to validate their assumptions, especially when they have different perspectives or results for one problem. The poster presents our design evaluation rubric to reflect the effectiveness of our practice, as well as the statistics about the students" achievements for the last three semesters.
Juan Chen 0001, Li Shen 0007, Jianping Yin, Chunyuan Zhang
SIGCSE1
2016 Detailed and clock-driven simulation for HPC interconnection network
Juan Chen 0001, Dezun Dong, Yuhua Tang
Frontiers Comput. Sci.2
2016 Reducing Static Energy in Supercomputer Interconnection Networks Using Topology-Aware Partitioning
abstract
The key to reducing static energy in supercomputers is switching off their unused components. Routers are the major components of a supercomputer. Whether routers can be effectively switched off or not has become the key to static energy management for supercomputers. For many typical applications, the routers in a supercomputer exhibit low utilization. However, there is no effective method to switch the routers off when they are idle. By analyzing the router occupancy in time and space, for the first time, we present a routing-policy guided topology partitioning methodology to solve this problem. We propose topology partitioning methods for three kinds of commonly used topologies (mesh, torus and fat-tree) equipped with the three most popular routing policies (deterministic routing, directionally adaptive routing and fully adaptive routing). Based on the above methods, we propose the key techniques required in this topology partitioning based static energy management in supercomputer interconnection networks to switch off unused routers in both time and space dimensions. Three topology-aware resource allocation algorithms have been developed to handle effectively different job-mixes running on a supercomputer. We validate the effectiveness of our methodology by using Tianhe-2 and a simulator for the aforementioned topologies and routing policies. The energy savings achieved on a subsystem of Tianhe-2 range from 3.8 to 79.7 percent. This translates into a yearly energy cost reduction of up to half a million US dollars for Tianhe-2.
Juan Chen 0001, Yuhua Tang, Yong Dong, Jingling Xue
IEEE Trans. Computers1
2015 GS-DMR: Low-overhead soft error detection scheme for stencil-based computation
Xiaoguang Ren, Xinhai Xu, Juan Chen 0001, Xuejun Yang
Parallel Comput.4
2011 Optimizing Linpack Benchmark on GPU-Accelerated Petascale Supercomputer
Feng Wang 0050, Canqun Yang, Yunfei Du 0001, Juan Chen 0001, Huizhan Yi, Weixia Xu 0001
J. Comput. Sci. Technol.4
2010 Adaptive Optimization for Petascale Heterogeneous CPU/GPU Computing
abstract
In this paper, we describe our experiment developing an implementation of the Linpack benchmark for TianHe-1, a petascale CPU/GPU supercomputer system, the largest GPU-accelerated system ever attempted before. An adaptive optimization framework is presented to balance the workload distribution across the GPUs and CPUs with the negligible runtime overhead, resulting in the better performance than the static or the training partitioning methods. The CPU-GPU communication overhead is effectively hidden by a software pipelining technique, which is particularly useful for large memory-bound applications. Combined with other traditional optimizations, the Linpack we optimized using the adaptive optimization framework achieved 196.7 GFLOPS on a single compute element of TianHe-1. This result is 70.1% of the peak compute capability and 3.3 times faster than the result using the vendor's library. On the full configuration of TianHe-1 our optimizations resulted in a Linpack performance of 0.563PFLOPS, which made TianHe-1 the 5th fastest supercomputer on the Top500 list released in November 2009.
Canqun Yang, Feng Wang 0050, Yunfei Du 0001, Juan Chen 0001, Jie Liu 0002, Huizhan Yi, Kai Lu 0001
CLUSTER4
2009 Solving 2D Nonlinear Unsteady Convection-Diffusion Equations on Heterogenous Platforms with Multiple GPUs
abstract
Solving complex convection-diffusion equations is very important to many practical mathematical and physical problems. After the finite difference discretization, most of the time for equations solution is spent on sparse linear equation solvers. In this paper, our goal is to solve 2D Nonlinear Unsteady Convection-Diffusion Equations by accelerating an iterative algorithm named Jacobi-preconditioned QMRCGSTAB on a heterogenous platform, which is composed of a multi-core processor and multiple GPUs. Firstly, a basic implementation and evaluation for adapting the problem to this kind of platform is given. Then, we propose two optimization methods to improve the performance: kernel merging method and matrix boundary data processing. Our experimental evaluation on an AMD Opteron(tm) quad-core processor 2380 linked to an NVIDIA Tesla S1070 platform with four GPUs delivers the peak performance of 33 GFLOPS (double precision), which is a speedup of close to a factor 32 compared to the same problem running on 4 cores of the same CPU.
Canqun Yang, Zhen Ge, Juan Chen 0001, Feng Wang 0050, Yunfei Du 0001
ICPADS3
2008 Energy-Constrained OpenMP Static Loop Scheduling
abstract
In high performance parallel computing, energy optimization for parallel loops becomes one key because the time of loops often takes a significant part of the whole execution time. Energy-constrained problem is one of the important research focuses. This paper studies energy-constrained problem based on OpenMP static loop scheduling. Firstly, we propose energy-constrained static scheduling algorithm (ECSS), which utilizes DVS to scale down voltage/frequency of the light-loaded processors in terms of energy constraint. Secondly, we propose Energy Constraint based Performance-Optimal Static Scheduling algorithm (ECPOSS), which combines loop rescheduling and DVS for the better performance under the same energy constraint. We prove ECPOSS can obtain the best performance under the same energy constraint. Through testing NPB3.2-OMP programs on 20-160 multiprocessor simulation environment, we evaluate the effectiveness of our algorithms. Experimental results show the performance of ECPOSS is better than that of ECSS by 4.81% under the 50% energy constraint on 100 processors.
Juan Chen 0001, Yong Dong, Xuejun Yang, Panfeng Wang
HPCC1
2008 Energy-Oriented OpenMP Parallel Loop Scheduling
abstract
In HPC, power-related concern becomes dominant aspects of hardware and software design. Significant research effort has been devoted towards the energy optimization of parallel loop. This article is focused on energy-oriented OpenMP static and dynamic parallel loop scheduling problem. Only DVS cannot obtain the maximum energy savings. It is necessary to combine parallel loop rescheduling and DVS. First, we propose an energy-saving static scheduling (ESSS) algorithm, which exploits the scheduling slack to save energy by DVS. Second, we propose an energy-saving optimal static scheduling (EOSS) algorithm, which obtains the maximum energy saving through combining loop rescheduling and DVS. Last, in order to reduce the energy of OpenMP dynamic scheduling, we shut down the processor when it is idle, which is called shut-down based dynamic scheduling (SBDS) algorithm. Finally, we demonstrate the effectiveness by experiments.
Yong Dong, Juan Chen 0001, Xuejun Yang, Xuemeng Zhang
ISPA2
2006 Compiler-Directed Energy-Time Tradeoff in MPI Programs on DVS-Enabled Parallel Systems
Huizhan Yi, Juan Chen 0001, Xue-Jun Yang
ISPA2
2005 Energy-Constrained Prefetching Optimization in Embedded Applications
Juan Chen 0001, Yong Dong, Huizhan Yi, Xuejun Yang
EUC1
2005 A Compiler-Directed Energy Saving Strategy for Parallelizing Applications in On-Chip Multiprocessors
abstract
As energy consumption becoming one of the key optimization objects in on-chip multiprocessor, compiling a parallelizing application combined with energy saving strategy is more significant. In this paper, we focus on an on-chip multiprocessors architecture, where each processor in on-chip multiprocessor can independently adjust its frequency and voltage for energy savings. Given an arrayintensive application, we simulate parallelizing application and analyze probable load imbalance; then our energy saving strategy determines each processor’s clock frequency and voltage level fit for each parallel fragment in terms of load imbalance. Here, parallel fragments mainly denote parallel loop nests. Further, we consider the serial code fragments as a severe load-unbalanced parallel partitioning when the redundant processors can be shut down. Initial experiment proves our energy saving strategy is successful in reducing the energy consumption of the parallel programs.
Juan Chen 0001, Yong Dong, Xuejun Yang
ISPDC1