EDBT 2026 Demo / reviewers in the wild / expert
Qiucheng Liao
dblp:259/7227 · also Qiu-Cheng Liao
· DBLP profile ↗
5ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TacVar: Tackling Variability in Short-Interval Timing Measurements on X86 ProcessorsabstractWhen timing short-interval loops, different timing methods may produce different results caused by timing fluctuations. These unstable results prevent us from drawing reliable conclusions in processor performance benchmark. To tackle this issue, we proposed the TacVar framework including three components: 1) we developed the TVkern benchmark to highlight misleading tail-end timing results caused by timing fluctuations; 2) we built the TVconv convolution model to describe the mechanism of coupling between timing fluctuations and performance variability; 3) we designed the TVfilt algorithm to filter out timing fluctuations from measurement results using in-situ timing fluctuation samples. We evaluated TacVar with three TVkern and two real-world stencil kernels on two mainstream X86 processors. The results showed TacVar can reduce measurement deviations caused by timing fluctuations by up to 99.0%. Qiucheng Liao, James Lin 0001 |
CCGrid | 1 |
| 2024 | FCUFS: Core-Level Frequency Tuning for Energy Optimization on Intel ProcessorsabstractSacrificing minor performance for better energy efficiency is effective in reducing the energy consumption of supercomputers. Recent studies have utilized some frequency tuning and power capping features of Intel processors to decrease energy consumption in supercomputers. However, two main issues persist: 1) current methods do not account for variability between cores, leading to insufficient energy savings in mixed workloads; and 2) they fail to control the extent of performance loss following frequency tuning. To address these issues, we developed the FCUFS framework, which includes two components: 1) a neural network for predicting performance and power; and 2) a strategy optimization algorithm for selecting frequencies with controllable performance loss. We evaluated FCUFS on Intel mainstream processors across 15 dedicated workloads and 5 mixed workloads. On a single dual-socket Ice Lake-SP server, at a 5% performance loss target, the average energy savings were 11.4% for dedicated workloads and 13.8% for mixed workloads, with average performance losses of 2.9% and 3.5%, respectively. At a 10% performance loss target, the energy savings increased to 14.2% for dedicated workloads and 14.4% for mixed workloads, with performance losses of 8.7% and 9.4%, respectively. When scaling up to 2,048 cores, the average energy savings were 9.7% with 4.4% performance loss. The results show that FCUFS achieves consistent energy savings across dedicated and mixed workloads while maintaining controllable performance loss. Hongjian Zhang, Akira Nukada, Qiucheng Liao |
CLUSTER | 3 |
| 2023 | Characterizing Performance Impacts of Subnormal Numbers on Vector Instructions and Transcendental FunctionsabstractWe investigated the performance impact of IEEE-754 double-precision floating-point subnormal numbers, focusing on vector arithmetic and transcendental functions across Intel, AMD, and HiSilicon CPUs. We developed a benchmark tool, SNBench, and uncovered that subnormal numbers can extend instruction latency and reciprocal throughput by up to 162 cycles on Intel, 13 cycles on AMD, and 2 cycles on HiSilicon processors. The impact is more pronounced in transcendental functions, with performance losses reaching up to 562 cycles on Intel, 404 cycles on AMD, and 62 cycles on HiSilicon CPUs. James Lin 0001, Qiucheng Liao |
ICPADS | 3 |
| 2020 | CUBE - Towards an Optimal Scaling of Cosmological N-body SimulationsabstractN-body simulations are essential tools in physical cosmology to understand the large-scale structure (LSS) formation of the universe. Large-scale simulations with high resolution are important for exploring the substructure of universe and for determining fundamental physical parameters like neutrino mass. However, traditional particle-mesh (PM) based algorithms use considerable amounts of memory, which limits the scalability of simulations. Therefore, we designed a two-level PM algorithm CUBE towards optimal performance in memory consumption reduction. By using the fixed-point compression technique, CUBE reduces the memory consumption per N-body particle to only 6 bytes, an order of magnitude lower than the traditional PM-based algorithms. We scaled CUBE to 512 nodes (20,480 cores) on an Intel Cascade Lake based supercomputer with ≃95% weak-scaling efficiency. This scaling test was performed in Cosmo-π - a cosmological LSS simulation using ≃4.4 trillion particles, tracing the evolution of the universe over ≃13.7 billion years. To our best knowledge, Cosmo-π is the largest completed cosmological N-body simulation. We believe CUBE has a huge potential to scale on exascale supercomputers for larger simulations. Shenggan Cheng, Hao-Ran Yu, Derek Inman, Qiucheng Liao, Qiaoya Wu, James Lin 0001 |
CCGRID | 4 |
| 2019 | An Empirical Study of HPC Workloads on Huawei Kunpeng 916 ProcessorabstractThe ARM-based server processors have been gaining momentum in high performance computing (HPC). While not designed specifically for HPC, Huawei Kunpeng 916 processor has 32 ARMv8 cores and is tempting for HPC workloads. However, its potential remains unknown. To throughly understand the potential, we conducted a systematic evaluation in three steps by using: 1) three well-known benchmarks (HPL, STREAM, and LMbench); 2) three typical scientific kernels (SpMV, N-body, and GEMM); 3) three widely used mini-apps (TeaLeaf, Neutral, and SNAP) and a real-world application GTC-P. We compared the performance results of Kunpeng 916 with that of Intel Xeon E5-2680v3/4 (Haswell/Broadwell). The evaluation results show that Kunpeng 916 has higher memory bandwidth than the two Intel processors, thus it can achieve compelling performance for running memory bound HPC applications. Yichao Wang 0001, Jin-Kun Chen, Bin-Rui Li, Sicheng Zuo, William Tang 0002, Bei Wang 0002, Qiucheng Liao, James Lin 0001 |
ICPADS | 7 |