EDBT 2026 Demo / reviewers in the wild / expert
Qinyun Tsai
dblp:304/6066 · also Qinyun Cai
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-7642-457XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DIDS: A distributed inference framework with dynamic scheduling capability
Yuwei Yan, Yikun Hu 0001, Qinyun Tsai, Wangdong Yang, Kenli Li 0001 |
Future Gener. Comput. Syst. | 3 |
| 2024 | COALA: A Compiler-Assisted Adaptive Library Routines Allocation Framework for Heterogeneous SystemsabstractExperienced developers often leverage well-tuned libraries and allocate their routines for computing tasks to enhance performance when building modern scientific and engineering applications. However, such well-tuned libraries are meticulously customized for specific target architectures or environments. Additionally, the performance of their routines is significantly impacted by the actual input data of computing tasks, which often remains uncertain until runtime. Accordingly, statically allocating these library routines may hinder the adaptability of applications and compromise performance, particularly in the context of heterogeneous systems. To address this issue, we propose the Compiler-Assisted Adaptive Library Routines Allocation (COALA) framework for heterogeneous systems. COALA is a fully automated mechanism that employs compiler assistance for dynamic allocation of the most suitable routine to each computing task on heterogeneous systems. It allows the deployment of varying allocation policies tailored to specific optimization targets. During the application compilation process, COALA reconstructs computing tasks and inserts a probe for each of these tasks. Probes serve the purpose of conveying vital information about the requirements of each task, including its computing objective, data size, and computing flops, to a user-level allocation component at runtime. Subsequently, the allocation component utilizes the probe information along with the allocation policy to assign the most optimal library routine for executing the computing tasks. In our prototype, we further introduce and deploy a performance-oriented allocation policy founded on a machine learning-based performance evaluation method for library routines. Experimental verification and evaluation on two heterogeneous systems reveal that COALA can significantly improve application performance, with gains of up to 4.3x for numerical simulation software and 4.2x for machine learning applications, and enhance system utilization by up to 27.8%. Qinyun Tsai, Guanghua Tan, Wangdong Yang, Xianhao He, Yuwei Yan, Keqin Li 0001, Kenli Li 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Parallel algorithm design and optimization of geodynamic numerical simulation application on the Tianhe new-generation high-performance computer
Wangdong Yang, Ruixuan Qi, Qinyun Tsai, Shengle Lin, Fengkun Dong, Kenli Li 0001, Keqin Li 0001 |
J. Supercomput. | 4 |
| 2021 | STM-multifrontal QR: streaming task mapping multifrontal QR factorization empowered by GCNabstractMultifrontal QR algorithm, which consists of symbolic analysis and numerical factorization, is a high-performance algorithm for orthogonal factorizing sparse matrix. In this work, a graph convolutional network (GCN) for adaptively selecting the optimal reordering algorithm is proposed in symbolic analysis. Using our GCN adaptive classifier, the average numerical factorization time is reduced by 20.78% compared with the default approach, and the additional memory overhead is approximately 4% higher than that of prior work. Moreover, for numerical factorization, an optimized tasks stream parallel processing strategy is proposed and a more efficient computing task mapping framework for NUMA architecture is adopted in this paper, which called STM-Multifrontal QR factorization. Numerical experiments on the TaiShan Server show average 1.22x performance gains over the original SuiteSparseQR. Nearly 80% of datasets have achieved better performance compared with the MKL sparse QR on Intel Xeon 6248. Shengle Lin, Wangdong Yang, Haotian Wang 0006, Qinyun Tsai, Kenli Li 0001 |
SC | 4 |