VLDB 2026 Research / reviewers in the wild / expert
Tao Liu 0029
dblp:43/656-29
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-9653-4108ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Heterogeneous Parallel Optimization Research of Plasma Guiding Center Orbit Simulation
Xianqian Meng, Tao Liu 0029, Baofeng Gao, Ying Guo 0028, Jingshan Pan |
ICA3PP (1) | 2 |
| 2025 | Parallel Optimization of Tokamak Guiding-Center Drift-Orbit Integration on the Sunway Bluelight II Supercomputer
Tao Liu 0029, Baofeng Gao, Ying Guo 0028, Jingshan Pan |
ICA3PP (3) | 2 |
| 2025 | Research on Dynamic Properties Evolution Algorithm of Polymer Nanomaterials Based on Heterogeneous Computing Platform
Xiaoman Zhang, Tao Liu 0029, Han Qin, Ying Guo 0028, Jingshan Pan |
ICA3PP (6) | 2 |
| 2025 | Leveraging Graph Analysis to Pinpoint Root Causes of Scalability Issues for Parallel ApplicationsabstractIt is challenging to scale parallel applications to modern supercomputers because of load imbalance, resource contention, and communications between processes. Profiling and tracing are two main performance analysis approaches for detecting these scalability bottlenecks. Profiling is low-cost but lacks detailed dependence for identifying root causes. Tracing records plentiful information but incurs significant overheads. To address these issues, we presentScalAna, which employs static analysis techniques to combine the benefits of profiling and tracing - it enables tracing's analyzability with overhead similar to profiling.ScalAnauses static analysis to capture program structures and data dependence of parallel applications, and leverages lightweight profiling approaches to record performance data during runtime. Then a parallel performance graph is generated with both static and dynamic data. Based on this graph, we design a backtracking detection approach to automatically pinpoint the root causes of scaling issues. We evaluate the efficacy and efficiency ofScalAnausing several real applications with up to 704K lines of code and demonstrate that our approach can effectively pinpoint the root causes of scaling loss with an average overhead of 5.65% for up to 16,384 processes. By fixing the root causes detected by our tool, it achieves up to 33.01% performance improvement. Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Zhenhua Guo 0003, Yaqian Zhao, Torsten Hoefler, Tao Liu 0029, Xu Liu 0001, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | SW-TRRM: Parallel Optimization Research of the Random Ray Method Based on Sunway Bluelight II Supercomputer
Zenghui Ren, Tao Liu 0029, Zhaoyuan Liu, Ying Guo 0028, Jingshan Pan, Meihong Yang |
ICA3PP (5) | 2 |
| 2023 | SW-LeNet: Implementation and Optimization of LeNet-1 Algorithm on Sunway Bluelight II Supercomputer
Zenghui Ren, Tao Liu 0029, Zhaoyuan Liu, Min Tian 0005, Ying Guo 0028, Jingshan Pan |
ICA3PP (5) | 2 |
| 2023 | Parallel optimization of method of characteristics based on Sunway Bluelight II supercomputer
Renjiang Chen, Tao Liu 0029, Zhaoyuan Liu, Min Tian 0005, Ying Guo 0028, Jingshan Pan, Meihong Yang |
J. Supercomput. | 2 |
| 2023 | swParaFEM: a highly efficient parallel finite element solver on Sunway many-core architecture
Jingshan Pan, Lei Xiao 0002, Min Tian 0005, Tao Liu 0029, Yinglong Wang 0001 |
J. Supercomput. | 4 |
| 2022 | swSuperLU: A highly scalable sparse direct solver on Sunway manycore architecture
Min Tian 0005, Zanjun Zhang, Jingshan Pan, Tao Liu 0029 |
J. Supercomput. | 6 |
| 2022 | iBalancer: Load-Aware in-Server Flow Scheduling for Sub-Millisecond Tail LatencyabstractAchieving microsecond-scale tail latency poses an extreme challenge to the conventional architecture of “NIC-OS-Application” in the face of high concurrent requests. Existing kernel-bypass network systems improve this situation significantly. Still, they cannot achieve load-aware in-server requests distribution, which in turn not only harms resource efficiency but, more importantly, beats the goal of squeezing tail latency. This paper proposes iBalancer, an in-server proactive load balancer for the kernel-bypass system, which aggressively handles NIC-side flow scheduling according to the load of threads on the processor-side. Furthermore, we propose a novel metric, “polling time interval (PTI),” to quantify the load of worker threads, which not only indicates utilization of the core bound to the worker thread but also reflects the differences in the processing time of different flows. By scheduling flows according to the metric PTI, iBalancer tends to average the queueing latencies of different flows, such as Set & Get operations for an in-memory key-value store. In addition, by decoupling flow scheduling from packet steering, iBalancer achieves a tail latency aware flow-to-core binding and preserves hardware-based request distribution among cores. The proposed system is evaluated and compared to mTCP and Shenango using two representative microsecond-scale network applications: Memcached KVS and a real-time deep-learning-based financial fraud identification application. Experimental results show that iBalancer can process up to 4.75$ \times $and 1.55$ \times \ $higher load over mTCP and Shenango under 500μs 99thpercentile tail latency limit on Memcached. For the financial fraud identification application, iBalancer is able to process 4.56$ \times $and 1.16$ \times $higher load than mTCP and Shenango considering 900μs tail latency. Qi Zhang 0108, Yi Liu 0013, Tao Liu 0029 |
IEEE Trans. Parallel Distributed Syst. | 3 |