EDBT 2026 Demo / reviewers in the wild / expert
Bo Yang 0023
dblp:46/999-23
· DBLP profile ↗
21ranked-venue papers
0as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | THAC: Unlocking Performance in Parallel HPC Applications via UQ-Aware Automated ApproximationabstractWhile approximate computing offers a promising paradigm for performance gains, significant technical challenges obstruct its practical adoption in High Performance Computing (HPC). The process of manually identifying approximable regions, quantifying the risk of cascading errors, and navigating the vast combinatorial tuning space is intricate and error-prone. To overcome these obstacles, we present Tianhe Approximate Computing (THAC), a framework that establishes a systematic, automated methodology. THAC implements a principled, three-stage workflow: (1) a hybrid AI-driven synthesizer that automatically identifies potential approximation candidates; (2) a principled Uncertainty Quantification screen that rigorously validates and prunes high-risk options through a principled risk-aware screening stage; and (3) a novel hierarchical Bayesian optimizer that efficiently navigates the search space. Although demonstrated on OpenMP as a representative case study, the framework is designed with extensibility for generic parallel patterns. Evaluation on a diverse suite of scientific benchmarks (including computational fluid dynamics and hydrodynamics) demonstrates the effectiveness of THAC, achieving a geometric mean speedup of 2.1× under strict quality budgets. These results confirm THAC as a robust solution that makes approximate computing in HPC practical and reliable. Zhenhao Zhao, Bo Yang 0023, Xinhai Chen 0001, Jie Liu 0002, Binglin Wang |
ICS | 3 |
| 2026 | LDNO: A low-power dynamic neural operator inspired by liquid state machines for solving partial differential equations
Chengxue Huang, Jie Liu 0002, Qingyang Zhang 0009, Xinhai Chen 0001, Bo Yang 0023 |
Neurocomputing | 9 |
| 2026 | GNNRL-smoothing: A prior-free reinforcement learning model for mesh optimization
Xinhai Chen 0001, Chunye Gong, Bo Yang 0023, Liang Deng, Yufei Pang, Xiang Zhang 0008, Jie Liu 0002 |
Neural Networks | 4 |
| 2025 | Bandwidth Optimized Scalable Designs with Inter-Layer Overlapping for MPI BroadcastabstractMPI (Massage Passing Interface) has been the dominant programming model for developing large-scale parallel applications. Existing work mainly focuses on the vast parallelism of modern multi-/many-cores architectures to parallelize MPI collectives. However, the abundant bandwidth provided by modern interconnects is either underutilized when processing small messages or overwhelmed by large messages. MPI_Bcast is one of the most widely used collective primitives in MPI, which broadcasts data from one process to all processes in the communication domain. In this paper, we address the issue of load imbalance arising from traditional tree-based designs in order to strike a better balance between bandwidth and latency of MPI Broadcast. By evenly distributing the broadcast load across lower layers, we effectively leverage the available resources at upper-layer nodes that recursively execute broadcast across different layers in a sequential and contention-free manner. This approach improves the scalability and performance of large-scale message broadcasts. Additionally, a generic inter-layer overlapped scheme is proposed to reduce overall broadcast latency by fully overlapping inter-layer and intra-layer data transmission. We further implement an online adaptive scheme for tree degree tuning to achieve optimal design at various message and system sizes. Extensive experiments are conducted to evaluate the performance of this Bandwidth-optimized Inter-layer Overlapping (BIO) design at both the microbenchmark and application levels. BIO-based MPI_Bcast designs demonstrate performance speedups of up to 2.71x compared to state-of-the-art MPI libraries. For application-level evaluation, BIO provides up to 165% acceleration for the initialization of the distributed deep learning model of Horovod with the PyTorch application. Qiang Wang 0006, Bo Yang 0023, Dongsheng Li 0001 |
ICDCS | 5 |
| 2025 | IA-Chol: Input-Aware Cholesky Decomposition on CPU and GPU
Jixiao Deng, Lin Chen 0028, Tun Li 0002, Bo Yang 0023, Xinhai Chen 0001, Jie Liu 0002 |
ICS | 5 |
| 2025 | Informative Discrimination Network for Efficient Single Image Super-ResolutionabstractDeploying convolutional neural networks on low-resource mobile devices for single image super-resolution (SISR) faces the issue of how to balance the parameter amount and performance. The default solution is simultaneously condensing both hierarchical representation and attention features into their respective light proxies. The insight underlying this solution lies in the fact that features are redundant since the super-resolution needs plenty of similar pixels. This work takes it to the next step from the viewpoint of informativeness and discrimination. In detail, we propose an informative disrcimination network (IDNet) for SISR. For informativeness, a multi-scale residual block (MRB) is explored to capture informative spatial details via the scale-in-scale structure. It mines rich intra-layer spatial details based on inter-layer ones of the default hierarchical representation. However, it also incurs feature redundancy. Though attention serves to reduce this redundancy, feature discrimination and pixel-wise structural preservation cannot be guaranteed. Here spatial discrimination attention behaves like the biased discriminant classifier to induce spatial discrimination, while the nuclear-norm regularization recovers the image low-rank structure to reduce artifacts or noises. Importantly, no extra network weights are introduced for model efficiency. Experiments show that IDNet delivers sound performance with fewer parameters, as compared to its cousins. Yuzheng Tu, Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Bo Yang 0023, Xiang Gao 0020, Xiang Zhang 0008 |
IJCNN | 5 |
| 2025 | GraphCom: Communication Hierarchy-aware Graph Engine for Distributed Model TrainingabstractEfficient processing of large-scale graphs with billions to trillions of edges is essential for training graph-based large language models (LLMs) in web-scale systems. The increasing complexity and size of these models create significant communication challenges due to the extensive message exchanges required across distributed nodes. Current graph engines struggle to effectively scale across hundreds of computing nodes because they often overlook variations in communication costs within the interconnection hierarchy. This paper presents GraphCom, a communication-efficient message graph engine for graph processing on supercomputers. Our key idea is to leverage the network topology information to perform communication hierarchy-aware message aggregation, where messages are (i) gathered to the responsible nodes (referred to as monitors) in the source domains, (ii) transferred between monitors, and (iii) scattered to the target nodes in the target domains. GraphCom's aggregation is more aggressive in that each source domain (instead of the source node). We have implemented GraphCom on top of MPI. We demonstrate GraphCom's effectiveness with synthetic benchmarks and real-world graphs, utilizing up to 79,024 nodes and over 1.2 million processor cores, demonstrating that GraphCom surpasses leading graph- parallel systems and state-of-the-art counterparts in both throughput and scalability. Moreover, we have deployed GraphCom on a production supercomputer, where it consistently outperforms the top solutions on the Graph500 list. These results highlight the potential GraphCom has to significantly improve the efficiency of distributed large-scale graph-based LLM training by optimizing communication between distributed systems, making it an invaluable graph engine for distributed training tasks on web-scale graphs. Xinbiao Gan, Qiang Zhang 0053, Lingyun Song, Bo Yang 0023, Jie Liu 0002, Kai Lu 0001 |
WWW | 6 |
| 2025 | GraphCSR: A Space and Time-Efficient Sparse Matrix Representation for Web-scale Graph ProcessingabstractGraph data processing is essential for web-scale applications, including social networks, recommendation systems, and web of things (WoT) systems, where large, sparsely connected graphs dominate. Traditional sparse matrix storage formats like compressed sparse row (CSR) face significant memory and performance bottlenecks in distributed, federated, and edge-based computing environments, which are increasingly central to the web. To address this challenge, we propose GraphCSR, a novel storage format that clusters vertices with identical edge degrees and stores only the starting index of each group. This approach minimizes memory overhead and facilitates batch memory access while enhancing overall performance, making it particularly suitable for federated systems and resource-constrained edge nodes. Our experiments across various graph operations and large datasets show that GraphCSR achieves considerable memory savings and performance gains of large-scale, distributed graph processing. When deployed GraphCSR on two production-grade supercomputers, demonstrating its potential for scaling web and WoT graph processing in large-scale distributed computing systems. Xinbiao Gan, Qiang Zhang 0053, Bo Yang 0023, Chunye Gong, Jie Liu 0002, Kai Lu 0001 |
WWW | 5 |
| 2025 | FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007 |
Sci. China Inf. Sci. | 7 |
| 2024 | TianheStar: Orchestrating SSSP Applications on Tianhe SupercomputerabstractComputing single-source shortest paths (SSSP) is one of the fundamental problems in graph theory and is also essential for data-intensive applications. As the potential of artificial intelligence (AI) continues to be explored, and with the advent of exascale supercomputing, there is a growing need for an extremely fast graph engine for SSSP applications. Current distributed SSSP engines for large-scale graph applications, unfortunately, often exhibit poor efficiency when running on supercomputers. In this paper, we introduce TianheStar, an ultra-fast SSSP engine designed specifically for graph search on the Tianhe supercomputer. TianheStar effectively minimizes communication costs and establishes a new balance between computation and communication. The key idea of TianheStar is to leverage network topology information for performing topology-aware message aggregation and architecture-aware group communication. These two techniques effectively reduce the number of messages and the average number of communication hops, respectively. We validate TianheStar using Graph500, a widely adopted benchmark for graph search on supercomputers. Extensive evaluation demonstrates that, compared to the state-of-the-art solutions, TianheStar achieves a remarkable performance improvement. We have deployed TianheStar on the latest Tianhe supercomputer and secured the top position in the latest Graph500. We achieved an outstanding performance of 23,021 GTEPS (Giga Traversed Edges Per Second) for SSSP using 4096 nodes. Furthermore, we have delved into real-world graphs representing the USA road networks and conducted computations to determine the shortest paths between vertices. Our experimental results demonstrate that TianheStar can traverse the USA road network, comprising over 58,333,344 edges, in less than 0.1 second on the Tianhe supercomputer. This performance represents a speedup of over a thousand times compared to parallel shortest-path graph computations on the Aziz supercomputer, a globally renowned high-performance computing system, using the same input data. Xinbiao Gan, Shijie Li 0002, Bo Yang 0023 |
CCGrid | 5 |
| 2024 | SuperCSR: A Space-Time-Efficient CSR Representation for Large-scale Graph Applications on SupercomputersabstractIt is widely accepted that graph representations such as the Compressed Sparse Row (CSR) format, directly affect the space and time complexities of graph processing. However, the standard CSR and its current variations are prone to high memory footprint and complicated calculations, which necessitates the development of more efficient graph processing techniques to save memory and reduce calculations. This paper presents SuperCSR, a more space-time-efficient CSR representation for fast graph processing. SuperCSR’s key idea is to leverage the law of sorted graphs, which would directly access adjacent vertex sets from an active vertex ID without complex indexing calculations and with a lower memory footprint. Xinbiao Gan, Qiang Zhang 0053, Bo Yang 0023, Xinhai Chen 0001, Jie Liu 0002 |
ICPP | 4 |
| 2024 | MST: Topology-Aware Message Aggregation for Exascale Graph Processing of Traversal-Centric AlgorithmsabstractThis article presents MST, a communication-efficient message library for fast graph traversal on exascale clusters. The key idea is to follow the multi-level network topology to perform topology-aware message aggregation, where small messages are gathered and scattered at each level of domain. To facilitate message aggregation, we equip MST with flexible buffer management including active buffer switching and dynamic buffer expansion. We implement MST on the newest-generation Tianhe supercomputer and evaluated its performance using various traversal-centric algorithms on both synthetic trillion-scale graphs and real-world big graphs. The results show that MST-based graph traversal is orders of magnitude faster than that based on Active Messages Library (AML). For the Graph500-BFS benchmark, MST-based Tianhe (with 77.2 K nodes) outperforms the Fugaku supercomputer (with 148.5 K nodes) by 18.53%, while Fugaku is ranked No. 1 in the latest Graph500-BFS ranking (June 2023). MST also greatly improves graph processing performance on other commercial large-scale computing systems at the National Supercomputing Center in Changsha (NSCC) and WuzhenLight. Xinbiao Gan, Bo Yang 0023, Xinhai Chen 0001, Chunye Gong, Shijie Li 0002, Kai Lu 0001, Qiao Li 0001, Yiming Zhang 0003 |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | Optimizing MPI Collectives on Shared Memory Multi-CoresabstractMessage Passing Interface (MPI) programs often experience performance slowdowns due to collective communication operations, like broadcasting and reductions. As modern CPUs integrate more processor cores, running multiple MPI processes on shared-memory machines to take advantage of hardware parallelism is becoming increasingly common. In this context, it is crucial to optimize MPI collective communications for shared-memory execution. However, existing MPI collective implementations on shared-memory systems have two primary drawbacks. The first is extensive redundant data movements when performing reduction collectives, and the second is the ineffective use of non-temporal instructions to optimize streamed data processing. To address these limitations, this paper proposes two optimization techniques that minimize data movements and enhance the use of non-temporal instructions. We evaluated our techniques by integrating them into the OpenMPI library and tested their performance using micro-benchmarks and real-world applications running on two multi-core clusters. Experimental results show that our approach significantly outperforms existing techniques, yielding a 1.2--6.4x performance improvement. Jintao Peng, Jianbin Fang, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Zheng Wang 0001 |
SC | 6 |
| 2023 | MT-office: parallel password recovery program for office on domestic heterogeneous multi-core processor
Yongtao Luo, Bo Yang 0023, Jie Liu 0002, Ruibo Wang, Jinmin Wen, Tiaojie Xiao, Xuguang Chen, Chunye Gong |
CCF Trans. High Perform. Comput. | 2 |
| 2022 | STEGNN: Spatial-Temporal Embedding Graph Neural Networks for Road Network ForecastingabstractAs intelligent transportation systems (ITS) are now being integrated into our everyday lives, it has been widely accepted that forecasting road networks is a promising killer engine for ITS with high social and economic benefits. However, current solutions ignore the heterogeneity of spatial-temporal traffic data and fail to capture hidden spatial-temporal correlations. This paper presents STEGNN: a novel spatial-temporal embedding graph neural network for road network forecasting. The key idea of STEGNN is utilizing Cosine Similarity to generate a high-quality temporal graph and thus fills the gap between the temporal-spatial correlations for traffic graph, which includes (i) a novel approach to construct temporal graph based on temporal-spatial similarity from traffic graphs, which is much more accurate on measured similarity of time series claimed by previous methods; (ii) an advanced spatial-temporal embedding model to exploit spatial-temporal dependencies by leveraging specific arrangements of temporal and spatial graphs; and (iii) an effective framework that gasps extensive spatial-temporal dependencies in the long-term by mixing multi-layer graph convolution with dilated convolution to understand wide-range spatial-temporal features. Extensive evaluations validate STEGNN by applying it to real-world traffic graphs and indicate that STEGNN outperforms state-of-the-art solutions with much more accurate forecasting of road networks. Jiaqi Si, Xinbiao Gan, Tiaojie Xiao, Bo Yang 0023, Dezun Dong, Zhengbin Pang |
ICPADS | 4 |
| 2022 | Improving the Performance of Lattice Boltzmann Method with Pipelined Algorithm on A Heterogeneous Multi-zone Processor
Qingyang Zhang 0009, Lei Xu 0035, Rongliang Chen, Lin Chen 0028, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023 |
PDCAT | 8 |
| 2021 | An efficient image to column algorithm for convolutional neural networksabstractConvolutional Neural Networks (CNNs) are a class of deep neural networks. The image to column (im2col) procedure is an important step for CNN and consumes about 28.8% of the whole inference time. In this paper, we present an efficient im2col algorithm, name im2cole (word “e” means efficient). The condition with different stride and pad in im2cole is well handled and the judgements in the innermost loop are removed. The procedure with pad = 1 is split into three conditions. This will reduce the pause of CPU instruction pipeline. The performances of the presented im2cole algorithm are reported with different inputs. Some discussion and performance issues are also reported. The experimental results show that the overall performance speedup of im2cole ranges from 2.12 to 4.33 compared with the original algorithm. The real application with Darknet shows that im2cole can get 20.75% whole performance improvement. Chunye Gong, Xinhai Chen 0001, Shuling Lv, Jie Liu 0002, Bo Yang 0023, Weimin Bao, Yufei Pang |
IJCNN | 5 |
| 2021 | DeepAtomicCharge: a new graph convolutional network-based architecture for accurate prediction of atomic chargesabstractAtomic charges play a very important role in drug-target recognition. However, computation of atomic charges with high-level quantum mechanics (QM) calculations is very time-consuming. A number of machine learning (ML)-based atomic charge prediction methods have been proposed to speed up the calculation of high-accuracy atomic charges in recent years. However, most of them used a set of predefined molecular properties, such as molecular fingerprints, for model construction, which is knowledge-dependent and may lead to biased predictions due to the representation preference of different molecular properties used for training. To solve the problem, we present a new architecture based on graph convolutional network (GCN) and develop a high-accuracy atomic charge prediction model named DeepAtomicCharge. The new GCN architecture is designed with only the atomic properties and the connection information between the atoms in molecules and can dynamically learn and convert molecules into appropriate atomic features without any prior knowledge of the molecules. Using the designed GCN architecture, substantial improvement is achieved for the prediction accuracy of atomic charges. The average root-mean-square error (RMSE) of DeepAtomicCharge is 0.0121 e, which is obviously more accurate than that (0.0180 e) reported by the previous benchmark study on the same two external test sets. Moreover, the new GCN architecture needs much lower storage space compared with other methods, and the predicted DDEC atomic charges can be efficiently used in large-scale structure-based drug design, thus opening a new avenue for high-performance atomic charge prediction and application. Jike Wang, Dong-Sheng Cao 0001, Cunchen Tang, Lei Xu 0035, Qiaojun He, Bo Yang 0023, Huiyong Sun, Tingjun Hou |
Briefings Bioinform. | 6 |
| 2020 | OHTMA: an optimized heuristic topology-aware mapping algorithm on the Tianhe-3 exascale supercomputer prototypeabstractWith the rapid increase of the size of applications and the complexity of the supercomputer architecture, topology-aware process mapping becomes increasingly important. High communication cost has become a dominant constraint of the performance of applications running on the supercomputer. To avoid a bad mapping strategy which can lead to terrible communication performance, we propose an optimized heuristic topology-aware mapping algorithm (OHTMA). The algorithm attempts to minimize the hop-byte metric that we use to measure the mapping results. OHTMA incorporates a new greedy heuristic method and pair-exchange-based optimization. It reduces the number of long-distance communications and effectively enhances the locality of the communication. Experimental results on the Tianhe-3 exascale supercomputer prototype indicate that OHTMA can significantly reduce the communication costs. Yishui Li, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023, Chunye Gong, Xinbiao Gan, Shengguo Li, Han Xu 0008 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2020 | VBSF: a new storage format for SIMD sparse matrix-vector multiplication on modern processors
Yishui Li, Peizhen Xie, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Chunye Gong, Xinbiao Gan, Han Xu 0008 |
J. Supercomput. | 5 |
| 2014 | An efficient parallel solution for Caputo fractional reaction-diffusion equation
Chunye Gong, Weimin Bao, Guojian Tang, Bo Yang 0023, Jie Liu 0002 |
J. Supercomput. | 4 |