Jianguo Liang

dblp:281/2331 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Torch-feat: GNN sampling training data loader based on feature data extraction operators
Jianzhi Yu, Zhencheng Liu, Guangjie Jin, Jianguo Liang
CCF Trans. High Perform. Comput.6
2026 Bridging the image-text gap: Reinforced Cross-modal Abnormality Driven Transformer for automatic chest X-ray report generation
Xiu-Long Yi, You Fu, Enxu Bi, Hao Zhang 0058, Jianguo Liang, Rong Hua
Eng. Appl. Artif. Intell.6
2026 Performance enhancement of CICE dynamics via data reconstruction and heterogeneous parallelization
Jianzhi Yu, Yan Yao 0001, Junyong Cao, Jianguo Liang
Future Gener. Comput. Syst.6
2026 Memory access optimization for the dynamics EVP model of the sea ice model on the SW39000 on-chip heterogeneous many-core processor
abstract
To improve the performance of the sea ice model in the Community Earth System Model (CESM) under a heterogeneous computing environment, this work conducts an in-depth study on the memory access optimization of the Elastic-Viscous-Plastic (EVP) dynamics model in Community Ice Code(CICE) on the SW39000 heterogeneous many-core processor, which is deployed in the new-generation Sunway supercomputer. The processor’s complex on-chip heterogeneous architecture, multi-level memory hierarchy, and unique inter-core communication mechanism present significant challenges for the parallel optimization of the sea ice dynamics simulation. To address the inefficiencies caused by diverse data access patterns, a differentiated processing strategy based on data read/write characteristics is proposed to reduce unnecessary data transfers. In addition, to alleviate load imbalance arising from the sparsity of sea ice boundary update data, a local dynamic compression method incorporating the probability density of data sparsity is designed. This method dynamically compresses data according to the probability density of Direct Memory Access (DMA) data transfers, thereby reducing communication volume and balancing the workload across slave cores. Finally, to enhance the computational intensity of the slave cores and reduce data dependencies between master and slave cores, an operator fusion algorithm based on Remote Memory Access (RMA) communication is introduced to achieve efficient data caching and transmission between operators. Experimental results demonstrate that, under the standard gx3 grid configuration, the optimized EVP model achieves a 27.54×speedup over the serial version running on a single master core when executed with a single core group. Multi-core group parallel tests validate the excellent scalability of the proposed optimization strategies, achieving up to a 123.93×speedup with a 10-core group, while also exhibiting effective load balancing in terms of both clock cycles and instruction counts across the slave-core array.
Jianzhi Yu, Jianguo Liang, You Fu, Ke-Kun Hu
Future Gener. Comput. Syst.3
2026 SDGraph: A scalable training system for GNNs with GPU sampling and parallel feature access
Jianzhi Yu, You Fu, Ke-Kun Hu, Jianguo Liang
Future Gener. Comput. Syst.6
2026 Radiology report generation via visual-semantic ambivalence-aware network and focal self-critical sequence training
Xiu-Long Yi, You Fu, Enxu Bi, Jianguo Liang, Hao Zhang 0058, Jianzhi Yu, Rong Hua
Neural Networks4
2026 RANS-KAN-DRIA: KAN-based relation-aware meta-learning with diffusion regularization for few-shot knowledge graph completion
Zhaoan Dong, Jiachen Gong, Jianguo Liang
J. Web Semant.4
2025 Parallel software design of large-scale diamond-structured crystals molecular dynamics simulation
Jianguo Liang, You Fu
Future Gener. Comput. Syst.1
2023 ESA: An efficient sequence alignment algorithm for biological database search on Sunway TaihuLight
Hao Zhang 0058, Zhiyi Huang 0001, Yawen Chen 0001, Jianguo Liang, Xiran Gao
Parallel Comput.4
2022 OpenACC + Athread collaborative optimization of Silicon-Crystal application on Sunway TaihuLight
Jianguo Liang, Rong Hua, Yuxi Ye, You Fu, Hao Zhang 0058
Parallel Comput.1
2022 A novel acceleration method for molecular dynamics of crystal silicon on GPUs using OpenACC
abstract
Abstract Compared with CUDA and OpenCL, OpenACC has the advantages of simple programming, openness, and good portability for GPU acceleration. An OpenMP/OpenACC implementation for molecular dynamics of silicon crystal on GPUs is proposed. First, to make effective use of vectorization and streaming, data structure conversion and data dependence elimination are designed. Second, the parallel version on the single GPU is realized by adding OpenACC guidance sentences, with very few modifications. Third, a patch block strategy is proposed to realize the parallel version on single machine multi‐GPUs using OpenMP+OpenACC, which greatly simplifies the construction of shadow area and the exchange of shadow area data. Experimental results show that 23 to 25 speedup is achieved for the single GPU at different scales over the serial program on Intel(R) Xeon(R) CPU E5‐2690 v4, and 6.37 speedup is achieved over the single GPU when the number of atoms reaches 2,097,152 on 8GPUs on single machine.
Jianguo Liang, You Fu, Rong Hua, Hao Zhang 0058, Yuxi Ye
Softw. Pract. Exp.1
2020 Accelerated molecular dynamics simulation of Silicon Crystals on TaihuLight using OpenACC
Jianguo Liang, Rong Hua, Hao Zhang 0058, You Fu
Parallel Comput.1