Tiaojie Xiao

dblp:275/3961 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-8378-5530ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating High-Frequency Electromagnetic Scattering Prediction with KNN-Augmented Radial Basis Function Networks
Chao Li 0002, Xinhai Chen 0001, Tiaojie Xiao, Jie Liu 0002
IPDPS5
2026 Flux-conserved physics-informed neural networks for electromagnetic scattering computation
Chenyu Peng, Tiaojie Xiao, Sifan Wang, Xinhai Chen 0001, Chunye Gong
Eng. Appl. Artif. Intell.2
2026 A Memory-Aware Sparse Matrix-Matrix Multiplication on Multicore Architectures
abstract
Sparse matrix–matrix multiplication (SpMM) is a fundamental operation in scientific computing with broad applications across numerous domains. Tiling is a key optimization technique for improving data locality and is widely adopted in high-performance computing. However, the irregular data access patterns inherent to SpMM make it challenging to exploit tiling effectively for data reuse. In this article, we propose MaSpMM , a memory-aware SpMM framework that integrates cache-aware tiling with a segment-oriented data layout. MaSpMM stores matrices as continuous segments to enhance data locality within each tile. Moreover, since many sparse matrices in real-world applications exhibit symmetry, we further develop MaSpMM-Sym, an extension that recursively partitions symmetric matrices to eliminate write conflicts and further improve locality. To adapt to diverse scenarios, we finally introduce MaSpMM-Adap, which adaptively selects the most suitable approach for each input matrix. Comprehensive evaluations on both x86 and ARM CPUs demonstrate that MaSpMM-Adap achieves average speedups of up to 1.86× over Intel oneMKL, 1.84× over ASpT, and 1.75× over J-Stream.
Deshun Bi, Shengguo Li, Haozhong Qiu, Chuanfu Xu, Xiaojian Yang, Dezun Dong, Tiaojie Xiao, Jie Liu 0002
ACM Trans. Archit. Code Optim.8
2024 A Hybrid Vectorized Merge Sort on ARM NEON
Jincheng Zhou, Jin Zhang 0018, Xiang Zhang 0008, Tiaojie Xiao, Chunye Gong
ICA3PP (6)4
2024 Gauss-Newton With Preconditioned Conjugate Gradient Magnetotelluric Inversion for 3-D Axial Anisotropic Conductivities
abstract
We present a regularized inversion method for three-dimensional (3D) magnetotelluric (MT) data with axial anisotropic conductivities based on the edge-based finite element (FE) method. The Gauss–Newton (GN) approach is used to minimize the inversion objective function, including data misfit and regularization penalties, considering both structural complexity and anisotropic penalties. The most time-intensive task in the 3D MT inversion process is solving the large sparse system of linear equations. To speed up the inversion calculation, a hybrid direct–iterative solver combined with a block-diagonal preconditioner that has not yet been applied in anisotropic inversion is developed to accelerate the solutions for the sparse linear system resulting from forward modeling and sensitivity computations. In each GN iteration, a preconditioned conjugate gradient (PCG) method is adopted to overcome the difficulty of the sensitivity matrix storage for the anisotropic scene and obtain a model update without explicitly calculating and storing the sensitivity matrix. Before the inversion test, we use a model to demonstrate that the hybrid solver is computationally beneficial in terms of memory usage and time spent as compared to the direct solver. The good convergence properties and efficiency of the GN–PCG inversion scheme are demonstrated by two synthetic models and USArray data. The proposed inversion scheme can be an important supplement to existing anisotropic inversion algorithms and provide technical support for MT data interpretation.
Junjun Zhou, Ningbo Bai, Xiangyun Hu, Tiaojie Xiao, Guoshu Huang
IEEE Trans. Geosci. Remote. Sens.5
2023 MT-office: parallel password recovery program for office on domestic heterogeneous multi-core processor
Yongtao Luo, Bo Yang 0023, Jie Liu 0002, Ruibo Wang, Jinmin Wen, Tiaojie Xiao, Xuguang Chen, Chunye Gong
CCF Trans. High Perform. Comput.6
2022 STEGNN: Spatial-Temporal Embedding Graph Neural Networks for Road Network Forecasting
abstract
As intelligent transportation systems (ITS) are now being integrated into our everyday lives, it has been widely accepted that forecasting road networks is a promising killer engine for ITS with high social and economic benefits. However, current solutions ignore the heterogeneity of spatial-temporal traffic data and fail to capture hidden spatial-temporal correlations. This paper presents STEGNN: a novel spatial-temporal embedding graph neural network for road network forecasting. The key idea of STEGNN is utilizing Cosine Similarity to generate a high-quality temporal graph and thus fills the gap between the temporal-spatial correlations for traffic graph, which includes (i) a novel approach to construct temporal graph based on temporal-spatial similarity from traffic graphs, which is much more accurate on measured similarity of time series claimed by previous methods; (ii) an advanced spatial-temporal embedding model to exploit spatial-temporal dependencies by leveraging specific arrangements of temporal and spatial graphs; and (iii) an effective framework that gasps extensive spatial-temporal dependencies in the long-term by mixing multi-layer graph convolution with dilated convolution to understand wide-range spatial-temporal features. Extensive evaluations validate STEGNN by applying it to real-world traffic graphs and indicate that STEGNN outperforms state-of-the-art solutions with much more accurate forecasting of road networks.
Jiaqi Si, Xinbiao Gan, Tiaojie Xiao, Bo Yang 0023, Dezun Dong, Zhengbin Pang
ICPADS3
2022 TianheGraph: Customizing Graph Search for Graph500 on Tianhe Supercomputer
abstract
As the era of exascale supercomputing is coming, it is vital for next-generation supercomputers to find appropriate applications with high social and economic benefit. In recent years, it has been widely accepted that extremely-large graph computation is a promising killer application for supercomputing. Although Tianhe series supercomputers are leading in the world-wide competition of supercomputing (ranked No. 1 in the Top500 list for six times), previously they had been inefficient in graph computation according to the Graph500 list. This is mainly because the previous graph processing system cannot leverage the advanced hardware features of Tianhe supercomputers. To address the problem, in this paper we present our integrated optimizations for improving the graph computation performance on our next-generation Tianhe supercomputing system, mainly including sorting with buffering for heavy vertices, vectorized searching with SVE (Scalable Vector Extension) on matrix2000+ CPUs, and group communication on the proprietary interconnection network. Performance evaluation on a subset of the Tianhe supercomputer (with 512 nodes and 196,608 cores) shows that our customized graph processing system effectively improves the graph search performance and achieves the BFS performance of 2131.98 GTEPS.
Xinbiao Gan, Yiming Zhang 0003, Ruibo Wang, Tiaojie Xiao, Ruigeng Zeng, Jie Liu 0002, Kai Lu 0001
IEEE Trans. Parallel Distributed Syst.5