Botao Wu

dblp:143/0219 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0001-1481-6399ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Parametric Mappings for Distributed-Memory Tensor Computations
abstract
Tensor computations are an important class of operations widely used in domains such as computational chemistry, machine learning, and various types of physical simulations that demand distributed-memory clusters. Recent work has shown that generating efficient mappings for multi-operator Directed Acyclic Graphs of distributed-memory tensor computations is possible by leveraging non-linear formulations underpinned by Satisfiability Modulo Theories (SMT) solvers. However, this approach is sensitive to the problem size, grid shape, and count of Processing Elements (PEs) given.
Botao Wu, Martin Kong
ICS1
2025 Generating Two-Level, GPU-Aware Mappings for Distributed Tensor Computations
abstract
We introduce a two-level scheme to generate GPUaware MPI/NCCL code for distributed tensor computations. Our generator takes the specification of a linearized Directed Acyclic Graph (DAG) of tensor operators and produces a global mapping solution that considers MPI communication (inter and intranode) and the local computation. The core of our generator is a new bit-vector representation that compactly models mappings as well as communication directions along the grid. We incorporate the 2-level mapping decisions into a non-linear formulation which is optimized in an iterative fashion with the Z3 SMT solver. The new mapper supports both NVIDIA NCCL, MVAPICH-gdr, allowing for better portability. We demonstrate the efficiency of our mapping generator on a set of matrix- and tensor- DAGs, on two multi-GPU clusters with NVLink or PCIe intra-node interconnect, and compare against the COSMA library and CTF framework, achieving speedups ranging from 2.6× (over COSMA) to 18× (over CTF).
Botao Wu, Martin Kong
PACT1
2025 TSIformer: Multi-Scale Dilation Transformer with Cross-variable and Cross-feature Dependency for Time Series Imputation
abstract
In the time series imputation task, most Transformer-based methods adopt the standard full attention mechanism, which not only has high complexity but also cannot aggregate semantic multi-scale information effectively. This paper proposes a Transformer-based model called TSIformer for time series imputation. TSIformer proposes a Sequence Multi-Scale Dilated Attention mechanism that assigns distinct dilation rates to different attention heads to enable the ability of multi-scale representation learning for series data. Self-attention is performed over temporal segments, through a dilated sliding window to capture dependencies within the sequence. TSIformer further introduces a Dual-Path Convolutional Feed-Forward Network to replace the traditional Feed-Forward Network in vanilla Transformer to better capture both the cross-feature and cross-variable dependency. Through extensive experiments on multiple benchmarks, TSIformer demonstrates its effectiveness and robustness.
Haozheng Yang, Xuelin Cheng, Runjie Zhao, Botao Wu, Xince Chen
ICASSP5
2025 WPMnet: Time Series Forecasting via Wavelet Patch Multi-Scale
abstract
Time series forecasting has important applications in fields like weather forecasting and traffic planning. Recent research questions the advantages of Transformers in long-term prediction, suggesting that simpler models, such as Multi-Layer Perceptrons (MLP), can perform similarly or even better. Inspired by image patch methods in the vision domain, we aim to extract richer semantic information from a broader perspective. However, existing approaches fail to capture multi-scale temporal features within image patches. This paper proposes WPMnet, a novel MLP-based model that combines multi-scale feature extraction and fusion techniques to predict both long-term trends and fine-grained details. We also introduce Discrete Wavelet Transform (DWT) to decompose time series data into components at different frequencies. This helps capture both the details and trends of the signal, improving the model’s ability to handle non-stationary signals. Experimental results on six real-world datasets show that WPMnet outperforms state-of-the-art methods, significantly improving prediction accuracy.
Botao Wu, Xuelin Cheng, Haozheng Yang, Runjie Zhao, Chang You
IJCNN1
2025 Automatic Generation of Mappings for Distributed Fourier Operations
abstract
The Fourier transform is an ubiquitous mathematical operation used in a multitude of scientific applications. Most distributed Fourier transform libraries provide rigid implementations that force developers of high performance applications to mold their code around the Fourier computation, omitting opportunities for minimizing communication across the Fourier transforms and the surrounding computation. In this work, we introduce a new automatic approach to generate distributed mappings for multi-dimensional Fourier operations, offering a solution to this problem. Our approach decides how to decompose, map, and schedule the computation as smaller and lower-dimensional parallel operations. We design and implement a novel non-linear iterative formulation that optimizes across Fourier and linear algebra operations. Our scheme leverages the Z3 SMT solver to minimize the number of communication steps across key MPI collectives, while selecting the grid shape. We evaluate the effectiveness of our new scheme and demonstrate 2 × -31 × speedups over coupled heFFTe and COSMA solutions.
Doru-Thom Popovici, Botao Wu, John Shalf, Martin Kong
SC2
2025 Denoising dual sparse graph attention model for session-based recommendation
abstract
Nowadays, session-based recommendation plays an increasingly important role in the e-commerce field, which predicts the item that a user may click next time based on the sequence of user clicks. However, in real-world scenarios, due to various factors, there is noise in the process of user click behavior. For example, an unexpected click may not be the user's true intention and thus affect the user's behavior prediction. The current denoising methods have the following challenges: denoising directly from a single user's click sequence is not sufficient, the impact of unexpected clicks is not fully considered, and the additional information of other users is not fully utilized. To address these challenges, a new denoising dual sparse graph attention model for session-based recommendation abbreviated as DDSG, which not only considers the information in the current session, but also utilizes the information outside the current session. In the current session, this paper uses position encoding and gated neural networks that are biased towards frequency information to obtain the initial embedding, and uses the self-attention mechanism to model the target representation. In terms of denoising, we first iteratively denoise the representation obtained in the session, and perform sparse self-attention denoising on the session representation based on the target representation. Outside the current session, we take other sessions with similar interests to the current session to enhance the current session. Finally, the user's next click item is predicted by combining internal and extra information of the session. Experiments on three e-commerce datasets demonstrate our model exceeded the optimal SOTA model by 107%, achieving the highest performance and verifying the effectiveness of our model.
Botao Wu
Discov. Comput.1
2022 A GPU-based multilevel additive schwarz preconditioner for cloth and deformable body simulation
abstract
In this paper, we wish to push the limit of real-time cloth and deformable body simulation to a higher level with 50K to 500K vertices, based on the development of a novel GPU-based multilevel additive Schwarz (MAS) pre-conditioner. Similar to other preconditioners under the MAS framework, our preconditioner naturally adopts multilevel and domain decomposition concepts. But contrary to previous works, we advocate the use of small, non-overlapping domains that can well explore the parallel computing power on a GPU. Based on this idea, we investigate and invent a series of algorithms for our preconditioner, including multilevel domain construction using Morton codes, low-cost matrix precomputation by one-way Gauss-Jordan elimination, and conflict-free symmetric-matrix-vector multiplication in runtime preconditioning. The experiment shows that our preconditioner is effective, fast, cheap to precompute and scalable with respect to stiffness and problem size. It is compatible with many linear and nonlinear solvers used in cloth and deformable body simulation with dynamic contacts, such as PCG, accelerated gradient descent and L-BFGS. On a GPU, our preconditioner speeds up a PCG solver by approximately a factor of four, and its CPU version outperforms a number of competitors, including ILU0 and ILUT.
Botao Wu, Zhendong Wang 0001, Huamin Wang 0001
ACM Trans. Graph.1
2021 A Safe and Fast Repulsion Method for GPU-based Cloth Self Collisions
abstract
Cloth dynamics and collision handling are the two most challenging topics in cloth simulation. While researchers have substantially improved the performances of cloth dynamics solvers recently, their success in fast collision detection and handling is rather limited. In this article, we focus our research on the safety, efficiency, and realism of the repulsion-based collision handling approach, which has demonstrated its potential in existing GPU-based simulators. Our first discovery is the necessary vertex distance conditions for cloth to enter self intersections, the negations of which can be viewed as vertex distance constraints continuous in time for sufficiently avoiding self collisions. Continuous constraints, however, cannot be enforced with ease. Our solution is to convert continuous constraints into three types of constraints: discrete edge length constraints, discrete vertex distance constraints, and vertex displacement constraints. Based on this solution, we develop a fast and safe collision handling process for enforcing constraints, a novel splitting method for integrating collision handling with dynamics solvers, and static and adaptive remeshing schemes to further improve the runtime performance. In summary, our cloth simulator is efficient, safe, robust, and parallelizable on a GPU. The experiment shows that it runs at least one order of magnitude faster than existing simulators.
Longhua Wu, Botao Wu, Yin Yang 0002, Huamin Wang 0001
ACM Trans. Graph.2