Baole Ai

dblp:199/8336 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0006-2497-7328ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 62% Language models and text generation · 27% Graph learning · 12%
Databases, data mining, and information retrieval
2 papers
Machine learning and data management · 29% Indexing and storage engines · 29% Query processing and optimization · 29%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 37% Parallel and multicore computing · 37% Storage systems · 22%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model fine-tuning
0.912025
SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning · AAAI 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning · AAAI 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning · AAAI 2025
Indexing and storage engines › caching
cache management
0.912025
CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution · ICDE 2025
Machine learning and data management › deep learning
graph neural network training
0.912025
CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution · ICDE 2025
Query processing and optimization › query execution
pipelining
0.912025
CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution · ICDE 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
graph neural network accelerator
0.912025
Voyager: Input-Adaptive Algebraic Transformations for High-Performance Graph Neural Networks · ASPLOS (3) 2025
Machine learning › Graph learning
graph neural network
0.412019
AliGraph: A Comprehensive Graph Neural Network Platform · Proc. VLDB Endow. 2019
Graph data management
graph storage
0.412019
AliGraph: A Comprehensive Graph Neural Network Platform · Proc. VLDB Endow. 2019
Machine learning › Efficient and distributed learning
distributed training
0.312025
SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning · AAAI 2025
Storage systems › i/o optimization
disk i/o optimization
0.312025
CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution · ICDE 2025
Storage systems
flash and SSD
0.312025
CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution · ICDE 2025
Distributed systems
distributed graph processing
0.112019
AliGraph: A Comprehensive Graph Neural Network Platform · Proc. VLDB Endow. 2019

Methods — techniques the papers use, named apart from their topics

pipelining · 1.7caching · 1.7auto-tuning · 1.7optimized sampling operators · 1.1distributed graph storage · 1.1caching strategy · 1.1quantization · 0.9parameter-efficient fine-tuning · 0.9operator reordering · 0.9operator fusion · 0.9multimodal training · 0.9algebraic transformation · 0.9
YearPublicationVenuePosition
2025 SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning
abstract
Recent development in Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) have achieved superior performance and generalization capabilities, covered extensive areas of traditional tasks. However, existing large model training frameworks support only a limited number of models and techniques, particularly lacking in support for new models, which makes fine-tuning LLMs challenging for most developers. Therefore, we develop SWIFT, a customizable one-stop infrastructure for large models. With support of over 350+ LLMs and 80+ MLLMs, SWIFT stands as the open-source framework that provide the most comprehensive support for fine-tuning large models. In particular, it is the first training framework that provides systematic support for MLLMs. Moreover, SWIFT integrates post-training processes such as inference, evaluation, and quantization, to facilitate fast adoptions of large models in various application scenarios, offering helpful utilities like benchmark comparisons among different training techniques.
Yuze Zhao, Jinghan Hu, Yunlin Mao, Daoze Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, Yingda Chen
AAAI9
2025 Voyager: Input-Adaptive Algebraic Transformations for High-Performance Graph Neural Networks
abstract
Graph neural networks (GNNs) are gaining popularity in diverse application domains and growing in complexity.As a result, it is crucial to achieve high-performance GNN execution.Among various techniques, algebraic transformations, including operator reordering and operator fusion, have been successfully applied to improve the computation and memory access efficiencies of DNN models.However,
Yangjie Zhou 0001, Wenting Shen, Jingwen Leng, Shuwen Lu, Zihan Liu 0002, Weihao Cui, Zhendong Zhang 0004, Wencong Xiao, Baole Ai, Yong Li 0045, Wei Lin 0016, Deze Zeng, Yun Liang 0001, Quan Chen 0001, Ning Liu 0007, Minyi Guo
ASPLOS (3)9
2025 CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and Execution
abstract
Graph neural networks (GNNs) have proven to be powerful tools for learning from graph-structured data and have achieved great success in many applications. As the sizes of real-world graphs continue to grow, traditional GNN training methods face significant scalability challenges. Recently, disks have gained attention as a cost-effective solution to store large-scale graphs, and several disk-based GNN systems have been proposed to train large-scale graphs on a single machine. However, these systems either overlook the unique data characteristics of GNN workloads when designing cache plans or fail to fully exploit the multilevel hierarchy of storage and computation in system execution, thus resulting in disk I/O bottleneck and resource under-utilization. To address these issues, we present CaliEX, an advanced disk-based GNN system that employs joint optimizations of caching and execution within and across different training stages. CaliEX first designs tailored cache plans and execution policy for both graph topology and features to accelerate neighborhood sampling and feature gathering. Since these two training stages work on different types of data, CaliEX further auto-tunes the cache allocation and pipelines the execution across different stages to improve resource utilization and overall training throughput. Evaluations on multiple GNN models and various large-scale datasets show that CaliEX achieves 3.28 × speedup on average compared to existing disk-based GNN training systems.
Can Su, Haipeng Zhang 0006, Wenting Shen, Baole Ai, Yong Li 0045, Kaigui Bian, Bin Cui 0001
ICDE5
2021 Graph Sampling with Fast Random Walker on HBM-enabled FPGA Accelerators
abstract
Graph neural networks (GNNs) have gained increasing popularity among researchers recently and have been employed in many applications. Training GNNs introduce a crucial stage called graph sampling. One of the most important sampling algorithms is Random Walk. However, Random Walk and many of its variants share and suffer from the same performance problem caused by random and fragmented memory access patterns, leading to significant system performance degradation. In this work, we present an efficient graph sampling engine on modern FPGAs integrated with in-package high bandwidth memory (HBM), which brings data closer and faster to the core logic. The hardware walker design is modular and easily scalable for massive parallelism, to fully utilize the available HBM channels. Our design also provides the flexibility to support Random Walk and two of its variants on both homogeneous and heterogeneous graphs. On real-world graph datasets, we achieve a 1.39 × -3.74 × speedup with a 2.42 × -6.69 × higher energy efficiency over highly optimized parallel baselines on a Xeon CPU. We also implement these algorithms on a NVIDIA Tesla VIOO GPU and achieve comparable dynamic energy consumption.
Chunyou Su, Hao Liang 0003, Wei Zhang 0012, Baole Ai, Wenting Shen, Zeke Wang
FPL5
2019 AliGraph: A Comprehensive Graph Neural Network Platform
abstract
An increasing number of machine learning tasks require dealing with large graph datasets, which capture rich and complex relationship among potentially billions of elements. Graph Neural Network (GNN) becomes an effective way to address the graph learning problem by converting the graph data into a low dimensional space while keeping both the structural and property information to the maximum extent and constructing a neural network for training and referencing. However, it is challenging to provide an efficient graph storage and computation capabilities to facilitate GNN training and enable development of new GNN algorithms. In this paper, we present a comprehensive graph neural network system, namely AliGraph , which consists of distributed graph storage, optimized sampling operators and runtime to efficiently support not only existing popular GNNs but also a series of in-house developed ones for different scenarios. The system is currently deployed at Alibaba to support a variety of business scenarios, including product recommendation and personalized search at Alibaba's E-Commerce platform. By conducting extensive experiments on a real-world dataset with 492.90 million vertices, 6.82 billion edges and rich attributes, AliGraph performs an order of magnitude faster in terms of graph building (5 minutes vs hours reported from the state-of-the-art PowerGraph platform). At training, AliGraph runs 40%-50% faster with the novel caching strategy and demonstrates around 12 times speed up with the improved runtime. In addition, our in-house developed GNN models all showcase their statistically significant superiorities in terms of both effectiveness and efficiency (e.g., 4.12%--17.19% lift by F1 scores).
Hongxia Yang, Wei Lin 0016, Chang Zhou 0005, Baole Ai, Yong Li 0045, Jingren Zhou 0001
Proc. VLDB Endow.6
2017 Human Pose Estimation Using Deep Structure Guided Learning
abstract
In this paper, we propose a novel approach to incorporate structure knowledge into Convolutional Neural Networks (CNNs) for articulated human pose estimation from a single still image. Recent research on pose estimation adopt CNNs as base blocks to combine with other graphical models. Different from existing methods using features from CNNs to model the tree structure, we directly use the structure pose prior to guide the learning of CNN. First, we introduce a deep CNN with effective receptive fields which capture the holistic context of the whole image. Second, limb loss is used as intermediate supervision of CNN to learn the correlations of joints. Both parts and joints features are extracted in the middle of neural network and then are used to guide the following network learning. The proposed framework can exploit an implicit structure model of human body. Only using one stage and without any complex post processing, our method achieves state-of-art results on both FLIC and LSP benchmarks.
Baole Ai, Yu Zhou 0007, Yao Yu 0001, Sidan Du
WACV1