VLDB 2026 Research / reviewers in the wild / expert
Shihao Bai
dblp:246/3127
· DBLP profile ↗
16ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-6885-8423ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Efficient and distributed learning · 29% Language models and text generation · 14% Image recognition and object detection · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Embedded and real-time systems · 38% Cloud and datacenter computing · 38% GPUs and heterogeneous computing · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 26 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.5 | 2 | 2025 | AtomNet: Designing Tiny Models from Operators Under Extreme MCU Constraints · AAAI 2025 Lossy and Lossless (L2) Post-training Model Size Compression · ICCV 2023 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model |
1.0 | 1 | 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
inference acceleration |
1.0 | 1 | 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model inference |
1.0 | 1 | 2026 | PiLLM: Resource-Efficient LLM Inference Using Workload Prediction · EuroSys 2026 |
Machine learning › Efficient and distributed learning
resource allocation |
1.0 | 1 | 2026 | PiLLM: Resource-Efficient LLM Inference Using Workload Prediction · EuroSys 2026 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.9 | 2 | 2020 | Few-shot Visual Learning with Contextual Memory and Fine-grained Calibration · IJCAI 2020 Transductive Relation-Propagation Network for Few-shot Learning · IJCAI 2020 |
Natural language and speech › Language models and text generation › text generation
structured generation |
0.9 | 1 | 2025 | Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation · ACL (1) 2025 |
Machine learning and data management
inference serving |
0.9 | 1 | 2025 | Past-Future Scheduler for LLM Serving under SLA Guarantees · ASPLOS (2) 2025 |
Compilers and program optimization
parsing |
0.9 | 1 | 2025 | Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation · ACL (1) 2025 |
Cloud and datacenter computing
inference serving |
0.9 | 1 | 2025 | Past-Future Scheduler for LLM Serving under SLA Guarantees · ASPLOS (2) 2025 |
Embedded and real-time systems › embedded software engineering › embedded software development
microcontroller deployment |
0.9 | 1 | 2025 | AtomNet: Designing Tiny Models from Operators Under Extreme MCU Constraints · AAAI 2025 |
Computer vision › Video understanding and tracking
spatiotemporal predictive learning |
0.8 | 1 | 2024 | SeeMore: a spatiotemporal predictive model with bidirectional distillation and level-specific meta-adaptation · Sci. China Inf. Sci. 2024 |
Computer vision › Image recognition and object detection › object detection
few-shot object detection |
0.7 | 1 | 2023 | Temporal Speciation Network for Few-Shot Object Detection · IEEE Trans. Multim. 2023 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | Temporal Speciation Network for Few-Shot Object Detection · IEEE Trans. Multim. 2023 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.7 | 1 | 2023 | A meaningful learning method for zero-shot semantic segmentation · Sci. China Inf. Sci. 2023 |
Computer vision › Segmentation and scene understanding › semantic segmentation › open-vocabulary segmentation
zero-shot semantic segmentation |
0.7 | 1 | 2023 | A meaningful learning method for zero-shot semantic segmentation · Sci. China Inf. Sci. 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › abstract reasoning
abstract visual reasoning |
0.5 | 1 | 2021 | Stratified Rule-Aware Network for Abstract Visual Reasoning · AAAI 2021 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot image classification |
0.4 | 1 | 2020 | Few-shot Visual Learning with Contextual Memory and Fine-grained Calibration · IJCAI 2020 |
Machine learning › Graph learning
graph neural network |
0.4 | 1 | 2020 | Transductive Relation-Propagation Network for Few-shot Learning · IJCAI 2020 |
Image and video processing › image restoration
image inpainting |
0.4 | 1 | 2019 | Coarse-to-Fine Image Inpainting via Region-wise Convolutions and Non-Local Correlation · IJCAI 2019 |
GPUs and heterogeneous computing
GPU resource management |
0.3 | 1 | 2026 | PiLLM: Resource-Efficient LLM Inference Using Workload Prediction · EuroSys 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.2 | 1 | 2024 | SeeMore: a spatiotemporal predictive model with bidirectional distillation and level-specific meta-adaptation · Sci. China Inf. Sci. 2024 |
Computer vision › Image recognition and object detection › object detection › object proposal generation
region proposal network |
0.2 | 1 | 2023 | Temporal Speciation Network for Few-Shot Object Detection · IEEE Trans. Multim. 2023 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2021 | Stratified Rule-Aware Network for Abstract Visual Reasoning · AAAI 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
relation embedding |
0.1 | 1 | 2020 | Few-shot Visual Learning with Contextual Memory and Fine-grained Calibration · IJCAI 2020 |
Image and video processing
image restoration |
0.1 | 1 | 2019 | Coarse-to-Fine Image Inpainting via Region-wise Convolutions and Non-Local Correlation · IJCAI 2019 |
Methods — techniques the papers use, named apart from their topics
workload prediction · 2.0dynamic resource allocation · 2.0past-future scheduling · 1.7operator profiling · 1.7neural architecture search · 1.7memory requirement estimation · 1.7constrained decoding · 1.7meta-learning · 1.2confidence-guided context focusing · 1.0bidirectional distillation · 0.8lossless compression · 0.7differentiable counter · 0.7region-wise convolution · 0.4non-local operation · 0.4coarse-to-fine framework · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context FocusingabstractLingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang, Ao Zhou, Jianlei Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang 0004, Jianlei Yang 0001 |
ACL (1) | 3 |
| 2026 | PiLLM: Resource-Efficient LLM Inference Using Workload PredictionabstractLLM inference demands substantial GPU resources, but its highly variable workload characteristics create efficiency challenges at both inter-GPU and intra-GPU levels. To meet Service Level Objectives (SLOs), existing systems typically overprovision resources in two ways: allocating excess GPUs to handle peak loads and reserving excessive memory per request to prevent out-of-memory during token generation. We introduce PiLLM (Predictable inference for LLMs), a system that addresses these inefficiencies through accurate workload prediction and dynamic resource allocation. Yunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang |
EuroSys | 2 |
| 2025 | AtomNet: Designing Tiny Models from Operators Under Extreme MCU ConstraintsabstractTiny machine learning (TinyML) has attracted heightened attention for its ability to provide low-cost and instantaneous performance on edge devices. Particularly, the commonly used microcontroller unit (MCU) imposes extreme constraints on peak memory (SRAM) and storage (Flash). Existing TinyML methods often rely on a customized and hard-to-obtain inference libraries, as well as necessitate a time-consuming search for a deployable architecture using advanced Neural Architecture Search (NAS) algorithms. To solve these problems, we fully exploit the resources on MCU and deduce hardware-oriented guidelines for designing models under extreme MCU constraints. In detail, we delve into thorough information about the atom operators by collecting the runtime data of Flash, SRAM, and latency to build a dataset named AtomDB. Based on AtomDB, several critical operator guidelines are established to fully utilize limited Flash and SRAM, while minimizing latency. By transferring the guidelines to analyze blocks, we propose a hybrid pattern that organizes appropriate blocks at different network stages to form the AtomNet, a more hardware-oriented architecture, to handle the former SRAM bottleneck and the latter Flash bottleneck. Extensive experiments demonstrate the effectiveness of the exploitation of the hardware characteristics. Remarkably, AtomNet pioneeringly achieve 3.5% accuracy enhancement and more than 15% latency reduction on 320KB MCU using readily available official inference libraries for ImageNet tasks, surpassing the current state-of-the-art method. Zhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei, Jinyang Guo 0002, Ruihao Gong, Song-Lu Chen, Xianglong Liu 0001, Xu-Cheng Yin |
AAAI | 3 |
| 2025 | Pre³: Enabling Deterministic Pushdown Automata for Faster Structured LLM GenerationabstractJunyi Chen, Shihao Bai, Zaijun Wang, Siyu Wu, Chuheng Du, Hailong Yang, Ruihao Gong, Shengzhong Liu, Fan Wu, Guihai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shihao Bai, Zaijun Wang, Siyu Wu 0001, Chuheng Du, Hailong Yang 0002, Ruihao Gong, Shengzhong Liu, Fan Wu 0006, Guihai Chen |
ACL (1) | 2 |
| 2025 | Past-Future Scheduler for LLM Serving under SLA GuaranteesabstractThe exploration and application of Large Language Models (LLMs) is thriving. To reduce deployment costs, continuous batching has become an essential feature in current service frameworks. The effectiveness of continuous batching relies on an accurate estimate of the memory requirements of requests. However, due to the diversity in request output lengths, existing frameworks tend to adopt aggressive or conservative schedulers, which often result in significant overestimation or underestimation of memory consumption. Consequently, they suffer from harmful request evictions or prolonged queuing times, failing to achieve satisfactory throughput under strict Service Level Agreement (SLA) guarantees (a.k.a. goodput), across various LLM application scenarios with differing input-output length distributions. To address this issue, we propose a novel Past-Future scheduler that precisely estimates the peak memory resources required by the running batch via considering the historical distribution of request output lengths and calculating memory occupancy at each future time point. It adapts to applications with all types of input-output length distributions, balancing the trade-off between request queuing and harmful evictions, thereby consistently achieving better goodput. Furthermore, to validate the effectiveness of the proposed scheduler, we developed a high-performance LLM serving framework, LightLLM, that implements the Past-Future scheduler. Compared to existing aggressive or conservative schedulers, LightLLM demonstrates superior goodput, achieving up to 2-3× higher goodput than other schedulers under heavy loads. LightLLM is open source to boost the research in such direction (https://github.com/ModelTC/lightllm). Ruihao Gong, Shihao Bai, Siyu Wu 0001, Yunqian Fan, Zaijun Wang, Hailong Yang 0002, Xianglong Liu 0001 |
ASPLOS (2) | 2 |
| 2025 | Inversed Pyramid Network with Spatial-adapted and Task-oriented Tuning for few-shot learning
Duorui Wang, Shihao Bai, Shuo Wang 0008, Yajun Gao, Yuqing Ma, Xianglong Liu 0001 |
Pattern Recognit. | 3 |
| 2024 | SeeMore: a spatiotemporal predictive model with bidirectional distillation and level-specific meta-adaptation
Yuqing Ma, Wei Liu 0005, Yajun Gao, Shihao Bai, Haotong Qin, Xianglong Liu 0001 |
Sci. China Inf. Sci. | 5 |
| 2023 | Lossy and Lossless (L2) Post-training Model Size CompressionabstractDeep neural networks have delivered remarkable performance and have been widely used in various visual tasks. However, their huge sizes cause significant inconvenience for transmission and storage. Many previous studies have explored model size compression. However, these studies often approach various lossy and lossless compression methods in isolation, leading to challenges in achieving high compression ratios efficiently. This work proposes a post-training model size compression method that combines lossy and lossless compression in a unified way. We first propose a unified parametric weight transformation, which ensures different lossy compression methods can be performed jointly in a post-training manner. Then, a dedicated differentiable counter is introduced to guide the optimization of lossy compression to arrive at a more suitable point for later lossless compression. Additionally, our method can easily control a desired global compression ratio and allocate adaptive ratios for different layers. Finally, our method can achieve a stable 10× compression ratio without sacrificing accuracy and a 20× compression ratio with minor accuracy loss in a short time. Our code is available at https://github.com/ModelTC/L2_Compression. Yumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong, Jianlei Yang 0001 |
ICCV | 2 |
| 2023 | A meaningful learning method for zero-shot semantic segmentation
Xianglong Liu 0001, Shihao Bai, Shan An, Shuo Wang 0008, Wei Liu 0005, Yuqing Ma |
Sci. China Inf. Sci. | 2 |
| 2023 | Regionwise Generative Adversarial Image Inpainting for Large Missing AreasabstractRecently, deep neural networks have achieved promising performance for in-filling large missing regions in image inpainting tasks. They have usually adopted the standard convolutional architecture over the corrupted image, leading to meaningless contents, such as color discrepancy, blur, and other artifacts. Moreover, most inpainting approaches cannot handle well the case of a large contiguous missing area. To address these problems, we propose a generic inpainting framework capable of handling incomplete images with both contiguous and discontiguous large missing areas. We pose this in an adversarial manner, deploying regionwise operations in both the generator and discriminator to separately handle the different types of regions, namely, existing regions and missing ones. Moreover, a correlation loss is introduced to capture the nonlocal correlations between different patches, and thus, guide the generator to obtain more information during inference. With the help of regionwise generative adversarial mechanism, our framework can restore semantically reasonable and visually realistic images for both discontiguous and contiguous large missing areas. Extensive experiments on three widely used datasets for image inpainting task have been conducted, and both qualitative and quantitative experimental results demonstrate that the proposed model significantly outperforms the state-of-the-art approaches, on the large contiguous and discontiguous missing areas. Yuqing Ma, Xianglong Liu 0001, Shihao Bai, Lei Wang 0018, Aishan Liu, Dacheng Tao, Edwin R. Hancock |
IEEE Trans. Cybern. | 3 |
| 2023 | Temporal Speciation Network for Few-Shot Object DetectionabstractRecently, few-shot object detection (FSOD) has become an increasing research focus, which can largely alleviate the heavy dependency on expensive annotations in the traditional object detection task. However, existing FSOD approaches fail to generate sufficient high-quality positive region proposals which are the key to detection performance, due to the lack of informative knowledge from base classes and non-specific alteration for novel classes. To address the problem, this paper presents a simple yet effective few-shot object detection framework referred to as Temporal Speciation Network (TeSNet) with an evolving training, which improves the diversity and rationality of positive proposal generation. Our TeSNet, imitating the natural evolution which relies on inheritation and mutation, correspondingly consists of two key components: a Selective Recombination Module (SRM) for effectively inheriting from base classes and a Mutational Region Proposal Network (MRPN) for flexibly mutating according to the unique traits of novel samples. Specifically, SRM selects and reorganizes relevant base categories, and further instantiates diverse individuals to ensure the diversity of positive proposals. MRPN adapts the parameters trained on base classes aiming for accurately locating positive proposals. Extensive experiments are conducted on several commonly-used datasets, in which our TeSNet achieves state-of-the-art results and outperforms baselines by large margin. Xianglong Liu 0001, Yuqing Ma, Shihao Bai, Zeyu Hao, Aishan Liu |
IEEE Trans. Multim. | 4 |
| 2022 | Transductive Relation-Propagation With Decoupling Training for Few-Shot LearningabstractFew-shot learning, aiming to learn novel concepts from one or a few labeled examples, is an interesting and very challenging problem with many practical advantages. Existing few-shot methods usually utilize data of the same classes to train the feature embedding module and in a row, which is unable to learn adapting to new tasks. Besides, traditional few-shot models fail to take advantage of the valuable relations of the support-query pairs, leading to performance degradation. In this article, we propose a transductive relation-propagation graph neural network (GNN) with a decoupling training strategy (TRPN-D) to explicitly model and propagate such relations across support-query pairs, and empower the few-shot module the ability of transferring past knowledge to new tasks via the decoupling training. Our few-shot module, namely TRPN, treats the relation of each support-query pair as a graph node, named relational node, and resorts to the known relations between support samples, including both intraclass commonality and interclass uniqueness. Through relation propagation, the model could generate the discriminative relation embeddings for support-query pairs. To the best of our knowledge, this is the first work that decouples the training of the embedding network and the few-shot graph module with different tasks, which might offer a new way to solve the few-shot learning problem. Extensive experiments conducted on several benchmark datasets demonstrate that our method can significantly outperform a variety of state-of-the-art few-shot learning methods. Yuqing Ma, Shihao Bai, Wei Liu 0005, Shuo Wang 0008, Yue Yu 0001, Xiao Bai 0001, Xianglong Liu 0001, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Stratified Rule-Aware Network for Abstract Visual ReasoningabstractAbstract reasoning refers to the ability to analyze information, discover rules at an intangible level, and solve problems in innovative ways. Raven's Progressive Matrices (RPM) test is typically used to examine the capability of abstract reasoning. The subject is asked to identify the correct choice from the answer set to fill the missing panel at the bottom right of RPM (e.g., a 3×3 matrix), following the underlying rules inside the matrix. Recent studies, taking advantage of Convolutional Neural Networks (CNNs), have achieved encouraging progress to accomplish the RPM test. However, they partly ignore necessary inductive biases of RPM solver, such as order sensitivity within each row/column and incremental rule induction. To address this problem, in this paper we propose a Stratified Rule-Aware Network (SRAN) to generate the rule embeddings for two input sequences. Our SRAN learns multiple granularity rule embeddings at different levels, and incrementally integrates the stratified embedding flows through a gated fusion module. With the help of embeddings, a rule similarity metric is applied to guarantee that SRAN can not only be trained using a tuplet loss but also infer the best answer efficiently. We further point out the severe defects existing in the popular RAVEN dataset for RPM test, which prevent from the fair evaluation of the abstract reasoning ability. To fix the defects, we propose an answer set generation algorithm called Attribute Bisection Tree (ABT), forming an improved dataset named Impartial-RAVEN (I-RAVEN for short). Extensive experiments are conducted on both PGM and I-RAVEN datasets, showing that our SRAN outperforms the state-of-the-art models by a considerable margin. Yuqing Ma, Xianglong Liu 0001, Yanlu Wei, Shihao Bai |
AAAI | 5 |
| 2020 | Transductive Relation-Propagation Network for Few-shot LearningabstractFew-shot learning, aiming to learn novel concepts from few labeled examples, is an interesting and very challenging problem with many practical advantages. To accomplish this task, one should concentrate on revealing the accurate relations of the support-query pairs. We propose a transductive relation-propagation graph neural network (TRPN) to explicitly model and propagate such relations across support-query pairs. Our TRPN treats the relation of each support-query pair as a graph node, named relational node, and resorts to the known relations between support samples, including both intra-class commonality and inter-class uniqueness, to guide the relation propagation in the graph, generating the discriminative relation embeddings for support-query pairs. A pseudo relational node is further introduced to propagate the query characteristics, and a fast, yet effective transductive learning strategy is devised to fully exploit the relation information among different queries. To the best of our knowledge, this is the first work that explicitly takes the relations of support-query pairs into consideration in few-shot learning, which might offer a new way to solve the few-shot learning problem. Extensive experiments conducted on several benchmark datasets demonstrate that our method can significantly outperform a variety of state-of-the-art few-shot learning methods. Yuqing Ma, Shihao Bai, Shan An, Wei Liu 0005, Aishan Liu, Xiantong Zhen, Xianglong Liu 0001 |
IJCAI | 2 |
| 2020 | Few-shot Visual Learning with Contextual Memory and Fine-grained CalibrationabstractFew-shot learning aims to learn a model that can be readily adapted to new unseen classes (concepts) by accessing one or few examples. Despite the successful progress, most of the few-shot learning approaches, concentrating on either global or local characteristics of examples, still suffer from weak generalization abilities. Inspired by the inverted pyramid theory, to address this problem, we propose an inverted pyramid network (IPN) that intimates the human's coarse-to-fine cognition paradigm. The proposed IPN consists of two consecutive stages, namely global stage and local stage. At the global stage, a class-sensitive contextual memory network (CCMNet) is introduced to learn discriminative support-query relation embeddings and predict the query-to-class similarity based on the contextual memory. Then at the local stage, a fine-grained calibration is further appended to complement the coarse relation embeddings, targeting more precise query-to-class similarity evaluation. To the best of our knowledge, IPN is the first work that simultaneously integrates both global and local characteristics in few-shot learning, approximately imitating the human cognition mechanism. Our extensive experiments on multiple benchmark datasets demonstrate the superiority of IPN, compared to a number of state-of-the-art approaches. Yuqing Ma, Wei Liu 0005, Shihao Bai, Aishan Liu, Weimin Chen 0002, Xianglong Liu 0001 |
IJCAI | 3 |
| 2019 | Coarse-to-Fine Image Inpainting via Region-wise Convolutions and Non-Local CorrelationabstractRecently deep neural networks have achieved promising performance for filling large missing regions in image inpainting tasks. They usually adopted the standard convolutional architecture over the corrupted image, where the same convolution filters try to restore the diverse information on both existing and missing regions, and meanwhile ignores the long-distance correlation among the regions. Only relying on the surrounding areas inevitably leads to meaningless contents and artifacts, such as color discrepancy and blur. To address these problems, we first propose region-wise convolutions to locally deal with the different types of regions, which can help exactly reconstruct existing regions and roughly infer the missing ones from existing regions at the same time. Then, a non-local operation is introduced to globally model the correlation among different regions, promising visual consistency between missing and existing regions. Finally, we integrate the region-wise convolutions and non-local correlation in a coarse-to-fine framework to restore semantically reasonable and visually realistic images. Extensive experiments on three widely-used datasets for image inpainting tasks have been conducted, and both qualitative and quantitative experimental results demonstrate that the proposed model significantly outperforms the state-of-the-art approaches, especially for the large irregular missing regions. Yuqing Ma, Xianglong Liu 0001, Shihao Bai, Lei Wang 0018, Dailan He, Aishan Liu |
IJCAI | 3 |