VLDB 2026 Research / reviewers in the wild / expert
Boyu Shi
dblp:311/9756
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0006-6412-2471ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GeoFA: A Geometric Finite Automaton Engine for Efficient Layout Pattern MatchingabstractAs Integrated Circuit (IC) design complexity increases exponentially, Layout Pattern Matching (LPM) has become a critical technique for physical verification, yet faces significant challenges in efficiency and accuracy. This paper presents a high-performance two-stage LPM framework centered around the innovative GeoFA (Geometric Finite Automaton) engine. The first stage successfully applies Deterministic Finite Automaton (DFA) theory for the first time to 2D layout pattern matching with geometric tolerances, enabling rapid candidate region screening with near-linear time complexity. The second stage then performs precise geometric verification, leveraging optimized spatial indexing, robust Boolean operations, and specific optimizations for cell arrays, to guarantee 100% sign-off accuracy. Experimental results demonstrate that the framework achieves 100% accuracy across diverse layouts. On standard benchmarks, it provides a speedup of approximately 1.7x over state-of-the-art academic work. Furthermore, compared to the industry-standard tool Calibre, it delivers average speedups of up to 201.7x and 22.9x on complex single-layer and cell-level matching tasks, respectively, demonstrating exceptional performance and scalability. Qingsheng Qiu, Ziwen Zheng, Boyu Shi |
ICCAD | 3 |
| 2024 | Building Variable-Sized Models via Learngene PoolabstractRecently, Stitchable Neural Networks (SN-Net) is proposed to stitch some pre-trained networks for quickly building numerous networks with different complexity and performance trade-offs. In this way, the burdens of designing or training the variable-sized networks, which can be used in application scenarios with diverse resource constraints, are alleviated. However, SN-Net still faces a few challenges. 1) Stitching from multiple independently pre-trained anchors introduces high storage resource consumption. 2) SN-Net faces challenges to build smaller models for low resource constraints. 3). SN-Net uses an unlearned initialization method for stitch layers, limiting the final performance. To overcome these challenges, motivated by the recently proposed Learngene framework, we propose a novel method called Learngene Pool. Briefly, Learngene distills the critical knowledge from a large pre-trained model into a small part (termed as learngene) and then expands this small part into a few variable-sized models. In our proposed method, we distill one pre-trained large model into multiple small models whose network blocks are used as learngene instances to construct the learngene pool. Since only one large model is used, we do not need to store more large models as SN-Net and after distilling, smaller learngene instances can be created to build small models to satisfy low resource constraints. We also insert learnable transformation matrices between the instances to stitch them into variable-sized models to improve the performance of these models. Exhaustive experiments have been implemented and the results validate the effectiveness of the proposed Learngene Pool compared with SN-Net. Boyu Shi, Shiyu Xia, Xu Yang 0021, Zhiqiang Kou, Xin Geng 0001 |
AAAI | 1 |
| 2024 | Exploiting Multi-Label Correlation in Label Distribution Learning
Zhiqiang Kou, Jing Wang 0113, Yuheng Jia, Boyu Shi, Xin Geng 0001 |
IJCAI | 5 |
| 2023 | Cutting Learned Index into Pieces: An In-depth Inquiry into Updatable Learned IndexesabstractNumerous high-performance updatable learned indexes have recently been designed to support the writing requirements in practical systems. Researchers have proposed various strategies to improve the availability of updatable learned indexes. However, it is unclear which strategy is more profitable. Therefore, we deconstruct the design of learned indexes into multiple dimensions and in-depth evaluate their impacts on the overall performance, respectively. Through the in-depth exploration of learned indexes, we reckon that the approximation algorithm is the most crucial design dimension for improving the performance of the learned indexes rather than the popular works that focus on the learned index structure. Moreover, this paper makes a comprehensive end-to-end evaluation based on a high-performance key-value store to answer people’s concerns about which learned index is better and whether learned indexes can outperform traditional ones. Finally, according to end-to-end and in-depth evaluation results, we give some constructive suggestions on designing a better learned index in these dimensions, especially how to design an excellent approximate algorithm to improve the lookup and insertion performance of learned indexes. Jiake Ge, Boyu Shi, Yanfeng Chai, Yuanhui Luo, Yunda Guo, Yinxuan He, Yunpeng Chai |
ICDE | 2 |
| 2023 | SALI: A Scalable Adaptive Learned Index Framework based on Probability ModelsabstractThe growth in data storage capacity and the increasing demands for high performance have created several challenges for concurrent indexing structures. One promising solution is the learned index, which uses a learning-based approach to fit the distribution of stored data and predictively locate target keys, significantly improving lookup performance. Despite their advantages, prevailing learned indexes exhibit constraints and encounter issues of scalability on multi-core data storage. This paper introduces SALI, the Scalable Adaptive Learned Index framework, which incorporates two strategies aimed at achieving high scalability, improving efficiency, and enhancing the robustness of the learned index. Firstly, a set of node-evolving strategies is defined to enable the learned index to adapt to various workload skews and enhance its concurrency performance in such scenarios. Secondly, a lightweight strategy is proposed to maintain statistical information within the learned index, with the goal of further improving the scalability of the index. Furthermore, to validate their effectiveness, SALI applied the two strategies mentioned above to the learned index structure that utilizes fine-grained write locks, known as LIPP. The experimental results have demonstrated that SALI significantly enhances the insertion throughput with 64 threads by an average of 2.04x compared to the second-best learned index. Furthermore, SALI accomplishes a lookup throughput similar to that of LIPP+. Jiake Ge, Huanchen Zhang, Boyu Shi, Yuanhui Luo, Yunda Guo, Yunpeng Chai, Yuxing Chen 0003, Anqun Pan |
Proc. ACM Manag. Data | 3 |
| 2022 | Edge-aware and spectral-spatial information aggregation network for multispectral image semantic segmentation
Di Zhang 0020, Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Rui Yao 0006 |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Multi-source collaborative enhanced for remote sensing images semantic segmentation
Jiaqi Zhao 0001, Di Zhang 0020, Boyu Shi, Yong Zhou 0003, Rui Yao 0006, Yong Xue |
Neurocomputing | 3 |
| 2021 | Multi-Stage Fusion and Multi-Source Attention Network for Multi-Modal Remote Sensing Image SegmentationabstractWith the rapid development of sensor technology, lots of remote sensing data have been collected. It effectively obtains good semantic segmentation performance by extracting feature maps based on multi-modal remote sensing images since extra modal data provides more information. How to make full use of multi-model remote sensing data for semantic segmentation is challenging. Toward this end, we propose a new network called Multi-Stage Fusion and Multi-Source Attention Network ((MS) 2 -Net) for multi-modal remote sensing data segmentation. The multi-stage fusion module fuses complementary information after calibrating the deviation information by filtering the noise from the multi-modal data. Besides, similar feature points are aggregated by the proposed multi-source attention for enhancing the discriminability of features with different modalities. The proposed model is evaluated on publicly available multi-modal remote sensing data sets, and results demonstrate the effectiveness of the proposed method. Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Jingsong Yang, Di Zhang 0020, Rui Yao 0006 |
ACM Trans. Intell. Syst. Technol. | 3 |