VLDB 2026 Research / reviewers in the wild / expert
Linxiao Bai
dblp:221/6002
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-9200-7060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | μScope: Evaluating storage stack robustness against SSD's latency variation
Linxiao Bai, Shanshan Li 0001, Zhouyang Jia, Yu Jiang 0001, Yuanliang Zhang, Zichen Xu 0001, Bin Lin 0011, Si Zheng 0003, Xiangke Liao |
J. Syst. Archit. | 1 |
| 2024 | ModelCS: A Two-Stage Framework for Model SearchabstractIn the open-source community, selecting models that meet user requirements and data distributions is essential due to numerous models with unique characteristics. However, existing model search methods often fail to meet diverse user requirements, varied data distributions, and have slow search speeds. To address these issues, we introduce ModelCS, a two-stage framework for model search based on recall-ranking. Its key idea is to preliminary screening of numerous models using representation learning and then precise ranking of selected ones. Specifically, we study model feature extraction and representation methods. We construct a dataset for this study and propose a rule-based data augmentation method to enhance its diversity. Based on the augmented dataset, we conduct an empirical study and propose the multidimensional feature representation, which influences the design of ModelCS. The recall stage of ModelCS involves a preliminary screening method based on the multidimensional feature representation, while the ranking stage of ModelCS involves a ranking method based on the extension to an existing method. We evaluate ModelCS on the multi-task model zoo in the PaddlePaddle framework. Experimental results indicate that ModelCS can reduce search time by up to 500 times and improve search effectiveness by up to 13.27 % compared to existing methods. Lingjun Zhao, Zhouyang Jia, Linxiao Bai |
APSEC | 5 |
| 2024 | LatVision: Modeling and Predicting Persisting Tail Latency in SSDsabstractAs Solid State Drives (SSDs) continue to evolve, the presence of tail latency within these devices remains a significant issue that can adversely affect overall performance. Various factors contribute to the emergence of tail latency spikes in SSDs. Current software-level management solutions primarily focus on the performance prediction of individual I/O operations, recognizing that persistent slow operations are prevalent in SSDs and tend to have a more pronounced impact. In this paper, we build a tool-LatVision to obtain I/O-related data directly from the kernel to predict persisting tail latency in SSDs by a neural network model. We conduct a comprehensive comparison and analysis of the input metrics and predictive models employed. Furthermore, we enhance LatVision’s performance through the application of heuristic algorithms. Through LatVision, we achieve real-time, lightweight, and high-accuracy performance prediction for low-latency SSDs. Linxiao Bai, Zhijie Jiang, Yuanliang Zhang, Xiangbing Huang, Wang Li 0003, Bin Lin 0011 |
HPCC | 1 |
| 2023 | deGraphCS: Embedding Variable-based Flow Graph for Neural Code SearchabstractWith the rapid increase of public code repositories, developers maintain a great desire to retrieve precise code snippets by using natural language. Despite existing deep learning-based approaches that provide end-to-end solutions (i.e., accept natural language as queries and show related code fragments), the performance of code search in the large-scale repositories is still low in accuracy because of the code representation (e.g., AST) and modeling (e.g., directly fusing features in the attention stage). In this paper, we propose a novel learnable de ep G raph for C ode S earch (called deGraphCS ) to transfer source code into variable-based flow graphs based on an intermediate representation technique, which can model code semantics more precisely than directly processing the code as text or using the syntax tree representation. Furthermore, we propose a graph optimization mechanism to refine the code representation and apply an improved gated graph neural network to model variable-based flow graphs. To evaluate the effectiveness of deGraphCS , we collect a large-scale dataset from GitHub containing 41,152 code snippets written in the C language and reproduce several typical deep code search methods for comparison. The experimental results show that deGraphCS can achieve state-of-the-art performance and accurately retrieve code snippets satisfying the needs of the users. Yue Yu 0001, Shanshan Li 0001, Xin Xia 0001, Mingyang Geng, Linxiao Bai, Wei Dong 0006, Xiangke Liao |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2018 | Does Reciprocal Gratefulness in Twitter Predict Neighborhood Safety?: Comparing 911 Calls Where Users Reside or Use Social Media
Ann Marie White, Linxiao Bai, Christopher Homan, Melanie Funchess, Catherine Cerulli, Amen Ptah, Deepak Pandita, Henry A. Kautz |
ICWSM | 2 |