VLDB 2026 Research / reviewers in the wild / expert
Jiaqi Dai
dblp:330/0988
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open-World Assembly of Damaged Fragments: An Algorithm-Driven Framework for Dunhuang Manuscripts
Ming-Kun Chen, Xiaokang Zhao, Yanping Xiang, Jiaqi Dai, Zeli Tong, Langtai Cheng, Yutong Zheng |
ICIC (21) | 6 |
| 2025 | Matching Ancient Dunhuang Manuscripts Based on Multi-dimensional Feature Fusion
Yanping Xiang, Jiaqi Dai, Ming-Kun Chen, Teer Song, Yutong Zheng |
KSEM (4) | 2 |
| 2024 | PURPLE: Making a Large Language Model a Better SQL WriterabstractLarge Language Model (LLM) techniques play an increasingly important role in Natural Language to SQL (NL2SQL) translation. LLMs trained by extensive corpora have strong natural language understanding and basic SQL generation abilities without additional tuning specific to NL2SQL tasks. Existing LLMs-based NL2SQL approaches try to improve the translation by enhancing the LLMs with an emphasis on user intention understanding. However, LLMs sometimes fail to generate appropriate SQL due to their lack of knowledge in organizing complex logical operator composition. A promising method is to input the LLMs with demonstrations, which include known NL2SQL translations from various databases. LLMs can learn to organize operator compositions from the input demonstrations for the given task. In this paper, we propose PURPLE (Pre-trained models Utilized to Retrieve Prompts for Logical Enhancement), which improves accuracy by retrieving demonstrations containing the requisite logical operator composition for the NL2SQL task on hand, thereby guiding LLMs to produce better SQL translation. PURPLE achieves a new state-of-the-art performance of 80.5% exact-set match accuracy and 87.8% execution match accuracy on the validation set of the popular NL2SQL benchmark Spider. PURPLE maintains high accuracy across diverse benchmarks, budgetary constraints, and various LLMs, showing robustness and cost-effectiveness. Tonghui Ren, Yuankai Fan, Zhenying He, Ren Huang, Jiaqi Dai, Can Huang 0003, Yinan Jing, Kai Zhang 0006, Yifan Yang 0001, Xiaoyang Sean Wang |
ICDE | 5 |
| 2024 | Personalized Image Aesthetics Assessment Based on Theme and Personality
Jiaqi Dai, Fanzhen Liu, Ronghua Huang |
KSEM (5) | 1 |
| 2023 | An automatic methodology for full dentition maturity staging from OPG images using deep learning
Wenxuan Dong, Meng You, Tao He 0016, Jiaqi Dai, Yueting Tang, Yuchao Shi, Jixiang Guo |
Appl. Intell. | 4 |
| 2022 | Performance evaluation of computational methods for splice-disrupting variants and improving the performance using the machine learning-based frameworkabstractA critical challenge in genetic diagnostics is the assessment of genetic variants associated with diseases, specifically variants that fall out with canonical splice sites, by altering alternative splicing. Several computational methods have been developed to prioritize variants effect on splicing; however, performance evaluation of these methods is hampered by the lack of large-scale benchmark datasets. In this study, we employed a splicing-region-specific strategy to evaluate the performance of prediction methods based on eight independent datasets. Under most conditions, we found that dbscSNV-ADA performed better in the exonic region, S-CAP performed better in the core donor and acceptor regions, S-CAP and SpliceAI performed better in the extended acceptor region and MMSplice performed better in identifying variants that caused exon skipping. However, it should be noted that the performances of prediction methods varied widely under different datasets and splicing regions, and none of these methods showed the best overall performance with all datasets. To address this, we developed a new method, machine learning-based classification of splice sites variants (MLCsplice), to predict variants effect on splicing based on individual methods. We demonstrated that MLCsplice achieved stable and superior prediction performance compared with any individual method. To facilitate the identification of the splicing effect of variants, we provided precomputed MLCsplice scores for all possible splice sites variants across human protein-coding genes (http://39.105.51.3:8090/MLCsplice/). We believe that the performance of different individual methods under eight benchmark datasets will provide tentative guidance for appropriate method selection to prioritize candidate splice-disrupting variants, thereby increasing the genetic diagnostic yield. Hao Liu 0089, Jiaqi Dai, Chunxia Zhao, Dao Wen Wang |
Briefings Bioinform. | 2 |